Device, method, and graphical user interface for displaying physical location view

By improving the user interface and methods, reducing user input, and combining visual and haptic feedback, the problems of low interaction efficiency and high energy consumption in virtual reality and augmented reality environments are solved, achieving more efficient user interaction and energy saving.

CN121532733APending Publication Date: 2026-02-13APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480037182.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-06-03
Filing Date
2024-05-31
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

In existing virtual reality and augmented reality environments, user interaction methods are cumbersome and inefficient, resulting in a heavy cognitive burden on users and high energy consumption, especially in battery-powered devices.

Method used

By providing improved user interfaces and methods, reducing the quantity and nature of user input, utilizing computer systems to detect user input and present virtual objects corresponding to their physical locations within a 3D environment, and combining visual, tactile, and audio feedback, interaction efficiency is improved and energy is saved.

Benefits of technology

It enables more intuitive and effective user interaction, reduces energy consumption, extends battery life, and improves device operability and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121532733A_ABST
    Figure CN121532733A_ABST
Patent Text Reader

Abstract

In some embodiments, a computer system displays a navigation user interface in a three-dimensional environment. The navigation user interface includes one or more first travel user interface elements displayed at a first immersion level corresponding to a first view of the first physical location, the one or more first travel user interface elements can be selected to change display of the navigation user interface from a first view of the first physical location to a second view of the second physical location.
Need to check novelty before this filing date? Find Prior Art

Description

Cross Reference to Related Applications

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 506,077, filed June 3, 2023, the contents of which are hereby incorporated by reference in their entirety for all purposes. TECHNICAL FIELD

[0002] The present disclosure relates generally to computer systems that provide computer-generated experiences, including but not limited to electronic devices that provide virtual reality and mixed reality experiences via a display. BACKGROUND

[0003] In recent years, there has been a significant increase in the development of computer systems for augmented reality. Example augmented reality environments include at least some virtual elements that replace or augment the physical world. Input devices for computer systems and other electronic computing devices, such as cameras, controllers, joysticks, touch-sensitive surfaces, and touch screen displays, are used to interact with virtual / augmented reality environments. Example virtual elements include virtual objects such as digital images, video, text, icons, and control elements such as buttons and other graphics. SUMMARY

[0004] Some methods and interfaces for interacting with environments (e.g., applications, augmented reality environments, mixed reality environments, and virtual reality environments) that include at least some virtual elements are cumbersome, inefficient, and limited. For example, systems that provide inadequate feedback for performing actions associated with virtual objects, systems that require a series of inputs to achieve a desired result in an augmented reality environment, and systems in which virtual object manipulation is complex, tedious, and error-prone, place a significant cognitive burden on users and detract from the experience of virtual / augmented reality environments. Moreover, these methods take longer than necessary, wasting computer system energy. This latter consideration is of increasing importance as battery-powered devices grow in popularity.

[0005] Accordingly, there is a need for computer systems with improved methods and interfaces for providing computer-generated experiences to users, such that the interaction between users and the computer systems is more efficient and more intuitive for the users. Such methods and interfaces optionally complement or replace conventional methods for providing extended reality experiences to users. Such methods and interfaces reduce the number, extent, and / or nature of the inputs from the user, thereby creating a more efficient human-machine interface.

[0006] The above-referenced deficiencies and other problems associated with user interfaces of computer systems are reduced or eliminated by the disclosed systems. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a notebook computer, a tablet computer, or a handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device such as a watch or a head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a display generation component (e.g., a display device such as a head-mounted (HMD), a display, a projector, a touch-sensitive display (also known as a “touch screen” or “touchscreen display”), or other device or component that presents visual content to a user, e.g., on or in the display generation component, or that produces visual content from the display generation component and visible elsewhere). In some embodiments, the computer system has one or more eye tracking components. In some embodiments, the computer system has one or more hand tracking components. In some embodiments, in addition to the display generation component, the computer system has one or more output devices including one or more haptic output generators and / or one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, memory, and one or more modules, programs or sets of instructions stored in memory for carrying out various functions. In some embodiments, a user interacts with the GUI by touch and gestures of a stylus and / or a finger on a touch-sensitive surface, movement of the user’s eyes and hands in space relative to the GUI (and / or the computer system) or the user’s body (as captured by cameras and other movement sensors), and / or voice input (as captured by one or more audio input devices). In some embodiments, the functions performed by the interactions include, optionally, image editing, drawing, presenting, word processing, spreadsheet making, game playing, phone calling, video conferencing, email sending / receiving, instant messaging, test support, digital photography, digital video recording, web browsing, digital music playing, note taking, and / or digital video playing. Executable instructions for performing these functions are optionally included in a transitory and / or non-transitory computer-readable storage medium or other computer program product configured to be executed by one or more processors.

[0007] Electronic devices with improved methods and interfaces are needed to interact with three-dimensional environments. Such methods and interfaces can supplement or replace conventional methods for interacting with three-dimensional environments. Such methods and interfaces reduce the quantity, degree, and / or nature of input from a user and result in a more efficient human-machine interface. For battery-powered computing devices, such methods and interfaces conserve power and increase the time between battery charges.

[0008] In some embodiments, the computer system presents a virtual object corresponding to a physical location within the three-dimensional environment in response to detecting user input indicative of an interaction with the virtual object. In some embodiments, the computer system presents a respective view of a respective physical location from a different perspective in response to detecting user input directed to the virtual object.

[0009] Note that the various embodiments described above can be combined with any of the other embodiments described herein. The features and advantages described in this specification are not all-inclusive. As will be recognized by one of ordinary skill in the art upon reading this disclosure, many additional features and advantages will be apparent. Furthermore, it should be noted that the language used in this specification has been principally selected for readability and instructional purposes and can not have been selected to delineate or circumscribe the subject of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0010] For a better understanding of the various described embodiments, reference should be made to the Detailed Description below, in conjunction with the following drawings, in which like reference numerals refer to corresponding parts throughout the figures. Accordingly, the description below is to be construed in conjunction with the drawings, and not in isolation.

[0011] FIG. 1A is a block diagram illustrating an operating environment of a computer system for providing an XR experience in accordance with some embodiments.

[0012] FIGS. 1B-1P is an example of a computer system for providing an XR experience in FIG. 1A the operating environment of

[0013] FIG. 2 is a block diagram illustrating a controller of a computer system configured to manage and coordinate an XR experience of a user in accordance with some embodiments.

[0014] FIG. 3 is a block diagram illustrating a display generation component of a computer system configured to provide visual components of an XR experience to a user in accordance with some embodiments.

[0015] FIG. 4 is a block diagram illustrating a hand tracking unit of a computer system configured to capture gesture input of a user in accordance with some embodiments.

[0016] is a block diagram illustrating a hand tracking unit of a computer system configured to capture gesture input of a user in accordance with some embodiments.FIG. 5 is a block diagram of an eye tracking unit configured to capture gaze input of a user of a computer system in accordance with some embodiments.

[0017] FIG. 6 is a flow diagram illustrating a flash-assisted gaze tracking pipeline in accordance with some embodiments.

[0018] FIGS. 7A-7Q An example is illustrated in which a computer system displays a navigation user interface having respective levels of immersion corresponding to respective views of respective physical locations in accordance with some embodiments.

[0019] FIG. 8 is a flow diagram illustrating an example method of displaying a navigation user interface having respective levels of immersion corresponding to respective views of respective physical locations in accordance with some embodiments. DETAILED DESCRIPTION

[0020] In accordance with some embodiments, the present disclosure relates to user interfaces for providing an extended reality (XR) experience to a user.

[0021] The systems, methods, and GUIs described herein improve user interface interactions with virtual / augmented reality environments in a variety of ways.

[0022] In some embodiments, a computer system displays a navigation user interface in a three-dimensional environment. The navigation user interface includes one or more first travel user interface elements displayed at a first level of immersion corresponding to a first view of a first physical location. In some embodiments, while displaying the navigation user interface including the one or more first travel user interface elements, the computer system detects a first input corresponding to a request to change the level of immersion. In response to detecting the first input, the computer system changes the display of the navigation user interface having the first level of immersion corresponding to the first view of the first physical location to a second level of immersion corresponding to the first view of the first physical location. The second level of immersion includes one or more second travel user interface elements different from the one or more first travel user interface elements that are selectable to change the display of the navigation user interface from the first view of the first physical location to a second view of a second physical location.

[0023] FIGS. 1 through FIG. 6 A description is provided of example computer systems for providing an XR experience to a user (such as described below with respect to method 800). FIGS. 7A-7Q An example is illustrated in which techniques for displaying a navigation user interface having respective levels of immersion corresponding to respective views of respective physical locations in accordance with some embodiments. FIG. 8A flowchart depicting an exemplary method of displaying a navigation user interface having respective levels of immersion corresponding to respective views of respective physical locations is described in accordance with some embodiments.

[0024] The processes described below enhance the operability of devices and make user-device interfaces more efficient (e.g., by helping users to provide appropriate inputs and reducing user mistakes when operating / interacting with devices) through various techniques, including by providing improved visual feedback to users, reducing the number of inputs needed to perform operations, providing additional control options without cluttering user interfaces with additional display controls, performing an operation when a set of conditions has been met without requiring further user input, improving privacy and / or security, providing richer, more detailed, and / or more realistic user experiences while conserving storage space, and / or additional techniques. These techniques also reduce power usage and improve battery life of devices by enabling users to use devices faster and more efficiently. Saving battery power, and thus weight, improves the ergonomics of the devices. These techniques also enable real-time communications, allow the use of less accurate and / or fewer sensors, resulting in more compact, lighter, and cheaper devices, and enable the devices to be used in various lighting conditions. These techniques reduce energy usage, which reduces the amount of heat emitted by the devices, which is particularly important for wearable devices, where the user can become uncomfortable wearing the device if the device generates too much heat within the operating parameters of the device components.

[0025] Further, in methods described herein in which one or more steps depend on one or more conditions having been met, it should be understood that the methods can be repeated in multiple iterations such that, over the course of the iterations, all of the conditions that determine steps in the method have been met in different iterations of the method. For example, if a method requires a first step to be taken if a condition is met, and a second step to be taken if the condition is not met, one of ordinary skill will appreciate that the recited steps can be repeated until both the condition is met and the condition is not met (in no particular order). Thus, a method that is described as having one or more steps that depend on one or more conditions having been met can be rewritten as a method that is repeated until every condition recited in the method has been met. However, this need not require a system or computer-readable medium to recite that the system or computer-readable medium contains instructions for taking conditional actions based on the satisfaction of the corresponding condition or conditions, and thus can determine whether the possible condition has been met without explicitly repeating the steps of the method until all of the conditions that determine steps in the method have been met. One of ordinary skill in the art will also appreciate that, similar to a method having conditional steps, a system or computer-readable storage medium can repeat the steps of a method as many times as necessary to ensure that all of the conditional steps have been taken.

[0026] In some implementation schemes, such as FIG. 1A As shown, an XR experience is provided to a user via an operating environment 100 including a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted display (HMD), a monitor, a projector, a touchscreen, etc.), one or more input devices 125 (e.g., an eye-tracking device 130, a hand-tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a haptic output generator 170, and other output devices 180), one or more sensors 190 (e.g., image sensors, light sensors, depth sensors, haptic sensors, orientation sensors, proximity sensors, temperature sensors, position sensors, motion sensors, speed sensors, etc.), and optionally one or more peripheral devices 195 (e.g., home appliances, wearable devices, etc.). In some embodiments, one or more of the input devices 125, output devices 155, sensors 190, and peripheral devices 195 are integrated with the display generation component 120 (e.g., in a head-mounted or handheld device).

[0027] In describing XR experiences, various terms are used to distinguish several related but different environments that a user can sense and / or interact with (e.g., interacting with inputs detected by the computer system 101 that generates the XR experience, causing the computer system to generate audio, visual, and / or haptic feedback corresponding to various inputs provided to the computer system 101). The following is a subset of these terms: Physical environment: The physical environment refers to the physical world that people can sense and / or interact with without the aid of electronic systems. Physical environments, such as physical parks, include physical objects such as physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment through senses such as sight, touch, hearing, taste, and smell.

[0028] Extended Reality: In contrast, an extended reality (XR) environment refers to a fully or partially simulated environment that people sense and / or interact with via electronic systems. In XR, a subset of a person’s physical motions, or representations thereof, are tracked, and, in response, one or more characteristics of one or more virtual objects simulated in the XR environment are adjusted in a manner that comports with at least one physical law. For example, an XR system can detect a person’s head turning, and, in response, adjust graphical content and an acoustic field presented to the person in a manner that parallels how such views and sounds would change in a physical environment. In some cases (e.g., for accessibility reasons), adjustments to characteristics of virtual objects in an XR environment can be made in response to representations of physical motions (e.g., voice commands). People can sense and / or interact with XR objects with any of their senses, including sight, hearing, touch, taste, and smell. For example, a person can sense and / or interact with an audio object that creates a 3D or spatial audio environment that provides a perception of a point audio source in 3D space. As another example, an audio object can enable audio transparency that selectively incorporates ambient sounds from a physical environment with or without computer-generated audio. In certain XR environments, a person can sense and / or interact with only audio objects.

[0029] Examples of XR include virtual reality and mixed reality.

[0030] Virtual Reality: A virtual reality (VR) environment refers to a simulated environment that is designed to be entirely based on computer-generated sensory inputs for one or more senses. A VR environment includes multiple virtual objects that a person can sense and / or interact with. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can sense and / or interact with virtual objects in a VR environment through a simulation of a person’s presence within the computer-generated environment and / or through a simulation of a subset of a person’s physical movements within the computer-generated environment.

[0031] Mixed reality: In contrast to VR environments, which are designed to be based entirely on computer-generated sensory inputs, a mixed reality (MR) environment refers to a simulated environment that is designed to incorporate sensory inputs from the physical environment, or representations thereof, in addition to including computer-generated sensory inputs (e.g., virtual objects). On a virtuality continuum, a mixed reality environment is any state in between a completely physical environment on one end and a virtual reality environment on the other end, but does not include either of the two extremes. In some MR environments, the computer-generated sensory inputs can respond to changes in the sensory inputs from the physical environment. Additionally, some electronic systems for presenting MR environments can track position and / or orientation with respect to the physical environment to enable virtual objects to interact with real objects (i.e., physical articles from the physical environment or representations thereof). For example, the system can cause motion such that a virtual tree appears to be stationary with respect to the physical ground.

[0032] Examples of mixed reality include augmented reality and augmented virtuality.

[0033] Augmented reality: An augmented reality (AR) environment refers to a simulated environment in which one or more virtual objects are superimposed over a physical environment or a representation of a physical environment. For example, an electronic system for presenting an AR environment can have a transparent or translucent display through which a person can directly view a physical environment. The system can be configured to present virtual objects on the transparent or translucent display so that a person, using the system, perceives the virtual objects as being overlaid on the physical environment. Alternatively, a system can have an opaque display and one or more imaging sensors that capture images or video of the physical environment, which are representations of the physical environment. The system combines the images or video with virtual objects and presents the combination on the opaque display. A person, using the system, indirectly views the physical environment via the images or video of the physical environment and perceives the virtual objects as being overlaid on the physical environment. As used herein, video of a physical environment displayed on an opaque display is referred to as “pass-through video,” meaning that the system uses one or more image sensors to capture images of the physical environment and uses those images in presenting the AR environment on the opaque display. Further alternatively, a system can have a projection system that projects virtual objects into the physical environment, for example, as a holograph or on a physical surface, so that a person, using the system, perceives the virtual objects as being overlaid on the physical environment. An augmented reality environment also refers to a simulated environment in which a representation of a physical environment is transformed by computer-generated sensory information. For example, in providing pass-through video, the system can transform one or more sensor images to impose a selected perspective (e.g., viewpoint) that is different from the perspective captured by the imaging sensors. As another example, a representation of a physical environment can be transformed by graphically modifying (e.g., enlarging) portions thereof so that the modified portions can be representative but not true versions of the originally captured images. As yet another example, a representation of a physical environment can be transformed by graphically eliminating or blurring portions thereof.

[0034] Augmented virtual: An augmented virtual (AV) environment refers to a simulated environment in which a virtual or computer-generated environment incorporates one or more sensory inputs from a physical environment. The sensory inputs can be representations of one or more characteristics of the physical environment. For example, an AV park can have virtual trees and virtual buildings, but a person’s face is realistically reproduced from images taken of a physical person. As another example, a virtual object can take on the shape or color of a physical article imaged by one or more imaging sensors. As yet another example, a virtual object can take on a shadow consistent with the positioning of the sun in the physical environment.

[0035] In augmented reality, mixed reality, or virtual reality environments, a view of the three-dimensional environment is visible to the user. This view is typically visible to the user via a virtual viewport through one or more display generating components (e.g., a display providing stereoscopic content to different eyes of the same user), which has a viewport boundary that defines the extent of the three-dimensional environment visible to the user via the one or more display generating components. In some embodiments, the area defined by the viewport boundary is smaller than the user's visual field in one or more dimensions (e.g., based on the user's visual field, the size of one or more display generating components, optical properties or other physical characteristics, and / or the position and / or orientation of one or more display generating components relative to the user's eyes). In some embodiments, the area defined by the viewport boundary is larger than the user's visual field in one or more dimensions (e.g., based on the user's visual field, the size of one or more display generating components, optical properties or other physical characteristics, and / or the position and / or orientation of one or more display generating components relative to the user's eyes). The viewport and viewport boundary typically move with the movement of one or more display generating components (e.g., with the user's head for head-mounted devices, or with the user's hand for handheld devices such as tablets or smartphones). The user's viewpoint determines what is visible within the viewport. The viewpoint typically specifies the position and orientation relative to the 3D environment, and as the viewpoint moves, the view of the 3D environment also moves within the viewport. For head-mounted devices, the viewpoint is usually based on the position and orientation of the user's head, face, and / or eyes to provide a perceptibly accurate view of the 3D environment that provides an immersive experience when the user is using the head-mounted device. For handheld or fixed devices, the viewpoint moves with the movement of the handheld or fixed device and / or with changes in the user's positioning relative to the handheld or fixed device (e.g., the user moves towards, away from, up, down, right, and / or left). For devices including display generation components with virtual pass-through, a portion of the physical environment visible (e.g., displayed and / or projected) via one or more display generation components is based on the field of view of one or more cameras communicating with the display generation components, which typically move with the movement of the display generation components (e.g., for head-mounted devices, moving with the movement of the user's head, or for handheld devices such as tablets or smartphones, moving with the movement of the user's hands), because the user's viewpoint moves with the movement of the field of view of the one or more cameras (and the appearance of one or more virtual objects displayed via one or more display generation components is updated based on the user's viewpoint (e.g., the display position and pose of the virtual objects are updated based on the movement of the user's viewpoint)).For display generation components with optical see-through, portions of the physical environment that are visible via the one or more display generation components (e.g., optically visible through one or more portions of the display generation component that are partially or fully transparent) are based on the user’s field of view through the partially or fully transparent portions of the display generation component (e.g., moving with the user’s head for head-mounted devices, or moving with the user’s hand for handheld devices such as tablets or smartphones) as the user’s point of view moves with the user’s movement through the user’s field of view of the partially or fully transparent portions of the display generation component (and the appearance of the one or more virtual objects is updated based on the user’s point of view).

[0036] In some embodiments, the representation of the physical environment (e.g., via virtual pass-through or optical pass-through display) can be partially or fully occluded by the virtual environment. In some embodiments, the amount of virtual environment displayed (e.g., the amount of physical environment that is not displayed) is based on the level of immersion of the virtual environment (e.g., relative to the representation of the physical environment). For example, increasing the level of immersion optionally causes more of the virtual environment to be displayed, replacing and / or occluding more of the physical environment, and decreasing the level of immersion optionally causes less of the virtual environment to be displayed, revealing portions of the physical environment that were previously not displayed and / or occluded. In some embodiments, at a particular level of immersion, one or more first background objects (e.g., in the representation of the physical environment) are more visually de-emphasized (e.g., darkened, blurred, displayed with increased transparency) than one or more second background objects, and one or more third background objects cease to be displayed. In some embodiments, the level of immersion includes an associated degree to which virtual content (e.g., virtual environment and / or virtual content) displayed by the computer system occludes background content (e.g., content other than the virtual environment and / or virtual content) behind / surrounding the virtual environment, optionally including a number of items of the displayed background content and / or displayed visual properties (e.g., color, contrast, and / or opacity) of the background content, an angular range of virtual content displayed via the display generation component (e.g., 60 degrees of content displayed at a low level of immersion, 120 degrees of content displayed at a medium level of immersion, or 180 degrees of content displayed at a high level of immersion), and / or a proportion of a field of view displayed via the display generation component that is occupied by the virtual content (e.g., 33% of the field of view occupied by the virtual content at a low level of immersion, 66% of the field of view occupied by the virtual content at a medium level of immersion, or 100% of the field of view occupied by the virtual content at a high level of immersion). In some embodiments, the background content is included in the background on which the virtual content is displayed (e.g., background content in the representation of the physical environment). In some embodiments, the background content includes a user interface (e.g., a user interface generated by the computer system corresponding to an application), virtual objects that are not associated with or included in the virtual environment and / or virtual content (e.g., representations of files or other users generated by the computer system, etc.), and / or real objects (e.g., pass-through objects representing real objects in the physical environment surrounding the user that are visible such that they are displayed via the display generation component and / or are visible via a transparent or semi-transparent component of the display generation component because the computer system does not occlude / impede their visibility through the display generation component). In some embodiments, at a low level of immersion (e.g., a first level of immersion), the background, virtual, and / or real objects are displayed in a manner that is not occluded. For example, a virtual environment with a low level of immersion is optionally displayed concurrently with the background content, which is optionally displayed at full brightness, color, and / or translucency.In some implementations, at higher immersion levels (e.g., a second immersion level above the first immersion level), background, virtual, and / or real objects are displayed in an occluded manner (e.g., dimmed, blurred, or removed from the display). For example, a corresponding virtual environment with a high immersion level is displayed without simultaneously displaying background content (e.g., in full-screen or fully immersive mode). As another example, a virtual environment displayed at a medium immersion level is displayed simultaneously with darkened, blurred, or otherwise de-emphasized background content. In some implementations, the visual characteristics of background objects differ between background objects. For example, at a particular immersion level, one or more first background objects are visually de-emphasized more than one or more second background objects (e.g., dimmed, blurred, and / or displayed with increased transparency), and one or more third background objects are stopped from being displayed. In some implementations, zero immersion or a zero immersion level corresponds to a virtual environment that is stopped from being displayed, and instead, a representation of the physical environment (optionally having one or more virtual objects, such as an application, window, or virtual 3D object) is displayed, and the representation of the physical environment is not occluded by the virtual environment. Using physical input elements to adjust immersion levels provides a quick and efficient way to adjust immersion, which enhances the operability of computer systems and makes user-device interfaces more efficient.

[0037] Viewpoint-locked virtual objects: When a computer system displays a virtual object at the same location and / or position within the user's viewpoint, the virtual object remains viewpoint-locked even if the user's viewpoint shifts (e.g., changes). In embodiments where the computer system is a head-mounted device, the user's viewpoint is locked to the direction forward of the user's head (e.g., when the user is looking straight ahead, the user's viewpoint is at least a portion of the user's field of view); therefore, the user's viewpoint remains fixed even when the user's gaze shifts without moving the user's head. In embodiments where the computer system has a display generating component (e.g., a display screen) that can be repositioned relative to the user's head, the user's viewpoint is the augmented reality view presented to the user on the computer system's display generating component. For example, a viewpoint-locked virtual object displayed in the upper left corner of the user's viewpoint when the user's viewpoint is in a first orientation (e.g., the user's head is facing north) continues to be displayed in the upper left corner of the user's viewpoint, even when the user's viewpoint changes to a second orientation (e.g., the user's head is facing west). In other words, the position and / or orientation of a viewpoint-locked virtual object displayed in the user's viewpoint is independent of the user's position and / or orientation in the physical environment. In an implementation where the computer system is a head-mounted device, the user's viewpoint is locked to the orientation of the user's head, so the virtual object is also referred to as a "head-locked virtual object".

[0038] Environment-locked visual objects: A virtual object is environment-locked (alternatively, "world-locked") when the computer system displays the virtual object at a position and / or location in the user's point of view that is based on a location and / or object in a three-dimensional environment (e.g., a physical environment or a virtual environment) that the virtual object is selected and / or anchored to with reference to. As the user's point of view moves, the location and / or object in the environment relative to the user's point of view changes, which causes the environment-locked virtual object to be displayed at a different position and / or location in the user's point of view. For example, an environment-locked virtual object that is locked to a tree immediately in front of the user is displayed at the center of the user's point of view. When the user's point of view shifts to the right (e.g., the user's head turns to the right) such that the tree is now to the left center of the user's point of view (e.g., the tree location in the user's point of view shifts), the environment-locked virtual object that is locked to the tree is displayed to the left center of the user's point of view. In other words, the position and / or location in the user's point of view at which the environment-locked virtual object is displayed depends on the location and / or orientation of the location and / or object in the environment that the virtual object is locked to. In some embodiments, the computer system uses a stationary frame of reference (e.g., a coordinate system that is anchored to a fixed location and / or object in the physical environment) in order to determine the location at which to display the environment-locked virtual object in the user's point of view. The environment-locked virtual object can be locked to a stationary portion of the environment (e.g., a floor, a wall, a table, or other stationary object), or can be locked to a movable portion of the environment (e.g., a vehicle, an animal, a person, or even a representation of a portion of the user's body such as the user's hand, wrist, arm, or foot that moves independently of the user's point of view) such that the virtual object moves with the portion of the environment to maintain a fixed relationship between the virtual object and the portion of the environment.

[0039] In some embodiments, an environment-locked or viewpoint-locked virtual object exhibits a lazy following behavior that reduces or delays movement of the environment-locked or viewpoint-locked virtual object relative to movement of a reference point that the virtual object is following. In some embodiments, when exhibiting a lazy following behavior, the computer system intentionally delays movement of the virtual object when movement of a reference point (e.g., a portion of the environment, a viewpoint, or a point fixed relative to a viewpoint, such as a point between 5 cm and 300 cm from the viewpoint) that the virtual object is following is detected. For example, when a reference point (e.g., a portion of the environment or a viewpoint) moves at a first speed, the virtual object is moved by the device to remain locked to the reference point, but moves at a second speed that is slower than the first speed (e.g., until the reference point stops moving or slows down, at which point the virtual object begins to catch up to the reference point). In some embodiments, when a virtual object exhibits a lazy following behavior, the device ignores small movements of the reference point (e.g., ignores movements of the reference point that are below a threshold amount of movement, such as movements of 0 to 5 degrees or movements of 0 to 50 cm). For example, when a reference point (e.g., a portion of the environment or a viewpoint to which the virtual object is locked) moves a first amount, the distance between the reference point and the virtual object increases (e.g., because the virtual object is being displayed so as to remain fixed or substantially fixed in position relative to a viewpoint or portion of the environment that is different from the reference point to which the virtual object is locked), and when the reference point (e.g., a portion of the environment or a viewpoint to which the virtual object is locked) moves a second amount that is greater than the first amount, the distance between the reference point and the virtual object first increases (e.g., because the virtual object is being displayed so as to remain fixed or substantially fixed in position relative to a viewpoint or portion of the environment that is different from the reference point to which the virtual object is locked), and then decreases when the amount of movement of the reference point increases above a threshold (e.g., a "lazy following" threshold) because the virtual object is moved by the computer system to remain fixed or substantially fixed in position relative to the reference point. In some embodiments, the virtual object remaining substantially fixed in position relative to the reference point includes the virtual object being displayed within a threshold distance (e.g., 1 cm, 2 cm, 3 cm, 5 cm, 15 cm, 20 cm, 50 cm) of the reference point in one or more dimensions (e.g., up / down, left / right, and / or forward / backward of the position of the reference point).

[0040] Hardware: There are many different types of electronic systems that enable a person to sense various XR environments and / or to interact with various XR environments. Examples include head-mounted systems, projection-based systems, heads-up displays (HUDs), vehicle windshields having integrated display capability, windows having integrated display capability, displays formed as lenses designed to be placed on a person's eyes (e.g., similar to contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system can have one or more speakers and an integrated opaque display. Alternatively, a head-mounted system can be configured to accept an external opaque display (e.g., a smartphone). A head-mounted system can incorporate one or more imaging sensors for capturing images or video of a physical environment and / or one or more microphones for capturing audio of the physical environment. A head-mounted system can have a transparent or semi-transparent display instead of an opaque display. A transparent or semi-transparent display can have a medium through which light representative of images is directed to a person's eyes. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium can be an optical waveguide, hologram medium, optical combiner, optical reflector, or any combination thereof. In one embodiment, a transparent or semi-transparent display can be configured to selectively become opaque. A projection-based system can employ retinal projection technology that projects graphical images onto a person's retinas. Projection systems can also be configured to project virtual objects into a physical environment, e.g., as a hologram or on a physical surface. In some embodiments, controller 110 is configured to manage and coordinate a person's XR experience. In some embodiments, controller 110 includes a suitable combination of software, firmware, and / or hardware. Reference is made to FIG. 1 for further details regarding the components of controller 110. FIG. 2Controller 110 is described in greater detail. In some embodiments, controller 110 is a computing device that is in a local or remote location relative to scene 105 (e.g., a physical environment). For example, controller 110 is a local server that is located within scene 105. As another example, controller 110 is a remote server (e.g., a cloud server, a central server, etc.) that is located outside of scene 105. In some embodiments, controller 110 is communicatively coupled with display generation component 120 (e.g., an HMD, a display, a projector, a touchscreen, etc.) via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802. l lx, IEEE 802.16x, IEEE 802.3x, etc.). As another example, controller 110 is included within a housing (e.g., a physical enclosure) of display generation component 120 (e.g., an HMD or a portable electronic device that includes a display and one or more processors, etc.), one or more input devices of input devices 125, one or more output devices of output devices 155, one or more sensors of sensors 190, and / or one or more peripheral devices of peripheral devices 195, or shares the same physical housing or support structure as one or more of the aforementioned devices.

[0041] In some embodiments, display generation component 120 is configured to provide an XR experience (e.g., at least a visual component of an XR experience) to a user. In some embodiments, display generation component 120 includes a suitable combination of software, firmware, and / or hardware. The following discussion of display generation component 120 is equally applicable to display generation component 120A. FIG. 3 Display generation component 120 is described in greater detail. In some embodiments, the functionality of controller 110 is provided by and / or in combination with display generation component 120.

[0042] According to some embodiments, display generation component 120 provides an XR experience to a user when the user is virtually and / or physically present within scene 105.

[0043] In some embodiments, the display generation component is worn on a portion of the user’s body (e.g., on his / her head, on his / her hand, etc.). Thus, the display generation component 120 includes one or more XR displays provided for displaying XR content. For example, in various embodiments, the display generation component 120 encloses the user’s field of view. In some embodiments, the display generation component 120 is a handheld device (such as a smartphone or tablet) configured to present XR content, and the user holds the device with a display facing the user’s field of view and a camera facing the scene 105. In some embodiments, the handheld device is optionally placed within a housing worn on the user’s head. In some embodiments, the handheld device is optionally placed on a support (e.g., a tripod) in front of the user. In some embodiments, the display generation component 120 is an XR chamber, housing, or room configured to present XR content, where the user does not wear or hold the display generation component 120. Many of the user interfaces described with reference to one type of hardware for displaying XR content (e.g., a handheld device or a device on a tripod) can be implemented on another type of hardware for displaying XR content (e.g., an HMD or other wearable computing device). For example, a user interface showing interactions with XR content triggered based on interactions occurring in a space in front of a handheld device or a tripod-mounted device can be similarly implemented with an HMD, where the interactions occur in a space in front of the HMD and responses to the XR content are displayed via the HMD. Similarly, a user interface showing interactions with XR content triggered based on movement of a handheld device or a tripod-mounted device relative to a physical environment (e.g., the scene 105 or a portion of the user’s body (e.g., the user’s eyes, head, or hand)) can be similarly implemented with an HMD, where the movement is caused by movement of the HMD relative to the physical environment (e.g., the scene 105 or a portion of the user’s body (e.g., the user’s eyes, head, or hand)).

[0044] Although relevant features of the operating environment 100 are shown in FIG. 1A the ordinary skill in the art will understand from the present disclosure that various other features have not been illustrated in the interest of conciseness and so as not to obscure more relevant aspects of the example embodiments disclosed herein.

[0045] FIGS. 1A-1PVarious examples of computer systems for performing the methods and providing audio, visual, and / or haptic feedback as part of the user interfaces described herein are illustrated. In some embodiments, the computer system includes one or more display generation components (e.g., first and second display components 1-120a, 1-120b and / or first and second optical modules 11.1.1-104a and 11.1.1-104b) for displaying a representation of virtual elements and / or a physical environment to a user of the computer system, the representation of the virtual elements and / or the physical environment optionally generated based on detected events and / or user input detected by the computer system. The user interfaces generated by the computer system are optionally corrected by one or more corrective lenses 11.3.2-216, optionally removably attached to one or more of the optical modules, to enable users who would otherwise use eyeglasses or contact lenses to correct their vision to more easily view the user interfaces. While many of the user interfaces illustrated herein illustrate a single view of the user interface, user interfaces in an HMD are optionally displayed using two optical modules (e.g., first and second display components 1-120a, 1-120b and / or first and second optical modules 11.1.1-104a and 11.1.1-104b), one optical module for the user’s right eye and a different optical module for the user’s left eye, and slightly different images are presented to the two different eyes to create the illusion of stereoscopic depth, the single view of the user interface typically being either a right-eye view or a left-eye view, the depth effects explained in the text or using other diagrams or views. In some embodiments, the computer system includes one or more external displays (e.g., display component 1-108) for displaying status information of the computer system to a user of the computer system (when the computer system is not being worn) and / or to other people in the vicinity of the computer system, the status information optionally generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more audio output components (e.g., electronic component 1-112) for generating audio feedback, the audio feedback optionally generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors (e.g., sensor components 1-356 and / or one or more sensors in FIG. 1I FIG. 1I ​The illuminators described in connection with FIGS. 6-124) to generate digital pass-through images, capture visual media (e.g., photos and / or videos) corresponding to a physical environment, or determine poses (e.g., positions and / or orientations) of physical objects and / or surfaces in the physical environment, such that virtual objects can be placed based on detected poses of the physical objects and / or surfaces. In some embodiments, the computer system includes one or more input devices to detect input, such as one or more sensors to detect hand positions and / or movements (e.g., sensor assembly 1-356 and / or FIG. 1I one or more sensors in FIG. 1-356), which can be used (optionally in combination with one or more illuminators, such as FIG. 1I the illuminators 6-124 described in connection with FIGS. 6-124) to determine when one or more air gestures are performed. In some embodiments, the computer system includes one or more input devices to detect input, such as one or more sensors to detect eye movements (e.g., eye tracking and gaze tracking sensors in FIG. 1I FIG. 1-356), which can be used (optionally in combination with one or more lights, such as FIG. 1OThe gaze and / or attention information is optionally combined with hand tracking information to determine user interaction with one or more user interfaces based on direct and / or indirect input, such as air gestures or input using one or more hardware input devices, such as one or more buttons (e.g., first button 1-128, button 11.1.1-114, second button 1-132, and / or dial or button 1-328), knobs (e.g., first button 1-128, button 11.1.1-114, and / or dial or button 1-328), digital crowns (e.g., first button 1-128, button 11.1.1-114, and / or dial or button 1-328 that can be pressed and that can be twisted or rotated), touchpads, touchscreens, keyboards, mice, and / or other input devices. One or more buttons (e.g., first button 1-128, button 11.1.1-114, second button 1-132, and / or dial or button 1-328) are optionally used to perform system operations, such as re-centering content in a three-dimensional environment that is visible to the user of the device, displaying a home user interface for launching applications, starting a live communication session, or initiating display of a virtual three-dimensional background. A knob or digital crown (e.g., first button 1-128, button 11.1.1-114, and / or dial or button 1-328 that can be pressed and that can be twisted or rotated) is optionally rotatable to adjust a parameter of visual content, such as an immersion level of a virtual three-dimensional environment (e.g., an extent to which virtual content occupies a user’s viewport in the three-dimensional environment) or other parameters associated with the three-dimensional environment and virtual content displayed via the optical modules (e.g., first display assembly 1-120a and second display assembly 1-120b and / or first optical module 11.1.1-104a and second optical module 11.1.1-104b).

[0046] FIG. 1BA front view, top view, perspective view of an example of a head-mountable display (HMD) device 1-100 configured to be worn by a user and provide virtual and altered / mixed reality (VR / AR) experiences is illustrated. The HMD 1-100 can include a display unit 1-102 or assembly, an electronic strap assembly 1-104 connected to and extending from the display unit 1-102, and a band assembly 1-106 secured to the electronic strap assembly 1-104 at either end. The electronic strap assembly 1-104 and the band 1-106 can be part of a retention assembly configured to wrap around a user’s head to hold the display unit 1-102 against the user’s face.

[0047] In at least one example, the band assembly 1-106 can include a first band 1-116 configured to wrap around a back side of the user’s head and a second band 1-117 configured to extend over a top of the user’s head. As shown, the second band can extend between a first electronic strap 1-105a and a second electronic strap 1-105b of the electronic strap assembly 1-104. The strap assembly 1-104 and the band assembly 1-106 can be part of a securing mechanism that extends rearward from the display unit 1-102 and is configured to hold the display unit 1-102 against the user’s face.

[0048] In at least one example, the securing mechanism includes a first electronic strap 1-105a that includes a first proximal end 1-134 coupled to the display unit 1-102 (e.g., a housing 1-150 of the display unit 1-102) and a first distal end 1-136 opposite the first proximal end 1-134. The securing mechanism can also include a second electronic strap 1-105b that includes a second proximal end 1-138 coupled to the housing 1-150 of the display unit 1-102 and a second distal end 1-140 opposite the second proximal end 1-138. The securing mechanism can also include a first band 1-116 that includes a first end 1-142 coupled to the first distal end 1-136 and a second end 1-144 coupled to the second distal end 1-140, and a second band 1-117 that extends between the first electronic strap 1-105a and the second electronic strap 1-105b. The straps 1-105a-b and the bands 1-116 can be coupled via a connection mechanism or assembly 1-114. In at least one example, the second band 1-117 includes a first end 1-146 coupled to the first electronic strap 1-105a between the first proximal end 1-134 and the first distal end 1-136 and a second end 1-148 coupled to the second electronic strap 1-105b between the second proximal end 1-138 and the second distal end 1-140.

[0049] In at least one example, the first and second electronic straps 1-105a-b comprise plastic, metal, or other structural material that forms a shape of the substantially rigid straps 1-105a-b. In at least one example, the first and second straps 1-116, 1-117 are formed of an elastically flexible material, including a woven textile, rubber, etc. The first and second straps 1-116, 1-117 can be flexible to conform to the shape of a user's head when wearing the HMD 1-100.

[0050] In at least one example, one or more of the first and second electronic straps 1-105a-b can define an internal strap volume and include one or more electronic components disposed in the internal strap volume. In one example, as shown, the first electronic strap 1-105a can include an electronic component 1-112. In one example, the electronic component 1-112 can include a speaker. In one example, the electronic component 1-112 can include a computing component, such as a processor. FIG. 1B

[0051] In at least one example, the housing 1-150 defines a first front-facing opening 1-152. The front-facing opening is labeled in dashed line as 1-152 in FIG. 1B because the display component 1-108 is disposed to occlude the first opening 1-152 from view when the HMD 1-100 is assembled. The housing 1-150 can also define a rear-facing second opening 1-154. The housing 1-150 also defines an internal volume between the first opening 1-152 and the second opening 1-154. In at least one example, the HMD 1-100 includes a display component 1-108 that can include a front cover and a display screen (shown in other figures) disposed in or across the front opening 1-152 to occlude the front opening 1-152. In at least one example, the display screen of the display component 1-108 and generally the display component 1-108 has a curvature configured to follow the curvature of a user's face. The display screen of the display component 1-108 can be curved as shown to complement the user's facial features and the overall curvature from side to side of the face, e.g., from left to right and / or from top to bottom, with the display unit 1-102 pressed.

[0052] ​In at least one example, the housing 1-150 can define a first aperture 1-126 between the first opening 1-152 and the second opening 1-154 and a second aperture 1-130 between the first opening 1-152 and the second opening 1-154. The HMD 1-100 can also include a first button 1-126 disposed in the first aperture 1-128 and a second button 1-132 disposed in the second aperture 1-130. The first button 1-128 and the second button 1-132 can be pressed through the respective apertures 1-126, 1-130. In at least one example, the first button 1-126 and / or the second button 1-132 can be a twistable dial and a pressable button. In at least one example, the first button 1-128 is a pressable and twistable dial button and the second button 1-132 is a pressable button.

[0053] FIG. 1C A rear perspective view of the HMD 1-100 is illustrated. The HMD 1-100 can include a light seal 1-110 extending rearward from the housing 1-150 of the display assembly 1-108 around a perimeter of the housing 1-150, as shown. The light seal 1-110 can be configured to extend from the housing 1-150 to a face of a user, around eyes of the user, to block external light from being visible. In one example, the HMD 1-100 can include a first display assembly 1-120a and a second display assembly 1-120b disposed at or in a second, rear-facing opening 1-154 defined by the housing 1-150 and configured to project light through the second opening 1-154. In at least one example, each display assembly 1-120a-b can include a respective display screen 1-122a, 1-122b configured to project light in a rearward direction through the second opening 1-154 toward eyes of a user.

[0054] In at least one example, with reference to FIG. 1B and FIG. 1C both, the display assembly 1-108 can be a front-facing, forward display assembly including a display screen configured to project light in a first, forward direction and the rear-facing display screens 1-122a-b can be configured to project light in a second, rearward direction opposite the first direction. As noted above, the light seal 1-110 can be configured to block light external to the HMD 1-100 from reaching the eyes of a user, including light projected by the forward display screen of the display assembly 1-108 shown in the front perspective view of FIG. 1B In at least one example, the HMD 1-100 can also include a curtain 1-124 occluding the second opening 1-154 between the housing 1-150 and the rear-facing display assemblies 1-120a-b. In at least one example, the curtain 1-124 can be elastic or at least partially elastic.

[0055] FIG. 1B and FIG. 1C Any of the features, components, and / or parts shown (including their arrangement and configuration) can be included alone or in any combination in FIGS. 1D-1F Any other examples of the devices, features, components, and parts shown and described herein. Likewise, any of the features, components, and / or parts shown or described can be included alone or in any combination in FIGS. 1D-1F Any of the features, components, and / or parts shown (including their arrangement and configuration) can be included alone or in any combination in FIG. 1B and FIG. 1C examples of the devices, features, components, and parts shown.

[0056] FIG. 1D An exploded view illustrating an example of an HMD 1-200 that includes individual portions or parts that are separated according to the modularization and selective coupling of these parts. For example, the HMD 1-200 can include a band 1-216 that is selectively coupleable to a first electronic strip 1-205a and a second electronic strip 1-205b. The first fixed strip 1-205a can include a first electronic component 1-212a and the second fixed strip 1-205b can include a second electronic component 1-212b. In at least one example, the first and second strips 1-205a-b are removably coupleable to a display unit 1-202.

[0057] Further, the HMD 1-200 can include a light seal 1-210 that is configured to be removably coupled to the display unit 1-202. The HMD 1-200 can also include a lens 1-218 that is removably coupleable to the display unit 1-202, for example, on a first display component and a second display component that include display screens. The lens 1-218 can include a custom prescription lens that is configured for correcting vision. As noted, in the exploded view of FIG. 1D Each of the parts described above and in the exploded view of FIG. 1-200 can be removably coupled, attached, reattached, and replaced to update the part or swap out the part for a different user. For example, the band such as the band 1-216, the light seal such as the light seal 1-210, the lens such as the lens 1-218, and the electronic strips such as the electronic strips 1-205a-b can be swapped out according to a user such that these portions are customized to fit and correspond to a single user of the HMD 1-200.

[0058] FIG. 1D Any of the features, components, and / or parts shown (including their arrangement and configuration) can be included alone or in any combination in FIG. 1B , FIG. 1C and FIGS. 1E-1FAny other examples of devices, features, components, and parts shown and described herein. Similarly, refer to... FIG. 1B , FIG. 1C and FIGS. 1E-1F Any of the features, components, and / or parts shown or described (including their arrangement and configuration) may be included individually or in any combination. FIG. 1D Examples of devices, features, components, and parts are shown.

[0059] FIG. 1E An exploded view illustrating an example of an HMD display unit 1-306 is shown. The display unit 1-306 may include a front display assembly 1-308, a frame / housing assembly 1-350, and a curtain assembly 1-324. The display unit 1-306 may also include a sensor assembly 1-356, a logic board assembly 1-358, and a cooling assembly 1-360 disposed between the frame assembly 1-350 and the front display assembly 1-308. In at least one example, the display unit 1-306 may also include a rear display assembly 1-320, which includes a first rear display screen 1-322a and a second rear display screen 1-322b disposed between the frame 1-350 and the curtain assembly 1-324.

[0060] In at least one example, the display unit 1-306 may further include a motor assembly 1-362 configured as an adjustment mechanism for adjusting the positioning of the display screens 1-322a to b of the display unit 1-320 relative to the frame 1-350. In at least one example, the display unit 1-320 is mechanically coupled to the motor assembly 1-362, and each display screen 1-322a to b has at least one motor, such that the motor is capable of translating the display screens 1-322a to b to match the interpupillary distance of the user's eyes.

[0061] In at least one example, display unit 1-306 may include a dial or button 1-328 that is pressable relative to frame 1-350 and accessible to a user outside frame 1-350. Button 1-328 may be electrically connected to motor assembly 1-362 via a controller, such that button 1-328 can be operated by a user to cause the motor of motor assembly 1-362 to adjust the positioning of display screens 1-322a to b.

[0062] FIG. 1E Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. FIGS. 1B-1D and FIG. 1F Any other examples of devices, features, components, and parts shown and described herein. Similarly, refer to... FIGS. 1B-1D and FIG. 1FAny of the features, components, and / or parts (including arrangements and configurations thereof) shown and described can be included alone or in any combination FIG. 1E Examples of the devices, features, components, and parts shown.

[0063] FIG. 1F An exploded view of another example of a display unit 1-406 of an HMD device similar to other HMD devices described herein is illustrated. The display unit 1-406 can include a front display assembly 1-402, a sensor assembly 1-456, a logic board assembly 1-458, a cooling assembly 1-460, a frame assembly 1-450, a rear display assembly 1-421, and a curtain assembly 1-424. The display unit 1-406 can also include a motor assembly 1-462 for adjusting the positioning of first and second display subassemblies 1-420a, 1-420b of the rear display assembly 1-421, including first and second respective display screens for inter-pupillary adjustment, as described above.

[0064] FIG. 1F The various parts, systems, and components shown in the exploded view of FIGS. 1B-1E are described in greater detail herein with reference to FIG. 1F The display unit 1-406 shown can be assembled and integrated with FIGS. 1B-1E a securing mechanism including an electronic band, a strap, and other components including light seals, connection assemblies, and the like.

[0065] FIG. 1F Any of the features, components, and / or parts (including arrangements and configurations thereof) shown and described can be included alone or in any combination FIGS. 1B-1E any other examples of devices, features, components, and parts shown and described herein. Likewise, reference is made to FIG. 1F Any of the features, components, and / or parts (including arrangements and configurations thereof) shown and described can be included alone or in any combination FIG. 1G Examples of the devices, features, components, and parts shown.

[0066] FIG. 1G A perspective exploded view of a front cover assembly 3-100 of an HMD device described herein is illustrated, for example FIG. 1G the front cover assembly 3-1 of the HMD 3-100 shown or any other HMD device shown and described herein. FIG. 1GThe illustrated front cover assembly 3-100 can include a transparent or translucent cover 3-102, a shroud 3-104 (or "cover"), an adhesive layer 3-106, a display assembly 3-108 including a lenticular lens panel or array 3-110, and a structural trim 3-112. The adhesive layer 3-106 can secure the shroud 3-104 and / or the transparent cover 3-102 to the display assembly 3-108 and / or the trim 3-112. The trim 3-112 can secure the various components of the front cover assembly 3-100 to a frame or chassis of the HMD device.

[0067] In at least one example, as illustrated, FIG. 1G The transparent cover 3-102, shroud 3-104, and display assembly 3-108 including the lenticular lens array 3-110 can be curved to accommodate the curvature of a user's face, in at least one example. The transparent cover 3-102 and shroud 3-104 can be curved in two or three dimensions, for example, vertically curved in the Z-direction within the Z-X plane, and horizontally curved in the X-direction within the Z-X plane. In at least one example, the display assembly 3-108 can include the lenticular lens array 3-110 and a display panel having pixels configured to project light through the shroud 3-104 and the transparent cover 3-102. The display assembly 3-108 can be curved in at least one direction (e.g., a horizontal direction) to accommodate the curvature of a user's face from one side of the face (e.g., the left side) to the other side (e.g., the right side). In at least one example, each layer or component of the display assembly 3-108 (which will be illustrated in subsequent figures and described in greater detail, but which can include the lenticular lens array 3-110 and a display layer) can be similarly or concentrically curved in the horizontal direction to accommodate the curvature of a user's face.

[0068] In at least one example, the shroud 3-104 can include a transparent or translucent material through which the display assembly 3-108 projects light. In one example, the shroud 3-104 can include one or more opaque portions, such as opaque ink printed portions or other opaque film portions on a back surface of the shroud 3-104. The back surface can be the surface of the shroud 3-104 that faces the eyes of a user when the HMD device is worn. In at least one example, the opaque portions can be on a front surface of the shroud 3-104 opposite the back surface. In at least one example, the one or more opaque portions of the shroud 3-104 can include a perimeter portion that visually hides any components surrounding an outer perimeter of a display screen of the display assembly 3-108. In this way, the opaque portions of the shroud hide any other components of the HMD device that would otherwise be visible through the transparent or translucent cover 3-102 and / or shroud 3-104, including electronic components, structural components, and the like.

[0069] In at least one example, the shroud 3-104 can define one or more apertured transparent portions 3-120 through which the sensor can transmit and receive signals. In one example, the portions 3-120 are apertures through which the sensor can extend or through which the sensor can transmit and receive signals. In one example, the portions 3-120 are transparent portions, or portions that are more transparent than the surrounding translucent or opaque portions of the shroud, through which the sensor can transmit and receive signals through the shroud and through the transparent cover 3-102. In one example, the sensor can include a camera, an IR sensor, a LUX sensor, or any other visual or non-visual environmental sensor of the HMD device.

[0070] FIG. 1G Any of the illustrated features, components, and / or parts, including their arrangement and configuration, can be included in any other example of the devices, features, components, and parts described herein, either alone or in any combination. Likewise, any of the features, components, and / or parts illustrated and described herein, including their arrangement and configuration, can be included in any example of the devices, features, components, and parts described herein, either alone or in any combination. FIG. 1H in the examples of the devices, features, components, and parts illustrated.

[0071] FIG. 1I An exploded view of an example of an HMD device 6-100 is illustrated. The HMD device 6-100 can include a sensor array or system 6-102 that includes one or more sensors, cameras, projectors, etc. mounted to one or more components of the HMD 6-100. In at least one example, the sensor system 6-102 can include a cradle 1-338 to which one or more sensors of the sensor system 6-102 can be secured / fastened.

[0072] FIG. 1J A portion of the HMD device 6-100 including a front transparent cover 6-104 and a sensor system 6-102 is illustrated. The sensor system 6-102 can include a number of different sensors, emitters, receivers, including cameras, IR sensors, projectors, etc. The transparent cover 6-104 is illustrated in front of the sensor system 6-102 to illustrate the relative positioning of the various sensors and emitters and the orientation of each sensor / emitter of the system 6-102. As referenced herein, "lateral," "side," "laterally," "horizontally," and other similar terms refer to the orientation or direction as indicated by the X-axis illustrated. FIG. 1J As referenced herein, "vertical," "up," "down," and other similar terms refer to the orientation or direction as indicated by the Z-axis illustrated. FIG. 1J As referenced herein, "vertical," "up," "down," and other similar terms refer to the orientation or direction as indicated by the Z-axis illustrated. FIG. 1I As referenced herein, "vertical," "up," "down," and other similar terms refer to the orientation or direction as indicated by the Z-axis illustrated.

[0073] In at least one example, the transparent cover 6-104 can define a front outer surface of the HMD device 6-100, and the sensor system 6-102 including various sensors and components thereof can be disposed behind the cover 6-104 in the Y axis / direction. The cover 6-104 can be transparent or translucent to allow light to pass through the cover 6-104, including both light detected by the sensor system 6-102 and light emitted thereby.

[0074] As noted elsewhere herein, the HMD device 6-100 can include one or more controllers including processors for electrically coupling the various sensors and emitters of the sensor system 6-102 with one or more motherboards, processing units, and other electronic devices such as display screens, etc. Further, as will be shown in greater detail below with reference to other figures, the various sensors, emitters, and other components of the sensor system 6-102 can be coupled to various structural frame members, brackets, etc. of the HMD device 6-100 not shown in FIG. 6-1. FIG. 1I For purposes of clarity, FIG. 1I The components of the sensor system 6-102 are illustrated as not attached and not electrically coupled with other components.

[0075] In at least one example, the device can include one or more controllers having processors configured to execute instructions stored on memory components electrically coupled to the processors. The instructions can include or cause the processors to execute one or more algorithms for self-correcting the angle and positioning of the various cameras described herein over time as the initial positioning, angle, or orientation of the cameras is impacted or distorted due to an accidental drop event or other event.

[0076] In at least one example, the sensor system 6-102 can include one or more scene cameras 6-106. The system 6-102 can include two scene cameras 6-102 disposed on either side of the bridge or arch structure of the HMD device 6-100 such that each of the two cameras 6-106 generally corresponds to the positioning of the left and right eyes of the user behind the cover 6-103. In at least one example, the scene cameras 6-106 are generally oriented forward in the Y direction to capture images of the user’s forward field of view during use of the HMD 6-100. In at least one example, the scene cameras are color cameras and provide image and content to display screens facing the eyes of the user for MR video pass-through when using the HMD device 6-100. The scene cameras 6-106 can also be used for environment and object reconstruction.

[0077] In at least one example, sensor system 6-102 can include a first depth sensor 6-108 that is generally directed forward in the Y direction. In at least one example, first depth sensor 6-108 can be used for environment and object reconstruction as well as hand and body tracking of a user. In at least one example, sensor system 6-102 can include a second depth sensor 6-110 that is centrally disposed along the width of HMD device 6-100 (e.g., along the X axis). For example, second depth sensor 6-110 can be disposed on a central nose bridge or on an adaptive structure above the nose when a user wears HMD 6-100. In at least one example, second depth sensor 6-110 can be used for environment and object reconstruction as well as hand and body tracking. In at least one example, the second depth sensor can include a LIDAR sensor.

[0078] In at least one example, sensor system 6-102 can include a depth projector 6-112 that is generally directed forward to project electromagnetic waves (e.g., in the form of a predetermined pattern of light dots) into or within a field of view of user and / or scene camera 6-106, or into or within a field of view that includes and extends beyond the field of view of user and / or scene camera 6-106. In at least one example, the depth projector is capable of projecting electromagnetic waves of light in the form of a pattern of dots that reflect off of objects and back into the aforementioned depth sensors, including depth sensors 6-108, 6-110. In at least one example, depth projector 6-112 can be used for environment and object reconstruction as well as hand and body tracking.

[0079] In at least one example, sensor system 6-102 can include downward facing cameras 6-114 that are generally directed downward in the Z axis relative to HMD device 6-100. In at least one example, downward facing cameras 6-114 can be disposed on the left and right sides of HMD device 6-100 as shown and used for hand and body tracking, headset tracking, and facial avatar detection and creation for display of a user avatar on a front-facing display screen of HMD device 6-100 as described elsewhere herein. For example, downward facing cameras 6-114 can be used to capture facial expressions and movements of a user’s face below HMD device 6-100, including cheeks, mouth, and chin.

[0080] In at least one example, the sensor system 6-102 can include a chin camera 6-116. In at least one example, the chin camera 6-116 can be disposed on the left and right sides of the HMD device 6-100 as shown and used for hand and body tracking, headset tracking, and facial avatar detection and creation for display of a user avatar on a front-facing display screen of the HMD device 6-100 as described elsewhere herein. For example, the chin camera 6-116 can be used to capture facial expressions and movements of a user's face below the HMD device 6-100, including the user's chin, cheeks, mouth, and jaw. The chin camera 6-116 can be used for hand and body tracking, headset tracking, and facial avatar detection and recreation. In at least one example, the sensor system 6-102 can include side cameras 6-118. The side cameras 6-118 can be oriented to capture left and right side views in the X-axis or direction relative to the HMD device 6-100. In at least one example, the side cameras 6-118 can be used for hand and body tracking, headset tracking, and facial avatar detection and recreation.

[0081] In at least one example, the sensor system 6-102 can include a plurality of eye tracking and gaze tracking sensors for determining identity, state, and gaze direction of the user's eyes during and / or prior to use. In at least one example, the eye / gaze tracking sensors can include a nose-eye camera 6-120 disposed on either side of the user's nose and adjacent to the user's nose when the HMD device 6-100 is worn. The eye / gaze sensors can also include bottom eye cameras 6-122 disposed below the respective user's eyes for capturing images of the eyes for facial avatar detection and creation, gaze tracking, and iris identification functions.

[0082] In at least one example, the sensor system 6-102 can include an infrared illuminator 6-124 that points outward from the HMD device 6-100 to illuminate the external environment and any objects therein with IR light for IR detection with one or more IR sensors of the sensor system 6-102. In at least one example, the sensor system 6-102 can include a flicker sensor 6-126 and an ambient light sensor 6-128. In at least one example, the flicker sensor 6-126 can detect a ceiling light refresh rate to avoid display flicker. In one example, the infrared illuminator 6-124 can include a light emitting diode and can be particularly used for low light environments for illuminating the user's hands and other objects in low light for detection by infrared sensors of the sensor system 6-102.

[0083] In at least one example, multiple sensors (including scene camera 6-106, downward camera 6-114, chin camera 6-116, side camera 6-118, depth projector 6-112, and depth sensors 6-108, 6-110) can be used in combination with an electrically coupled controller to combine depth data with camera data for hand tracking and for sizing, thereby improving the hand tracking and object recognition and tracking functions of the HMD device 6-100. In at least one example, as described above and FIG. 1I The downward-facing camera 6-114, the chin camera 6-116, and the side camera 6-118 shown can be wide-angle cameras capable of operating in both the visible and infrared spectra. In at least one example, these cameras 6-114, 6-116, and 6-118 can operate solely in black-and-white light detection to simplify image processing and achieve sensitivity.

[0084] FIGS. 1J-1L Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. FIGS. 1J-1L Any other examples of devices, features, components, and parts shown and described herein. Similarly, refer to... FIG. 1I Any of the features, components, and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. FIG. 1J Examples of devices, features, components, and parts are shown.

[0085] FIG. 1I A lower perspective view of an example HMD 6-200 including a cover or shield 6-204 fixed to a frame 6-230 is shown. In at least one example, a sensor 6-203 of a sensor system 6-202 may be disposed around the periphery of the HMD 6-200 such that the sensor 6-203 is disposed outwardly around the periphery of the display area or region 6-232 so as not to obstruct the view of the displayed light. In at least one example, the sensor may be disposed behind the shield 6-204 and aligned with a transparent portion of the shield, thereby allowing light to pass back and forth through the shield 6-204 by the sensor and the projector. In at least one example, an opaque ink or other opaque material or film / layer may be disposed on the shield 6-204 around the display area 6-232 to conceal the components of the HMD 6-200 outside the display area 6-232 rather than through a transparent portion defined by the opaque portion through which the sensor and the projector transmit and receive light and electromagnetic signals during operation. In at least one example, the shield 6-204 allows light to pass through the display (e.g., within the display area 6-232), but does not allow light to pass radially outward from the display area surrounding the periphery of the display and the shield 6-204.

[0086] In some examples, the shroud 6-204 includes a transparent portion 6-205 and an opaque portion 6-207, as described above and elsewhere herein. In at least one example, the opaque portion 6-207 of the shroud 6-204 can define one or more transparent regions 6-209 through which the sensors 6-203 of the sensor system 6-202 can transmit and receive signals. In the illustrated example, the sensors 6-203 of the sensor system 6-202 that transmit and receive signals through the shroud 6-204, or more specifically through the transparent regions 6-209 of (or defined by) the opaque portion 6-207 of the shroud 6-204, can include the same or similar sensors as those illustrated in the example of FIG. 1K FIG. 6-1, such as the depth sensors 6-108 and 6-110, the depth projector 6-112, the first and second scene cameras 6-106, the first and second downward cameras 6-114, the first and second side cameras 6-118, and the first and second infrared illuminators 6-124. These sensors are also illustrated in the example of FIG. 1L FIG. 6-2, and FIG. 1J FIG. 6-3. Other sensors, sensor types, numbers of sensors, and their relative positioning can be included in one or more other examples of an HMD.

[0087] FIG. 1I Any of the features, components, and / or parts illustrated in the examples of FIGS. 1K-1L FIG. 6-1 and FIG. 1I FIG. 6-2, and described herein, can be included individually or in any combination in any other example of an apparatus, feature, component, and part described herein. Likewise, any of the features, components, and / or parts illustrated in or described with respect to FIGS. 1K-1L FIG. 6-2 and FIG. 1J FIG. 6-3 can be included individually or in any combination in the examples of an apparatus, feature, component, and part illustrated in FIG. 1K FIG. 6-1.

[0088] FIG. 1K An elevation view of a portion of an example of an HMD apparatus 6-300 is illustrated, including a display 6-334, brackets 6-336, 6-338, and a frame or housing 6-330. FIG. 1J The example illustrated in FIG. 1K FIG. 6-2 includes a shroud 6-204 that includes an opaque portion 6-207 that would visually cover / block viewing of anything outside (e.g., radially / peripherally outward) of the display / display area 6-334, including the sensors 6-303 and the brackets 6-338.

[0089] In at least one example, various sensors of the sensor system 6-302 are coupled to the cradle 6-336, 6-338. In at least one example, the scene cameras 6-306 include tight tolerances for the angles relative to one another. For example, the tolerance for the mounting angle between two scene cameras 6-306 can be 0.5 degrees or less, such as 0.3 degrees or less. To achieve and maintain such tight tolerances, in one example, the scene cameras 6-306 can be mounted to the cradle 6-338 rather than the shroud. The cradle can include a cantilever on which the scene cameras 6-306, as well as other sensors of the sensor system 6-302, can be mounted to remain fixed in position and orientation in the event of a drop event that causes other deformation of the cradle 6-226, the housing 6-330, and / or the shroud by a user.

[0090] FIGS. 1I-1J Any of the features, components, and / or parts shown (including their arrangement and configuration) can be included individually or in any combination in FIG. 1L and FIGS. 1I-1J any other examples of devices, features, components, and parts shown and described herein. Likewise, any of the features, components, and / or parts shown or described (including their arrangement and configuration) with reference to FIG. 1L and FIG. 1K may be included individually or in any combination in FIG. 1L examples of devices, features, components, and parts shown.

[0091] FIGS. 1I-1K A bottom view of an example of an HMD 6-400 including a front display / cover assembly 6-404 and a sensor system 6-402 is illustrated. The sensor system 6-402 can be similar to other sensor systems described above and elsewhere herein, including with reference to FIG. 1L In at least one example, the chin camera 6-416 can face downward to capture images of the lower facial features of the user. In one example, the chin camera 6-416 can be directly coupled to the frame or housing 6-430 or one or more internal cradles that are directly coupled to the frame or housing 6-430 shown. The frame or housing 6-430 can include one or more holes / openings 6-415 through which the chin camera 6-416 can send and receive signals.

[0092] FIGS. 1I-1K Any of the features, components, and / or parts shown (including their arrangement and configuration) can be included individually or in any combination in FIGS. 1I-1K any other examples of devices, features, components, and parts shown and described herein. Likewise, any of the features, components, and / or parts shown or described (including their arrangement and configuration) with reference to FIG. 1LAny of the features, components, and / or parts shown and described (including arrangements and configurations thereof) can be included alone or in any combination FIG. 1M Examples of the devices, features, components, and parts shown.

[0093] FIG. 1M A rear perspective view of an interpupillary distance (IPD) adjustment system 11.1.1-102 is illustrated, including first and second optical modules 11.1.1-104a-b that are slidably engaged / coupled to respective guide rods 11.1.1-108a-b and motors 11.1.1-110a-b of left and right adjustment subsystems 11.1.1-106a-b. The IPD adjustment system 11.1.1-102 can be coupled to a cradle 11.1.1-112 and include a button 11.1.1-114 in electrical communication with the motors 11.1.1-110a-b. In at least one example, the button 11.1.1-114 can be in electrical communication with the first and second motors 11.1.1-110a-b via a processor or other circuitry component to cause the first and second motors 11.1.1-110a-b to activate and cause the first and second optical modules 11.1.1-104a-b to change positioning relative to one another, respectively.

[0094] In at least one example, the first and second optical modules 11.1.1-104a-b can include respective display screens configured to project light toward a user’s eyes when wearing the HMD 11.1.1-100. In at least one example, a user can manipulate (e.g., press and / or rotate) the button 11.1.1-114 to activate position adjustment of the optical modules 11.1.1-104a-b to match the interpupillary distance of the user’s eyes. The optical modules 11.1.1-104a-b can also include one or more cameras or other sensors / sensor systems for imaging and measuring the IPD of a user, such that the optical modules 11.1.1-104a-b can be adjusted to match the IPD.

[0095] In one example, the user manipulates the button 11.1.1-114 to cause automatic position adjustment of the first and second optical modules 11.1.1-104a-b. In one example, the user manipulates the button 11.1.1-114 to cause manual adjustment such that the optical modules 11.1.1-104a-b move further apart or closer together (e.g., when the user rotates the button 11.1.1-114 one way or another) until the user visually matches her / his own IPD. In one example, the manual adjustment is communicated electronically via one or more circuits, and power to move the optical modules 11.1.1-104a-b via the motors 11.1.1-110a-b is provided by a power source. In one example, the adjustment and movement of the optical modules 11.1.1-104a-b via the manipulation of the button 11.1.1-114 is mechanically actuated via the movement of the button 11.1.1-114.

[0096] FIG. 1M Any of the features, components, and / or parts shown, including their arrangement and configuration, can be included individually or in any combination in any other example of the devices, features, components, and parts shown in any other figure and described herein. Likewise, any of the features, components, and / or parts shown or described with reference to any other figure can be included individually or in any combination in any example of the devices, features, components, and parts shown FIG. 1N in the examples of the devices, features, components, and parts shown.

[0097] FIG. 1N A front perspective view illustrates a portion of an HMD 11.1.2-100, including an outer structural frame 11.1.2-102 and an inner or intermediate structural frame 11.1.2-104 that define a first aperture 11.1.2-106a and a second aperture 11.1.2-106b. The apertures 11.1.2-106a-b are shown in dashed lines in FIG. 1N as the view of the apertures 11.1.2-106a-b can be obstructed by one or more other components of the HMD 11.1.2-100 that are coupled to the inner frame 11.1.2-104 and / or the outer frame 11.1.2-102, as shown. In at least one example, the HMD 11.1.2-100 can include a first mounting bracket 11.1.2-108 that is coupled to the inner frame 11.1.2-104. In at least one example, the mounting bracket 11.1.2-108 is coupled to the inner frame 11.1.2-104 between the first and second apertures 11.1.2-106a-b.

[0098] Mounting brackets 11.1.2-108 may include intermediate or central portions 11.1.2-109 coupled to the inner frame 11.1.2-104. In some examples, the intermediate or central portions 11.1.2-109 may not be the geometric center or middle of the brackets 11.1.2-108. Instead, the intermediate / central portions 11.1.2-109 may be positioned between a first cantilever extension arm and a second cantilever extension arm extending away from the intermediate portions 11.1.2-109. In at least one example, mounting bracket 108 includes first cantilever arms 11.1.2-112 and second cantilever arms 11.1.2-114 extending away from the intermediate portions 11.1.2-109 of the mounting brackets 11.1.2-108 coupled to the inner frame 11.1.2-104.

[0099] like FIG. 1N As shown, the outer frame 11.1.2-102 may define a curved geometry on its underside to adapt to the user's nose when the user wears the HMD 11.1.2-100. This curved geometry may be referred to as the bridge of the nose 11.1.2-111 and is centrally located on the underside of the HMD 11.1.2-100 as shown. In at least one example, the mounting bracket 11.1.2-108 may be connected to the inner frame 11.1.2-104 between holes 11.1.2-106a and b, such that the cantilever 11.1.2-112, 11.1.2-114 extend downward and laterally outward away from the central portion 11.1.2-109 to complement the nose bridge geometry of the outer frame 11.1.2-102. In this way, the mounting bracket 11.1.2-108 is configured to adapt to the user's nose, as described above. The geometry of the bridge of the nose 11.1.2-111 adapts to the nose, as it provides a curvature that conforms to the shape of the user's nose, offering a comfortable fit from above, above, and around.

[0100] The first cantilever 11.1.2-112 can extend away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-108 in a first direction, and the second cantilever 11.1.2-114 can extend away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-108 in a second direction opposite the first direction. The first cantilever 11.1.2-112 and the second cantilever 11.1.2-114 are referred to as“cantilevered” or“cantilever” arms because each arm 11.1.2-112, 11.1.2-114 includes a free distal end 11.1.2-116, 11.1.2-118, respectively, that is not attached to the inner frame 11.1.2-102 and the outer frame 11.1.2-104. In this way, the arms 11.1.2-112, 11.1.2-114 cantilever from the middle portion 11.1.2-109, which can be connected to the inner frame 11.1.2-104, while the distal ends 11.1.2-102, 11.1.2-104 are unattached.

[0101] In at least one example, the HMD 11.1.2-100 can include one or more components coupled to the mounting bracket 11.1.2-108. In one example, the components include a plurality of sensors 11.1.2-110a-f. Each of the plurality of sensors 11.1.2-110a-f can include various types of sensors, including cameras, IR sensors, and the like. In some examples, one or more of the sensors 11.1.2-110a-f can be used for object recognition in three-dimensional space, such that it is important to maintain the precise relative positioning of two or more of the plurality of sensors 11.1.2-110a-f. The cantilevered nature of the mounting bracket 11.1.2-108 can protect the sensors 11.1.2-110a-f from damage and altered positioning in the event of an accidental drop by the user. Because the sensors 11.1.2-110a-f are overhanging on the arms 11.1.2-112, 11.1.2-114 of the mounting bracket 11.1.2-108, stresses and deformations of the inner and / or outer frames 11.1.2-104, 11.1.2-102 are not transmitted to the cantilevers 11.1.2-112, 11.1.2-114, and thus do not affect the relative positions of the sensors 11.1.2-110a-f coupled / mounted to the mounting bracket 11.1.2-108.

[0102] FIG. 1NAny of the illustrated features, components, and / or parts, including their arrangement and configuration, can be included in any other example of an apparatus, feature, component described herein, either alone or in any combination. Likewise, any of the features, components, and / or parts illustrated and described herein, including their arrangement and configuration, can be included in any example of an apparatus, feature, component FIG. 1O Examples of the illustrated apparatus, features, components, and parts.

[0103] FIG. 1O Examples of optical modules 11.3.2-100 for use in electronic devices, such as HMDs, including the HDM apparatus described herein, are illustrated. As shown in one or more other examples described herein, optical module 11.3.2-100 can be one of two optical modules within an HMD, with each optical module aligned to project light toward an eye of a user. In this way, a first optical module can project light toward a first eye of a user via a display screen, and a second optical module of the same device can project light toward a second eye of the user via another display screen.

[0104] In at least one example, optical module 11.3.2-100 can include an optical frame or housing 11.3.2-102, which can also be referred to as a barrel or optical module barrel. Optical module 11.3.2-100 can also include a display 11.3.2-104 coupled to housing 11.3.2-102, which includes one or more display screens. Display 11.3.2-104 can be coupled to housing 11.3.2-102 such that display 11.3.2-104 is configured to project light toward an eye of a user when wearing an HMD to which display module 11.3.2-100 belongs during use. In at least one example, housing 11.3.2-102 can surround display 11.3.2-104 and provide connection features for coupling other components of an optical module described herein.

[0105] In one example, optical module 11.3.2-100 can include one or more cameras 11.3.2-106 coupled to housing 11.3.2-102. Cameras 11.3.2-106 can be positioned relative to display 11.3.2-104 and housing 11.3.2-102 such that cameras 11.3.2-106 are configured to capture one or more images of a user’s eyes during use. In at least one example, optical module 11.3.2-100 can also include a light bar 11.3.2-108 that surrounds display 11.3.2-104. In one example, light bar 11.3.2-108 is disposed between display 11.3.2-104 and cameras 11.3.2-106. Light bar 11.3.2-108 can include a plurality of lights 11.3.2-110. The plurality of lights can include one or more light-emitting diodes (LEDs) or other lights configured to project light toward a user’s eyes when the HMD is worn. Individual lights 11.3.2-110 in light bar 11.3.2-108 can be spaced apart around light bar 11.3.2-108, and thus, evenly or unevenly, around display 11.3.2-104 at various locations on light bar 11.3.2-108 and around display 11.3.2-104.

[0106] In at least one example, housing 11.3.2-102 defines a viewing opening 11.3.2-101 through which a user can view display 11.3.2-104 when the HMD device is worn. In at least one example, the LEDs are configured and arranged to emit light through viewing opening 11.3.2-101 onto a user’s eyes. In one example, cameras 11.3.2-106 are configured to capture one or more images of a user’s eyes through viewing opening 11.3.2-101.

[0107] As noted above, FIG. 1O Each of the components and features of optical module 11.3.2-100 shown can be replicated in another (e.g., second) optical module provided with the HMD to interact with (e.g., project light and capture images of) the other eye of the user.

[0108] FIG. 1P Any of the features, components, and / or parts shown (including their arrangement and configuration) can be included, alone or in any combination, in FIG. 1P Any other example of the devices, features, components, and parts shown or otherwise described herein. Likewise, reference to FIG. 1O Any of the features, components, and / or parts shown or otherwise described herein (including their arrangement and configuration) can be included, alone or in any combination, in FIG. 1PExamples of the devices, features, components, and parts shown.

[0109] FIG. 1P A cross-sectional view illustrating an example of an optical module 11.3.2-200 is shown, including a housing 11.3.2-202, a display assembly 11.3.2-204 coupled to the housing 11.3.2-202, and a lens 11.3.2-216 coupled to the housing 11.3.2-202. In at least one example, the housing 11.3.2-202 defines a first aperture or channel 11.3.2-212 and a second aperture or channel 11.3.2-214. The channels 11.3.2-212, 11.3.2-214 can be configured to slidably engage corresponding rails or guide rods of an HMD device to allow the optical module 11.3.2-200 to adjust positioning relative to a user’s eyes to match the user’s interpupillary distance (IPD). The housing 11.3.2-202 can slidably engage the guide rods to secure the optical module 11.3.2-200 in place within the HMD.

[0110] In at least one example, the optical module 11.3.2-200 can also include a lens 11.3.2-216 coupled to the housing 11.3.2-202 and disposed between the display assembly 11.3.2-204 and the user’s eyes when the HMD is worn. The lens 11.3.2-216 can be configured to direct light from the display assembly 11.3.2-204 to the user’s eyes. In at least one example, the lens 11.3.2-216 can be part of a lens assembly, including a corrective lens removably attached to the optical module 11.3.2-200. In at least one example, the lens 11.3.2-216 is disposed above the light bar 11.3.2-208 and the one or more eye tracking cameras 11.3.2-206, such that the cameras 11.3.2-206 are configured to capture images of the user’s eyes through the lens 11.3.2-216, and the light bar 11.3.2-208 includes lights configured to project light to the user’s eyes through the lens 11.3.2-216 during use.

[0111] FIG. 1P Any of the features, components, and / or parts shown (including arrangements and configurations thereof) can be included in any other example of a device, feature, component, and part described herein, alone or in any combination. Likewise, any of the features, components, and / or parts shown and described herein (including arrangements and configurations thereof) can be included in any other example of a device, feature, component, and part described herein, alone or in any combination. FIG. 2 Examples of the devices, features, components, and parts shown.

[0112] FIG. 1Ais a block diagram of an example of a controller 110 according to some embodiments. While certain specific features are illustrated, those skilled in the art will appreciate from the present disclosure that various other features have not been illustrated for the sake of brevity and so as not to obscure more pertinent aspects of the embodiments disclosed herein. To that end, as a non-limiting example, in some embodiments, the controller 110 includes one or more processing units 202 (e.g., microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), graphics processing units (GPUs), central processing units (CPUs), processing cores, and / or the like), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., universal serial bus (USB), FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, global system for mobile communications (GSM), code division multiple access (CDMA), time division multiple access (TDMA), global positioning system (GPS), infrared (IR), BLUETOOTH, ZIGBEE, and / or the like type of interface), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these and various other components.

[0113] In some embodiments, the one or more communication buses 204 include circuitry that interconnects and controls communications between system components. In some embodiments, the one or more I / O devices 206 include at least one of a keyboard, mouse, touchpad, joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.

[0114] The memory 220 includes high-speed random access memory, such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate random access memory (DDR RAM), or other random access solid state memory devices. In some embodiments, the memory 220 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid state storage devices. The memory 220 optionally includes one or more storage devices remotely located from the one or more processing units 202. The memory 220 comprises a non-transitory computer readable storage medium. In some embodiments, the memory 220, or the non-transitory computer readable storage medium of the memory 220, stores the following programs, modules, and data structures, or a subset thereof, including an optional operating system 230 and an XR experience module 240.

[0115] The operating system 230 includes instructions for handling various basic system services and for executing hardware-related tasks. In some embodiments, the XR experience module 240 is configured to manage and coordinate single or multiple XR experiences for one or more users (e.g., a single XR experience for one or more users, or multiple XR experiences for respective groups of one or more users). To this end, in various embodiments, the XR experience module 240 includes a data acquisition unit 241, a tracking unit 242, a coordination unit 246, and a data transmission unit 248.

[0116] In some embodiments, the data acquisition unit 241 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the display generation component 120, and optionally from one or more of the input devices 125, the output devices 155, the sensors 190, and / or the peripheral devices 195. To this end, in various embodiments, the data acquisition unit 241 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics. FIG. 1A

[0117] In some embodiments, the tracking unit 242 is configured to map the scene 105, and to track the positioning / location of at least the display generation component 120 relative to the scene 105, and optionally to track the position of one or more of the input devices 125, the output devices 155, the sensors 190, and / or the peripheral devices 195. To this end, in various embodiments, the tracking unit 242 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics. In some embodiments, the tracking unit 242 includes a hand tracking unit 244 and / or an eye tracking unit 243. In some embodiments, the hand tracking unit 244 is configured to track the positioning / location of one or more portions of a user’s hand, and / or the motion of one or more portions of a user’s hand relative to the scene 105, relative to the display generation component 120, and / or relative to a coordinate system that is defined relative to the user’s hand. FIG. 1A FIG. 4 The hand tracking unit 244 is described in greater detail below with respect to the hand tracking unit 244. In some embodiments, the eye tracking unit 243 is configured to track the positioning or movement of a user’s gaze (or more broadly, the user’s eyes, face, or head) relative to the scene 105 (e.g., relative to the physical environment and / or relative to the user (e.g., the user’s hand)) or relative to XR content displayed via the display generation component 120. The eye tracking unit 243 is described in greater detail below with respect to the eye tracking unit 243. FIG. 5 FIG. 2

[0118] ​​​​In some embodiments, the coordination unit 246 is configured to manage and coordinate the XR experience presented to the user by the display generation component 120, and optionally one or more of the input devices 125, the output devices 155, and / or the peripheral devices 195. To this end, in various embodiments, the coordination unit 246 includes instructions and / or logic therefor, and heuristics and metadata therefor.

[0119] In some embodiments, the data sending unit 248 is configured to send data (e.g., presentation data, position data, etc.) to at least the display generation component 120, and optionally to one or more of the input devices 125, the output devices 155, the sensors 190, and / or the peripheral devices 195. To this end, in various embodiments, the data sending unit 248 includes instructions and / or logic therefor, and heuristics and metadata therefor.

[0120] Although the data acquisition unit 241, the tracking unit 242 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the coordination unit 246, and the data sending unit 248 are shown as residing on a single device (e.g., the controller 110), it will be appreciated that, in other embodiments, any combination of the data acquisition unit 241, the tracking unit 242 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the coordination unit 246, and the data sending unit 248 can reside in separate computing devices.

[0121] Further, FIG. 2 More functionally described as various features that can be present in particular implementations, as distinct from structural illustrations of embodiments described herein. As will be appreciated by those of ordinary skill in the art, items shown separately could be combined and some items could be separated. For example, FIG. 3 Some of the functional modules shown separately in the can be implemented in a single module, and various functions of a single functional block can be implemented by one or more functional blocks in various embodiments. The actual number of modules and the division of particular functions between the modules can vary from embodiment to embodiment depending on the hardware, software, and / or firmware chosen for a particular implementation, and in some embodiments, depends in part on the particular combination of hardware, software, and / or firmware selected for the particular implementation.

[0122] FIG. 1Ais a block diagram of an example of a display generation component 120 according to some embodiments. While certain specific features are illustrated, one of ordinary skill in the art will appreciate from the present disclosure that the application also contemplates various other features. Hence, as a non-limiting example, in some embodiments, the display generation component 120 (e.g., an HMD) includes one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.1 lx, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, BLUETOOTH, ZIGBEE, and / or the like type of interface), one or more programming (e.g., I / O) interfaces 310, one or more XR displays 312, one or more optional inward- and / or outward-facing image sensors 314, memory 320, and one or more communication buses for interconnecting these and various other components, and

[0123] In some embodiments, the one or more communication buses 304 include circuitry that interconnects and controls communications between system components. In some embodiments, the one or more I / O devices and sensors 306 include an inertial measurement unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., blood pressure monitor, heart rate monitor, blood oxygen sensor, blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptics engine, and / or one or more depth sensors (e.g., structured light, time-of-flight, etc.), among others.

[0124] In some embodiments, the one or more XR displays 312 are configured to provide an XR experience to a user. In some embodiments, the one or more XR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transitory (OLET), organic light-emitting diode (OLED), surface-conduction electron-emitter display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), micro-electromechanical system (MEMS), and / or similar display types. In some embodiments, the one or more XR displays 312 correspond to diffractive, reflective, polarized, holographic, etc. waveguide displays. For example, the display generation component 120 (e.g., HMD) includes a single XR display. As another example, the display generation component 120 includes an XR display for each eye of the user. In some embodiments, the one or more XR displays 312 are capable of presenting MR and VR content. In some embodiments, the one or more XR displays 312 are capable of presenting MR or VR content.

[0125] In some embodiments, the one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user’s face, including the user’s eyes (and can be referred to as eye tracking cameras). In some embodiments, the one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user’s hands and, optionally, the user’s arms (and can be referred to as hand tracking cameras). In some embodiments, the one or more image sensors 314 are configured to face forward so as to acquire image data corresponding to a scene that the user would see in the absence of the display generation component 120 (e.g., HMD) (and can be referred to as a scene camera). The one or more optional image sensors 314 can include one or more RGB cameras (e.g., with a complementary metal-oxide-semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), one or more infrared (IR) cameras, and / or one or more event-based cameras, etc.

[0126] Memory 320 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM or other random access solid state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid state storage devices. Memory 320 optionally includes one or more storage devices remotely located from the one or more processing units 302. Memory 320 comprises a non-transitory computer readable storage medium. In some embodiments, memory 320 or the non-transitory computer readable storage medium of memory 320 stores the following programs, modules, and data structures, among others:

[0127] Operating system 330 includes instructions for handling various basic system services and for performing hardware dependent tasks. In some embodiments, XR presentation module 340 is configured to present XR content to a user via one or more XR displays 312. To do so, in various embodiments, XR presentation module 340 includes a data acquisition unit 342, an XR presentation unit 344, an XR mapping generation unit 346, and a data transmission unit 348.

[0128] In some embodiments, data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least FIG. 1A controller 110. To do so, in various embodiments, data acquisition unit 342 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.

[0129] In some embodiments, XR presentation unit 344 is configured to present XR content via one or more XR displays 312. To do so, in various embodiments, XR presentation unit 344 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.

[0130] In some embodiments, XR mapping generation unit 346 is configured to generate an XR map (e.g., a 3D map of a mixed reality scene or a map of a physical environment in which computer-generated objects can be placed to generate an extended reality) based on media content data. To do so, in various embodiments, XR mapping generation unit 346 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.

[0131] In some embodiments, data sending unit 348 is configured to send data (e.g., presentation data, location data, etc.) to at least controller 110, and optionally to one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, data sending unit 348 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.

[0132] Although data acquisition unit 342, XR presentation unit 344, XR mapping generation unit 346, and data sending unit 348 are shown as residing on a single device (e.g., display generation component 120 of system 100), FIG. 3 in other embodiments, any combination of data acquisition unit 342, XR presentation unit 344, XR mapping generation unit 346, and data sending unit 348 can be located in separate computing devices.

[0133] Further, FIG. 3 More functionally described as various features that can be present in a particular implementation, rather than as structural illustrations of embodiments described herein. As will be appreciated by one of ordinary skill in the art, items shown separately could be combined, and items shown separately could be divided. For example, FIG. 4 Some of the functional modules shown separately in FIG. 10 can be implemented in a single module, and various functions of a single functional block can be implemented by one or more functional blocks in various embodiments. The actual number of modules and the division of particular functions between them, and how features are allocated among them, will vary from one implementation to another and, in some embodiments, depends in part on the particular combination of hardware, software, and / or firmware chosen to implement the particular implementation.

[0134] FIG. 1A is a schematic illustration of an example implementation of hand tracking device 140. In some embodiments, hand tracking device 140 is controlled by hand tracking unit 244 to track the positioning / location of one or more portions of a user’s hand, and / or the orientation of one or more portions of a user’s hand relative to FIG. 2 FIG. 1A FIG. 4 ​​motion of the scene 105 relative to a portion of the physical environment surrounding the user, relative to the display generation component 120, or relative to a portion of the user (e.g., the user’s face, eyes, or head), and / or relative to a coordinate system that is defined relative to the user’s hand. In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).

[0135] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras, etc.) that captures three-dimensional scene information including at least a human user’s hand 406. The image sensor 404 captures hand images with sufficient resolution to enable the fingers and their respective positions to be distinguished. The image sensor 404 typically captures images of other portions of the user’s body, also or possibly all portions of the body, and can have zoom capability or a dedicated sensor with increased magnification to capture images of the hand with the desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with other image sensors to capture the physical environment of the scene 105, or as the image sensor that captures the physical environment of the scene 105. In some embodiments, the image sensor 404 or a portion thereof is positioned relative to the user or the user’s environment in a manner that uses the field of view of the image sensor 404 for defining an interaction space in which movement of the hand captured by the image sensor is treated as input to the controller 110.

[0136] In some embodiments, the image sensor 404 outputs a sequence of frames containing 3D map data (and, in addition, possibly color image data) to the controller 110, which extracts high-level information from the map data. This high-level information is typically provided to applications running on the controller via an application program interface (API), which in turn drives the display generation component 120. For example, a user can interact with software running on the controller 110 by moving his hand 406 and changing his hand pose.

[0137] In some embodiments, image sensor 404 projects a pattern of dots onto a scene containing hand 406 and captures an image of the projected pattern. In some embodiments, controller 110 computes 3D coordinates of points in the scene (including points on the surface of the user's hand) based on lateral shifts of the dots in the pattern by triangulation. This approach is advantageous because it does not require the user to hold or wear any kind of beacon, sensor, or other marker. It gives depth coordinates of points in the scene at a particular distance from image sensor 404 relative to a predetermined reference plane. In this disclosure, it is assumed that image sensor 404 defines an orthogonal set of x, y, z axes so that the depth coordinates of points in the scene correspond to the z component measured by the image sensor. Alternatively, image sensor 404 (e.g., a hand tracking device) can use other 3D mapping methods such as stereo imaging or time-of-flight measurements based on a single or multiple cameras or other types of sensors.

[0138] In some embodiments, hand tracking device 140 captures and processes a time series of depth maps containing the user's hand as he moves his hand (e.g., the entire hand or one or more fingers). Software running on a processor in image sensor 404 and / or controller 110 processes the 3D map data to extract image patch descriptors of the hand in these depth maps. The software can match these descriptors to image patch descriptors stored in database 408 based on a previous learning process in order to estimate the pose of the hand in each frame. The pose typically includes the 3D positions of the user's hand joints and finger tips.

[0139] The software can also analyze the trajectories of the hand and / or fingers over multiple frames in the sequence to identify gestures. The pose estimation functionality described herein can alternate with motion tracking functionality so that image patch based pose estimation is performed only once every two (or more) frames, while tracking is used to find changes in pose that occur over the remaining frames. The pose, motion, and gesture information is provided to applications running on controller 110 via the API described above. The program can move and modify images presented on display generation component 120, for example, in response to the pose and / or gesture information, or perform other functions.

[0140] In some embodiments, a gesture includes an in-air gesture. An in-air gesture is a gesture that is detected without the user touching an input element that is part of a device (e.g., computer system 101, one or more input devices 125, and / or hand tracking device 140) (or independent of an input element that is part of a device) and is based on detected movement of a portion of the user’s body (e.g., head, one or both arms, one or both hands, one or more fingers, and / or one or both legs) through the air (including movement of the user’s body relative to an absolute reference (e.g., angle of a user’s arm relative to the ground or distance of a user’s hand from the ground), relative to another portion of the user’s body (e.g., movement of a user’s hand relative to a user’s shoulder, movement of one of a user’s hands relative to the other of a user’s hands, and / or movement of a user’s finger relative to another finger or portion of a hand), and / or absolute movement of a portion of the user’s body (e.g., a tap gesture including a hand moving a predetermined amount and / or speed in a predetermined pose, or a shake gesture including a predetermined speed or amount of rotation of a portion of the user’s body)).

[0141] In some embodiments, input gestures used in various examples and embodiments described herein include in-air gestures performed by movement of a user’s finger relative to other fingers or portions of the user’s hand for interacting with an XR environment (e.g., a virtual or mixed reality environment), in accordance with some embodiments. In some embodiments, an in-air gesture is a gesture that is detected without the user touching an input element that is part of a device (or independent of an input element that is part of a device) and is based on detected movement of a portion of the user’s body through the air (including movement of the user’s body relative to an absolute reference (e.g., angle of a user’s arm relative to the ground or distance of a user’s hand from the ground), relative to another portion of the user’s body (e.g., movement of a user’s hand relative to a user’s shoulder, movement of one of a user’s hands relative to the other of a user’s hands, and / or movement of a user’s finger relative to another finger or portion of a hand), and / or absolute movement of a portion of the user’s body (e.g., a tap gesture including a hand moving a predetermined amount and / or speed in a predetermined pose, or a shake gesture including a predetermined speed or amount of rotation of a portion of the user’s body).

[0142] In some embodiments in which the input gesture is an in-air gesture (e.g., in the absence of physical contact with an input device that provides information to the computer system about which user interface element is the target of the user input, such as contact with a user interface element displayed on a touchscreen, or contact with a mouse or trackpad to move a cursor to a user interface element), the gesture takes into account the user's attention (e.g., gaze) to determine the target of the user input (e.g., for direct input, as described below). Thus, in implementations involving in-air gestures, for example, the input gesture is combined with (e.g., simultaneous with) detection of attention (e.g., gaze) toward a user interface element to perform pinch and / or tap input, as described in greater detail below.

[0143] In some embodiments, an input gesture directed at a user interface object is performed directly or indirectly with reference to the user interface object. For example, an input gesture is performed directly on a user interface object in accordance with performing the input gesture at a location corresponding to the location of the user interface object in a three-dimensional environment (e.g., as determined based on the user's current viewpoint). In some embodiments, an input gesture is performed indirectly on a user interface object in accordance with the location of the user's hand not being at the location corresponding to the location of the user interface object in a three-dimensional environment at the same time the user performs the input gesture when the user's attention (e.g., gaze) is detected to be on the user interface object. For example, for a direct input gesture, the user can direct the user's input to a user interface object by initiating the gesture at or near a location corresponding to the displayed location of the user interface object (e.g., within 0.5 cm, 1 cm, 5 cm, or a distance between 0 and 5 cm, measured from the outer edge of the option or the center portion of the option). For an indirect input gesture, the user can direct the user's input to a user interface object by focusing on the user interface object (e.g., by gazing at the user interface object), and while focusing on the option, the user initiates the input gesture (e.g., at any location that the computer system can detect) (e.g., at a location that does not correspond to the displayed location of the user interface object).

[0144] In some embodiments, input gestures (e.g., in-air gestures) used in various examples and embodiments described herein include pinch input and tap input for interacting with virtual or mixed reality environments, in accordance with some embodiments. For example, the pinch input and tap input described below are performed as in-air gestures.

[0145] In some embodiments, the pinch input is part of an in-air gesture that includes one or more of: a pinch gesture, a long pinch gesture, a pinch-and-drag gesture, or a double pinch gesture. For example, a pinch gesture as an in-air gesture includes movement of two or more fingers of a hand into contact with each other, i.e., optionally followed by an immediate (e.g., within 0-1 seconds) break in contact with each other. A long pinch gesture as an in-air gesture includes movement of two or more fingers of a hand into contact with each other for at least a threshold amount of time (e.g., at least 1 second) before a break in contact with each other is detected. For example, a long pinch gesture includes a user holding a pinch gesture (e.g., with two or more fingers in contact), and the long pinch gesture continues until a break in contact between the two or more fingers is detected. In some embodiments, a double pinch gesture as an in-air gesture includes two (e.g., or more) pinch inputs (e.g., performed by the same hand) that are detected consecutively immediately (e.g., within a predefined time period) of each other. For example, a user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., breaks contact between the two or more fingers), and performs a second pinch input within a predefined time period (e.g., within 1 second or within 2 seconds) after releasing the first pinch input.

[0146] In some embodiments, a pinch-and-drag gesture as an in-air gesture includes a pinch gesture (e.g., a pinch gesture or a long pinch gesture) performed in conjunction with (e.g., following) a drag input that changes a position of a user’s hand from a first position (e.g., a start position of the drag) to a second position (e.g., an end position of the drag). In some embodiments, a user holds the pinch gesture while performing the drag input, and releases the pinch gesture (e.g., opens their two or more fingers) to end the drag gesture (e.g., at the second position). In some embodiments, the pinch input and the drag input are performed by the same hand (e.g., a user pinches two or more fingers into contact with each other and moves the same hand into the air to a second position with a drag gesture). In some embodiments, the pinch input is performed by a first hand of a user, and the drag input is performed by a second hand of the user (e.g., while the user continues the pinch input with the first hand of the user, the second hand of the user moves in the air from the first position to the second position. In some embodiments, an input gesture as an in-air gesture includes input (e.g., pinch and / or tap input) performed using both hands of a user. For example, an input gesture includes two (e.g., or more) pinch inputs performed in conjunction (e.g., simultaneously or within a predefined time period) of each other. For example, a first pinch gesture (e.g., a pinch input, a long pinch input, or a pinch-and-drag input) is performed using a first hand of a user, and a second pinch input is performed using another hand (e.g., a second hand of the two hands of the user) in conjunction with the pinch input performed using the first hand.

[0147] In some embodiments, a tap input (e.g., a pointing to a user interface element) performed as an air gesture includes movement of a user's finger toward the user interface element, movement of the user's hand toward the user interface element (optionally, the user's finger extending toward the user interface element), a downward motion of the user's finger (e.g., mimicking a mouse click motion or a tap on a touch screen), or other predefined movement of the user's hand. In some embodiments, a tap input performed as an air gesture is detected based on movement characteristics of a finger or hand performing a tap gesture movement that is a finger or hand moving away from the user's point of view and / or toward an object that is the target of the tap input, followed by an end of the movement. In some embodiments, the end of the movement is detected based on a change in movement characteristics of the finger or hand performing the tap gesture (e.g., an end of movement away from the user's point of view and / or toward an object that is the target of the tap input, a reversal of the direction of movement of the finger or hand, and / or a reversal of the direction of acceleration of the movement of the finger or hand).

[0148] In some embodiments, a user's attention is determined to be directed to a portion of a three-dimensional environment based on detection of a gaze directed to the portion of the three-dimensional environment (optionally, without requiring other conditions). In some embodiments, a user's attention is determined to be directed to a portion of a three-dimensional environment based on detection of a gaze directed to the portion of the three-dimensional environment with one or more additional conditions, such as requiring the gaze to be directed to the portion of the three-dimensional environment for at least a threshold duration (e.g., a dwell duration) and / or requiring the gaze to be directed to the portion of the three-dimensional environment while the user's point of view is within a distance threshold from the portion of the three-dimensional environment for the device to determine that the user's attention is directed to the portion of the three-dimensional environment, where if one of these additional conditions is not met, the device determines that the attention is not directed to the portion of the three-dimensional environment that the gaze is directed to (e.g., until the one or more additional conditions are met).

[0149] In some embodiments, detection of the readiness state configuration of the user or a portion of the user is detected by the computer system. Detection of the readiness state configuration of the hand is used by the computer system as an indication that the user can be preparing to use one or more mid-air hand gesture inputs performed by the hand (e.g., pinch, tap, pinch-and-drag, double pinch, long pinch, or other mid-air hand gestures described herein) to interact with the computer system. For example, the readiness state of the hand is determined based on whether the hand has a predetermined hand shape (e.g., a pre-pinch shape in which the thumb and one or more fingers are extended and spaced apart in preparation to make a pinch or grasp gesture, or a pre-tap in which one or more fingers are extended and the palm is facing away from the user), based on whether the hand is in a predetermined position relative to the user's point of view (e.g., below the user's head and above the user's waist and extending at least 15 cm, 20 cm, 25 cm, 30 cm, or 50 cm from the body), and / or based on whether the hand has moved in a particular manner (e.g., moved toward an area in front of the user above the user's waist and below the user's head or moved away from the user's body or legs). In some embodiments, the readiness state is used to determine whether an interactive element of a user interface is responsive to attention (e.g., gaze) inputs.

[0150] In scenarios where inputs are described with reference to in-air gestures, it will be understood that similar gestures can be detected using a hardware input device attached to or held by one or both hands of a user, where the position of the hardware input device in space can be tracked using optical tracking, one or more accelerometers, one or more gyroscopes, one or more magnetometers, and / or one or more inertial measurement units, and the position and / or movement of the hardware input device is used in place of the position and / or movement of the one or both hands in the corresponding in-air gesture. In scenarios where inputs are described with reference to in-air poses, it will be understood that similar poses can be detected using a hardware input device attached to or held by one or both hands of a user. User inputs can be detected with controls contained in the hardware input device, such as one or more touch-sensitive input elements, one or more pressure-sensitive input elements, one or more buttons, one or more knobs, one or more dials, one or more joysticks, one or both hands or finger coverings that can detect the position or change in position of the hands and / or fingers relative to each other, relative to the user’s body, and / or relative to the user’s physical environment, and / or other hardware input device controls, where user inputs made with controls contained in the hardware input device are used in place of hand and / or finger gestures such as in-air taps or in-air pinches in the corresponding in-air gesture. For example, a selection input described as being performed with an in-air tap or in-air pinch can alternatively be detected with a button press, a tap on a touch-sensitive surface, a press on a pressure-sensitive surface, or other hardware input. As another example, a movement input described as being performed with an in-air pinch and drag can alternatively be detected based on interaction with a hardware input control such as a button press and hold, a touch on a touch-sensitive surface, a press on a pressure-sensitive surface, or other hardware input following movement of the hardware input device (e.g., along with the hand associated with the hardware input device) through space. Similarly, two-handed inputs that include movement of the hands relative to each other can be performed with one in-air gesture and one hardware input device held in a hand that is not performing the in-air gesture, two hardware input devices held in different hands, or two in-air gestures performed by different hands and / or various combinations of inputs detected by one or more of the above hardware input devices.

[0151] In some embodiments, the software can be downloaded to the controller 110 in electronic form, over a network, for example, or it can alternatively be provided on a tangible, non-transitory medium, such as optical, magnetic or electronic memory media. In some embodiments, the database 408 is likewise stored in memory associated with the controller 110. Alternatively or additionally, some or all of the described functionality of the computer can be implemented in dedicated hardware, such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP), for example. Although the described functionality is implemented in one or more computers in the example of FIG. 4, it will be appreciated that the functionality can be implemented in other ways, such as in one or more special-purpose hardware components, for example. FIG. 4The controller 110 is shown, but for example, as a separate unit from the image sensor 404, some or all of the controller's processing functions may be performed by a suitable microprocessor and software, or by dedicated circuitry within the housing of the image sensor 404 (e.g., a hand-tracking device), or by other devices associated with the image sensor 404. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with the display generation component 120 (e.g., in a television receiver, handheld device, or head-mounted device) or with any other suitable computerized device (such as a game console or media player). The sensing function of the image sensor 404 may also be integrated into a computer or other computerized device controlled by the sensor output.

[0152] FIG. 4 It also includes a schematic diagram of a depth map 410 captured by image sensor 404 according to some embodiments. As described above, the depth map comprises a matrix of pixels with corresponding depth values. Pixel 412 corresponding to hand 406 has been segmented from the background and wrist in the map. The brightness of each pixel within the depth map 410 is inversely proportional to its depth value (i.e., the measured z-distance from image sensor 404), where gray shadows become darker as depth increases. Controller 110 processes these depth values ​​to identify and segment components of the image that have human hand characteristics (i.e., a group of adjacent pixels). These characteristics may include, for example, overall size, shape, and frame-to-frame motion from the depth map sequence.

[0153] FIG. 4 The controller 110 also schematically illustrates, according to some embodiments, the hand skeleton 414 ultimately extracted from the depth map 410 of the hand 406. FIG. 5 In this configuration, the hand skeleton 414 is superimposed on the hand background 416, which has already been segmented from the original depth map. In some embodiments, key feature points of the hand, and optionally on the wrist or arm connected to the hand (e.g., points corresponding to knuckles, fingertips, the center of the palm, the end of the hand connecting to the wrist, etc.), are identified and located on the hand skeleton 414. In some embodiments, the controller 110 uses the position and movement of these key feature points across multiple image frames to determine, according to some embodiments, the gesture performed by the hand or the current state of the hand.

[0154] FIG. 1A An eye-tracking device 130 is illustrated. FIG. 2 Example implementation of ). In some implementations, the eye-tracking device 130 consists of an eye-tracking unit 243 ( FIG. 5) control to track the position and movement of the user’s gaze relative to the scene 105 or relative to XR content displayed via the display generation component 120. In some embodiments, the eye tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, when the display generation component 120 is a head-mounted device, such as a headset, helmet, goggles, or glasses, or a handheld device that is placed in a wearable frame, the head-mounted device includes both components that generate XR content for the user to view and components to track the user’s gaze relative to the XR content. In some embodiments, the eye tracking device 130 is separate from the display generation component 120. For example, when the display generation component is a handheld device or an XR room, the eye tracking device 130 is optionally a device that is separate from the handheld device or the XR room. In some embodiments, the eye tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, the head-mounted eye tracking device 130 is used in combination with a display generation component that is also head-mounted or a display generation component that is not head-mounted. In some embodiments, the eye tracking device 130 is not a head-mounted device and is used in combination with a head-mounted display generation component. In some embodiments, the eye tracking device 130 is not a head-mounted device and is part of a non-head-mounted display generation component.

[0155] In some embodiments, the display generation component 120 uses display mechanisms (e.g., left and right near-eye display panels) to display frames including left and right images in front of the user’s eyes, providing a 3D virtual view to the user. For example, a head-mounted display generation component can include left and right optical lenses (referred to herein as eye lenses) positioned between the display and the user’s eyes. In some embodiments, the display generation component can include or be coupled to one or more external cameras that capture video of the user’s environment for display. In some embodiments, a head-mounted display generation component can have a transparent or semi-transparent display and display virtual objects on the transparent or semi-transparent display through which the user can directly view the physical environment. In some embodiments, the display generation component projects virtual objects into the physical environment. The virtual objects may, for example, be projected on a physical surface or as a hologram, such that an individual observing the virtual objects using the system overlaps above the physical environment. In this case, separate display panels and image frames for the left and right eyes can not be needed.

[0156] As FIG. 5As shown, in some embodiments, eye tracking device 130 (e.g., a gaze tracking device) includes at least one eye tracking camera (e.g., an infrared (IR) or near infrared (NIR) camera), and an illumination source (e.g., an IR or NIR light source, such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user’s eyes. The eye tracking camera can be pointed at the user’s eyes to receive IR or NIR light that is directly reflected from the eyes by the light source, or alternatively can be pointed at “hot” mirrors that are positioned between the user’s eyes and the display panel that reflect IR or NIR light from the eyes to the eye tracking camera while allowing visible light to pass through. Eye tracking device 130 optionally captures images of the user’s eyes (e.g., as a video stream captured at 60-120 frames per second (fps)), analyzes the images to generate gaze tracking information, and communicates the gaze tracking information to controller 110. In some embodiments, both of the user’s eyes are tracked separately by respective eye tracking cameras and illumination sources. In some embodiments, only one of the user’s eyes is tracked by a respective eye tracking camera and illumination source.

[0157] In some embodiments, a device-specific calibration procedure is used to calibrate eye tracking device 130 to determine parameters of the eye tracking device for a particular operating environment 100, such as 3D geometry and parameters of the LEDs, cameras, hot mirrors (if present), eye lenses, and display screen. The device-specific calibration procedure can be performed at a factory or another facility prior to delivery of the AR / VR equipment to an end user. The device-specific calibration procedure can be an automatic calibration procedure or a manual calibration procedure. According to some embodiments, a user-specific calibration procedure can include estimation of eye parameters for a particular user, such as pupil position, fovea position, optical axis, visual axis, interpupillary distance, etc. According to some embodiments, once the device-specific parameters and user-specific parameters are determined for eye tracking device 130, a glint-assisted method can be used to process images captured by the eye tracking camera to determine a current visual axis and a gaze point of the user relative to the display.

[0158] As FIG. 5As shown, eye tracking device 130 (e.g., 130A or 130B) includes an eye lens 520 and a gaze tracking system that includes at least one eye tracking camera 540 (e.g., an infrared (IR) or near-IR (NIR) camera) positioned on the side of the user’s face on which eye tracking is performed, and an illumination source 530 (e.g., an IR or NIR light source such as an array or ring of NIR light-emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user’s eye 592. Eye tracking camera 540 can be directed toward a mirror 550 (which mirrors IR or NIR light from eye 592 while allowing visible light to pass) located between user’s eye 592 and display 510 (e.g., a left display panel or a right display panel of a head-mounted display, or a display of a handheld device, a projector, etc.) (e.g., as shown in the top portion of FIG. 6A), or alternatively can be directed toward user’s eye 592 to receive reflected IR or NIR light from eye 592 (e.g., as shown in the bottom portion of FIG. 6A). FIG. 5 FIG. 5

[0159] In some embodiments, controller 110 renders AR or VR frames 562 (e.g., left and right frames for left and right display panels) and provides frames 562 to display 510. Controller 110 uses gaze tracking input 542 from eye tracking camera 540 for various purposes, e.g., for processing frames 562 for display. Controller 110 optionally estimates a gaze point of the user on display 510 based on gaze tracking input 542 acquired from eye tracking camera 540 using a glint-assisted method or other suitable method. The gaze point estimated from gaze tracking input 542 is optionally used to determine a direction in which the user is currently looking.

[0160] ​​The following describes several possible use cases for the user's current gaze direction and is not intended to be limiting. As an example use case, controller 110 may render virtual content differently based on the determined direction of the user's gaze. For example, controller 110 may generate virtual content at a higher resolution in the concave region determined according to the user's current gaze direction than in the peripheral region. As another example, the controller may position or move virtual content in the view based at least partially on the user's current gaze direction. As yet another example, the controller may display specific virtual content in the view based at least partially on the user's current gaze direction. As another example use case in an AR application, controller 110 may guide an external camera used to capture the physical environment of an XR experience to focus in the determined direction. The external camera's autofocus mechanism may then focus on an object or surface in the environment that the user is currently looking at on display 510. As another example use case, eye lens 520 may be a focusable lens, and the controller uses gaze tracking information to adjust the focus of eye lens 520 so that the virtual object the user is currently looking at has appropriate convergence / divergence to match the convergence of the user's eyes 592. The controller 110 can use gaze tracking information to guide the eye lens 520 to adjust its focus so that the nearby object that the user is looking at appears at the correct distance.

[0161] In some embodiments, the eye-tracking device is part of a head-mounted device that includes a display (e.g., display 510), two eye lenses (e.g., eye lens 520), an eye-tracking camera (e.g., eye-tracking camera 540), and a light source (e.g., illumination source 530 (e.g., IR or NIR LED)). The light source emits light (e.g., IR or NIR light) toward the user's eyes 592. In some embodiments, the light source may be arranged in a ring or circle around each lens in the head-mounted device, such as... FIG. 5 As shown. In some embodiments, for example, eight light sources 530 (e.g., LEDs) are arranged around each lens 520. However, more or fewer light sources 530 may be used, and other arrangements and positions of the light sources 530 may be used.

[0162] In some embodiments, the display 510 emits light in the visible light range and does not emit light in the IR or NIR range, and thus does not introduce noise in the gaze tracking system. Note that the position and angle of the eye tracking cameras 540 are given by way of example and are not intended to be limiting. In some embodiments, a single eye tracking camera 540 is positioned on each side of the user’s face. In some embodiments, two or more NIR cameras 540 can be used on each side of the user’s face. In some embodiments, a camera 540 with a wider field of view (FOV) and a camera 540 with a narrower FOV can be used on each side of the user’s face. In some embodiments, a camera 540 operating at one wavelength (e.g., 850 nm) and a camera 540 operating at a different wavelength (e.g., 940 nm) can be used on each side of the user’s face.

[0163] As FIG. 6 embodiments of gaze tracking systems as exemplified in

[0164] FIG. 1A A flash-assisted gaze tracking pipeline is exemplified in accordance with some embodiments. In some embodiments, the gaze tracking pipeline is implemented by a flash-assisted gaze tracking system (e.g., an eye tracking device 130 as shown in FIG. 5 and FIG. 6 The flash-assisted gaze tracking system can maintain a tracking state. Initially, the tracking state is off or “no”. When in the tracking state, the flash-assisted gaze tracking system uses prior information from a previous frame when analyzing a current frame to track the pupil outline and glint in the current frame. When not in the tracking state, the flash-assisted gaze tracking system attempts to detect the pupil and glint in the current frame, and if successful, initializes the tracking state to “yes” and continues in the tracking state for the next frame.

[0165] As FIG. 6 indicated by the arrow returning to element 600, the gaze tracking system can continue to capture images of the user’s eyes, for example, at a rate of 60 to 120 frames per second. In some embodiments, each set of captured images can be input to the pipeline for processing. However, in some embodiments or under some conditions, not all captured frames are processed by the pipeline.

[0166] At 610, for the current captured image, if the tracking status is yes, the method proceeds to element 640. At 610, if the tracking status is no, the image is analyzed to detect the user’s pupils and glints in the image as indicated at 620. At 630, if the pupils and glints are successfully detected, the method proceeds to element 640. Otherwise, the method returns to element 610 to process the next image of the user’s eye.

[0167] At 640, if proceeding from element 610, the current frame is analyzed to track the pupils and glints based in part on previous information from previous frames. At 640, if proceeding from element 630, the tracking status is initialized based on the pupils and glints detected in the current frame. The results of the processing at element 640 are checked to verify that the results of the tracking or detection can be trusted. For example, the results can be checked to determine whether the pupils and a sufficient number of glints for performing gaze estimation were successfully tracked or detected in the current frame. At 650, if the results can not be trusted, at element 660, the tracking status is set to no, and the method returns to element 610 to process the next image of the user’s eye. At 650, if the results are trusted, the method proceeds to element 670. At 670, the tracking status is set to yes (if it is not already yes), and the pupil and glint information is passed to element 680 to estimate the user’s gaze point.

[0168] User interface and associated processes It is intended to serve as one example of an eye tracking technique that can be used in particular implementations. As will be recognized by one of ordinary skill in the art, in accordance with various implementations, other eye tracking techniques that are currently existing or are developed in the future can be used in place of or in combination with the glint-assisted eye tracking technique described herein in computer systems 101 for providing XR experiences to users.

[0169] In some implementations, portions of the captured real-world environment 602 are used to provide an XR experience to the user, such as a mixed reality environment in which one or more virtual objects are overlaid above a representation of the real-world environment 602.

[0170] Accordingly, the description herein describes some embodiments of a three-dimensional environment (e.g., an XR environment) that includes representations of real-world objects and representations of virtual objects. For example, the three-dimensional environment optionally includes a representation of a table that is present in a physical environment, which is captured and displayed in the three-dimensional environment (e.g., actively displayed via a camera and display of the computer system or passively displayed via a transparent or semi-transparent display of the computer system). As previously described, the three-dimensional environment is optionally a mixed reality system in which the three-dimensional environment is based on a physical environment that is captured by one or more sensors of the computer system and displayed via a display generation component. As a mixed reality system, the computer system is optionally able to selectively display portions and / or objects of the physical environment such that the respective portions and / or objects of the physical environment appear as if they are present in the three-dimensional environment displayed by the computer system. Similarly, the computer system is optionally able to display virtual objects in the three-dimensional environment at respective locations that have corresponding locations in the real world (e.g., the physical environment) by placing the virtual objects in the three-dimensional environment at the respective locations to appear as if the virtual objects are present in the real world. For example, the computer system optionally displays a vase such that the vase appears as if a real vase is placed on top of a table in the physical environment. In some embodiments, the respective locations in the three-dimensional environment have corresponding locations in the physical environment. Accordingly, when the computer system is described as displaying a virtual object at a respective location relative to a physical object (e.g., a location such as at or near a hand of the user or at or near a physical table), the computer system displays the virtual object at a particular location in the three-dimensional environment such that it appears as if the virtual object is at or near the physical object in the physical environment (e.g., the virtual object is displayed in the three-dimensional environment at a location that corresponds to a location in the physical environment where the virtual object would be displayed if the virtual object were a real object at the particular location).

[0171] In some embodiments, real-world objects that are present in a physical environment that are displayed in the three-dimensional environment (e.g., and / or are visible via the display generation component) can interact with virtual objects that are only present in the three-dimensional environment. For example, the three-dimensional environment can include a table and a vase placed on top of the table, where the table is a view (or representation) of a physical table in a physical environment, and the vase is a virtual object.

[0172] In three-dimensional environments (e.g., real environments, virtual environments, or environments that include a mix of real and virtual objects), objects are sometimes referred to as having depth or simulated depth, or objects are referred to as being visible, displayed, or placed at different depths. In this context, depth refers to a dimension that is different from height or width. In some embodiments, depth is defined relative to a fixed set of coordinates (e.g., where a room or object has a height, depth, and width defined relative to a fixed set of coordinates). In some embodiments, depth is defined relative to a user’s position or viewpoint, in which case the depth dimension varies based on the user’s position and / or the position and angle of the user’s viewpoint. In some embodiments where depth is defined relative to a user’s position that is positioned relative to a surface of the environment (e.g., a floor or ground surface of the environment), objects that are farther away from the user along a line that extends parallel to the surface are considered to have greater depth in the environment, and / or the depth of an object is measured along an axis that extends outward from the user’s position and parallel to the surface of the environment (e.g., depth is defined in a cylindrical or substantially cylindrical coordinate system, where the user’s position is at the center of a cylinder that extends from the user’s head toward the user’s feet). In some embodiments where depth is defined relative to a user’s viewpoint (e.g., relative to a direction of a point in space that determines which portion of the environment is visible via a head-mounted device or other display), objects that are farther away from the user’s viewpoint along a line that extends parallel to the direction of the user’s viewpoint are considered to have greater depth in the environment, and / or the depth of an object is measured along an axis that extends outward from the user’s viewpoint and parallel to the direction of the user’s viewpoint (e.g., depth is defined in a spherical or substantially spherical coordinate system, where the origin of the viewpoint is at the center of a sphere that extends outward from the user’s head). In some embodiments, depth is defined relative to a user interface container (e.g., a window or application in which application programs and / or system content are displayed), where the user interface container has a height and / or width, and the depth is a dimension that is orthogonal to the height and / or width of the user interface container. In some embodiments where depth is defined relative to a user interface container, when the container is placed in a three-dimensional environment or initially displayed (e.g., such that the depth dimension of the container extends outward away from the user or the user’s viewpoint), the height and / or width of the container is typically orthogonal or substantially orthogonal to a straight line that extends from the user’s position (e.g., the user’s viewpoint or the user’s position) to the user interface container (e.g., the center of the user interface container or another characteristic point of the user interface container). In some embodiments where depth is defined relative to a user interface container, the depth of an object relative to the user interface container refers to the positioning of the object along the depth dimension of the user interface container. In some embodiments, multiple different containers can have different depth dimensions (e.g., different depth dimensions that extend in different directions and / or from different starting points away from the user or the user’s viewpoint).In some embodiments, when defining depth relative to a user interface container, the direction of the depth dimension remains constant for the user interface container as the location of the user interface container, the user, and / or the user’s point of view changes (e.g., or when multiple different viewers are viewing the same container in a three-dimensional environment, such as during a physical collaboration session and / or when multiple participants are in a live communication session with shared virtual content including the container). In some embodiments, for a curved container (e.g., a container that includes a curved surface or a curved content area), the depth dimension optionally extends into the surface of the curved container. In some cases, a z-separation (e.g., the separation of two objects in the depth dimension), a z-height (e.g., the distance of one object from another object in the depth dimension), a z-position (e.g., the positioning of one object in the depth dimension), a z-depth (e.g., the positioning of one object in the depth dimension), or a simulated z-dimension (e.g., a depth used as a dimension of an object, a dimension of an environment, a direction in space, and / or a simulated direction in space) is used to refer to the concept of depth as described above.

[0173] In some embodiments, a user optionally is able to interact with virtual objects in a three-dimensional environment using one or both hands as if the virtual objects were real objects in a physical environment. For example, as described above, one or more sensors of a computer system optionally capture a user’s one or both hands and display a representation of the user’s hands in the three-dimensional environment (e.g., in a manner similar to displaying real-world objects in a three-dimensional environment as described above), or in some embodiments, the user’s hands are viewable via the display generation component, via the ability to see the physical environment through the user interface due to the transparency / semi- transparency of the portion of the user interface being displayed by the display generation component, or due to the projection of the user interface onto a transparent / semi-transparent surface or onto or into the field of view of the user’s eyes. Thus, in some embodiments, the user’s hands are displayed at their respective locations in the three-dimensional environment and are treated as if they were objects in the three-dimensional environment that are able to interact with virtual objects in the three-dimensional environment as if the virtual objects were physical objects in the physical environment. In some embodiments, the computer system is able to update the display of the representation of the user’s hands in the three-dimensional environment in conjunction with the movement of the user’s hands in the physical environment.

[0174] In some of the embodiments described below, the computer system is optionally capable of determining an "effective" distance between a physical object in the physical world and a virtual object in the three-dimensional environment, e.g., for determining whether a physical object is directly interacting with a virtual object (e.g., whether a hand is touching, grabbing, holding, etc. a virtual object or within a threshold distance of a virtual object). For example, a hand directly interacting with a virtual object optionally includes one or more of: a finger of the hand pressing a virtual button, a hand of the user grabbing a virtual vase, the user's hands coming together and pinching / holding a user interface of an application, and two fingers of the user's hand doing any other type of interaction described herein. For example, when determining whether a user is interacting with a virtual object and / or how a user is interacting with a virtual object, the computer system optionally determines a distance between a hand of the user and the virtual object. In some embodiments, the computer system determines a distance between a hand of the user and a virtual object by determining a distance between a location of the hand in the three-dimensional environment and a location of the virtual object of interest in the three-dimensional environment. For example, the hand or hands of the user are at a particular location in the physical world, the computer system optionally captures the hand or hands and displays the hand or hands at a particular corresponding location in the three-dimensional environment (e.g., if the hand is a virtual hand rather than a physical hand, the hand will be displayed at a location in the three-dimensional environment). The location of the hand in the three-dimensional environment is optionally compared to the location of the virtual object of interest in the three-dimensional environment to determine a distance between the hand or hands of the user and the virtual object. In some embodiments, the computer system optionally determines a distance between a physical object and a virtual object by comparing locations in the physical world (e.g., rather than comparing locations in the three-dimensional environment). For example, when determining a distance between a hand or hands of the user and a virtual object, the computer system optionally determines a corresponding location of the virtual object in the physical world (e.g., if the virtual object is a physical object rather than a virtual object, the location at which the virtual object would be located in the physical world), and then determines a distance between the corresponding physical location and the hand or hands of the user. In some embodiments, the same techniques are optionally used to determine a distance between any physical object and any virtual object. Thus, as described herein, when determining whether a physical object is in contact with a virtual object or whether a physical object is within a threshold distance of a virtual object, the computer system optionally performs any of the techniques described above to map a location of the physical object to the three-dimensional environment and / or to map a location of the virtual object to the physical environment.

[0175] In some embodiments, the same or similar techniques are used to determine where and what a user’s gaze is directed at, and / or where and what a physical stylus held by the user is directed at. For example, if a user’s gaze is directed at a particular location in the physical environment, the computer system optionally determines a corresponding location in the three-dimensional environment (e.g., a virtual location of the gaze), and if a virtual object is located at that corresponding virtual location, the computer system optionally determines that the user’s gaze is directed at that virtual object. Similarly, the computer system optionally is able to determine a direction in which a physical stylus is directed in the physical environment based on the orientation of the stylus. In some embodiments, based on the determination, the computer system determines a corresponding virtual location in the three-dimensional environment that corresponds to the location in the physical environment at which the stylus is directed, and optionally determines that the stylus is directed at the corresponding virtual location in the three-dimensional environment.

[0176] Similarly, embodiments described herein can refer to a location of a user (e.g., a user of a computer system) in a three-dimensional environment and / or a location of a computer system in a three-dimensional environment. In some embodiments, a user of a computer system is holding, wearing, or otherwise located at or near the computer system. Thus, in some embodiments, the location of the computer system is used as a proxy for the location of the user. In some embodiments, the location of the computer system and / or the user in the physical environment corresponds to a respective location in the three-dimensional environment. For example, the location of the computer system would be a location in the physical environment (and its corresponding location in the three-dimensional environment) from which the user would see the physical environment in the same location, orientation, and / or size (e.g., in absolute terms and / or relative to each other) of objects in the physical environment as the objects are displayed in the three-dimensional environment by the display generation component of the computer system or are visible in the three-dimensional environment via the display generation component of the computer system if the user were standing in that location facing the respective portion of the physical environment that is visible via the display generation component. Similarly, if the virtual objects displayed in the three-dimensional environment are physical objects in the physical environment (e.g., physical objects placed in the physical environment at the same locations as the virtual objects in the three-dimensional environment, and physical objects in the physical environment having the same size and orientation as when in the three-dimensional environment), the location of the computer system and / or the user is a location from which the user would see the physical environment in the same location, orientation, and / or size (e.g., in absolute terms and / or relative to each other and real-world objects) of the virtual objects as the virtual objects are displayed in the three-dimensional environment by the display generation component of the computer system.

[0177] In this disclosure, various input methods are described with respect to interactions with a computer system. When one input device or input method is used to provide an example, and another input device or input method is used to provide another example, it will be understood that each example can be compatible with and optionally utilize the input device or input method described with respect to the other example. Similarly, various output methods are described with respect to interactions with a computer system. When one output device or output method is used to provide an example, and another output device or output method is used to provide another example, it will be understood that each example can be compatible with and optionally utilize the output device or output method described with respect to the other example. Similarly, various methods are described with respect to interactions with a virtual environment or a mixed reality environment by a computer system. When an interaction with a virtual environment is used to provide an example, and a mixed reality environment is used to provide another example, it will be understood that each example can be compatible with and optionally utilize the methods described with respect to the other example. Accordingly, this disclosure discloses embodiments that are combinations of features of multiple examples without the need to exhaustively list all features of an embodiment in the description of each example embodiment.

[0178] FIGS. 7A-7Q Attention is now directed to embodiments of user interfaces (“UIs”) and associated processes that can be implemented on a computer system having a display generation component, one or more input devices, and (optionally) one or more cameras, such as a portable multifunction device or a head-mounted device.

[0179] In some embodiments, the computer system presents virtual objects that correspond to physical locations within a three-dimensional environment. The virtual objects described below include images and / or videos taken from and / or of the physical locations. In some embodiments, the virtual objects provide respective views of respective physical locations from different perspectives without requiring subsequent input to manipulate (e.g., rotate, pan, and / or zoom in or out) the map, thereby enabling the user to easily view and experience the physical locations from a variety of perspectives. Enhancing interactions with the computer system reduces the time required for the user to perform operations, thereby reducing power consumption of the computer system and increasing battery life for battery-powered computer systems.

[0180] FIG. 7A Examples are illustrated of a computer system updating, based on an immersion level, a display of a navigation user interface that includes a travel user interface element and / or one or more navigation user interface elements, in accordance with some embodiments. In some embodiments, the travel user interface element is selectable to update, in accordance with some embodiments, the display of the navigation user interface with a respective immersion level corresponding to a respective view of a respective physical location.

[0181] FIG. 6 The computer system 101 is illustrated as displaying the three-dimensional environment 702 from the user's point of view via a display generation component (e.g., the display generation component 120 of FIG. 1). As noted above with reference to FIGS. 1-2, the computer system 101 optionally includes a display generation component (e.g., a touchscreen) and a plurality of image sensors (e.g., the image sensors 314 of FIG. 1). The image sensors optionally include one or more of: a visible light camera; an infrared camera; a depth sensor; or any other sensor that the computer system 101 is able to use to capture one or more images of the user or a portion of the user (e.g., one or both hands of the user) as the user interacts with the computer system 101. In some embodiments, the user interfaces illustrated and described below can also be implemented on a head-mounted display that includes a display generation component that displays the user interfaces or three-dimensional environments to the user, and sensors (e.g., external sensors facing outward from the user's face) that detect movement of the physical environment and / or the user's hands (such as movement that the computer system interprets as gestures, such as air gestures), and / or sensors (e.g., internal sensors facing inward toward the user's face) that detect the user's gaze. FIG. 3 FIG. 7A As illustrated, the computer system 101 captures one or more images of the physical environment (e.g., the operating environment 100) around the computer system 101, including one or more objects in the physical environment around the computer system 101. In some embodiments, the computer system 101 displays a representation of the physical environment in the three-dimensional environment 702, or portions of the physical environment are visible via the display generation component 120 of the computer system 101. For example, the three-dimensional environment 2302 includes portions of the left and back walls and floor in the user's physical environment.

[0182] As illustrated, the computer system 101 captures one or more images of the physical environment (e.g., the operating environment 100) around the computer system 101, including one or more objects in the physical environment around the computer system 101. In some embodiments, the computer system 101 displays a representation of the physical environment in the three-dimensional environment 702, or portions of the physical environment are visible via the display generation component 120 of the computer system 101. For example, the three-dimensional environment 2302 includes portions of the left and back walls and floor in the user's physical environment. FIG. 7A

[0183] FIG. 7A ​​​In particular embodiments, three-dimensional environment 702 also includes virtual objects, such as virtual objects 704a and 714. Virtual objects 704a and 714 are optionally one or more of a user interface of an application (e.g., a navigation user interface of a map application), a three-dimensional object (e.g., a virtual three-dimensional map element of a navigation user interface), or any other element displayed by computer system 101 that is not included in the physical environment of computer system 101. In some embodiments, virtual object 704a is a navigation user interface that includes virtual content and / or imagery (e.g., captured photos and / or videos) associated with various geographic locations, points of interest, etc. For example, computer system 101 displays navigation user interface 704a in a partially immersive manner that corresponds to a respective view of a respective physical location. Navigation user interface 704a also includes selectable travel user interface objects 708a, 708b, 708c, and 708d that, when selected, cause computer system 101 to display content with a different respective view of the same respective physical location or a different physical location, as will be described below. Navigation user interface 704a also includes a selectable user interface element 706 that, when selected, causes the computer system to increase or decrease the size of navigation user interface 604a. Navigation user interface 704a also includes a selectable user interface element 710a that includes selectable navigation user interface objects 710b, 710c, 710d, and 710e that, when selected, cause computer system 101 to display a navigation user interface element associated with a respective physical location, as will be described below. In some embodiments, and as will be described below, physical buttons 778a and 778b of computer system 101 change the level of immersion of navigation user interface 704a in response to manipulation of physical buttons 778a and / or 778b.

[0184] In FIG. 7A In particular embodiments, computer system 101 displays navigation user interface element 714 that includes representations of physical objects (e.g., buildings, landmarks, buildings, landmarks, parks, and / or road lanes, trees). Navigation user interface element 714 also includes an indication 712 (e.g., a cloud) of weather at a respective physical location that corresponds to a physical location that is experiencing the weather represented by indication 712 (e.g., cloudy weather). Navigation user interface element 714 also includes a selectable user interface control element 716 that, when selected (optionally by hand 718a), causes the computer system to adjust the positioning and / or orientation of navigation user interface element 714, as will be described below.

[0185] In some embodiments, the computer system displays a navigation user interface 704a with different levels of immersion as described herein with respect to presenting different levels of immersion, where the VR and / or AR content occupies a greater or lesser portion of the three-dimensional environment 702 relative to content corresponding to the physical environment of the computer system 101 (e.g., pass-through video). For example, the computer system detects a user input (e.g., a finger of the hand 718a touching the physical button 778a and / or 778b). Additionally or alternatively, the computer system detects a user input when the attention / gaze of the user of the computer system 101, if present, is directed to a settings menu user interface element for changing the level of immersion (e.g., a finger of the hand 718a touching a trackpad, an air pinch gesture from the hand 718a, and / or an air tap gesture). For example, in response to the computer system detecting a user input as described herein, such as an input directed to a settings menu user interface element or an input manipulating the physical button 778a and / or 778b, the computer system changes the level of immersion. In some embodiments, the computer system 101 detects other types of user inputs directed to the selectable user interface element 706, such as via an input device in communication with the computer system 101 and / or voice input from the user.

[0186] In response to the hand 718a manipulating the physical button 778a and / or 778b, the computer system changes the level of immersion in which the navigation user interface 704a occupies a greater or lesser portion of the three-dimensional environment 702 relative to content corresponding to the physical environment of the computer system 101. For example, in response to the computer system detecting a user input directed to a settings menu user interface element for changing the level of immersion or an input for manipulating the physical button 778a and / or 778b, the computer changes the display of the navigation user interface 704a from a partially immersive (e.g., with a first level of immersion) display as shown in FIG. 7B corresponding to the respective view of the respective physical location to a display of a fully immersive (e.g., with a second level of immersion) navigation user interface 704b as shown in FIG. 7B corresponding to the respective view of the respective physical location to a display of a fully immersive (e.g., with a second level of immersion) navigation user interface 704b as shown in FIG. 7A corresponding to the respective view of the respective physical location to a display of a fully immersive (e.g., with a second level of immersion) navigation user interface 704b as shown in FIG. 7B In some embodiments, the navigation user interface 704b includes a collapsed view of the selectable user interface element 710a that includes the selectable user interface element 710f. In some embodiments, in response to detecting a selection of the selectable user interface element 710a or 710f, the computer system expands the selectable user interface element 710a to display the selectable navigation user interface objects 710b, 710c, 710d, and 710e shown in

[0187] In some embodiments, the navigation user interface 704b includes more selectable travel user interface objects as the computer system 101 increases the level of immersion. For example, in FIG. 7A , the navigation user interface 704b displayed by the computer system 101 includes additional travel user interface objects 708e and 708f that were not included when the computer system 101 displayed the navigation user interface 704a in the partial immersion mode as illustrated in FIG. 7B

[0188] In some embodiments, the computer system displays the travel user interface objects in the navigation user interface 704b at respective locations that correspond to respective views and / or respective physical locations associated with the respective travel user interface objects. For example, in FIG. 7B , the navigation user interface 704b includes the travel user interface object 708f that, when selected, causes the computer system 101 to update the display of the navigation user interface 704b to correspond to the respective viewpoint and / or physical location associated with the travel user interface object 708f (e.g., a viewpoint from the top of a building). As another example, the navigation user interface 704b includes the travel user interface object 708e that, when selected, causes the computer system 101 to update the display of the navigation user interface 704b to correspond to the respective viewpoint and / or physical location associated with the travel user interface object 708e (e.g., a viewpoint from the Golden Gate Bridge, which is a different physical location than FIG. 7B . As another example, the computer system detects the user input 718b directed to the travel user interface object 708b (e.g., an input similar to or corresponding to the inputs described above). In response to receiving the user input 718b directed to the travel user interface object 708b, the computer system 101 changes the display of the navigation user interface 704b. For example, when the computer system 101 displays the navigation user interface 704b to correspond to a first view of a first physical location as illustrated in FIG. 7C , the computer system 101 receives the user input 718b directed to the travel user interface object 708b. In response to receiving the user input 718b directed to the travel user interface object 708b, the computer system 101 displays the navigation user interface 704b corresponding to a second view of a second physical location that is different than the first view of the first physical location as illustrated in FIG. 7C . For example, the second view of the second physical location includes a street view image from a location in the northwest direction of the building, while the first view of the first physical location includes a street view image from a location in the southeast direction of the building.

[0189] In FIG. 7B ​In some embodiments, computer system 101 displays navigation user interface 704b in response to receiving FIG. 7B the illustrated input. For example, the second view of the second physical location includes travel user interface objects 720a, 720b, 720c, 720f, 720e, and 720f, which are different from the travel user interface objects displayed by navigation user interface 704a in FIG. 7C In some embodiments, computer system 101 displays navigation user interface 704b that includes travel user interface objects corresponding to a second view of a first physical location that are different from the travel user interface objects displayed by navigation user interface 704b in

[0190] In some embodiments, computer system 101 displays navigation user interface 704b that includes a representation 722a of a business located within the respective physical location. For example, in FIG. 7C In some embodiments, computer system 101 displays navigation user interface 704b that includes a representation 722a of a business located within the respective physical location. For example, in FIG. 7D In some embodiments, computer system 101 displays navigation user interface 704b that includes a representation 722a of a business located within the respective physical location. For example, in FIG. 7D In some embodiments, computer system 101 displays navigation user interface 704b that includes a representation 722a of a business located within the respective physical location. For example, in FIG. 7B In some embodiments, computer system 101 displays navigation user interface 704b that includes a representation 722a of a business located within the respective physical location. For example, in FIG. 7D In some embodiments, computer system 101 displays navigation user interface 704b that includes a representation 722a of a business located within the respective physical location. For example, in

[0191] In some embodiments, the computer system 101 displays the navigation user interface 704b corresponding to the respective view of the respective physical location that includes an annotation, such as text, an image, and / or a reference to other information (e.g., a link). In some embodiments, the computer system 101 displays the annotation within the respective view of the respective physical location proximate to the corresponding physical location of the annotation. In some embodiments, the annotation is created by a user via the computer system 101 to describe or provide useful information about the physical location. For example, in FIG. 7E In some embodiments, the computer system 101 detects a voice input 724 from the user that corresponds to a request to add an annotation to the respective view of the respective physical location. In response to receiving the voice input 724, the computer system 101 generates and displays a representation 726 of the annotation, as shown in FIG. 7P In some embodiments, when the computer system 101 is physically located at the respective location corresponding to the annotation, the computer system 101 displays a notification to the user about the annotation, as will be described below with reference to FIG. 7E In some embodiments, the computer system 101 detects a voice input 724 from the user that corresponds to a request to add an annotation to the respective view of the respective physical location. In response to receiving the voice input 724, the computer system 101 generates and displays a representation 726 of the annotation, as shown in

[0192] As previously described herein, the navigation user interface 704b includes a collapsed view of the selectable user interface element 710a that includes FIG. 7E In some embodiments, the computer system 101 detects a user input 718e (e.g., an input similar to or corresponding to the inputs described above) directed to the selectable user interface element 710f in FIG. 7F In response to receiving the user input 718e directed to the selectable user interface element 710f in FIG. 7F the computer system 101 changes the display of the navigation user interface 704b to include the expanded selectable user interface element 710a as shown in FIG. 7N In some embodiments, the expanded selectable user interface element 710a includes the selectable navigation user interface objects 710b, 710c, 710d, and 710e.

[0193] In some embodiments, in response to detecting an input selecting a respective one of the navigation user interface objects 710b, 710c, 710d, the computer system 101 displays a respective navigation user interface element associated with the respective physical location displayed by the navigation user interface 704b. For example, in response to receiving an input selecting the navigation user interface object 710b, the computer system 101 displays a navigation user interface element corresponding to the globe described below with reference to FIG. 7F Returning to FIGS. 7G-7MThe selectable user interface element 710a includes a selectable user interface element 710c that, when selected, causes the computer system 101 to display a navigation user interface element 714, as described in greater detail below. In some embodiments, the selectable user interface element 710a includes a navigation user interface object 710d that, when selected, causes the computer system 101 to display a navigation user interface 704b that corresponds to a respective view of the respective physical location, such as a street view of the respective physical location, as described herein. FIG. 7F In some embodiments, the selectable user interface element 710a includes a navigation user interface object 710e that, when selected, causes the computer system 101 to display video content associated with the respective physical location, such as an aerial tour video of the respective physical location. FIG. 7F In some embodiments, the selectable user interface element 710a includes a navigation user interface object 710e that, when selected, causes the computer system 101 to display video content associated with the respective physical location, such as an aerial tour video of the respective physical location.

[0194] In some embodiments, the navigation user interface 704a includes similar user interface objects and / or elements displayed in the navigation user interface 704b in FIG. 7F In some embodiments, the navigation user interface 704a includes similar user interface objects and / or elements displayed in the navigation user interface 704b in FIG. 7G In some embodiments, the navigation user interface 704a includes similar user interface objects and / or elements displayed in the navigation user interface 704b in FIG. 7F In some embodiments, the navigation user interface 704a includes similar user interface objects and / or elements displayed in the navigation user interface 704b in FIG. 7G In some embodiments, the navigation user interface 704a includes similar user interface objects and / or elements displayed in the navigation user interface 704b in FIG. 7F In some embodiments, the navigation user interface 704a includes similar user interface objects and / or elements displayed in the navigation user interface 704b in FIG. 7G In some embodiments, the navigation user interface 704a includes similar user interface objects and / or elements displayed in the navigation user interface 704b in FIG. 7G In some embodiments, the navigation user interface 704a includes similar user interface objects and / or elements displayed in the navigation user interface 704b in FIG. 7G In some embodiments, the navigation user interface 704a includes similar user interface objects and / or elements displayed in the navigation user interface 704b in FIG. 7A In some embodiments, the navigation user interface 704a includes similar user interface objects and / or elements displayed in the navigation user interface 704b in FIG. 7A In some embodiments, the navigation user interface 704a includes similar user interface objects and / or elements displayed in the navigation user interface 704b in FIG. 7Gsimilar user interface objects and / or elements shown in FIGS. 7A-7D.

[0195] In some embodiments, in response to receiving user input 718f directed to selectable user interface element 710c, computer system 101 displays FIG. 7F the navigation user interface element 714 and maintains a high level of immersion. For example, displaying a three-dimensional environment with a high level of immersion includes displaying the navigation user interface 704b and the navigation user interface element 714 with a greater visual salience (e.g., size, brightness, sharpness, opacity, etc.) than the visual salience of the representations of real objects in the three-dimensional environment. In some embodiments, the pass-through video (e.g., a representation of the physical environment of computer system 101) is hidden. In another example, displaying a three-dimensional environment with a high level of immersion includes displaying the navigation user interface 704b with a fully immersive view as shown in FIG. 7G the navigation user interface element 714 overlaid on the fully immersive view. In FIG. 7E In some embodiments, computer system displays a navigation user interface 704a corresponding to a first view of the first physical location that is different from a second view of the first physical location as shown in FIG. 7G In some embodiments, computer system displays a navigation user interface 704a corresponding to a first view of the first physical location that is different from a second view of the first physical location as shown in FIG. 7A Also included is the navigation user interface element 714. In some embodiments, the navigation user interface element 714 has one or more characteristics and / or includes FIG. 7G similar user interface objects and / or elements shown in FIGS. 7A-7D.

[0196] In some embodiments, computer system 101 detects a pose of a user of the computer system and determines that the pose of the user indicates an interaction with the navigation user interface element 714. For example, in response to detecting that the head of the user is tilted toward the navigation user interface element 714, computer system 101 displays the navigation user interface element 714 with a certain level of zoom according to the pose of the user. FIG. 7H depicted as positioning the navigation user interface element 714 on top of the selectable user interface control element 716. In some embodiments, computer system 101 detects a gaze 728 of a user of computer system 101 directed at the navigation user interface element 714. In some embodiments, in response to computer system detecting the gaze 728 of the user of computer system 101 directed at the navigation user interface element 714, computer system 101 outputs spatial audio 730 corresponding to environmental noise at a respective physical location corresponding to the location of the user’s gaze 728. In some embodiments, computer system displays the navigation user interface element 714 with a level of zoom and / or perspective corresponding to the pose of the user 734 as depicted in the legend 732.

[0197] As shown in FIG. 7H , in response to the computer system detecting the gaze 728 of the user of the computer system 101 directed at the navigation user interface element 714, the computer system 101 zooms in on the display of the navigation user interface element 714. In some embodiments, in response to detecting the gaze of the user, the computer system 101 ceases to display the respective content corresponding to the respective view of the respective physical location and displays a representation 736 of the content related to the respective physical location, such as text, images, and / or video related to the respective physical location, as shown in FIG. 7H . In some embodiments, the computer system 101 changes the spatial audio 730 that is output to correspond to the zoomed-in view of the respective physical location. For example, in FIG. 7G , the computer system 101 displays the navigation user interface element 714 representing a smaller area of the respective physical location with more detail than the area of the respective location displayed in FIGS. 7G-7H . From FIG. 7H , in response to the computer system 101 zooming in on the navigation user interface element, the computer system 101 increases the volume of the spatial audio 730.

[0198] In some embodiments, the computer system 101 continues to detect zooming input from the user of the computer system 101. For example, in FIG. 7I , the legend 732 depicts that the pose of the user 734 has changed, corresponding to leaning further down to look at the navigation user interface element 714 while maintaining the gaze at the navigation user interface element 714, particularly the representation of the building. In some embodiments, in response to detecting that the pose of the user has changed and that their gaze is maintained on the particular building represented by the navigation user interface element 714, the computer system 101 displays a representation of the building with more clarity than the representation of the area outside of a predetermined distance from the building and the area immediately adjacent to the building, as shown in FIG. 7I .

[0199] In addition to zooming in and / or out on the navigation user interface element 714, the computer system 101 also changes the positioning and / or orientation of the navigation user interface element in response to receiving input directed at the selectable user interface control element 716. For example, in FIG. 7J , the computer system 101 detects input 738 (e.g., input corresponding to or similar to the input described above) directed at the selectable user interface control element 716 corresponding to a request to rotate (e.g., 740i) and zoom out on the navigation user interface element 714. In response to the input 738, the computer system 101 displays the navigation user interface element 714 representing a rotated and zoomed-out view of the respective physical location. In some embodiments, the computer system 101 changes the spatial audio that is output accordingly, as shown inFIG. 7J For example, in FIG. 7J , the computer system outputs spatial audio 742 corresponding to the ambient noise at the respective physical location and decreases the volume of the spatial audio 730 in accordance with the request to zoom out and rotate the navigation user interface element 714. In some embodiments, the computer system 101 changes the representation 736 to include content related to the zoomed-out view of the respective physical location, as shown in FIG. 7J .

[0200] In some embodiments, as shown in FIG. 7K , while displaying the zoomed-in and rotated view of the respective location, the computer system 101 detects another input 738 (e.g., input similar to or corresponding to the input described above) directed to the selectable user interface control element 716 corresponding to a request to further zoom out and rotate (e.g., 740j) the navigation user interface element 714. In response to the input 738, the computer system 101 displays FIG. 7K the navigation user interface element 714 representing a further zoomed-out and rotated view of the respective physical location in FIG. 7K . In FIG. 7K , the computer system further decreases the volume of the spatial audio 730 in accordance with the zoom-out request and increases the spatial audio 742 corresponding to the ambient noise of the respective physical location now represented by the navigation user interface element 714. In some embodiments and as shown in FIG. 7J , the computer system determines that the user input 738 is directed to a first portion (e.g., an outer portion) of the selectable user interface control element 716k that is different from the portion (e.g., an inner portion) of the selectable user interface control element 716j shown in FIG. 7J . In some embodiments, the computer system rotates the navigation user interface element 714 at a faster speed when the user input 738 is directed to the first portion of the selectable user interface control element 716j than when the user input 738 is directed to a portion of the selectable user interface control element 716k that is different from the first portion, as shown in FIG. 7K . In some embodiments, the computer system scales both the navigation user interface element 714 and the selectable user interface control element 716 in response to receiving input directed to the selectable user interface control element 716. In some embodiments, the computer system scales the navigation user interface element 714 but not the selectable user interface control element 716 in response to receiving input directed to the navigation user interface element 714.

[0201] In some embodiments, while computer system 101 displays navigation user interface element 714 representing a further zoomed-out and rotated view of the respective physical location, computer system 101 detects user input 738 (e.g., input similar to or corresponding to the input described above) directed to selectable user interface control element 716 corresponding to a request to zoom in and rotate on navigation user interface element 714 (e.g., 740k), as shown in FIG. 7O. In response, computer system 101 displays navigation user interface element 714 representing a zoomed-in and rotated view of the respective physical location, as shown in FIG. 7P. In some embodiments, in response to the request to zoom in and rotate, computer system 101 outputs spatial audio 730 at an even lower volume in accordance with rotating navigation user interface element 714 because the portion of navigation user interface element 714 associated with spatial audio 730 is no longer (or only partially) in view. In some embodiments, computer system 101 increases spatial audio 742 in accordance with zooming in on the portion of navigation user interface element 714 associated with spatial audio 742, as shown in FIG. 7Q. In some embodiments, computer system 101 changes representation 736 to include content related to the zoomed-in view of the respective physical location, as shown in FIG. 7R. FIG. 7L FIG. 7L FIG. 7L FIG. 7L

[0202] In some embodiments, while computer system 101 displays navigation user interface element 714 representing a zoomed-in and rotated view of the respective physical location, computer system 101 detects user input 738 (e.g., input similar to or corresponding to the input described above) directed to selectable user interface control element 716 corresponding to a request to zoom out navigation user interface element 714 (e.g., 740), as shown in FIG. 7S. In response to receiving input as shown in FIG. 7T, computer system 101 displays navigation user interface element 714 representing a zoomed-out view of the respective physical location, and outputs spatial audio 730 at a similar lower volume, and slightly decreases spatial audio 742 to correspond to the zoomed-out view of the respective physical location, as shown in FIG. 7U. In some embodiments, computer system 101 changes representation 736 to include content related to the zoomed-out view of the respective physical location, as shown in FIG. 7V. FIG. 7L FIG. 7M FIG. 7M FIG. 7M

[0203] In some embodiments, while computer system 101 displays navigation user interface element 714 representing a zoomed-in and rotated view of the respective physical location, computer system 101 detects user input 738 (e.g., input similar to or corresponding to the input described above) directed to selectable user interface control element 716 corresponding to a request to zoom out navigation user interface element 714 (e.g., 740), as shown in FIG. 7S. In response to receiving input as shown in FIG. 7T, computer system 101 displays navigation user interface element 714 representing a zoomed-out view of the respective physical location, and outputs spatial audio 730 at a similar lower volume, and slightly decreases spatial audio 742 to correspond to the zoomed-out view of the respective physical location, as shown in FIG. 7U. In some embodiments, computer system 101 changes representation 736 to include content related to the zoomed-out view of the respective physical location, as shown in FIG. 7V. FIG. 7L FIG. 7M In response to receiving input​​​​​​​​​FIG. 7N The input in the computer system 101 is displayed and FIG. 7N The globe in the navigation user interface element 754 corresponds to a globe that is positioned and / or oriented such that the corresponding physical location is centered in the user's field of view. FIG. 7C In the navigation user interface element 754, two-dimensional representations of landmasses (e.g., 758), bodies of water, and / or points of interest (e.g., 746, 748, 750, and 752) are located on the navigation user interface element 754 at positions corresponding to their respective physical locations. In some embodiments, in response to receiving input that selects a representation of a point of interest (e.g., 746, 748, 750, and / or 752), the computer system 101 displays information related to the corresponding point of interest, similar to... FIG. 7N The representation is 722a. In some embodiments, the navigation user interface element 754 also includes a representation 744 of the weather at the corresponding physical location.

[0204] In some implementations, similar to navigation user interface element 714, navigation user interface element 754 is interactive. For example, in FIG. 7O In this process, the computer system detects user input 718 (e.g., input similar to or corresponding to the input described above) that points to the navigation user interface element 754 and corresponds to a request to rotate the navigation user interface element 754. In response to detecting user input 718, the computer system 101 rotates the navigation user interface element 754 to display, as shown below. FIG. 7N The navigation user interface element 754 shown was previously not... FIG. 7O The navigation user interface element 754 displays different parts. For example, in FIG. 7O In the navigation user interface element 754, landmass 760 with terrain indication is included. FIG. 7P In this process, the computer system also displays representations of points of interest (e.g., 746, 748, 750, and 752), which are located on the navigation user interface element 754 at positions corresponding to their respective physical locations. Reference method 800 also anticipates and describes additional or alternative representations of additional or alternative map virtual objects, user interface elements, information, and / or features.

[0205] In some implementations, and as described above, when computer system 101 is physically located at a corresponding physical location corresponding to the annotation, computer system 101 displays a notification about the annotation to the user. In some implementations, another computer system associated with the same user account as computer system 101 displays the annotation. For example, in FIG. 7QIn particular, the second computer system 101 presents an immersive augmented reality 762 and a representation 766 of the annotation as described above. In some embodiments, in response to detecting an input that selects the representation 766, the computer system 101p displays information related to the annotation.

[0206] In some embodiments, the computer system 101 shares the navigation user interface 704a with a second computer system that is associated with another user account that is different from the user account of the computer system 101. For example, FIG. 7Q Two different computer systems in a shared communication session are illustrated: a computer system 101a that is associated with user Mary and a computer system 101b that is associated with user Bob. A legend 768 depicts, for example, user Mary (e.g., representation 770) and user Bob (representation 772) and their respective computer systems in different physical environments. In FIG. 7Q In particular, the computer system 101a that is associated with user Mary displays a representation of the navigation user interface with a fully immersive view that includes selectable travel user interface objects 774a, 774b, and 774c. In some embodiments, in response to receiving an input that selects the selectable travel user interface object 774b, the computer system 101a displays a view of the navigation user interface that corresponds to the viewpoint of user Bob (e.g., displays the navigation user interface associated with user Bob from the viewpoint of user Bob). In FIG. 8 In particular, the computer system 101b that is associated with user Bob displays a representation of the navigation user interface with a fully immersive view that includes selectable travel user interface objects 776a, 776b, and 776c. In some embodiments, in response to receiving an input that selects the selectable travel user interface object 776a, the computer system 101b displays a view of the navigation user interface that is displayed by the computer system 101a that is associated with user Mary from the viewpoint of user Mary.

[0207] FIG. 3 is a flowchart illustrating an example method of displaying a navigation user interface with respective levels of immersion corresponding to respective views of respective physical locations, in accordance with some embodiments. In some embodiments, the method 800 is performed at a computer system (e.g., computer system 101 in Figure 1, such as a tablet, a smartphone, a wearable computer, or a head-mounted device) that includes a display generation component (e.g., display generation component 170 in Figure 1, display generation component 270 in Figure 2, display generation component 370 in Figure 3, display generation component 470 in Figure 4, display generation component 570 in Figure 5, display generation component 670 in Figure 6, or display generation component 770 in Figure 7) and a position determination component (e.g., position determination component 160 in Figure 1, position determination component 260 in Figure 2, position determination component 360 in Figure 3, position determination component 460 in Figure 4, position determination component 560 in Figure 5, position determination component 660 in Figure 6, or position determination component 760 in Figure 7). FIG. 4 and FIG. 1Aa display generation component (e.g., 120) (e.g., a heads-up display, a display, a touchscreen, a projector, etc.) and one or more cameras (e.g., a camera (e.g., a color sensor, an infrared sensor, and other depth-sensing cameras) pointed down at a user's hand or a camera pointed forward from a user's head). In some embodiments, method 800 is managed by instructions executed by a control unit (e.g., 110) in the non-transitory computer-readable storage medium and one or more processors of a computer system, such as one or more processors 202 of computer system 101 (e.g., FIG. 7A Some operations in method 800 are, optionally, combined and / or the order of some operations is, optionally, changed.

[0208] In some embodiments, method 800 is performed at a computer system (e.g., 101) in communication with a display generation component (e.g., 120) and one or more input devices (e.g., 314). For example, the computer system is or includes a mobile device (e.g., a tablet, a smartphone, a media player, or a wearable device) or a computer. In some embodiments, the display generation component is a display (optionally a touchscreen display) integrated with the computer system, an external display (such as a monitor, a projector, a television), or a hardware component (optionally integrated or external) for projecting or causing a user interface to be visible to one or more users. In some embodiments, the one or more input devices include an electronic device or component capable of receiving (e.g., capturing or detecting) user input and sending information associated with the user input to the computer system. Examples of input devices include a touchscreen, a mouse (e.g., external), a trackpad (optionally integrated or external), a trackpad (optionally integrated or external), a remote control device (e.g., external), another mobile device (e.g., separate from the computer system), a handheld device (e.g., external), a controller (e.g., external), a camera, a depth sensor, an eye tracking device, and / or a motion sensor (e.g., a hand tracking device or a hand motion sensor). In some embodiments, the computer system is in communication with a hand tracking device (e.g., one or more cameras, depth sensors, proximity sensors, touch sensors (e.g., a touchscreen, a trackpad)). In some embodiments, the hand tracking device is a wearable device, such as a smart glove. In some embodiments, the hand tracking device is a handheld input device, such as a remote control or a stylus.

[0209] In some embodiments, the computer system displays (802a), via the display generation component, a navigation user interface within a three-dimensional environment, such as FIG. 7Aa navigation user interface 704a within a three-dimensional environment 702 in FIG. 7B. For example, the three-dimensional environment is generated, displayed, or otherwise made viewable by a computer system (e.g., an extended reality (XR) environment, such as a virtual reality (VR) environment, a mixed reality (MR) environment, or an augmented reality (AR) environment). In some embodiments, the physical environment surrounding the display generation component is visible through a transparent portion of the display generation component (e.g., real or reality pass-through). In some embodiments, a representation of the physical environment is displayed in the three-dimensional environment via the display generation component (e.g., virtual or video pass-through). In some embodiments, the navigation user interface includes one or more first travel user interface elements displayed at a first level of immersion corresponding to a first view of the first physical location, such as FIG. 7A selectable travel user interface objects 708a, 708b, 708c, and 708d in FIG. 7B.

[0210] In some embodiments, the computer system displays the navigation user interface in a three-dimensional environment that is within the field of view of the user of the computer system from the viewpoint of the user. In some embodiments, the user interface is a user interface of a map application. In some embodiments, the navigation user interface is a user interface of an application other than a map application, such as a travel guide application. In some embodiments, the navigation user interface includes respective content corresponding to a first view of a first physical location, such as one or more images or videos taken from and / or of the first physical location (e.g., one or more street view images and / or videos from the first physical location). In some embodiments, the first respective content is a (e.g., live) video recorded at the first physical location. The navigation user interface including the respective content corresponding to the first view of the first physical location is optionally presented from a first-person perspective of the viewpoint of the user associated with the computer system at a respective location in the three-dimensional environment (e.g., corresponding to the location of the computer system). In some embodiments, the computer system is configured to increase or decrease an immersion level of the navigation user interface (e.g., increase or decrease the number of virtual elements included in the navigation user interface that include a travel user interface element, and / or increase or decrease the portion of the three-dimensional environment that the navigation user interface occupies). In some embodiments, the immersion level includes an associated degree to which the navigation user interface and / or virtual content displayed by the computer system occludes background content (e.g., a three-dimensional environment including a physical environment) around behind the virtual content, optionally including a number of items of the displayed background content and visual characteristics (e.g., color, contrast, and / or opacity) with which the background content is displayed, and / or an angular range of virtual content displayed via the display generation component (e.g., 60 degrees of content displayed with a low immersion, 120 degrees of content displayed with a medium immersion, or 180 degrees of content displayed with a high immersion), and / or a proportion of a field of view displayed via the display generation display that is consumed by the virtual content (e.g., 33% of the field of view consumed by the virtual content with a low immersion, 66% of the field of view consumed by the virtual environment with a medium immersion, and / or 100% of the field of view consumed by the virtual content with a high immersion).

[0211] In some embodiments, at a first (e.g., low) level of immersion, background, virtual, and / or real objects are displayed in an occluded manner (e.g., dimmed, blurred, and / or removed from display). For example, virtual content with a low level of immersion is optionally displayed concurrently with background content, which is optionally displayed at full brightness, color, and / or translucency. In some embodiments, at a second (e.g., high) level of immersion, background, virtual, and / or real objects are displayed in an occluded manner. For example, respective virtual content with a high level of immersion is displayed without concurrent display of background content (e.g., in a full screen or fully immersive mode). As another example, virtual content displayed with a medium level of immersion is optionally displayed concurrently with background content that is dimmed, blurred, or otherwise de-emphasized. In some embodiments, the visual properties of background objects differ between background objects. For example, at a particular level of immersion, one or more first background objects are more visually de-emphasized (e.g., dimmed, blurred, and / or displayed with increased transparency) than one or more second background objects, and one or more third background objects cease to be displayed. In some embodiments, a navigation user interface with a first level of immersion includes one or more first travel user interface elements that, when selected, cause the computer system to change the display of the navigation user interface from a first view of a first physical location to a view of the selected first travel user interface element, such as a second view of a second physical location as described herein.

[0212] In some embodiments, while displaying, via the display generation component, a navigation user interface that includes one or more first travel user interface elements, the computer system detects (802b), via the one or more input devices, a first input corresponding to a request to change the level of immersion, such as FIG. 7A In some embodiments, the computer system detects (802c) a first user input corresponding to a request to change the level of immersion of the navigation user interface. For example, the computer system detects a first user input corresponding to a request to change the level of immersion of the navigation user interface. In some embodiments, the computer system detects (802c) a first user input corresponding to a request to change the level of immersion of the navigation user interface, such as

[0213] In response to detecting the first input (802c), the computer system changes (802d) the display of the navigation user interface with a first level of immersion corresponding to the first view of the first physical location, such as FIG. 7B In some embodiments, the computer system changes (802d) the display of the navigation user interface 704a in FIG. 7A to a second level of immersion corresponding to the first view of the first physical location, such as the navigation user interface 704b in FIG. 7B, where the second level of immersion includes one or more second travel user interface elements that are different from the one or more first travel user interface elements, such as FIG. 7B In some embodiments, the computer system changes (802d) the display of the navigation user interface 704a in FIG. 7A to a second level of immersion corresponding to the first view of the first physical location, such as the navigation user interface 704b in FIG. 7B, where the second level of immersion includes one or more second travel user interface elements that are different from the one or more first travel user interface elements, such asFIG. 7C selectable travel user interface objects 708a, 708b, 708c, 708d, 708e, and 708f in the navigation user interface 704b in FIG. 7B, which can be selected to change the display of the navigation user interface from the first view of the first physical location to the second view of the second physical location, such as FIG. 7B the second view of the second physical location in the navigation user interface 704b in FIG. 7B. In some embodiments, the navigation user interface having the second level of immersion that includes the one or more second travel user interface elements is more (e.g., in number) than the one or more first travel user interface elements associated with the navigation user interface having the first level of immersion. In some embodiments, the navigation user interface having the second level of immersion that includes the one or more second travel user interface elements is less (e.g., in number) than the one or more first travel user interface elements associated with the navigation user interface having the first level of immersion. In some embodiments, the one or more second travel user interface elements are of a second type of travel user interface element that is different from a first type of travel user interface element associated with the one or more first travel user interface elements. For example, the first type of travel user interface element optionally includes views from a street level, while the second type of travel user interface element includes views from the street level and other levels, such as an aerial view. As another example, the first type of travel user interface element optionally includes views from physical locations within a first predetermined distance (e.g., 0.02 km, 0.04 km, 0.06 km, 0.08 km, 0.2 km, 0.4 km, 0.6 km, 0.8 km, 1 km, 1.5 km, 2 km, 3 km, 4 km, 5 km, 6 km, 7 km, 8 km, 9 km, or 10 km) from the first physical location, while the second type of travel user interface element optionally includes views from physical locations more than the first predetermined distance from the first physical location.

[0214] In some embodiments, changing the display of the navigation user interface having the second level of immersion to a second view corresponding to a second physical location associated with the second travel user interface element includes the computer system displaying an image taken from the second physical location, which is different from the first physical location. In some embodiments, changing the display of the navigation user interface having the second level of immersion to a second view corresponding to a second physical location associated with the second travel user interface element includes the computer system displaying a (e.g., live) video recorded at the second physical location. In some embodiments, changing the display of the navigation user interface having the second level of immersion to a second view corresponding to a second physical location associated with the second travel user interface element includes updating the user’s point of view to be at the location of the second travel user interface element within the three-dimensional environment. In some embodiments, the first view is different from the second view, as described below, regardless of whether associated with the same physical location. In some embodiments, the one or more first travel user interface elements are selectable to change the display of the navigation user interface having the second level of immersion corresponding to the first view of the first physical location to a third view corresponding to a third physical location associated with the respective one or more first travel user interface elements. In some embodiments, changing the display of the navigation user interface having the second level of immersion to a third view corresponding to the third physical location corresponding to the first travel user interface element includes the computer system displaying an image taken from the third physical location, which is different from the second physical location. In some embodiments, changing the display of the navigation user interface having the second level of immersion to a third view corresponding to the third physical location associated with the first travel user interface element includes the computer system displaying a (e.g., live) video recorded at the third physical location. In some embodiments, changing the display of the navigation user interface having the second level of immersion to a third view corresponding to the third physical location associated with the first travel user interface element includes updating the user’s point of view to be at the location of the first travel user interface element within the three-dimensional environment. In some embodiments, the computer system displays the navigation user interface having the second level of immersion including the one or more first travel user interface elements and the one or more second travel user interface elements at respective locations of the navigation user interface (e.g., at different depths and / or heights that do not conflict).changing the number of displayed travel user interface elements based on the level of immersion, where a respective travel user interface element, when selected, causes the computer system to change the display of the navigation user interface having a first view of a first physical location to a second view corresponding to a second physical location associated with the respective travel user interface element, provides a quick display of the respective location without requiring the user to traverse the map, and reduces confusion due to displaying too many travel user interface elements, thereby reducing the number of inputs and providing a more efficient interaction between the user and the computer system.

[0215] In some embodiments, the first view of the first physical location includes a first perspective from a simulated camera, such as FIG. 7C In some embodiments, the first perspective from a simulated camera optionally includes a first simulated camera angle that is relative to a normal (perpendicular) to the ground (e.g., the simulated camera is pointed at an angle to the scene (or environment) of the first physical location). In some embodiments, the first perspective from a simulated camera includes a first angle that is normal or deviates from normal to the ground, such as 60 degrees, 65 degrees, 70 degrees, 75 degrees, 80 degrees, 85 degrees, 90 degrees, 100 degrees, 110 degrees, 120 degrees, 130 degrees, 140 degrees, 150 degrees, or 160 degrees, and / or a first distance from the ground, such as 50 centimeters, 60 centimeters, 70 centimeters, 80 centimeters, 90 centimeters, 100 centimeters, 150 centimeters, 200 centimeters, 250 centimeters, 300 centimeters, 400 centimeters, or 500 centimeters from the ground. For example, the first perspective from a simulated camera includes a street view perspective of the first physical location.

[0216] In some embodiments, the second view of the second physical location is from a second perspective of a simulated camera that is different from the first perspective of the simulated camera, such as FIG. 7Aa second view of a second physical location in the navigation user interface 704b in Figure 7B. In some embodiments, the second perspective from the simulated camera includes a second angle normal or offset from the ground that is equal to, less than, or greater than the first angle associated with the first perspective, such as 2 degrees, 5 degrees, 10 degrees, 20 degrees, 30 degrees, 40 degrees, 50 degrees, 60 degrees, 70 degrees, 80 degrees, or 90 degrees. In some embodiments, the second perspective from the simulated camera includes a second distance from the ground that is equal to, less than, or greater than the first distance associated with the first perspective, such as 5 meters, 10 meters, 50 meters, 100 meters, 200 meters, 300 meters, 400 meters, 500 meters, 600 meters, 700 meters, 800 meters, 900 meters, or 1000 meters from the ground. For example, the second perspective from the simulated camera is optionally above the street level (e.g., on a rooftop, balcony, or other overhead view) or below the street level (e.g., in a public transit tunnel, on a boat, or other underground view). In some embodiments, the second perspective includes different map information and / or satellite imagery, as described below with respect to different zoom levels. Providing respective views of respective physical locations from different perspectives via travel user interface elements enables the user to easily view and experience the first physical location from a variety of perspectives, thereby reducing the need for subsequent input by the user to zoom, rotate, and / or pan respective views of respective locations, which reduces power consumption and improves battery life of the computer system by enabling the user to use the computer system faster and more efficiently.

[0217] In some embodiments, when displaying the navigation user interface corresponding to the first view of the first physical location, the computer system displays, via the display generation component, a first navigation user interface element within the three-dimensional environment, where the first navigation user interface element represents a third view of the first physical location, such as FIG. 7AThe first navigation user interface element optionally includes a three-dimensional topographical map of the physical location associated with the first location experience. In some embodiments, the three-dimensional topographical map includes satellite image data of different resolutions, as described in more detail below with reference to zoom levels. For example, the first navigation user interface element includes a three-dimensional topographical map of a city that includes three-dimensional representations of buildings, streets, and other landmarks, with signs, placemarker, or other visual indications displayed at first locations that correspond to addresses, landmarks, or coordinates in the city. In some embodiments, the computer system displays the first navigation user interface element representing the third view of the first physical location as oriented along a horizontal surface in the three-dimensional environment or floating along a horizontal plane, and respective content corresponding to the first view of the first physical location is oriented vertically. In some embodiments, the first navigation user interface element is positioned between the user’s viewpoint and the respective content in the three-dimensional environment. In some embodiments, the computer system changes the display of the first navigation user interface element (e.g., rotates, resizes, and / or tilts) and / or renders respective portions of the first navigation user interface element at a zoom level to focus on the area and / or location of the map associated with the first physical location. For example, when the first navigation user interface element represents the third view of the first physical location, the computer system optionally displays the first navigation user interface element at a first rotation, a first size, a first zoom level, and / or a first tilt such that the first navigation user interface element is centered on the area and / or portion associated with the first physical location (e.g., without changing the viewpoint of the user of the computer system). In some embodiments, the first navigation user interface element representing the third view of the first physical location includes an area of the first physical location that is the same, larger, or smaller than an area of the respective physical location associated with the display of the respective content having a respective level of immersion. In some embodiments, the respective content is displayed concurrently with the first navigation user interface element. In some embodiments, and as will be described below, the computer system does not display the respective content concurrently with the first navigation user interface element. In some embodiments, the first navigation user interface element is a two-dimensional map of the first physical location.

[0218] In some embodiments, the first navigation user interface element includes an indication (e.g., a visual indication) of the position of the viewpoint displayed at the respective location within the first navigation user interface element. For example, the computer system optionally displays, within the first navigation user interface element, a visual indication (e.g., a binoculars icon or other graphical representation) corresponding to the position of the first view of the first physical location (e.g., the position from which the first view of the first physical location is captured by a camera or simulated camera).

[0219] In some embodiments, the first navigation user interface element includes an indication of a field of view displayed at a respective orientation relative to the first navigation user interface element. For example, the computer system optionally displays, within the first navigation user interface element, a visual indication of a field of view corresponding to an orientation of the first view of the first physical location (e.g., relative to a fixed coordinate system of the first physical location). In some embodiments, the indication of the field of view is displayed adjacent to or in association with (e.g., adjacent to or incorporated into) an indication of a location of the view corresponding to a location of the first view of the first physical location described herein. In some embodiments, the indication of the field of view indicates a boundary of the first physical location displayed within the respective content. In some embodiments, the computer system updates the indication of the location of the view and the indication of the field of view in accordance with a changed location and / or orientation of the respective content. It will be appreciated that although embodiments described herein relate to respective content corresponding to a first view of a first physical location, such indications and associated functionality and / or features optionally apply to other views and / or physical locations, including a second view of a second physical location. Providing an indication of a viewpoint location and / or a field of view in accordance with displayed respective content provides an effective way of indicating a portion of a first physical location represented by the respective content, which reduces power consumption and improves battery life of the computer system by enabling a user to use the computer system more quickly and efficiently.

[0220] In some embodiments, the navigation user interface having the first level of immersion occupies a first portion of the three-dimensional environment, such as FIG. 7B The navigation user interface 704a occupies a portion of the three-dimensional environment 702. In some embodiments, the navigation user interface having the second level of immersion occupies a second portion of the three-dimensional environment, which is greater than the first portion of the three-dimensional environment, such as FIG. 7AThe navigation user interface 704b occupies a second portion of the three- dimensional environment 702. In some embodiments, the computer system increases or decreases the size of the navigation user interface displayed by the display generation component in the three-dimensional environment. For example, in response to detecting a first input corresponding to a request to change the level of immersion (e.g., capturing and / or receiving user input), the computer system optionally increases the display area of the navigation user interface that is occupied by respective content corresponding to the first view of the first physical location. As described above with respect to the level of immersion, when displaying the navigation user interface with the second level of immersion, the computer system optionally replaces, occludes, or blocks the view of physical objects or surfaces (e.g., the front wall, the front wall and ceiling, the front wall and floor, the front wall and side walls, and / or a combination of any of the foregoing) with respective content including newly displayed virtual elements (e.g., the travel user interface elements) or newly displayed portions of the respective content that were not previously displayed when displaying the navigation user interface with the first level of immersion. As another example, when displaying the navigation user interface with the second level of immersion, the computer system optionally replaces, occludes, or blocks a portion of fewer physical objects (or a subset thereof) or surfaces (e.g., half of the front wall, the ceiling, or the floor) with respective content including a subset of virtual elements (e.g., the travel user interface elements) or a less portion of the respective content than was previously displayed when displaying the navigation user interface with the second level of immersion. In some embodiments, displaying and / or changing the navigation user interface with respective levels of immersion is performed by the computer system without requiring user input to change the level of immersion of the navigation user interface. Displaying the navigation user interface with different levels of immersion that occupy a greater or lesser portion of the three-dimensional environment provides an efficient way of presenting a smaller or greater view of the first physical location, thereby reducing the need for subsequent input by the user to manipulate the view of the first physical location, which reduces power usage and improves battery life of the computer system by enabling the user to use the computer system more quickly and efficiently.

[0221] In some embodiments, displaying, via the display generation component, the navigation user interface including respective content with respective levels of immersion includes displaying the navigation user interface with a first level of immersion that simulates the navigation user interface being along a physical object (such as a wall) within a physical environment surrounding the computer system FIG. 7CThe navigation user interface is displayed in a manner that mimics the orientation of the physical wall in the physical environment. For example, the computer system optionally displays the navigation user interface overlaid on a physical wall or any surface as described above, or no surface at all, such as overlaid on a virtual object. In some embodiments, the computer system displays the navigation user interface as flat or curved around a user sphere of the computer system. In some embodiments, the navigation user interface is displayed as curved or any shape to map to any physical surface or virtual object. In some embodiments, the physical wall, the physical table, and other physical objects described herein surround the display generation component and are visible through transparent portions of the display generation component (e.g., real or real-tight pass-through). For example, the representation of the physical environment, including the representation of the physical wall and the representation of the physical table, is displayed in a three-dimensional environment via the display generation component (e.g., virtual or video pass-through). Displaying the navigation user interface in a manner that mimics the orientation of the navigation user interface along the physical objects within the physical environment around the computer system provides an effective way of presenting a view of the first physical location as if it were part of the physical environment, taking advantage of the physical environment of the computer system, reducing the need for subsequent input by the user to manipulate the view of the first physical location, which reduces power usage and improves battery life of the computer system by enabling the user to use the computer system faster and more efficiently.

[0222] In some embodiments, while displaying the navigation user interface with the second level of immersion including the one or more second travel user interface elements, the computer system detects, via the one or more input devices, a second input directed to a respective travel user interface element, such as FIG. 7B In some embodiments, the user input 718c is directed to the travel user interface object 720a. For example, the computer system optionally detects a second input directed to a respective travel user interface element (e.g., a gaze of the user, a contact on a touch-sensitive surface, an actuation of a physical input device, a predefined gesture (e.g., a pinch gesture or an air tap gesture), and / or a voice input from the user). In some embodiments, the respective travel user interface element corresponds to a request to change a view of the respective physical location and / or the respective physical location.

[0223] In some embodiments, in response to detecting the second input, and in accordance with a determination that the respective travel user interface element is a first travel user interface element of the one or more second travel user interface elements, the computer system changes the display of the navigation user interface with the second level of immersion corresponding to the first view of the first physical location to the second level of immersion corresponding to a second view of a second physical location, such as from FIG. 7Cthe navigation user interface 704b in FIG. 7B changes to FIG. 7C the navigation user interface 704b in FIG. 7B. In some embodiments, changing the display of the navigation user interface having the second level of immersion to the second view corresponding to the second physical location includes one or more characteristics of the display of the navigation user interface having the second level of immersion corresponding to the second view of the second physical location described above.

[0224] In some embodiments, in response to detecting the second input, and in accordance with a determination that the respective travel user interface element is a second travel user interface element of the one or more second travel user interface elements that is different from the first travel user interface element, the computer system changes the display of the navigation user interface having the second level of immersion corresponding to the first view of the first physical location to the second level of immersion corresponding to a third view of a third physical location that is different from the second view of the second physical location. In some embodiments, changing the display of the navigation user interface having the second level of immersion to the third view corresponding to the third physical location includes one or more characteristics of the display of the navigation user interface having the second level of immersion corresponding to the third view of the third physical location described above. In some embodiments, the computer system automatically (or in response to a user input) moves or repositions the navigation user interface including the display of the respective content from a first location within the physical environment surrounding the computer system to a second location within the physical environment surrounding the computer system (e.g., from an orientation along a first physical object within the physical environment surrounding the computer system to an orientation along a second physical object within the physical environment surrounding the computer system). In some embodiments, the second location is automatically selected by the computer system or selected by a user of the computer system. In some embodiments, the size of the display of the respective content optionally does not change (e.g., remains the same). In some embodiments, the second location is optionally further away (relative to the viewpoint of the user) than the first location, such that the display of the respective content appears smaller relative to the perspective of the user than when the display of the respective content is displayed at the first location. In other embodiments, the second location is optionally closer (relative to the viewpoint of the user) than the first location, such that the display of the respective content appears larger relative to the perspective of the user than when the display of the respective content is displayed at the first location. In some embodiments, the computer system automatically (or in response to a user input) adjusts the size of the display of the respective content and / or changes the size of the display of the respective content. Changing the display of the navigation user interface corresponding to the first view of the first physical location to the respective view corresponding to the respective physical location associated with the respective travel user interface element provides for a quick display of the respective physical location without requiring the user to traverse the map, thereby reducing the number of inputs and providing for more effective interactions between the user and the computer system.

[0225] In some embodiments, where the navigation user interface is displayed with a second level of immersion corresponding to a second view of a second physical location includes displaying a travel option that, when selected, causes the computer system to display a previous view of the respective physical location, such as FIG. 7C the travel user interface object 720a in FIG. 7B. In some embodiments, the travel option is displayed as part of the navigation user interface at a location that does not disrupt the respective level of immersion of the navigation user interface. In some embodiments, the computer system displays the travel option in response to changing the display of the navigation user interface to a view and / or physical location that is different from the first view of the first physical location. In some embodiments, when the computer system determines that the display of the navigation user interface with the second level of immersion corresponding to the second view of the second physical location has not changed (e.g., no respective travel user interface element has been selected since the navigation user interface display was initiated), the computer system ceases to display the travel option.

[0226] In some embodiments, while displaying the navigation user interface with the second level of immersion corresponding to the second view of the second physical location, the computer system detects, via the one or more input devices, a third input directed to the travel option, such as FIG. 7C the user input 718c directed to the travel user interface object 720a in FIG. 7B. For example, the third input includes the following directed to the travel option: a gaze of the user, a contact on the touch- sensitive surface, an actuation of a physical input device, a predefined gesture (e.g., a pinch gesture or an air tap gesture), and / or a voice input from the user.

[0227] In some embodiments, in response to detecting the third input, the computer system changes the display of the navigation user interface with the second level of immersion corresponding to the second view of the second physical location to the first view corresponding to the first physical location, such as from the navigation user interface 704b in FIG. 7B to FIG. 7D the navigation user interface 704a in FIG. 7A. FIG. 7CThe computer system displays the navigation user interface 704b corresponding to the first view of the first physical location. In some embodiments, the computer system maintains the level of immersion. In some embodiments, displaying the navigation user interface 704b corresponding to the first view of the first physical location includes displaying a first travel option that, when selected, causes the computer system to display a previously displayed view of the first physical location. For example, selecting the first travel option causes the computer system to optionally change from displaying the navigation user interface 704b with the first level of immersion corresponding to the first view of the first physical location to displaying the navigation user interface 704b with the first level of immersion corresponding to the previously displayed view of the first physical location. Providing the option to display the previously displayed view of the first physical location provides a quick display of the previously displayed view of the first physical location without requiring the user to traverse the map, thereby reducing the number of inputs and providing a more efficient interaction between the user and the computer system.

[0228] In some embodiments, while displaying the navigation user interface corresponding to the first view of the first physical location, the computer system displays, via the display generation component, a first navigation user interface element within the three-dimensional environment, where the first navigation user interface element represents a third view of the first physical location and includes an indication of a second computer system (e.g., associated with another user different from the user of the computer system) at a respective location within the first physical location corresponding to a current location of the second computer system within the first physical location, such as the travel user interface object 720e, within the first navigation user interface element. In some embodiments, the first navigation user interface element includes one or more characteristics of the first navigation user interface elements described above. In some embodiments, the computer system receives an indication of the current location of the second computer system within the first physical location (e.g., from the second computer system, from a server). For example, the user of the computer system and the user of the second computer system are connected via a service that presents the location of the second computer system to the computer system and the location of the computer system to the second computer system. In some embodiments, as the second computer system moves such that the current location of the second computer system changes, the computer system updates the display of the first navigation user interface element to move and / or change the indication of the second computer system to correspond to the changed current location of the second computer system.

[0229] In some embodiments, while displaying the first navigation user interface element including the indication of the second computer system within the first navigation user interface element, the computer system detects, via the one or more input devices, a second input directed to the indication of the second computer system, such as FIG. 7CThe middle finger points to the user input 718c of the user interface object 720e. For example, the second input includes the following: indications to the second computer system: the user's gaze, contact on a touch-sensitive surface, actuation of a physical input device, predefined gestures (e.g., pinch gestures or air tap gestures) and / or voice input from the user.

[0230] In some implementations, in response to detecting a second input, the computer system changes the display of a navigation user interface having a second level of immersion corresponding to a first view of a first physical location to a corresponding view corresponding to the current location from the second computer system, such as from... FIG. 7D The navigation user interface in 704b has been changed to FIG. 7C The navigation user interface 704b is described. In some embodiments, changing the display of the navigation user interface with a second level of immersion to correspond to a corresponding view of the current location of the second computer system includes displaying an image taken from the current location of the second computer system, which is different from the first physical location. In some embodiments, the image includes multiple images and / or video content (e.g., captured by the second computer system or by different electronic devices) associated with the current location of the second computer system. In some embodiments, changing the display of the navigation user interface with a second level of immersion to correspond to a corresponding view of the current location of the second computer system includes the computer system displaying (e.g., live) video recorded at the current location of the second computer system. For example, the recorded video is captured by the second computer system (e.g., one or more cameras and / or microphones communicating with the second computer system). Changing the display of the navigation user interface to correspond to a corresponding view from the current location of the second computer system provides an efficient way to view a corresponding view of the location of the second computer system without leaving the navigation user interface, thereby reducing the need for subsequent input to search for the current location of the second computer system. This simplifies the interaction between the user and the computer system, enhances the operability of the computer system, and makes the user-computer system interface more efficient.

[0231] In some implementations, when displaying a navigation user interface corresponding to a first view of a first physical location, the computer system displays a first navigation user interface element within a three-dimensional environment via a display generation component. This first navigation user interface element represents a third view of the first physical location and includes additional content corresponding to the first physical location, such as… FIG. 7Ga representation 722a of a merchant within a third view of the first physical location. In some embodiments, the first navigation user interface element includes one or more features of the first navigation user interface element described above. In some embodiments, the additional content corresponding to the first physical location includes information corresponding to one or more points of interest (e.g., landmarks, parks, buildings, merchants, or other entities of interest to the user). For example, when the point of interest corresponds to a merchant, the additional content includes the name of the merchant, the address of the merchant, the category of the merchant, the distance to the merchant, and / or other information associated with the merchant. In some embodiments, the additional content includes a virtual object that, when selected, initiates a communication with the merchant and / or navigates to an application or website associated with the merchant. In some embodiments, the computer system displays the additional content at a respective location of the first navigation user interface element (e.g., within, proximate to, overlaid on, etc. the respective physical location of the merchant). In some embodiments, the computer system overlays the additional content on a virtual object or a physical object within the physical environment surrounding the computer system, such as a wall or any other surface. In some embodiments, the computer system presents the additional content in response to user input corresponding to selection of a virtual object within the first navigation user interface element corresponding to the merchant. In some embodiments, the virtual object is displayed proximate to a respective location within the first navigation user interface element corresponding to the physical location of the merchant within the first physical location. In some embodiments, the computer system displays the additional content within the navigation user interface. In some embodiments, the quantity and / or amount of additional content displayed within the navigation user interface is based on the level of immersion. For example, the second level of immersion includes more, less, or the same amount of additional content than the first level of immersion. Displaying additional content corresponding to the first physical location provides an efficient way to view additional content corresponding to the first physical location without having to exit the navigation user interface, thereby reducing the need for subsequent input to search for additional content corresponding to the first physical location, which simplifies the interaction between the user and the computer system and enhances the operability of the computer system and the user-computer system interface, making it easier to use.

[0232] In some embodiments, while displaying the navigation user interface corresponding to the first view of the first physical location, the computer system displays, via the display generation component, a first navigation user interface element within the three-dimensional environment, the first navigation user interface element representing a third view of the first physical location and including additional content corresponding to weather conditions at the first physical location, such as FIG. 7Ian indication of the weather at the first physical location 712. In some embodiments, the first navigation user interface element includes one or more characteristics of the first navigation user interface elements described above. In some embodiments, the additional content corresponding to the weather conditions at the first physical location includes an indication (and / or representation) of precipitation, clouds, sunshine, wind, fog, and / or haze that is occurring, has occurred, or is forecasted to occur at the first physical location. In some embodiments, the additional content indication is a graphical representation of the weather conditions displayed at a location in the first navigation user interface element. For example, when the first physical location is experiencing rain, the computer system displays a representation of rain clouds and rain at a corresponding location in the first navigation user interface element that corresponds to the first physical location that is experiencing rain. In some embodiments, in addition to displaying the graphical representation, the computer system displays a textual description of the weather conditions, such as forecasted conditions, temperature, UV index, wind information, and / or other weather measurements. In some embodiments, the computer system displays the additional content within the navigation user interface. Displaying the additional content corresponding to the weather conditions of the first physical location provides an efficient way to view the additional content corresponding to the first physical location without leaving the navigation user interface, thereby reducing the need for subsequent input to search for the weather conditions of the physical location, which simplifies the interaction between the user and the computer system and enhances the operability of the computer system and the user-computer system interface.

[0233] In some embodiments, while displaying the navigation user interface corresponding to the first view of the first physical location, the computer system detects a second input corresponding to a request to change a zoom level of the first view of the first physical location, such as FIG. 7J In some embodiments, the second input corresponds to a request 740i to change the zoom level. For example, the second input includes the following corresponding to a request to change the zoom level: a gaze of the user, a contact on the touch- sensitive surface, an actuation of a physical input device, a predefined gesture (e.g., a pinch gesture or an air tap gesture), and / or a voice input from the user. In some embodiments, the second input is directed to the navigation user interface. In some embodiments, the second input is directed to a user interface element that, when selected, initiates a process to change the zoom level as described herein. In some embodiments, the computer system initially displays respective content of the navigation user interface corresponding to the first view of the first physical location with one or more characteristics from the first perspective of the simulated camera described above.

[0234] In some embodiments, in response to detecting the second input, and in accordance with a determination that the request to change the zoom level corresponds to changing the zoom level by a first amount, the computer system, via the display generation component, displays, within the three-dimensional environment, a first navigation user interface element representing a third view of the first physical location, such asFIG. 7J In some embodiments, the zoom level is changed by a first amount based on a magnitude of the second input (e.g., a magnitude of velocity, distance, and / or duration). For example, the second input is optionally a pinch zoom gesture, when the user’s attention is directed to the navigation user interface (e.g., a predetermined portion of the respective content and / or a virtual object within the navigation user interface that is selectable to change the zoom level) while the user’s hand performs a pinch air gesture that includes the thumb and index finger of the hand contacting and pinching together, or the index finger directly interacts with (e.g., air tap or air touch) a virtual object within the navigation user interface that is selectable to change the zoom level, or the user’s two hands are brought closer together (e.g., to zoom out the respective content of the navigation user interface) or further apart (e.g., to zoom in the respective content of the navigation user interface). In some embodiments, the first navigation user interface element includes one or more characteristics of the first navigation user interface elements described above. For example, the computer system initially displays the first navigation user interface element at a first zoom level that is appropriate for the first physical location, such that the first navigation user interface element is centered on the region and / or portion associated with the first physical location (e.g., without changing the viewpoint of the user of the computer system). In some embodiments, the computer system changes the zoom level in response to a voice input from the user that corresponds to a request to change the zoom level of the first view of the first physical location.

[0235] In some embodiments, and in accordance with a determination that the request to change the zoom level corresponds to a change in the zoom level by a second amount that is different from the first amount, the computer system changes the display of the first navigation user interface element from the third view representing the first physical location to a fourth view of the first physical location, such as from FIG. 7K In some embodiments, and in accordance with a determination that the request to change the zoom level corresponds to a change in the zoom level by a second amount that is different from the first amount, the computer system changes the display of the first navigation user interface element from the third view representing the first physical location to a fourth view of the first physical location, such as from FIG. 7Nthe first navigation user interface element. In some embodiments, the view of the first navigation user interface element changes in accordance with a zoom level (e.g., an amount of zoom) applied to the first navigation user interface element in response to input as described herein. In some embodiments, the second amount of zoom level is greater than the first amount of zoom level. In some embodiments, the second amount of zoom level is less than the first amount of zoom level. For example, when the computer system determines that the amount of zoom is such that the zoom level is below a first threshold, the computer system optionally displays a second navigation user interface element described below. As another example, when the zoom level is between the first threshold and a second threshold that is greater than the first threshold, the computer system displays the first navigation user interface element. In some embodiments, when the zoom level is greater than the second threshold, the computer system displays the respective content. In some embodiments, when the first navigation user interface element represents the fourth view of the first physical location, the first navigation user interface element is centered on a second region and / or a second portion associated with the first physical location that is the same, larger, or smaller than the region and / or portion associated with the first physical location when the first navigation user interface element represents the third view of the first physical location (e.g., prior to the second input that changes the zoom level by the second amount). For example, when the first navigation user interface element represents the fourth view of the first physical location, the first navigation user interface element includes natural features (e.g., terrain, vegetation, etc.) that are optionally not shown when the first navigation user interface element represents the third view of the first physical location.

[0236] In some embodiments, and in accordance with a determination that the request to change the zoom level corresponds to changing the zoom level by a third amount that is different from the first amount and different from the second amount, the computer system, via the display generation component, displays a second navigation user interface element within the three-dimensional environment, where the second navigation user interface element represents a first portion corresponding to the first physical location, such as FIG. 7Kthe first navigation user interface element in FIG. 7A. In some embodiments, the view of the first navigation user interface element changes according to a zoom level (e.g., an amount of zoom) applied to the first navigation user interface element in response to input as described herein. In some embodiments, the third amount is greater than or less than the second amount. For example, when the computer system determines that the third amount of zooming causes the zoom level to be below the first threshold, the computer system optionally displays a second navigation user interface element representing a first portion corresponding to the first physical location. In some embodiments, the second navigation user interface element is a three-dimensional globe centered on the region and / or portion associated with the first physical location (e.g., without changing the viewpoint of the user of the computer system). For example, the second navigation user interface element representing the first portion corresponding to the first physical location optionally includes a view of North America and South America. Providing the ability to zoom in or out to view a more detailed and / or larger region corresponding to the first physical location provides an efficient way for viewing respective views corresponding to the first physical location, which simplifies the interaction between the user and the computer system and enhances the operability of the computer system and makes the user-computer system interface more efficient.

[0237] In some embodiments, while displaying the second navigation user interface element including the first portion corresponding to the first physical location, the computer system detects a third input corresponding to a request to change the zoom level of the first portion, such as FIG. 7O In some embodiments, the third input has one or more characteristics of the second input described above.

[0238] In some embodiments, in response to detecting the third input, and in accordance with a determination that the request to change the zoom level corresponds to changing the zoom level by a fourth amount, the computer system changes the display of the second navigation user interface element from representing the first portion of the first physical location to a second portion of the first physical location, such as, for example, zooming in FIG. 7Ma portion of the second navigation user interface element 754. In some embodiments, the view of the second navigation user interface element changes in accordance with a zoom level (e.g., an amount of zoom) that is applied to the second navigation user interface element in response to input as described herein. In some embodiments, the fourth amount is greater than the third amount. In some embodiments, the fourth amount is less than the third amount. In some embodiments, when the second navigation user interface element represents the second portion of the first physical location, the second navigation user interface element is centered on a second region and / or a second portion associated with the first physical location that is the same, larger, or smaller than a region and / or portion associated with the first physical location when the second navigation user interface element represents the first portion of the first physical location (e.g., prior to the third input that changes the zoom level by the fourth amount). For example, the second portion optionally includes a view of North America, in contrast to a view that optionally includes North America and South America when the second navigation user interface element represents the first portion corresponding to the first physical location. In some embodiments, in response to detecting the third input, the computer system ceases to display the second navigation user interface element.

[0239] In some embodiments, in accordance with a determination that the request to change the zoom level corresponds to changing the zoom level by a fifth amount that is different from the fourth amount, the computer system, via the display generation component, displays, within the three-dimensional environment, a first navigation user interface element representing a third view of the first physical location, such as FIG. 7F the navigation user interface element 714 in FIG. 7M. In some embodiments, the fifth amount is greater than or less than the fourth amount. In some embodiments, the first navigation user interface element representing the third view of the first physical location within the three-dimensional environment includes one or more characteristics of the first navigation user interface element representing the third view of the first physical location within the three-dimensional environment described above.

[0240] In some embodiments, and in accordance with a determination that the request to change the zoom level corresponds to changing the zoom level by a sixth amount that is different from the fourth amount and that is different from the fifth amount, the computer system displays respective content corresponding to the first view of the first physical location, such as FIG. 7Ithe navigation user interface 704b in FIG. 7B. In some embodiments, the sixth amount is greater than or less than the fifth amount. In some embodiments, the respective content corresponding to the first view of the first physical location includes one or more characteristics of the respective content corresponding to the first view of the first physical location described above. In some embodiments, while displaying the respective content corresponding to the first view of the first physical location, the computer system ceases to display the first navigation user interface element and / or the second navigation user interface element. Providing the ability to zoom in or out to view a more detailed and / or larger area corresponding to the first physical location in response to user input provides an efficient way to view respective views corresponding to the first physical location, which simplifies the interaction between the user and the computer system and enhances the operability of the computer system and makes the user-computer system interaction more efficient.

[0241] In some embodiments, while displaying the first navigation user interface element representing the third view of the first physical location, the computer system displays the first navigation user interface element within a first (e.g., simulated physical) volume of the three-dimensional environment, such as FIG. 7J the navigation user interface element 714 in FIG. 7B. For example, the first navigation user interface element representing the third view of the first physical location encompasses a region of the three-dimensional environment having a predefined first volume appropriate for the first physical location, such that the first navigation user interface element is centered on the region and / or portion associated with the first physical location and includes all map elements (e.g., artificial and / or natural features described above). In some embodiments, the computer system changes the volume of the first navigation user interface element to display more or fewer map elements and / or at a higher or lower level of detail than when the first navigation user interface element is displayed within the first amount, as described below.

[0242] In some embodiments, while displaying the respective content corresponding to the first view of the first physical location, the computer system displays the respective content in a second (e.g., simulated physical) volume of the three-dimensional environment that is larger than the first volume in at least one dimension, such as FIG. 7G the navigation user interface element 714 in FIG. 7B. For example, the second volume is larger than the first volume in height and / or width. In some embodiments, the computer system changes the volume of the respective content as described above with respect to the level of immersion. Displaying respective content of the navigation user interface in a volume larger than the volume associated with the first navigation user interface element in the three-dimensional environment frees up space for display in the navigation user interface and reduces clutter.

[0243] In some embodiments, in which detecting the second input includes detecting that the pose of the user of the computer system has changed, such as FIG. 7Ha change in the user's 734 pose (e.g., the user tilting toward or away from the first navigation user interface element) that satisfies one or more criteria described below with respect to at least a portion of the user's pose (e.g., the positioning of the user's head indicates that the user's head is within a predetermined distance, such as 1 centimeter, 2 centimeters, 3 centimeters, 4 centimeters, 5 centimeters, 10 centimeters, 20 centimeters, 30 centimeters, 40 centimeters, 50 centimeters, 60 centimeters, 70 centimeters, 80 centimeters, 90 centimeters, 100 centimeters, or 200 centimeters, from the first navigation user interface element). For example, the computer system optionally detects a magnitude and / or degree of the change in the user's pose (e.g., the amount of head movement closer to the first navigation user interface element or back away from the first navigation user interface element) and, in response, the computer system updates the zoom level of the first navigation user interface element accordingly. For example, when the user's pose indicates that the user's head is moved closer to the first navigation user interface element, the computer system zooms in more than it did prior to detecting the change in pose. As another example, when the user's pose indicates that the user's head is moved further away from the first navigation user interface element, the computer system zooms out (or zooms in less) than it did when the user's pose indicated that the user's head was moved closer to the first navigation user interface element. In some embodiments, a change in pose that does not satisfy the one or more criteria does not cause the computer system to change the zoom level of the first navigation user interface element. In some embodiments, a change in pose other than tilting toward or away from the first navigation user interface element does not cause the computer system to change the zoom level of the first navigation user interface element. Changing the zoom level in response to detecting a change in the user's pose provides an intuitive way to zoom in and out of the first navigation user interface element (e.g., without having to manipulate a physical input device), which additionally reduces power consumption and improves battery life of the computer system by enabling the user to use the computer system faster and more efficiently.

[0244] In some embodiments, while displaying the first navigation user interface element representing the third view of the first physical location, the computer system detects a change in the user's pose of the computer system that satisfies the one or more criteria, such as FIG. 7H a change in the user's pose. In some embodiments, detecting a change in the user's pose of the computer system that satisfies the one or more criteria includes one or more of the characteristics of detecting a change in the user's pose of the computer system that satisfies the one or more criteria described above.

[0245] In some embodiments, in response to detecting a change in the pose of the user of the computer system that satisfies the one or more criteria, and in accordance with a determination that the change in the pose of the user of the computer system includes a first predetermined pose, the computer system scales the third view of the first physical location by a first amount, such as the first navigation user interface element 714 in Figure G. For example, the first predetermined pose includes a positioning of the user’s head within a second predetermined distance from the first navigation user interface element, such as 50 centimeters, 60 centimeters, 70 centimeters, 80 centimeters, 90 centimeters, 100 centimeters, or 200 centimeters. In some embodiments, scaling the third view of the first physical location by the first amount includes zooming the third view in by a first zoom value (e.g., 110%, 120%, 130%, 140%, or 150% of its original size), the original size being the size prior to detecting the change in the pose of the user by the first amount. In some embodiments, scaling the third view of the first physical location by the first amount includes displaying the map element in a larger or smaller size in accordance with the change in the pose of the user of the computer system that satisfies the one or more criteria.

[0246] In some embodiments, in accordance with a determination that the change in the pose of the user of the computer system includes a second predetermined pose that is different from the first predetermined pose, the computer system scales the third view of the first physical location by a second amount that is different from the first amount, such as the first navigation user interface element 714 in Figure G. FIG. 7H In some embodiments, in accordance with a determination that the change in the pose of the user of the computer system includes a second predetermined pose that is different from the first predetermined pose, the computer system scales the third view of the first physical location by a second amount that is different from the first amount, such as the first navigation user interface element 714 in Figure G.

[0247] In some embodiments, while displaying the first navigation user interface element representing the third view of the first physical location, the computer system detects a change in the pose of the user of the computer system that satisfies the one or more criteria, such as the first navigation user interface element 714 in Figure G. FIG. 7Hthe user of the computer system. In some embodiments, detecting the change in the pose of the user of the computer system that satisfies the one or more criteria includes one or more characteristics of detecting the change in the pose of the user of the computer system that satisfies the one or more criteria described above.

[0248] In some embodiments, in response to detecting the change in the pose of the user of the computer system that satisfies the one or more criteria, and in accordance with a determination that the change in the pose of the user of the computer system includes a first predetermined pose, the computer system displays a third view of the first physical location with a first level of detail, such as FIG. 7I the navigation user interface element 714 in FIG. 7B. In some embodiments, the first predetermined pose includes one or more characteristics of the first predetermined pose described above. In some embodiments and as described above, scaling the third view by the respective zoom value includes displaying a larger or smaller first navigation user interface element. In some embodiments, displaying the third view of the first physical location with the first level of detail includes a more detailed first physical area than a level of detail of the third view when the first predetermined pose is not detected. For example, displaying the third view of the first physical location with the first level of detail includes buildings, landmarks, parks, and / or road lanes that are optionally not included in the third view when the first predetermined pose is not detected.

[0249] In some embodiments, in accordance with a determination that the change in the pose of the user of the computer system includes a second predetermined pose that is different from the first predetermined pose, the computer system displays the third view of the first physical location with a second level of detail that is different from the first level of detail, such as FIG. 7G the level of detail shown by the navigation user interface element 714 in FIG. 7B. In some embodiments, the second predetermined pose includes one or more characteristics of the second predetermined pose described above. In some embodiments, displaying the third view of the first physical location with the second level of detail includes displaying a more detailed first physical area than a level of detail of the third view with the first level of detail. For example, displaying the third view of the first physical location with the second level of detail includes buildings, landmarks, parks, and / or road lanes, trees, vegetation, sidewalks, bike lanes, medians, and / or crosswalks that are optionally not included in the first level of detail. Changing the level of detail of the first physical location in response to detecting a predetermined pose of the user provides an intuitive way of zooming in and out of the first navigation user interface element (e.g., without having to manipulate a physical input device), which additionally reduces power consumption and improves battery life of the computer system by enabling the user to use the computer system faster and more efficiently.

[0250] In some implementations, detecting the second input corresponding to the zoom request includes detecting the user's viewpoint of the computer system (such as...). FIG. 7G The computer system detects a change in the user's viewpoint (734 in the system) that satisfies one or more criteria. For example, it detects that the user of the computer system has moved position within the physical environment. In some embodiments, the computer system determines the viewing angle (e.g., the viewing angle) of the user's viewpoint, which is formed by a first vector extending from the normal of a corresponding portion (e.g., the center) of a first navigation user interface element and a second different vector extending from the corresponding portion of the first navigation user interface element toward the user's viewpoint (e.g., the user's field of view, the center of the user's head, and / or the center of the computer system). For example, the viewing angle is the angle between the user's viewpoint and the first navigation user interface element. In some embodiments, the computer system detects a change in the viewing angle of the user's viewpoint that satisfies one or more criteria (e.g., within an angular range, such as 5 degrees, 10 degrees, 15 degrees, 20 degrees, 25 degrees, 30 degrees, 35 degrees, 40 degrees, 45 degrees, 50 degrees, 55 degrees, 60 degrees, 65 degrees, 70 degrees, or 75 degrees relative to the normal of the first navigation user interface element).

[0251] In some implementations, in response to detecting a change in the user's viewpoint on the computer system that satisfies one or more criteria, and based on determining that the change in the user's viewpoint includes movement along a first axis, the computer system changes zoom along a second axis corresponding to the first axis, such as... FIG. 7H The navigation user interface element 714 is shown in the diagram. In some embodiments, the first axis is the same as the second axis. In some embodiments, the first axis is different from the second axis. For example, a first predetermined viewpoint optionally includes a first viewing angle (e.g., at 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, or 75 degrees with respect to the normal) and a zoom axis corresponding to that first viewing angle. In some embodiments, changing the zoom along the second axis corresponding to the first axis includes one or more of the features described above for displaying a third view from a first zoom level.

[0252] In some implementations, based on determining changes in the user's viewpoint, including movement along a second axis, the computer system changes zoom along a first axis corresponding to the second axis, such as... FIG. 7Gcorresponding to the second viewing angle. In some embodiments, changing the zoom along the first axis corresponding to the second axis includes displaying one or more characteristics of the third view from the second zoom level as described above. In another example, while the computer system displays a first portion of the first navigational user interface element from a first zoom level corresponding to a first viewing angle relative to the first navigational user interface element, the computer system determines a change from the first viewing angle to a second viewing angle relative to the first navigational user interface element that is different from the first viewing angle. In response to the change from the first viewing angle to the second viewing angle, the computer system displays the same portion of the first navigational user interface element from a second zoom level that is different from the first zoom level. In some embodiments, the computer system zooms in on a given location in the first navigational user interface element— optionally for the same amount of head movement— when the viewing angle towards that given location during the head movement is different (e.g., when a first viewing angle towards that location from a first viewpoint when the head movement begins results in zooming along a first axis (e.g., perpendicular to the first viewing angle), while a second viewing angle towards that location from a second viewpoint when the head movement begins results in zooming along a second axis (e.g., perpendicular to the second viewing angle) that is different from the first axis). Changing the zoom level of the third view of the first physical location in accordance with a determination that the change in the viewpoint of the user includes a respective predetermined viewpoint improves the visual feedback regarding the positioning of the user relative to the first navigational user interface element, informing the user how subsequent changes will affect the level of detail of the respective view of the first physical location, which reduces errors in interactions with the computer system, enhances the operability of the computer system, and makes the user-computer system interface more efficient, which, additionally, saves power and reduces the possibility of user mistakes.

[0253] In some embodiments, in response to detecting the second input, and in accordance with a determination that a position of a gaze of a user of the computer system is directed at the first portion of the first navigational user interface element, such as while the second input is detected FIG. 7G In some embodiments, in response to detecting the second input, and in accordance with a determination that a position of a gaze of a user of the computer system is directed at the first portion of the first navigational user interface element, such as while the second input is detected FIG. 7H In some embodiments, in response to detecting the second input, and in accordance with a determination that a position of a gaze of a user of the computer system is directed at the first portion of the first navigational user interface element, such as while the second input is detected FIG. 7HFor example, the computer system determines that the user’s gaze has been directed at the first navigation user interface element for an amount of time that exceeds a predetermined time threshold (e.g., 0.05 seconds, 0.1 seconds, 0.2 seconds, 0.3 seconds, 0.5 seconds, 1 second, 2 seconds, 5 seconds, 10 seconds, 20 seconds, or 30 seconds), while detecting a second input that corresponds to a request to change the zoom level, and in response, the computer system zooms in on the first portion of the first navigation user interface element. In some embodiments, when the computer system zooms in on the first portion of the first navigation user interface element (e.g., the portion of the first navigation user interface element that the user’s gaze is directed at), the first portion of the first navigation user interface element remains at a fixed position in the three-dimensional environment (and the computer system optionally moves other portions of the first navigation user interface element in the three-dimensional environment to achieve this result as needed). In some embodiments, the computer system continues to increase the zoom level in accordance with the duration that the user’s gaze is directed at the first portion of the first navigation user interface element. In some embodiments, the first portion that the user’s gaze is directed at is rendered at a higher quality, clarity / resolution than other portions of the first navigation user interface element that are rendered at a lower quality. In some embodiments, the computer system continues to zoom in on the first portion of the first navigation user interface element as the computer system detects that the user’s gaze continues to be directed at the first portion of the first navigation user interface element. In some embodiments, as previously described, zooming in or out on the first portion requires head movement.

[0254] In some embodiments, in accordance with a determination that the position of the user’s gaze of the computer system when the second input is detected is directed at a second portion of the first navigation user interface element that is different from the first portion of the first navigation user interface element, the computer system changes the zoom level in accordance with the second portion of the first navigation user interface element along a second axis that is different from the first axis, such as FIG. 7I In some embodiments, the computer system changes the zoom level in accordance with the first portion of the first navigation user interface element along a first axis that is different from the second axis, such as FIG. 7Ithe first navigation user interface element. For example, the computer system determines that the user’s gaze has moved away from the first portion to a second portion of the first navigation user interface element, and in response, the computer system zooms in on the second portion of the first navigation user interface element. In some embodiments, the computer system changes the positioning of the first navigation user interface element so that the first navigation user interface element is centered on the second portion. In some embodiments, when the computer system zooms in on the second portion of the first navigation user interface element (e.g., the portion of the first navigation user interface element that the user’s gaze is directed towards), the second portion of the first navigation user interface element remains at a fixed position in the three-dimensional environment (and the computer system optionally moves other portions of the first navigation user interface element in the three-dimensional environment to achieve this result as needed). In some embodiments, the computer system continues to increase the level of zoom according to the duration of time that the user’s gaze is directed at the second portion of the first navigation user interface element. In some embodiments, the second portion that is the target of the user’s gaze is rendered at a higher quality, clarity / resolution than the first portion that is no longer the target of the user’s gaze, such that the first portion is rendered at a lower quality. In some embodiments, the computer system continues to zoom in on the second portion of the first navigation user interface element when the computer system detects that the user’s gaze continues to be directed at the second portion of the first navigation user interface element. In some embodiments, as previously described, zooming in or out on the second portion requires head movement. Changing the level of zoom in response to detecting a change in the location of the user’s gaze provides an intuitive way of zooming in and out on the first navigation user interface element (e.g., without needing to manipulate a physical input device), which additionally reduces power usage and improves battery life of the computer system by enabling the user to use the computer system faster and more efficiently.

[0255] In some embodiments, while displaying the first navigation user interface element representing the third view of the first physical location, the computer system detects a third input corresponding to a request to update the positioning and / or orientation of the first navigation user interface element within the three-dimensional environment, such as FIG. 7Icorresponding to a request 740i to update the position and / or orientation of the first navigation user interface element. In some embodiments, the third input includes the following corresponding to a request to update the position and / or orientation of the first navigation user interface element within the three-dimensional environment: a gaze of the user, a contact on the touch- sensitive surface, an actuation of a physical input device, a predefined gesture (e.g., a pinch gesture or an air tap gesture), and / or a voice input from the user. For example, the computer system initially displays the first navigation user interface element with a first position and / or orientation that is appropriate for the first physical location, such that the first navigation user interface element is centered on the area and / or portion associated with the first physical location. In some embodiments, the third input is directed to the first navigation user interface element. In some embodiments, the third input is directed to a control element for adjusting the position and / or orientation of the first navigation user interface element, as described in further detail below.

[0256] In some embodiments, in response to detecting the third input, the computer system changes the position and / or orientation of the first navigation user interface element within the three-dimensional environment based on the third input, such as from ​ the navigation user interface element 714 in FIG. 7M to FIG. 7Jthe first navigation user interface element in the three-dimensional environment, the computer system displays the first navigation user interface element with a second position and / or orientation that is appropriate for a second physical location that is different from the first physical location, such that the first navigation user interface element is rotated and / or positioned to be centered on a region and / or portion associated with the second physical location. In some embodiments, changing the position and / or orientation of the first navigation user interface element is based on one or more characte...

Claims

1. A method, the method comprising: At the computer system that communicates with the display generation components and one or more input devices: The navigation user interface is displayed in a three-dimensional environment via the display generation component, wherein the navigation user interface includes one or more first travel user interface elements displayed at a first immersion level corresponding to a first view of a first physical location; When the navigation user interface including the one or more first navigating user interface elements is displayed via the display generation component, a first input corresponding to a request to change the immersion level is detected via the one or more input devices; as well as In response to the detection of the first input: The display of the navigation user interface having a first immersion level corresponding to the first view of the first physical location is changed to a second immersion level corresponding to the first view of the first physical location, wherein the second immersion level includes one or more second navigation user interface elements that are different from the one or more first navigation user interface elements, the one or more second navigation user interface elements being selectable to change the display of the navigation user interface from the first view of the first physical location to a second view of the second physical location.

2. The method according to claim 1, wherein: The first view at the first physical location includes a first perspective from an analog camera; and The second view from the second physical location is from a second view of the analog camera, which is different from the first view of the analog camera.

3. The method according to any one of claims 1 to 2, further comprising: When the navigation user interface corresponding to the first view of the first physical location is displayed, a first navigation user interface element is displayed in the three-dimensional environment via the display generation component, wherein the first navigation user interface element represents a third view of the first physical location and includes: An indication of the viewpoint's location displayed at the corresponding position within the first navigation user interface element; or An indication of the field of view displayed relative to the corresponding orientation of the first navigation user interface element.

4. The method according to any one of claims 1 to 3, wherein: The navigation user interface having the first immersion level occupies a first portion of the three-dimensional environment; as well as The navigation user interface with the second immersion level occupies a second portion of the three-dimensional environment, which is larger than the first portion of the three-dimensional environment.

5. The method according to any one of claims 1 to 4, wherein displaying the navigation user interface via the display generation component comprises displaying the navigation user interface in a manner that simulates the orientation of physical objects within the physical environment surrounding the computer system.

6. The method according to any one of claims 1 to 5, further comprising: When the navigation user interface having the second immersion level, including the one or more second navigation user interface elements, is displayed, a second input pointing to the corresponding navigation user interface element is detected via the one or more input devices; as well as In response to the detection of the second input: Based on determining that the corresponding navigation user interface element is a first navigation user interface element among one or more second navigation user interface elements, the display of the navigation user interface having a second immersion level corresponding to the first view of the first physical location is changed to a second immersion level corresponding to the second view of the second physical location; as well as Based on the determination that the corresponding navigation user interface element is a second navigation user interface element that is different from the first navigation user interface element among one or more second navigation user interface elements, the display of the navigation user interface having a second immersion level corresponding to the first view of the first physical location is changed to a second immersion level corresponding to a third view of a third physical location, the third view being different from the second view of the second physical location.

7. The method of claim 6, wherein displaying the navigation user interface having a second immersion level corresponding to the second view of the second physical location includes displaying a travel option, which, when selected, causes the computer system to display a previous view of the corresponding physical location, the method further comprising: When the navigation user interface with a second immersion level corresponding to the second view of the second physical location is displayed, a third input pointing to the travel option is detected via the one or more input devices; as well as In response to the detection of the third input, the display of the navigation user interface having a second immersion level corresponding to the second view of the second physical location is changed to the first view corresponding to the first physical location.

8. The method according to any one of claims 1 to 7, further comprising: When the navigation user interface corresponding to the first view of the first physical location is displayed, a first navigation user interface element is displayed in the three-dimensional environment via the display generation component, wherein the first navigation user interface element represents a third view of the first physical location and includes an indication of the second computer system at a corresponding position within the first navigation user interface element that corresponds to the current position of the second computer system in the first physical location. When the first navigation user interface element, which includes the instruction of the second computer system, is displayed within the first navigation user interface element, a second input pointing to the instruction of the second computer system is detected via the one or more input devices; as well as In response to detecting the second input, the display of the navigation user interface having a second immersion level corresponding to the first view of the first physical location is changed to a corresponding view corresponding to the current location from the second computer system.

9. The method according to any one of claims 1 to 8, further comprising: When the navigation user interface corresponding to the first view of the first physical location is displayed, a first navigation user interface element is displayed in the three-dimensional environment via the display generation component, wherein the first navigation user interface element represents a third view of the first physical location and includes additional content corresponding to the first physical location.

10. The method according to any one of claims 1 to 9, further comprising: When the navigation user interface corresponding to the first view of the first physical location is displayed, a first navigation user interface element is displayed in the three-dimensional environment via the display generation component. The first navigation user interface element represents a third view of the first physical location and includes additional content corresponding to the weather conditions at the first physical location.

11. The method according to any one of claims 1 to 10, further comprising: When the navigation user interface corresponding to the first view of the first physical location is displayed, a second input corresponding to a request to change the zoom level of the first view of the first physical location is detected; as well as In response to the detection of the second input: Based on the request to change the zoom level, corresponding to changing the zoom level by a first amount, a first navigation user interface element representing a third view of the first physical location is displayed in the three-dimensional environment via the display generation component; According to the determination that the request to change the zoom level corresponds to changing the zoom level by a second amount different from the first amount, the display of the first navigation user interface element is changed from the third view representing the first physical location to a fourth view representing the first physical location. as well as Based on the request to change the zoom level, which corresponds to changing the zoom level by a third amount, the third amount being different from the first amount and different from the second amount, a second navigation user interface element is displayed in the three-dimensional environment via the display generation component, wherein the second navigation user interface element represents a first portion corresponding to the first physical location.

12. The method according to claim 11, further comprising: When the second navigation user interface element, which includes the first portion corresponding to the first physical location, is displayed, a third input corresponding to a request to change the zoom level of the first portion is detected; as well as In response to the detection of the third input: The request to change the zoom level corresponds to changing the zoom level by a fourth amount, changing the display of the second navigation user interface element from the first portion representing the first physical location to the second portion representing the first physical location; The request to change the zoom level corresponds to changing the zoom level by a fifth amount, which is different from the fourth amount, and the first navigation user interface element representing the third view of the first physical location is displayed in the three-dimensional environment via the display generation component. as well as The request to change the zoom level corresponds to changing the zoom level by a sixth amount, which is different from the fourth amount and different from the fifth amount, and displays the corresponding content corresponding to the first view at the first physical location.

13. The method according to claim 12, further comprising: When the first navigation user interface element representing the first physical location is displayed in the third view, the first navigation user interface element is displayed within a first volume of the three-dimensional environment; as well as When the corresponding content corresponding to the first view of the first physical location is displayed, the corresponding content is displayed in a second volume of the three-dimensional environment, the second volume being larger than the first volume in at least one dimension.

14. The method of any one of claims 11 to 13, wherein detecting the second input includes detecting a change in the posture of the user of the computer system that satisfies one or more criteria.

15. The method according to claim 14, further comprising: When the first navigation user interface element representing the third view of the first physical location is displayed, the computer system detects changes in the user's posture that satisfy one or more criteria; as well as In response to detecting a change in the user's posture on the computer system that satisfies one or more of the criteria: Based on the determination of the change in the posture of the user of the computer system, including a first predetermined pose, the third view of the first physical location is scaled by a first amount; as well as Based on the determination that the change in the user's posture in the computer system includes a second predetermined pose different from the first predetermined pose, the third view of the first physical location is scaled by a second amount different from the first amount.

16. The method according to any one of claims 14 to 15, further comprising: When the first navigation user interface element representing the third view of the first physical location is displayed, the computer system detects changes in the user's posture that satisfy one or more criteria; as well as In response to detecting a change in the user's posture on the computer system that satisfies one or more of the criteria: Based on the determination of the change in the posture of the user of the computer system, including a first predetermined pose, the third view of the first physical location with a first level of detail is displayed; as well as Based on the determination that the user's pose changes in the computer system include a second predetermined pose different from the first predetermined pose, the third view displays the first physical location having a second level of detail different from the first level of detail.

17. The method of any one of claims 11 to 16, wherein detecting the second input corresponding to the zoom request includes detecting a change in the viewpoint of the user of the computer system satisfying one or more criteria; and In response to detecting a change in the user's viewpoint on the computer system that satisfies one or more of the criteria: The change in the user's viewpoint is determined to include movement along a first axis, and the scaling is changed along a second axis corresponding to the first axis; and the change in the user's viewpoint is determined to include movement along the second axis, and the scaling is changed along the first axis corresponding to the second axis.

18. The method of any one of claims 11 to 17, further comprising responding to detecting the second input: Based on determining that the user's gaze on the computer system points to a first portion of the first navigation user interface element when the second input is detected, the zoom level is adjusted along a first axis according to the first portion of the first navigation user interface element; and Based on determining that when the second input is detected, the location of the user's gaze on the computer system is pointed to a second part of the first navigation user interface element that is different from the first part of the first navigation user interface element, the zoom level is changed along a second axis different from the first axis based on the second part of the first navigation user interface element.

19. The method according to any one of claims 11 to 18, the method further comprising: When the first navigation user interface element representing the first physical location is displayed in the third view, a third input corresponding to a request to update the positioning and / or orientation of the first navigation user interface element in the three-dimensional environment is detected. as well as In response to the detection of the third input, the positioning and / or orientation of the first navigation user interface element within the three-dimensional environment is changed based on the third input.

20. The method of claim 19, wherein the first navigation user interface element is displayed in a corresponding volume within the three-dimensional environment, and includes a control element that is interactive to rotate the first navigation user interface element.

21. The method according to claim 20, further comprising: When the first navigation user interface element representing the first physical location is displayed and when the control element is displayed at a first size, a third input pointing to the control element is detected; as well as In response to the detection of the third input, a fourth view in which the corresponding volume of the first navigation user interface element is displayed to represent the first physical location is scaled according to the third input, and the control element is scaled to a second size different from the first size according to the third input.

22. The method according to any one of claims 20 to 21, further comprising: When the first navigation user interface element representing the first physical location is displayed and when the control element is displayed at the first size, a third input pointing to the first navigation user interface element is detected; as well as In response to the detection of the third input, the display of the third view of the first physical location is scaled according to the third input to represent a fourth view of the first physical location, without scaling the view volume associated with the first navigation user interface element.

23. The method according to any one of claims 19 to 22, further comprising: When the first navigation user interface element representing the third view of the first physical location is displayed, a fourth input corresponding to a request to change the zoom level of the third view of the first physical location is detected. as well as In response to the detection of the fourth input: The request to change the zoom level corresponds to changing the zoom level to a first level, changing the display of the first navigation user interface element representing the third view of the first physical location to a fourth view representing the first physical location, and outputting spatial audio having one or more characteristics corresponding to the fourth view of the first physical location via one or more output devices communicating with the computer system. as well as The request to change the zoom level corresponds to changing the zoom level to a second level different from the first level, changing the display of the first navigation user interface element representing the third view of the first physical location to a fifth view of the first physical location different from the fourth view of the first physical location, and outputting the spatial audio having the one or more characteristics corresponding to the fifth view of the first physical location via the one or more output devices.

24. The method according to any one of claims 19 to 23, further comprising: When the first navigation user interface element, representing the third view of the first physical location, is displayed, and when the position of the user's gaze on the first navigation user interface element is detected via the one or more input devices: Based on determining that the user's gaze position points to a first position, spatial audio with one or more characteristics corresponding to the first position is output via one or more output devices communicating with the computer system; as well as Based on the determination that the user's gaze position points to a second position, spatial audio with one or more characteristics corresponding to the second position is output via the one or more output devices.

25. The method according to any one of claims 19 to 24, further comprising: When the first navigation user interface, representing the third view indicating the first physical location, is displayed, changes in the user's viewpoint in the three-dimensional environment are detected; and In response to detecting the change in the user's viewpoint: Based on the determination that the user's viewpoint has changed from a first viewpoint to a second viewpoint different from the first viewpoint, the display of the first navigation user interface element representing the third view of the first physical location is changed to a fourth view representing the first physical location, and spatial audio with one or more characteristics corresponding to the fourth view of the first physical location is presented via one or more output devices communicating with the computer system. as well as Based on the determination that the user's viewpoint changes from the first viewpoint to a third viewpoint, which is different from both the first and second viewpoints, the display of the first navigation user interface element representing the third view of the first physical location is changed to a fifth view of the first physical location, which is different from the fourth view of the first physical location, and spatial audio with one or more characteristics corresponding to the fifth view of the first physical location is presented via the one or more output devices, which is different from spatial audio with the one or more characteristics corresponding to the fourth view of the first physical location.

26. The method according to any one of claims 19 to 25, further comprising: In response to the detection of the third input: Based on the determination that the third input corresponds to a request to update the positioning and / or orientation of the first navigation user interface element in a first manner, spatial audio with one or more characteristics corresponding to the positioning and / or orientation of the first navigation user interface element updated in the first manner is output via one or more output devices communicating with the computer system. as well as Based on the determination that the third input corresponds to a request to update the positioning and / or orientation of the first navigation user interface element in a second manner different from the first manner, spatial audio having one or more characteristics corresponding to the positioning and / or orientation of the first navigation user interface element updated in the second manner is output via one or more output devices communicating with the computer system.

27. The method according to any one of claims 12 to 26, further comprising: When the second navigation user interface element is displayed: Displaying multiple two-dimensional representations corresponding to one or more physical locations represented by the second navigation user interface elements; detecting a third input that selects a corresponding two-dimensional representation from the multiple two-dimensional representations; and In response to the detection of the third input: Based on the determination that the corresponding two-dimensional representation is a first two-dimensional representation corresponding to the third physical location, the second navigation user interface element is rotated such that the first portion of the second navigation user interface element representing the third physical location is visible in the three-dimensional environment; and Based on the determination that the corresponding two-dimensional representation is a second two-dimensional representation corresponding to the fourth physical location, and the second two-dimensional representation is different from the first two-dimensional representation corresponding to the third physical location, the second navigation user interface element is rotated such that the second part of the second navigation user interface element representing the fourth physical location is visible in the three-dimensional environment.

28. The method of claim 27, further comprising: In response to the detection of the third input, the positioning of the plurality of two-dimensional representations in the three-dimensional environment is updated according to the rotation of the second navigation user interface element.

29. The method of any one of claims 12 to 28, wherein displaying the second navigation user interface element includes displaying a visual indication of weather conditions at a corresponding physical location represented by the second navigation user interface element at a corresponding location in the three-dimensional environment relative to the second navigation user interface element, and the method further comprises: When the second navigation user interface element, which includes the visual indication of the weather conditions, is displayed, the computer system detects changes in the user's viewpoint in the three-dimensional environment; and In response to detecting a change in the user's viewpoint, the visual indication of the weather conditions is updated relative to the corresponding position of the second navigation user interface element in the three-dimensional environment in a manner that simulates a parallax effect.

30. The method of any one of claims 1 to 29, wherein the navigation user interface displaying the second immersion level corresponding to the second view of the second physical location includes outputting spatial audio associated with the second view of the second physical location via one or more output devices in communication with the computer system.

31. The method according to any one of claims 1 to 30, further comprising: When the navigation user interface with a second immersion level corresponding to the second view of the second physical location is displayed, a second input corresponding to a request to add virtual objects to the corresponding content of the navigation user interface is detected via the one or more input devices; as well as In response to detecting the second input, the display of the navigation user interface having a second immersion level corresponding to the second view of the second physical location is changed to include the virtual object.

32. The method according to any one of claims 1 to 31, further comprising: When a second three-dimensional environment including the transparent video of the second physical location is displayed via the display generation component, the virtual object is displayed via the display generation component.

33. The method according to any one of claims 1 to 32, wherein the three-dimensional environment is shared with the second computer system, and displaying the navigation user interface includes displaying a representation of the user of the second computer system at a location in the three-dimensional environment corresponding to the viewpoint of the user of the second computer system via the display generation component.

34. A computer system in communication with a display generation component and one or more input devices, the computer system comprising: One or more processors; Memory; as well as One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: A navigation user interface is displayed in a three-dimensional environment via a display generation component, wherein the navigation user interface includes one or more first travel user interface elements displayed at a first immersion level corresponding to a first view of a first physical location; When the navigation user interface including the one or more first navigating user interface elements is displayed via the display generation component, a first input corresponding to a request to change the immersion level is detected via one or more input devices; as well as In response to the detection of the first input: The display of the navigation user interface having a first immersion level corresponding to the first view of the first physical location is changed to a second immersion level corresponding to the first view of the first physical location, wherein the second immersion level includes one or more second navigation user interface elements that are different from the one or more first navigation user interface elements, the one or more second navigation user interface elements being selectable to change the display of the navigation user interface from the first view of the first physical location to a second view of the second physical location.

35. A non-transitory computer-readable storage medium storing one or more programs, said one or more programs comprising instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform a method comprising: A navigation user interface is displayed in a three-dimensional environment via a display generation component, wherein the navigation user interface includes one or more first travel user interface elements displayed at a first immersion level corresponding to a first view of a first physical location; When the navigation user interface, including the one or more first navigating user interface elements, is displayed via the display generation component, a first input corresponding to a request to change the immersion level is detected via one or more input devices; and In response to the detection of the first input: The display of the navigation user interface having a first immersion level corresponding to the first view of the first physical location is changed to a second immersion level corresponding to the first view of the first physical location, wherein the second immersion level includes one or more second navigation user interface elements that are different from the one or more first navigation user interface elements, the one or more second navigation user interface elements being selectable to change the display of the navigation user interface from the first view of the first physical location to a second view of the second physical location.

36. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; Memory; Components for the following operations: displaying a navigation user interface in a three-dimensional environment via a display generation component, wherein the navigation user interface includes one or more first travel user interface elements displayed at a first immersion level corresponding to a first view of a first physical location; Components for the following operations: when the navigation user interface including the one or more first navigating user interface elements is displayed via the display generation component, detecting a first input corresponding to a request to change the immersion level via one or more input devices; as well as Components used for the following operation: in response to detecting the first input: Components for the following operation: changing the display of the navigation user interface having a first immersion level corresponding to the first view of the first physical location to a second immersion level corresponding to the first view of the first physical location, wherein the second immersion level includes one or more second moving user interface elements different from the one or more first moving user interface elements, the one or more second moving user interface elements being selectable to change the display of the navigation user interface from the first view of the first physical location to a second view of the second physical location.

37. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; Memory; as well as One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any one of the methods according to claims 1 to 33.

38. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform any one of the methods according to claims 1 to 33.

39. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; Memory; as well as Components for performing any one of the methods according to claims 1 to 33.