Devices, methods, and graphical user interfaces for interacting with three-dimensional environments

The computer system improves interaction with virtual and augmented reality environments by dynamically adjusting user interface objects based on user attention and movement, reducing complexity and conserving energy.

JP2025124679APending Publication Date: 2025-08-26APPLE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025080557
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-19
Filing Date
2025-05-13
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

Existing methods and interfaces for interacting with virtual and augmented reality environments are cumbersome, inefficient, and complex, leading to a significant cognitive burden on users and inefficient energy usage, particularly in battery-operated devices.

Method used

A computer system with improved methods and interfaces that include displaying user interface objects with modified appearances based on user attention and movement, providing visual, audio, and tactile feedback to enhance interaction efficiency and reduce the number of required inputs.

Benefits of technology

The system enhances user interaction efficiency, reduces errors, and conserves energy by minimizing unnecessary inputs and optimizing user interface responses to user attention and movement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025124679000001_ABST
    Figure 2025124679000001_ABST
Patent Text Reader

Abstract

To solve the problem that interfaces for interacting with environments including virtual elements are inefficient and limited.SOLUTION: A system displays a first user interface object in a first view of a three-dimensional environment at a first position in the three-dimensional environment and with a first spatial arrangement relative to an individual portion of a user. The system detects movement of a viewpoint of the user from a first location to a second location in a physical environment. In accordance with determination that the movement does not satisfy a threshold amount of movement, the system maintains display of the first user interface object at the first position. In accordance with determination that the movement satisfies the threshold amount of movement, the system ceases to display the first user interface object at the first position in the three-dimensional environment and displays the first user interface object at a second position having the first spatial arrangement relative to the individual portion of the user.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (Related Applications) This application is a continuation of U.S. Patent Application No. 17 / 948,096, filed September 19, 2022, which claims priority to U.S. Provisional Patent Application No. 63 / 247,241, filed September 22, 2021, each of which is incorporated by reference in its entirety.

[0002] The present disclosure generally relates to computer systems having a display generation component, including but not limited to electronic devices that provide virtual reality and mixed reality experiences via the display generation component, and one or more input devices that provide computer-generated augmented reality (XR) experiences.

[0003] background The development of computer systems for virtual reality, augmented reality, and augmented reality has progressed significantly in recent years. Exemplary augmented reality and augmented reality environments include at least some virtual elements that replace or enhance the physical world. Input devices such as cameras, controllers, joysticks, touch-sensitive surfaces, and touchscreen displays for computer systems and other electronic computing devices are used to interact with the virtual / augmented / augmented reality environment. Exemplary virtual elements include virtual objects, including digital images, video, text, icons, and control elements such as buttons and other graphics.

[0004] However, methods and interfaces for interacting with environments that include at least some virtual elements (e.g., applications, augmented reality environments, mixed reality environments, virtual reality environments, and augmented reality environments) are cumbersome, inefficient, and limited. For example, systems that provide insufficient feedback for performing actions associated with virtual objects, systems that require a series of inputs to achieve a desired result in a virtual / augmented / augmented reality environment, and systems in which manipulating virtual objects is complex and error-prone create a significant cognitive burden for users and detract from the experience of the virtual / augmented / augmented reality environment. In addition, these methods are unnecessarily time-consuming, thereby wasting energy. This latter consideration is particularly important in battery-operated devices. Summary of the Invention

[0005] Therefore, there is a need for a computer system with improved methods and interfaces for providing users with computer-generated experiences that make interacting with the computer system more efficient and intuitive for the user. The above-mentioned deficiencies and other problems associated with user interfaces for computer systems having a display generation component and one or more input devices are reduced or eliminated by the disclosed systems, methods, and user interfaces. Such systems, methods, and interfaces optionally complement or replace conventional systems, methods, and user interfaces for providing users with augmented reality experiences. Such methods and interfaces reduce the number, extent, and / or type of inputs from the user by helping the user understand the connection between the input provided and the device response to that input, thereby creating a more efficient human-machine interface.

[0006] According to some embodiments, a method is performed on a computer system in communication with a first display generation component and one or more input devices. The method includes displaying a first user interface object in a first view of a three-dimensional environment via the first display generation component. The method further includes detecting, via the one or more input devices, whether a user satisfies an attention criterion for the first user interface object while displaying the first user interface object. The method further includes displaying the first user interface object with a modified appearance in response to detecting that the user does not satisfy the attention criterion for the first user interface object, wherein displaying the first user interface object with the modified appearance includes de-emphasizing the first user interface object relative to one or more other objects in the three-dimensional environment. The method further includes detecting, via the one or more input devices, a first movement of a user's viewpoint relative to the physical environment while displaying the first user interface object with the modified appearance. The method further includes detecting, after detecting the first movement of the user's viewpoint relative to the physical environment, that the user satisfies the attention criterion for the first user interface object. The method further includes, in response to detecting that the user meets an attention criterion, displaying the first user interface object in a second view of the three-dimensional environment that is different from the first view of the three-dimensional environment, where displaying the first user interface object in the second view of the three-dimensional environment includes displaying the first user interface object in an appearance that highlights the first user interface object relative to one or more other objects in the three-dimensional environment more than when the first user interface object is displayed with the modified appearance.

[0007] In some embodiments, a method is performed on a computer system in communication with a first display generating component and one or more input devices. The method includes displaying, via the first display generating component, a first user interface object in a first view of a three-dimensional environment at a first position within the three-dimensional environment and at a first spatial arrangement relative to a distinct portion of a user. The method further includes detecting, while displaying the first user interface object, movement of a user's viewpoint from a first location to a second location within the physical environment via the one or more input devices. In response to detecting movement of the user's viewpoint from the first location to the second location, the method further includes maintaining the display of the first user interface object at the first position within the three-dimensional environment in accordance with a determination that the movement of the user's viewpoint from the first location to the second location does not satisfy a threshold amount of movement. The method further includes, in response to detecting movement of the user's viewpoint from the first location to the second location, ceasing to display the first user interface object at a first position within the three-dimensional environment in accordance with a determination that the movement of the user's viewpoint from the first location to the second location satisfies a threshold amount of movement, and displaying the first user interface object at a second position within the three-dimensional environment, the second position within the three-dimensional environment having a first spatial arrangement with respect to individual portions of the user.

[0008] According to some embodiments, a computer system includes or is in communication with a display generation component (e.g., a display, projector, or head-mounted display), one or more input devices (e.g., one or more cameras, a touch-sensitive surface, optionally one or more sensors for detecting intensity of contact with the touch-sensitive surface), optionally one or more audio output components, optionally one or more tactile output generators, one or more processors, and memory storing one or more programs, the one or more programs configured to be executed by the one or more processors, the one or more programs including instructions to perform or cause to be performed any of the operations of the methods described herein. According to some embodiments, a non-transitory computer-readable storage medium has instructions stored therein that, when executed by a computer system having a display generation component, one or more input devices (e.g., one or more cameras, a touch-sensitive surface, optionally one or more sensors for detecting intensity of contact with the touch-sensitive surface), optionally one or more audio output components, and optionally one or more tactile output generators, cause the device to perform or cause to be performed any of the operations of the methods described herein. According to some embodiments, a graphical user interface of a computer system having a display generation component, one or more input devices (e.g., one or more cameras, a touch-sensitive surface, optionally one or more sensors that detect the intensity of contact with the touch-sensitive surface), optionally one or more audio output components, optionally one or more tactile output generators, memory, and one or more processors executing one or more programs stored in the memory, includes one or more of the elements displayed in any of the ways described herein, and these elements are updated in response to input as described in any of the ways described herein.According to some embodiments, a computer system includes a display generation component, one or more input devices (e.g., one or more cameras, a touch-sensitive surface, optionally one or more sensors that detect the intensity of contact with the touch-sensitive surface), optionally one or more audio output components, optionally one or more tactile output generators, and means for performing or causing to be performed any of the operations of the methods described herein. According to some embodiments, an information processing apparatus for use in a computer system having a display generation component, one or more input devices (e.g., one or more cameras, a touch-sensitive surface, optionally one or more sensors that detect the intensity of contact with the touch-sensitive surface), optionally one or more audio output components, and optionally one or more tactile output generators includes means for performing or causing to be performed any of the operations of the methods described herein.

[0009] Thus, computer systems having display generation components are provided with improved methods and interfaces for interacting with three-dimensional environments and facilitating a user's use of the computer systems when interacting with the three-dimensional environments, thereby increasing the effectiveness, efficiency, and safety and satisfaction of such computer systems. Such methods and interfaces may complement or replace conventional methods for interacting with three-dimensional environments and facilitating a user's use of the computer systems when interacting with the three-dimensional environments.

[0010] It should be noted that the various embodiments described above can be combined with any other embodiment described herein. The features and advantages described herein are not exhaustive, and many additional features and advantages will become apparent to those skilled in the art, particularly in light of the drawings, specification, and claims. Furthermore, it should be noted that the language used in this specification has been selected solely for the purposes of readability and explanation, and not to define or limit the subject matter of the present invention. [Brief explanation of the drawings]

[0011] For a better understanding of the various described embodiments, reference should be made to the following Detailed Description of the Invention in conjunction with the following drawings, in which like reference numerals refer to corresponding parts throughout:

[0012] [Figure 1] FIG. 1 is a block diagram illustrating an operating environment for a computer system for providing an extended reality (XR) experience, according to some embodiments.

[0013] [Figure 2] FIG. 1 is a block diagram illustrating a controller of a computer system configured to manage and coordinate a user's XR experience, according to some embodiments.

[0014] [Figure 3] FIG. 1 is a block diagram illustrating display generation components of a computer system configured to provide a visual component of an XR experience to a user, according to some embodiments.

[0015] [Figure 4] FIG. 1 is a block diagram illustrating a hand tracking unit of a computer system configured to capture a user's gesture input, according to some embodiments.

[0016] [Figure 5] FIG. 1 is a block diagram illustrating an eye-tracking unit of a computer system configured to capture a user's gaze input, according to some embodiments.

[0017] [Figure 6] 1 is a flowchart illustrating a glint-assisted gaze tracking pipeline, according to some embodiments.

[0018] [Figure 7A]FIG. 1 is a block diagram illustrating displaying user interface objects at distinct positions in a three-dimensional environment, according to some embodiments. [Figure 7B] FIG. 1 is a block diagram illustrating displaying user interface objects at distinct positions in a three-dimensional environment, according to some embodiments. [Figure 7C] FIG. 1 is a block diagram illustrating displaying user interface objects at distinct positions in a three-dimensional environment, according to some embodiments. [Figure 7D] FIG. 1 is a block diagram illustrating displaying user interface objects at distinct positions in a three-dimensional environment, according to some embodiments. [Figure 7E] FIG. 1 is a block diagram illustrating displaying user interface objects at distinct positions in a three-dimensional environment, according to some embodiments. [Figure 7F] FIG. 1 is a block diagram illustrating displaying user interface objects at distinct positions in a three-dimensional environment, according to some embodiments. [Figure 7G] FIG. 1 is a block diagram illustrating displaying user interface objects at distinct positions in a three-dimensional environment, according to some embodiments. [Figure 7H] FIG. 1 is a block diagram illustrating displaying user interface objects at distinct positions in a three-dimensional environment, according to some embodiments. [Figure 7I] FIG. 1 is a block diagram illustrating displaying user interface objects at distinct positions in a three-dimensional environment, according to some embodiments. [Figure 7J] FIG. 1 is a block diagram illustrating displaying user interface objects at distinct positions in a three-dimensional environment, according to some embodiments.

[0019] [Figure 8]1 is a flowchart of a method for visually de-emphasizing a user interface element in a three-dimensional environment while a user is not focusing on the user interface element, according to some embodiments.

[0020] [Figure 9] 1 is a flowchart of a method for updating the display of user interface elements in a three-dimensional environment to follow a user as the user changes their current view of the three-dimensional environment, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0021] The present disclosure relates to a user interface that provides a computer-generated augmented reality (XR) experience to a user, according to some embodiments.

[0022] The systems, methods, and GUIs described herein improve user interface interaction with virtual / augmented reality environments in several ways.

[0023] In some embodiments, the computer system displays the user interface element visually de-emphasized while the user is not attending to the user interface element, the user interface element remains de-emphasized as the user moves about in the physical environment, and pursuant to a determination that the user is attending to the user interface element, the user interface element is no longer visually de-emphasized and is displayed for the user in a position within the three-dimensional environment based on the user's current view of the three-dimensional environment.

[0024] In some embodiments, a computer system is provided that displays user interface elements within a three-dimensional environment, where the display of the user interface elements is updated to follow the user as the user changes their current view of the three-dimensional environment (e.g., by moving around the physical environment). The user interface elements initially do not move as the user's view changes until the user's view changes by more than a threshold amount. After the user's view changes by more than the threshold amount, the user interface elements follow the user (e.g., delay following the user, at a slower rate of movement than the user's movement).

[0025] Figures 1-6 illustrate an exemplary computer system for providing an XR experience to a user. The user interfaces of Figures 7A-7J are used to illustrate the processes of Figures 8-9, respectively.

[0026] The processes described below improve the usability of a device and make the user's interface with the device more efficient (e.g., by assisting the user in providing appropriate inputs and reducing user errors when operating / interacting with the device) through various techniques, including providing improved visual, audio, and / or tactile feedback to the user, reducing the number of inputs required to perform an operation, providing additional control options without cluttering the user interface with additional displayed controls, performing an operation when a set of conditions is met without requiring further user input, and / or additional techniques. These techniques also reduce power usage and improve the device's battery life by allowing the user to use the device more quickly and efficiently.

[0027] 1, an XR experience is provided to a user via an operating environment 100 that includes a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted device (HMD), a display, a projector, and / or a touchscreen), one or more input devices 125 (e.g., an eye-tracking device 130, a hand-tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a tactile output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a tactile sensor, an orientation sensor, a proximity sensor, a temperature sensor, a location sensor, a motion sensor, and / or a speed sensor), and optionally one or more peripheral devices 195 (e.g., a consumer electronics device and / or a wearable device). In some embodiments, one or more of the input device 125, the output device 155, the sensor 190, and the peripheral device 195 are integrated with the display generation component 120 (e.g., within a head-mounted or handheld device).

[0028] When describing an XR experience, various terms are used to individually refer to several related, but distinct, environments that a user can sense and / or interact with (e.g., using inputs detected by the computer system 101 generating the XR experience that cause the computer system generating the XR experience to generate audio, visual, and / or haptic feedback corresponding to various inputs provided to the computer system 101). The following is a subset of these terms:

[0029] Physical Environment: The physical environment refers to the physical world that people can sense and / or interact with without the aid of electronic systems. A physical environment, such as a physical park, includes physical objects such as physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment through their senses, such as sight, touch, hearing, taste, and smell.

[0030] Augmented reality: In contrast, an extended reality (XR) environment refers to a wholly or partially mimicked environment that people sense and / or interact with through electronic systems. In XR, a subset of a person's physical movements or representations thereof are tracked, and one or more properties of one or more simulated virtual objects within the XR environment are adjusted accordingly to behave according to at least one law of physics. For example, an XR system may detect a person's head rotation and adjust the graphical content and sound field presented to the person accordingly, in a manner similar to how such views and sounds change in a physical environment. In some circumstances (e.g., for accessibility reasons), adjustments to the property(ies) of virtual object(s) in the XR environment may be made in response to representations of physical movements (e.g., voice commands). A person may sense and / or interact with an XR object using any one of these senses, including sight, hearing, touch, taste, and smell. For example, a person may sense and / or interact with audio objects that create a 3D or spatially expansive audio environment that provides the perception of a point sound source in 3D space. In another example, audio objects may enable audio transparency that selectively incorporates ambient sounds from the physical environment, with or without computer-generated audio. In some XR environments, a person may sense and / or interact with only audio objects.

[0031] Examples of XR include virtual reality and mixed reality.

[0032] Virtual Reality: A virtual reality (VR) environment refers to an emulated environment designed to be based entirely on computer-generated sensory input for one or more senses. A VR environment includes multiple virtual objects that a person can sense and / or interact with. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can sense and / or interact with virtual objects in the VR environment through a simulation of the person's presence in the computer-generated environment and / or through a simulation of a subset of the person's physical movement within the computer-generated environment.

[0033] Mixed Reality: A mixed reality (MR) environment refers to a mimetic environment designed to incorporate sensory input from or representations of a physical environment in addition to including computer-generated sensory input (e.g., virtual objects), as opposed to a VR environment designed to be based entirely on computer-generated sensory input. On the virtual continuum, a mixed reality environment is anywhere between, but not including, a fully physical environment at one end and a virtual reality environment at the other. In some MR environments, computer-generated sensory input may respond to changes in sensory input from the physical environment. Some electronic systems for presenting MR environments may also track location and / or orientation relative to the physical environment to allow virtual objects to interact with real objects (i.e., physical items from the physical environment or representations thereof). For example, the system may account for movement so that a virtual tree appears stationary relative to the physical ground.

[0034] Examples of mixed reality include augmented reality and augmented virtuality.

[0035] Augmented reality: An augmented reality (AR) environment refers to a simulated environment in which one or more virtual objects are superimposed on a physical environment or a representation thereof. For example, an electronic system for presenting an AR environment may have a transparent or translucent display through which a person can directly view the physical environment. The system may be configured to present virtual objects on the transparent or translucent display, whereby a person using the system perceives the virtual objects superimposed on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture images or videos of the physical environment, which are representations of the physical environment. The system composites the images or videos with virtual objects and presents the composite on the opaque display. The person uses the system to indirectly view the physical environment through the images or videos of the physical environment and perceive the virtual objects superimposed on the physical environment. As used herein, video of a physical environment shown on an opaque display is referred to as "pass-through video," meaning that the system captures images of the physical environment using one or more image sensors and uses those images in presenting the AR environment on the opaque display. Alternatively, the system may include a projection system that projects virtual objects, e.g., as holograms, into a physical environment or onto a physical surface, such that a person using the system perceives the virtual objects superimposed on the physical environment. Augmented reality environments also refer to mimic environments in which a representation of a physical environment is transformed by computer-generated sensory information. For example, when providing pass-through video, a system may distort one or more sensor images to impose a selected perspective (e.g., viewpoint) different from the perspective captured by the imaging sensor. As another example, a representation of a physical environment may be distorted by graphically modifying (e.g., enlarging) portions thereof, thereby rendering the modified portions a non-photorealistic, altered version of the originally captured image. As a further example, a representation of a physical environment may be distorted by graphically removing or obscuring portions thereof.

[0036] Augmented Virtual: An augmented virtual (AV) environment refers to a mimicking environment in which a virtual or computer-generated environment incorporates one or more sensory inputs from a physical environment. The sensory inputs may be representations of one or more characteristics of the physical environment. For example, an AV park may have virtual trees and virtual buildings, while people with faces are realistically recreated from images taken of physical people. As another example, virtual objects may adopt the shape or color of physical items imaged by one or more imaging sensors. As a further example, virtual objects may adopt shadows that match the position of the sun in the physical environment.

[0037] Hardware: There are many different types of electronic systems that enable a person to sense and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, heads-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be placed over a person's eyes (e.g., similar to contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system may have one or more speaker(s) and an integrated opaque display. Alternatively, a head-mounted system may be configured to accept an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing images or video of the physical environment and / or one or more microphones for capturing audio of the physical environment. A head-mounted system may have a transparent or translucent display rather than an opaque display. The transparent or translucent display may have a medium through which light representing an image is directed to a person's eyes. The display may utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser-scanned light source, or any combination of these technologies. The medium may be a light guide, a holographic medium, an optical combiner, an optical reflector, or any combination thereof. In one embodiment, the transparent or translucent display may be configured to be selectively opaque. The projection-based system may employ retinal projection technology to project a graphical image onto a person's retina. The projection system may also be configured to project virtual objects into the physical environment, for example, as holograms or as physical surfaces.In some embodiments, controller 110 is configured to manage and coordinate the user's XR experience. In some embodiments, controller 110 includes a suitable combination of software, firmware, and / or hardware. Controller 110 is described in more detail below with reference to FIG. 2. In some embodiments, controller 110 is a computing device that is local or remote to scene 105 (e.g., physical setting / environment). For example, controller 110 is a local server located within scene 105. In another example, controller 110 is a remote server (e.g., a cloud server, a central server, and / or another server) located outside of scene 105. In some embodiments, controller 110 is communicatively coupled to display generation component 120 (e.g., HMD, display, projector, and / or touchscreen) via one or more wired or wireless communication channels 144 (e.g., BLUETOOTH, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). In another example, the controller 110 is contained within the housing (e.g., physical housing) of, or shares the same physical housing or support structure as, one or more of the display generation component 120 (e.g., an HMD or a portable electronic device including a display and one or more processors), one or more of the input devices 125, one or more of the output devices 155, one or more of the sensors 190, and / or one or more of the peripheral devices 195.

[0038] In some embodiments, display generation component 120 is configured to provide an XR experience (e.g., at least a visual component of an XR experience) to a user. In some embodiments, display generation component 120 includes a suitable combination of software, firmware, and / or hardware. Display generation component 120 is described in more detail below with reference to FIG. 3. In some embodiments, functionality of controller 110 is provided by and / or combined with display generation component 120.

[0039] According to some embodiments, the display generation component 120 provides an XR experience to the user while the user is virtually and / or physically present in the scene 105.

[0040] In some embodiments, the display generation component is worn on a part of the user's body (e.g., on their head and / or their hand). Thus, display generation component 120 includes one or more XR displays provided for displaying XR content. For example, in various embodiments, display generation component 120 surrounds the user's field of view. In some embodiments, display generation component 120 is a handheld device (such as a smartphone or tablet) configured to present XR content, where the user holds the device with a display pointed toward the user's field of view and a camera pointed toward scene 105. In some embodiments, the handheld device is optionally located within a housing worn on the user's head. In some embodiments, the handheld device is optionally located on a support (e.g., a tripod) in front of the user. In some embodiments, display generation component 120 is an XR chamber, housing, or room configured to present XR content without the user wearing or holding display generation component 120. Many user interfaces described with reference to one type of hardware for displaying XR content (e.g., a handheld device or a device on a tripod) may be implemented on another type of hardware for displaying XR content (e.g., an HMD or other wearable computing device). For example, a user interface illustrating interactions with XR content that are triggered based on interactions occurring in the space in front of a handheld or tripod-mounted device may be implemented similarly to an HMD in which the interactions occur in the space in front of the HMD and the XR content responses are displayed via the HMD. Similarly, a user interface illustrating interactions with XR content that are triggered based on movement of a handheld or tripod-mounted device relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eye(s), head, or hands)) may be implemented similarly to an HMD in which the movement is caused by movement of the HMD relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eye(s), head, or hands)).

[0041] While relevant features of operating environment 100 are shown in FIG. 1, those skilled in the art will understand from this disclosure that various other features have not been shown for the sake of brevity so as not to obscure more pertinent aspects of the exemplary embodiments disclosed herein.

[0042] 2 is a block diagram of an example controller 110, according to some embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that various other features are not shown for the sake of brevity so as not to obscure more pertinent aspects of the embodiments disclosed herein. Thus, by way of non-limiting example, in some embodiments, the controller 110 includes one or more processing units 202 (e.g., a microprocessor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a graphics processing unit (GPU), a central processing unit (CPU), a processing core, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), BLUETOOTH, ZIGBEE, or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these and various other components.

[0043] In some embodiments, one or more communication buses 204 include circuitry that interconnects and controls communications between system components. In some embodiments, one or more I / O devices 206 include at least one of a keyboard, a mouse, a touchpad, a joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.

[0044] Memory 220 includes high-speed random-access memory, such as dynamic random-access memory (DRAM), static random-access memory (SRAM), double-data-rate random-access memory (DDRRAM), or other random-access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 220 optionally includes one or more storage devices located remotely from the one or more processing units 202. Memory 220 includes a non-transitory computer-readable storage medium. In some embodiments, memory 220, or its non-transitory computer-readable storage medium, stores the following programs, modules, and data structures, or a subset thereof, including an optional operating system 230 and an XR experience module 240:

[0045] Operating system 230 includes instructions for handling various basic system services and for performing hardware-dependent tasks. In some embodiments, XR experience module 240 is configured to manage and coordinate one or more XR experiences for one or more users (e.g., a single XR experience for one or more users, or multiple XR experiences for respective groups of one or more users). To that end, in various embodiments, XR experience module 240 includes a data acquisition unit 242, a tracking unit 244, an adjustment unit 246, and a data transmission unit 248.

[0046] 1 , and optionally one or more of input devices 125, output devices 155, sensors 190, and / or peripheral devices 195. To that end, in various embodiments, data acquisition unit 242 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0047] In some embodiments, tracking unit 244 is configured to map scene 105 and track the position / location of at least display generation component 120 relative to scene 105 of FIG. 1 , and optionally relative to one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. To that end, in various embodiments, tracking unit 244 includes instructions and / or logic therefor, as well as heuristics and metadata therefor. In some embodiments, tracking unit 244 includes hand tracking unit 245 and / or eye tracking unit 243. In some embodiments, hand tracking unit 245 is configured to track the position / location of one or more parts of a user's hand and / or the movement of one or more parts of a user's hand relative to scene 105 of FIG. 1 , relative to display generation component 120, and / or relative to a coordinate system defined relative to the user's hand. Hand tracking unit 245 is described in more detail below with respect to FIG. 4. In some embodiments, eye tracking unit 243 is configured to track the position and movement of the user's gaze (or, more broadly, the user's eyes, face, or head) relative to scene 105 (e.g., relative to the physical environment and / or the user (e.g., the user's hands)), or relative to XR content displayed via display generation component 120. Eye tracking unit 243 is described in more detail below with respect to FIG. 5.

[0048] In some embodiments, coordination unit 246 is configured to manage and coordinate the XR experience presented to the user by display generation component 120 and, optionally, by one or more of output devices 155 and / or peripheral devices 195. To that end, in various embodiments, coordination unit 246 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0049] In some embodiments, data transmission unit 248 is configured to transmit data (e.g., presentation data and / or location data) to at least display generation component 120, and optionally to one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. To that end, in various embodiments, data transmission unit 248 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0050] Although the data acquisition unit 242, the tracking unit 244 (e.g., including the eye tracking unit 243 and the hand tracking unit 245), the adjustment unit 246, and the data transmission unit 248 are shown as being present on a single device (e.g., the controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 242, the tracking unit 244 (e.g., including the eye tracking unit 243 and the hand tracking unit 245), the adjustment unit 246, and the data transmission unit 248 may be located within separate computing devices.

[0051] Furthermore, Figure 2 is intended more to illustrate the functionality of various features that may be present in particular embodiments, as opposed to a structural overview of the embodiments described herein. As will be recognized by those skilled in the art, items shown separately may be combined and some items may be separated. For example, some functional modules shown separately in Figure 2 may be implemented within a single module, and various functions of a single functional block may be performed by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of specific functions and how functions are allocated among them, will vary depending on implementation and, in some embodiments, will depend in part on the particular combination of hardware, software, and / or firmware selected for a particular implementation.

[0052] 3 is a block diagram of an example of a display generation component 120, according to some embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that, for the sake of brevity, various other features are not shown so as not to obscure more pertinent aspects of the embodiments disclosed herein. To that end, by way of non-limiting example, in some embodiments, the HMD 120 includes one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, infrared, BLUETOOTH, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more XR displays 312, one or more optional inward-facing and / or outward-facing image sensors 314, memory 320, and one or more communication buses 304 for interconnecting these and various other components.

[0053] In some embodiments, the one or more communication buses 304 include circuitry that interconnects and controls communications between system components. In some embodiments, the one or more I / O devices and sensors 306 include at least one of an inertial measurement unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, and / or a blood glucose sensor), one or more microphones, one or more speakers, a haptic engine, one or more depth sensors (e.g., structured light, time of flight, etc.), etc.

[0054] In some embodiments, the one or more XR displays 312 are configured to provide an XR experience to a user. In some embodiments, the one or more XR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conduction electron-emissive element display (SED), field-emission display (FED), quantum dot light-emitting diode (QD-LED), MEMS, and / or similar display types. In some embodiments, the one or more XR displays 312 correspond to a waveguide display, such as a diffractive, reflective, polarized, holographic, etc. For example, the HMD 120 includes a single XR display. In another example, the HMD 120 includes an XR display for each eye of the user. In some embodiments, the one or more XR displays 312 are capable of presenting mixed reality (MR) or virtual reality (VR) content. In some embodiments, the one or more XR displays 312 are capable of presenting mixed reality (MR) or virtual reality (VR) content.

[0055] In some embodiments, the one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face, including the user's eyes (and may be referred to as eye-tracking cameras). In some embodiments, the one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hand(s) and optionally the user's arm(s) (and may be referred to as hand-tracking cameras). In some embodiments, the one or more image sensors 314 are configured to face forward to acquire image data corresponding to a scene viewed by the user when the HMD 120 is not present (and may be referred to as scene cameras). The one or more optional image sensors 314 may include one or more RGB cameras (e.g., with a complementary metal-oxide semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), one or more infrared (IR) cameras, one or more event-based cameras, and / or the like.

[0056] Memory 320 includes high-speed random-access memory, such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located remotely from the one or more processing units 302. Memory 320 includes a non-transitory computer-readable storage medium. In some embodiments, memory 320, or its non-transitory computer-readable storage medium, stores the following programs, modules, and data structures, or a subset thereof, including an optional operating system 330 and an XR presentation module 340:

[0057] The operating system 330 includes instructions for handling various basic system services and for performing hardware-dependent tasks. In some embodiments, the XR presentation module 340 is configured to present XR content to a user via one or more XR displays 312. To that end, in various embodiments, the XR presentation module 340 includes a data acquisition unit 342, an XR presentation unit 344, an XR map generation unit 346, and a data transmission unit 348.

[0058] In some embodiments, the data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, and / or location data) from at least the controller 110 of Figure 1. To that end, in various embodiments, the data acquisition unit 342 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0059] In some embodiments, the XR presentation unit 344 is configured to present XR content via one or more XR displays 312. To that end, in various embodiments, the XR presentation unit 344 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0060] In some embodiments, the XR map generation unit 346 is configured to generate an XR map (e.g., a 3D map of a mixed reality scene or a map of a physical environment in which computer-generated objects can be placed to generate an augmented reality) based on the media content data. To that end, in various embodiments, the XR map generation unit 346 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0061] In some embodiments, data transmission unit 348 is configured to transmit data (e.g., presentation data and / or location data) to at least controller 110, and optionally to one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. To that end, in various embodiments, data transmission unit 348 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0062] Although the data acquisition unit 342, the XR presentation unit 344, the XR map generation unit 346, and the data transmission unit 348 are shown as residing on a single device (e.g., the display generation component 120 of FIG. 1), it should be understood that in other embodiments, any combination of the data acquisition unit 342, the XR presentation unit 344, the XR map generation unit 346, and the data transmission unit 348 may be located in separate computing devices.

[0063] Furthermore, Figure 3 is intended more to illustrate the functionality of various features that may be present in particular implementations, as opposed to a structural overview of the embodiments described herein. As will be recognized by those skilled in the art, items shown separately can be combined and some items can be separated. For example, some functional modules shown separately in Figure 3 can be implemented within a single module, and various functions of a single functional block can be performed by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of specific functions and how functions are allocated among them, will vary from implementation to implementation and, in some embodiments, will depend in part on the particular combination of hardware, software, and / or firmware selected for a particular implementation.

[0064] 4 is a schematic diagram of an example embodiment of a hand tracking device 140. In some embodiments, hand tracking device 140 (FIG. 1) is controlled by hand tracking unit 245 (FIG. 2) to track the position / location of one or more parts of a user's hand and / or the movement of one or more parts of a user's hand relative to scene 105 of FIG. 1 (e.g., relative to a portion of the physical environment surrounding the user, relative to display generation component 120, or relative to a portion of the user (e.g., the user's face, eyes, or head), and / or relative to a coordinate system defined relative to the user's hand). In some embodiments, hand tracking device 140 is part of display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, hand tracking device 140 is separate from display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).

[0065] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras) that captures three-dimensional scene information including at least the hand 406 of a human user. The image sensor 404 captures hand images with sufficient resolution to allow for differentiation of the fingers and their respective positions. The image sensor 404 typically captures images of other parts of the user's body, or all of the body, and can have either zoom capabilities or a dedicated sensor with high magnification to capture hand images at a desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with or functions as an image sensor that captures the physical environment of the scene 105. In some embodiments, the image sensor 404 is positioned relative to the user or the user's environment such that the field of view of the image sensor, or a portion thereof, is used to define an interaction space in which hand movements captured by the image sensor are processed as inputs to the controller 110.

[0066] In some embodiments, image sensor 404 outputs a sequence of frames containing 3D map data (and possibly color image data) to controller 110, which extracts high-level information from the map data. This high-level information is provided, typically via an application program interface (API), to an application running on the controller, which drives display generation component 120 accordingly. For example, a user can interact with software running on controller 110 by moving their hand 408 and changing the posture of their hand.

[0067] In some embodiments, the image sensor 404 projects a spot pattern onto a scene including the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral shift of the pattern's spots. This approach is advantageous in that it does not require the user to hold or wear any type of beacon, sensor, or other marker. This provides depth coordinates of points in the scene relative to a predetermined reference plane at a specific distance from the image sensor 404. In this disclosure, the image sensor 404 is assumed to define a set of orthogonal x, y, and z axes such that the depth coordinate of a point in the scene corresponds to the z-component measured by the image sensor. Alternatively, the hand tracking device 440 can use other 3D mapping methods, such as stereoscopic imaging or time-of-flight measurements, based on single or multiple cameras or other types of sensors.

[0068] In some embodiments, the hand tracking device 140 captures and processes a time sequence of depth maps containing the user's hand while the user moves the hand (e.g., the entire hand or one or more fingers). Software running on the image sensor 404 and / or a processor in the controller 110 processes the 3D map data to extract patch descriptors of the hand in these depth maps. The software matches these descriptors with patch descriptors stored in the database 408, based on a previous learning process, to estimate the pose of the hand in each frame. The pose typically includes the 3D locations of the user's wrist joints and fingertips.

[0069] The software can also analyze hand and / or finger trajectories across multiple frames in a sequence to identify gestures. The pose estimation functionality described herein may be interleaved with motion tracking functionality, whereby patch-based pose estimation is performed only once every two (or more) frames, while tracking is used to discover pose changes that occur across the remaining frames. The pose, motion, and gesture information is provided to an application program running on controller 110 via the API described above. This program can, for example, move and modify an image presented on display generation component 120 or perform other functions in response to the pose and / or gesture information.

[0070] In some embodiments, the gesture includes an air gesture, which is detected without (or independent of) the user touching an input element that is part of a device (e.g., computer system 101, one or more input devices 125, and / or hand tracking device 140) and is based on detected movement of a part of the user's body in the air (e.g., head, one or more arms, one or more hands, one or more fingers, and / or one or more legs), including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement of the user's body relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to another of the user's hands, and / or movement of a user's finger relative to another finger or part of the user's hand), and / or absolute movement of the user's body part (e.g., a tap gesture involving movement of a hand in a predetermined posture by a predetermined amount and / or speed, or a shake gesture involving a predetermined speed or amount of rotation of the user's body part).

[0071] In some embodiments, input gestures used in various examples and embodiments described herein include air gestures performed by movement of a user's finger(s) relative to other finger(s) or part(s) of the user's hand to interact with an XR environment (e.g., a virtual or mixed reality environment), according to some embodiments. In some embodiments, an air gesture is a gesture that is detected without the user touching an input element that is part of the device (or independent of an input element that is part of the device) and is based on detected movement of a part of the user's body, including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground, or the distance of the user's hand relative to the ground), movement of the user's body relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of the user's other hand relative to one of the user's hands, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tap gesture that includes movement of the hand in a predetermined pose by a predetermined amount and / or speed, or a shake gesture that includes rotation of a part of the user's body at a predetermined speed or amount).

[0072] In some embodiments where the input gesture is an air gesture (e.g., in the absence of physical contact with an input device that provides a computer system with information about which user interface element is the target of the user input, such as contact with a user interface element displayed on a touchscreen or contact with a mouse or trackpad to move a cursor to a user interface element), the gesture takes into account the user's attention (e.g., gaze) to determine the target of the user input (e.g., in the case of direct input, as described below). Thus, in implementations that include air gestures, the input gesture is detected attention (e.g., gaze) to a user interface element in combination with (e.g., simultaneous with) movement of the user's finger(s) and / or hand to perform pinch and / or tap input, as described in more detail below.

[0073] In some embodiments, an input gesture directed at a user interface object is performed directly or indirectly with reference to the user interface object. For example, user input is performed directly at a user interface object in response to performing an input gesture with the user's hand at a position corresponding to the user interface object's position in the three-dimensional environment (e.g., as determined based on the user's current viewpoint). In some embodiments, an input gesture is performed indirectly at a user interface object in response to detecting the user's attention (e.g., gaze) to the user interface object while performing the input gesture while the user's hand position is not at a position corresponding to the user interface object's position in the three-dimensional environment. For example, for a direct input gesture, a user can direct the user's input at a user interface object by initiating the gesture at or near a position corresponding to the user interface object's displayed position (e.g., within a distance of 0.5 cm, 1 cm, 5 cm, or 0-5 cm, measured from an outer edge of the option or a central portion of the option). For indirect input gestures, a user can direct their input to a user interface object by paying attention to the user interface object (e.g., by gazing at the user interface object), and while paying attention to the option, the user initiates an input gesture (e.g., at any position detectable by the computer system) (e.g., at a position that does not correspond to the displayed position of the user interface object).

[0074] In some embodiments, input gestures (e.g., air gestures) used in various examples and embodiments described herein include pinch inputs and tap inputs for interacting with a virtual or mixed reality environment, according to some embodiments. For example, pinch inputs and tap inputs, as described below, are performed as air gestures.

[0075] In some embodiments, the pinch input is part of an air gesture, including one or more of a pinch gesture, a long pinch gesture, a pinch-and-drag gesture, or a double pinch gesture. For example, a pinch gesture that is an air gesture includes moving two or more fingers of a hand to contact each other, i.e., optionally with a short break (e.g., within 0-1 second) after contact with each other. A long pinch gesture that is an air gesture includes moving two or more fingers of a hand to contact each other for at least a threshold amount of time (e.g., at least 1 second) before detecting a break in contact with each other. For example, a long pinch gesture includes a user holding a pinch gesture (e.g., when two or more fingers are in contact), and the long pinch gesture continues until a break in contact between the two or more fingers is detected. In some embodiments, a double pinch gesture that is an air gesture includes two (e.g., or more) pinch inputs (e.g., performed by the same hand) that are detected immediately in succession (e.g., within a predetermined period of time) after each other. For example, a user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., breaking contact between two or more fingers), and performs a second pinch input within a predetermined period of time (e.g., within 1 second or 2 seconds) after releasing the first pinch input.

[0076] In some embodiments, a pinch-and-drag gesture that is an air gesture includes a pinch gesture (e.g., a pinch gesture or a long pinch gesture) performed in conjunction with (e.g., followed by) a drag input that changes the position of a user's hand from a first position (e.g., a start position of the drag) to a second position (e.g., an end position of the drag). In some embodiments, a user maintains the pinch gesture while performing the drag input and releases the pinch gesture (e.g., spreading two or more fingers apart) to end the drag gesture (e.g., at the second position). In some embodiments, the pinch input and the drag input are performed by the same hand (e.g., a user pinches two or more fingers together and moves the same hand to a second position in the air with a drag gesture). In some embodiments, the pinch input is performed by a user's first hand and the drag input is performed by the user's second hand (e.g., the user's second hand moves from a first position to a second position in the air while the user continues the pinch input with the user's first hand). In some embodiments, an input gesture that is an air gesture includes an input (e.g., a pinch input and / or a tap input) performed using both of a user's hands. For example, the input gesture includes two (e.g., or more) pinch inputs performed in conjunction with each other (e.g., simultaneously or within a predetermined period of time). For example, a first pinch gesture (e.g., a pinch input, a long pinch input, or a pinch and drag input) performed using a first hand of the user and a second pinch input performed using the other hand (e.g., a second of the user's hands) in conjunction with performing the pinch input using the first hand. In some embodiments, a movement between a user's hands (e.g., to increase and / or decrease the distance or relative orientation between the user's hands).

[0077] In some embodiments, a tap input (e.g., directed toward a user interface element) performed as an air gesture includes movement(s) of a user's finger(s) toward the user interface element, movement of a user's hand toward a user interface element, optionally with the user's finger(s) extended toward the user interface element, a downward movement of a user's finger (e.g., mimicking a mouse click action or a tap on a touchscreen), or other predefined movement of the user's hand. In some embodiments, a tap input performed as an air gesture is detected based on movement characteristics of the finger or hand performing the tap gesture, moving the finger or hand away from the user's viewpoint and / or toward the object that is the target of the tap input followed by an end of the movement. In some embodiments, an end of the movement is detected based on a change in movement characteristics of the finger or hand performing the tap gesture (e.g., an end of movement away from the user's viewpoint and / or toward the object that is the target of the tap input, a reversal of the direction of movement of the finger or hand, and / or a reversal of the direction of acceleration of the movement of the finger or hand).

[0078] In some embodiments, the user's attention is determined to be directed to a portion of the three-dimensional environment based on detecting a gaze directed to the portion of the three-dimensional environment (optionally, without requiring other conditions). In some embodiments, the device determines that the user's attention is directed to the portion of the three-dimensional environment based on detecting a gaze directed to the portion of the three-dimensional environment with one or more additional conditions, such as requiring the gaze to be directed to the portion of the three-dimensional environment for at least a threshold duration (e.g., dwell time) while the user's viewpoint is within a distance threshold from the portion of the three-dimensional environment, and / or requiring the gaze to be directed to the portion of the three-dimensional environment, and if one of the additional conditions is not met, the device determines that the user's attention is not directed to the portion of the three-dimensional environment to which the gaze is directed (e.g., until one or more additional conditions are met).

[0079] In some embodiments, detection of a ready configuration of a user or a portion of a user is detected by a computer system, and detection of a ready configuration of the hands is used by the computer system as an indication that the user is likely preparing to interact with the computer system using one or more air gesture inputs performed with the hands (e.g., pinch, tap, pinch and drag, double pinch, long pinch, or other air gestures described herein). For example, the ready state of a hand is determined based on whether the hand has a predetermined hand geometry (e.g., a pre-pinch geometry with the thumb and one or more fingers extended and spaced apart, ready to perform a pinch or grab gesture, or a pre-tap geometry with one or more fingers extended and the palm facing away from the user), whether the hand is in a predetermined position relative to the user's viewpoint (e.g., below the user's head, above the user's waist, extended at least 15 cm, 20 cm, 25 cm, 30 cm, or 50 cm from the body), and / or whether the hand has moved in a particular manner (e.g., above the user's waist, moved toward an area in front of the user below the user's head, or away from the user's body or legs). In some embodiments, the ready state is used to determine whether an interactive element of a user interface is responsive to attentional (e.g., gaze) input.

[0080] In some embodiments, the software may be downloaded to the controller 110 in electronic form, for example, over a network, or alternatively may be provided on a tangible, non-transitory medium, such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is similarly stored in memory associated with the controller 110. Alternatively, or additionally, some or all of the described functionality of the computer may be implemented in dedicated hardware, such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP). While the controller 110 is shown in FIG. 4 as, by way of example, a separate unit from the image sensor 440, some or all of the processing functionality of the controller may be implemented by a suitable microprocessor and software, or by dedicated circuitry within the housing of the hand tracking device 402, or otherwise associated with the image sensor 404. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with the display generation component 120 (e.g., in a television set, handheld device, or head-mounted device) or using any other suitable computerized device, such as a game console or media player. The sensing function of the image sensor 404 may likewise be integrated into a computer or other computerized device that is controlled by the sensor output.

[0081] FIG. 4 also includes a schematic diagram of a depth map 410 captured by the image sensor 404, according to some embodiments. The depth map includes a matrix of pixels having respective depth values, as described above. A pixel 412 corresponding to the hand 406 is segmented from the background and wrist in this map. The intensity of each pixel in the depth map 410 is inversely proportional to the depth value, i.e., the measured z-distance from the image sensor 404, with increasing gray levels as depth increases. The controller 110 processes these depth values ​​to identify and segment components of the image (i.e., groups of adjacent pixels) that have characteristics of a human hand. These characteristics can include, for example, the overall size, shape, and frame-to-frame motion of the depth map sequence.

[0082] 4 also schematically illustrates a hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406, according to some embodiments. In FIG. 4, the skeleton 414 is superimposed on a hand background 416 that was segmented from the original depth map. In some embodiments, key feature points on the hand (e.g., knuckles, fingertips, palm center, and / or end of the hand connecting to the wrist), and optionally the wrist or arm connected to the hand, are identified and positioned on the hand skeleton 414. In some embodiments, the location and movement of these key feature points over multiple image frames are used by the controller 110 to determine hand gestures performed by the hand or the current state of the hand, according to some embodiments.

[0083] FIG. 5 shows an exemplary embodiment of eye tracking device 130 ( FIG. 1 ). In some embodiments, eye tracking device 130 is controlled by eye tracking unit 243 ( FIG. 2 ) to track the position and movement of a user's gaze relative to scene 105 or relative to XR content displayed via display generation component 120. In some embodiments, eye tracking device 130 is integrated with display generation component 120. For example, in some embodiments, if display generation component 120 is a head-mounted device such as a headset, helmet, goggles, or glasses, or a handheld device disposed in a wearable frame, the head-mounted device includes both components for generating XR content for viewing by the user and components for tracking the user's gaze relative to the XR content. In some embodiments, eye tracking device 130 is separate from display generation component 120. For example, if the display generation component is a handheld device or an XR chamber, eye tracking device 130 is optionally a device separate from the handheld device or the XR chamber. In some embodiments, eye tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, head-mounted eye tracking device 130 is optionally used in conjunction with head-mounted or non-head-mounted display generation components. In some embodiments, eye tracking device 130 is not a head-mounted device, and is optionally used in conjunction with head-mounted display generation components. In some embodiments, eye tracking device 130 is not a head-mounted device, and is optionally part of non-head-mounted display generation components.

[0084] In some embodiments, the display generation component 120 uses a display mechanism (e.g., left and right near-eye display panels) that displays frames including left and right images in front of the user's eyes to provide the user with a 3D virtual view. For example, a head-mounted display generation component may include left and right optical lenses (referred to herein as eyepieces) positioned between the display and the user's eyes. In some embodiments, the display generation component may include or be coupled to one or more external video cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display that allows the user to view the physical environment directly and display virtual objects on the transparent or translucent display. In some embodiments, the display generation component projects virtual objects into the physical environment. The virtual objects are projected, for example, onto a physical surface or as a hologram, allowing an individual using the system to observe the virtual objects superimposed on the physical environment. In such cases, separate display panels and image frames for the left and right eyes may not be required.

[0085] As shown in FIG. 5 , in some embodiments, the gaze tracking device 130 includes at least one eye tracking camera (e.g., an infrared (IR) or near-IR (NIR) camera) and an illumination source (e.g., an IR or NIR light source such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eyes. The eye tracking camera may be aimed at the user's eyes to receive reflected IR or NIR light from the light source directly from the eyes, or alternatively, may be aimed at a “hot” mirror positioned between the user's eyes and a display panel that reflects the IR or NIR light from the eyes to the eye tracking camera while allowing visual light to pass through. The gaze tracking device 130 optionally captures images of the user's eyes (e.g., as a video stream captured at 60-120 frames per second (fps)), analyzes the images, generates gaze tracking information, and communicates the gaze tracking information to the controller 110. In some embodiments, the user's eyes are tracked separately by their respective eye tracking cameras and illumination sources. In some embodiments, only one eye of the user is tracked by a separate eye-tracking camera and lighting source.

[0086] In some embodiments, the eye tracking device 130 is calibrated using a device-specific calibration process to determine the eye tracking device's parameters for the particular operating environment 100, such as the 3D geometric relationships and parameters of the LEDs, camera, hot mirror (if present), eyepiece, and display screen. The device-specific calibration process may be performed at a factory or another facility before delivery of the AR / VR equipment to the end user. The device-specific calibration process may be an automatic or manual calibration process. The user-specific calibration process may include estimation of a particular user's eye parameters, such as pupil location, central visual location, optical axis, visual axis, eye spacing, etc. According to some embodiments, once the device-specific and user-specific parameters for the eye tracking device 130 have been determined, images captured by the eye tracking camera can be processed using glint-assisted methods to determine the user's current visual axis and viewpoint relative to the display.

[0087] As shown in FIG. 5, eye tracking device 130 (e.g., 130A or 130B) includes an eyepiece(s) 520 and a gaze tracking system including at least one eye tracking camera 540 (e.g., an infrared (IR) or near-IR (NIR) camera) positioned on the side of the user's face where eye tracking occurs and an illumination source 530 (e.g., an IR or NIR light source such as an array or ring of NIR light emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eye(s) 592. The eye tracking camera 540 may be positioned between the user's eye(s) 592 and the display 510 (e.g., the left or right display panel of a head-mounted display, or the display of a handheld device, and / or a projector) and may be directed at a mirror 550 that reflects IR or NIR light from the eye(s) 592 while transmitting visible light (e.g., as shown at the top of FIG. 5), or may be directed at the user's eye(s) 592 to receive reflected IR or NIR light from the eye(s) 592 (e.g., as shown at the bottom of FIG. 5).

[0088] In some embodiments, controller 110 renders AR or VR frames 562 (e.g., left and right frames for left and right display panels) and provides frames 562 to display 510. Controller 110 uses gaze tracking input 542 from eye tracking camera 540 for various purposes, such as in processing frames 562 for display. Controller 110 optionally estimates the user's viewpoint on display 510 based on gaze tracking input 542 obtained from eye tracking camera 540, using a glint-assisted method or other suitable method. The viewpoint estimated from gaze tracking input 542 is optionally used to determine the direction the user is currently looking.

[0089] Some possible use cases of the user's current gaze direction are described below, but are not intended to be limiting. As an exemplary use case, the controller 110 can render virtual content differently based on the determined user's gaze direction. For example, the controller 110 may generate virtual content with higher resolution in a central visual area determined from the user's current gaze direction than in a peripheral area. As another example, the controller may position or move virtual content within a view based at least in part on the user's current gaze direction. As another example, the controller may display particular virtual content within a view based at least in part on the user's current gaze direction. As another exemplary use case in an AR application, the controller 110 can orient an external camera to capture the physical environment of the XR experience and focus in the determined direction. The external camera's autofocus mechanism can then focus on an object or surface within the environment the user is currently viewing on the display 510. As another exemplary use case, eyepiece 520 may be a focusable lens, and eye-tracking information is used by the controller to adjust the focus of eyepiece 520 so that the virtual object the user is currently looking at has the proper binocular coordination to match the convergence of the user's eyes 592. Controller 110 can utilize the eye-tracking information to orient and focus eyepiece 520 so that close objects the user is looking at appear at the correct distance.

[0090] In some embodiments, the eye tracking device is part of a head-mounted device that includes a display (e.g., display 510), two eyepieces (e.g., eyepiece(s) 520), an eye tracking camera (e.g., eye tracking camera(s) 540), and a light source (e.g., light source 530 (e.g., IR LED or NIR LED)) attached to the wearable housing. The light source emits light (e.g., IR or NIR light) toward the user's eye(s) 592. In some embodiments, the light sources may be arranged in a ring or circle around each lens, as shown in FIG. 5. In some embodiments, eight light sources 530 (e.g., LEDs) are arranged around each lens 520, as an example. However, more or fewer light sources 530 may be used, and other arrangements and locations of the light sources 530 may be used.

[0091] In some embodiments, the display 510 emits light in the visible light range and not in the IR or NIR range, and therefore does not introduce noise into the gaze tracking system. Note that the location and angle of the eye tracking camera(s) 540 are given by way of example and are not intended to be limiting. In some embodiments, a single eye tracking camera 540 is located on each side of the user's face. In some embodiments, two or more NIR cameras 540 may be used on each side of the user's face. In some embodiments, a camera 540 with a wider field of view (FOV) and a camera 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, a camera 540 operating at one wavelength (e.g., 850 nm) and a camera 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.

[0092] Embodiments of an eye tracking system such as that shown in FIG. 5 may be used, for example, in augmented reality (e.g., including virtual reality and / or mixed reality) applications to provide a user with an augmented reality (e.g., including virtual reality, augmented reality, and / or augmented virtuality) experience.

[0093] Figure 6 shows a glint-assisted gaze tracking pipeline according to some embodiments. In some embodiments, the gaze tracking pipeline is implemented by a glint-assisted gaze tracking system (e.g., eye tracking device 130 as shown in Figures 1 and 5). The glint-assisted gaze tracking system can maintain a tracking state. Initially, the tracking state is off or "no." When in the tracking state, the glint-assisted gaze tracking system tracks the pupil contour and glint in the current frame using prior information from the previous frame when analyzing the current frame. When not in the tracking state, the glint-assisted gaze tracking system attempts to detect the pupil and glint in the current frame, and if successful, initializes the tracking state to "yes" and continues to the next frame in the tracking state.

[0094] As shown in FIG. 6, an eye-tracking camera can capture left and right images of a user's left and right eyes. The captured images are then input into an eye-tracking pipeline for processing beginning at 610. As indicated by the arrow returning to element 600, the eye-tracking system can continue to capture images of the user's eyes at a rate of, for example, 60-120 frames per second. In some embodiments, each set of captured images may be input into the pipeline for processing. However, in some embodiments, or under some conditions, not all captured frames are processed by the pipeline.

[0095] At 610, if the tracking status is yes for the currently captured image, the method proceeds to element 640. If the tracking status is no at 610, the image is analyzed to detect the user's pupil and glint in the image, as shown at 620. If the pupil and glint are successfully detected at 630, the method proceeds to element 640. If not, the method returns to element 610 to process the next image of the user's eyes.

[0096] At 640, proceeding from element 410, the current frame is analyzed to track pupils and glints based in part on previous information from the previous frame. At 640, proceeding from element 630, a tracking state is initialized based on the detected pupils and glints in the current frame. The results of the processing at element 640 are checked to ensure that the tracking or detection results are reliable. For example, the results can be checked to determine whether a sufficient number of glints are successfully tracked or detected in the current frame to perform pupil and gaze estimation. At 650, if the results are not reliable, the tracking state is set to no and the method returns to element 610 to process the next image of the user's eyes. At 650, if the results are reliable, the method proceeds to element 670. At 670, the tracking state is set to yes (if not already yes) and the pupil and glint information is passed to element 680 to estimate the user's gaze point.

[0097] 6 is intended to serve as an example of eye-tracking technology that may be used in particular implementations. As will be recognized by those skilled in the art, other eye-tracking technologies, now existing or developed in the future, may be used in place of or in combination with the glint-assisted eye-tracking technology described herein in computer system 101 to provide a user with an XR experience according to various embodiments.

[0098] In this disclosure, various input methods are described with respect to interaction with a computer system. Where one example is provided using one input device or input method and another example is provided using a different input device or input method, it should be understood that each example may be compatible with, and optionally utilize, the input device or input method described with respect to the other example. Similarly, various output methods are described with respect to interaction with a computer system. Where one example is provided using one output device or output method and another example is provided using a different output device or output method, it should be understood that each example may be compatible with, and optionally utilize, the output device or output method described with respect to the other example. Similarly, various methods are described with respect to interaction with a virtual environment or a mixed reality environment via a computer system. Where one example is provided using interaction with a virtual environment and another example is provided using a mixed reality environment, it should be understood that each example may be compatible with, and optionally utilize, the method described with respect to the other example. Thus, this disclosure discloses embodiments that are combinations of features of multiple examples, without exhaustively listing all features of the embodiments in the description of each exemplary embodiment. User Interface and Related Processes

[0099] We now turn our attention to embodiments of user interfaces (“UIs”) and associated processes that may be executed in a computer system, such as a portable multifunction device or a head-mounted device, equipped with a display generation component, one or more input devices, and (optionally) one or more cameras.

[0100] 7A-7J illustrate a three-dimensional environment displayed via a display generation component (e.g., display generation component 7100 or display generation component 120) and interactions occurring in the three-dimensional environment caused by user input directed to the three-dimensional environment and / or input received from other computer systems and / or sensors. In some embodiments, input is directed to a virtual object in the three-dimensional environment by a user's gaze detected within an area occupied by the virtual object and / or by a hand gesture performed at a location in the physical environment corresponding to the area of ​​the virtual object. In some embodiments, input is directed to a virtual object in the three-dimensional environment (e.g., optionally at a location in the physical environment independent of the area of ​​the virtual object in the three-dimensional environment) by a hand gesture performed while the virtual object has an input focus (e.g., while the virtual object is being selected by simultaneously and / or previously detected gaze input, while being selected by simultaneously and / or previously detected pointer input, while being selected by simultaneously and / or previously detected gesture input). In some embodiments, input is directed to a virtual object in the three-dimensional environment by an input device that places a focus selector object (e.g., a pointer object or a selector object) at the position of the virtual object. In some embodiments, input is directed to a virtual object in the three-dimensional environment via other means (e.g., voice and / or control buttons). In some embodiments, input is directed to a physical object or a representation of a virtual object corresponding to a physical object by a user's hand movements (e.g., whole hand movements, whole hand movements in a discrete pose, movement of one part of the hand relative to another part of the hand, and / or relative movement between two hands) and / or manipulation of the physical object (e.g., touching, swiping, tapping, opening, moving towards, and / or moving relative to).In some embodiments, the computer system modifies the display in the three-dimensional environment (e.g., displaying additional virtual content, ceasing to display existing virtual content, and / or transitioning between different immersion levels displaying visual content) according to input from sensors (e.g., image sensors, temperature sensors, biometric sensors, motion sensors, and / or proximity sensors) and contextual conditions (e.g., location, time, and / or the presence of others in the environment). In some embodiments, the computer system modifies the display in the three-dimensional environment (e.g., displaying additional virtual content, ceasing to display existing virtual content, or transitioning between different immersion levels displaying visual content) according to input from other computers used by other users sharing the computer-generated environment with the user of the computer system (e.g., in a shared computer-generated experience, shared virtual environment, or shared virtual or augmented reality environment of a communication session). In some embodiments, the computer system displays some changes in the three-dimensional environment (e.g., displaying movement, deformation, changes in visual characteristics, etc. of the user interface, virtual surfaces, user interface objects, and / or virtual scenery) according to input from sensors that detect the movement of other people and objects, as well as user movements that may not be of high quality as recognized gesture input for triggering related actions of the computer system.

[0101] In some embodiments, the three-dimensional environment displayed via the display generation components described herein is a virtual three-dimensional environment that includes virtual objects and content at different virtual positions within the three-dimensional environment without a representation of the physical environment. In some embodiments, the three-dimensional environment is a mixed reality environment that displays virtual objects at different virtual positions within the three-dimensional environment constrained by one or more physical aspects of the physical environment (e.g., the position and orientation of walls, floors, surfaces, the direction of gravity, the time of day, and / or the spatial relationships between physical objects). In some embodiments, the three-dimensional environment is an augmented reality environment that includes a representation of the physical environment. In some embodiments, the representation of the physical environment includes respective representations of physical objects and surfaces at different positions within the three-dimensional environment, such that the spatial relationships between the different physical objects and surfaces in the physical environment are reflected by the spatial relationships between the representations of the physical objects and surfaces in the three-dimensional environment. In some embodiments, when the virtual objects are positioned relative to the positions of the representations of the physical objects and surfaces in the three-dimensional environment, they appear to have corresponding spatial relationships with the physical objects and surfaces in the physical environment. In some embodiments, the computer system transitions between displaying different types of environments (e.g., transitioning between presenting computer-generated environments or experiences at different levels of immersion, or adjusting the relative prominence of audio / visual sensory inputs from the virtual content and from the representation of the physical environment) based on user input and / or contextual conditions.

[0102] In some embodiments, the display generation component includes a pass-through portion in which a representation of the physical environment is displayed. In some embodiments, the pass-through portion of the display generation component is a transparent or translucent (e.g., see-through) portion of the display generation component that surrounds the user's field of view and reveals at least a portion of the physical environment within the field of view. For example, the pass-through portion is a translucent (e.g., less than 50%, 40%, 30%, 20%, 15%, 10%, or 5% opacity) or transparent portion of a head-mounted or head-up display, allowing the user to view the real world surrounding the user through it without removing the head-mounted display or moving away from the head-up display. In some embodiments, the pass-through portion gradually transitions from translucent or transparent to fully opaque when displaying a virtual or mixed reality environment. In some embodiments, the pass-through portion of the display generation component displays a live feed of images or video of at least a portion of the physical environment captured by one or more cameras (e.g., rear-facing camera(s) of a mobile device or associated with a head-mounted display, or other cameras that provide image data to a computer system). In some embodiments, one or more cameras are aimed at a portion of the physical environment that is directly in front of the user (e.g., relative to the user of the display generation component, behind the display generation component). In some embodiments, one or more cameras are aimed at a portion of the physical environment that is not directly in front of the user (e.g., in a different physical environment, or to the side or behind the user).

[0103] In some embodiments, when displaying virtual objects in positions corresponding to the locations of one or more physical objects in a physical environment (e.g., positions in a virtual reality environment, a mixed reality environment, or an augmented reality environment), at least some of the virtual objects are displayed in place of (e.g., replace the display of) a portion of a camera's live view (e.g., a portion of the physical environment captured in the live view). In some embodiments, at least some of the virtual objects and content are projected onto physical surfaces or open space in the physical environment and are visible through a pass-through portion of the display generation component (e.g., as part of the camera view of the physical environment or visible through a transparent or semi-transparent portion of the display generation component). In some embodiments, at least some of the virtual objects and virtual content are displayed to overlay a portion of the display, blocking the view of at least a portion of the physical environment that is visible through the transparent or semi-transparent portion of the display generation component.

[0104] In some embodiments, the display generation component displays different views of the three-dimensional environment according to user input or movement to change the virtual position of the viewpoint of the currently displayed view of the three-dimensional environment relative to the three-dimensional environment. In some embodiments, if the three-dimensional environment is a virtual environment, the viewpoint moves according to a navigation or movement request (e.g., an air hand gesture and / or a gesture performed by movement of one part of a hand relative to another part of the hand) without requiring movement of the user's head, torso, and / or display generation component in the physical environment. In some embodiments, movement of the user's head and / or torso relative to the physical environment (e.g., by the user holding the display generation component or wearing an HMD) and / or movement of the display generation component or other location-sensing element of the computer system, etc., causes a corresponding movement of the viewpoint relative to the three-dimensional environment (e.g., with a corresponding change in direction, distance, speed, and / or orientation of movement), resulting in a corresponding change in the currently displayed view of the three-dimensional environment. In some embodiments, when a virtual object has a predetermined spatial relationship to the viewpoint (e.g., is fixed to the viewpoint), movement of the viewpoint relative to the three-dimensional environment causes movement of the virtual object relative to the three-dimensional environment while the position of the virtual object within the field of view is maintained (e.g., the virtual object is said to be head-locked). In some embodiments, the virtual object is body-locked to the user, moving relative to the three-dimensional environment as the user moves as a whole within the physical environment (e.g., carrying or wearing the display generating components and / or other location sensing components of the computer system), but does not move within the three-dimensional environment solely in response to movement of the user's head (e.g., the display generating components and / or other location sensing components of the computer system rotating around the user's fixed location within the physical environment).In some embodiments, the virtual object is optionally locked to another part of the user, such as the user's hand or wrist, and moves within the three-dimensional environment according to movements of the part of the user in the physical environment, maintaining a preset spatial relationship between the position of the virtual object and the virtual position of the part of the user in the three-dimensional environment. In some embodiments, the virtual object is locked to a preset portion of the field of view provided by the display generation component, and moves within the three-dimensional environment according to movements of the field of view, regardless of user movements that do not cause a change in the field of view.

[0105] In some embodiments, as shown in FIGS. 7B-7J, the view of the three-dimensional environment may not include representation(s) of the user's hand(s), arm(s), and / or wrist(s). In some embodiments, representation(s) of the user's hand(s), arm(s), and / or wrist(s) are included in the view of the three-dimensional environment. In some embodiments, representation(s) of the user's hand(s), arm(s), and / or wrist(s) are included in the view of the three-dimensional environment as part of the representation of the physical environment provided via the display generation component. In some embodiments, the representations are not part of the representation of the physical environment, but are captured separately (e.g., by one or more cameras directed at the user's hand(s), arm(s), and wrist(s)) and displayed in the three-dimensional environment independently of the currently displayed view of the three-dimensional environment. In some embodiments, the representation(s) include a stylized version of the arm(s), wrist(s), and / or hand(s) based on camera images captured by one or more cameras of the computer system(s) or information captured by various sensors. In some embodiments, the representation(s) replace the display of, overlap with, or block the view of a portion of the representation of the physical environment. In some embodiments, if the display generation component does not provide a view of the physical environment but provides an entirely virtual environment (e.g., no camera view and no transparent pass-through portions), a real-time visual representation (e.g., a stylized representation or segmented camera image) of one or both of the user's arms, wrists, and / or hands is optionally still displayed in the virtual environment. In some embodiments, if a representation of the user's hand is not provided within the view of the three-dimensional environment, a position corresponding to the user's hand is optionally indicated within the three-dimensional environment, for example, by a change in appearance of virtual content (e.g., through a change in translucency and / or simulated reflectivity) at a position in the three-dimensional environment that corresponds to the location of the user's hand in the physical environment.In some embodiments, a representation of the user's hand or wrist is outside the currently displayed view of the three-dimensional environment while a virtual position in the three-dimensional environment corresponding to the location of the user's hand or wrist is outside the current field of view provided via the display generation component, and the representation of the user's hand or wrist is made visible within the view of the three-dimensional environment in response to the virtual position corresponding to the location of the user's hand or wrist being moved within the current field of view due to movement of the display generation component, the user's hand or wrist, the user's head, and / or the entire user, etc.

[0106] 7A-7J are block diagrams illustrating user interaction with user interface objects displayed in a three-dimensional environment, according to some embodiments. In some embodiments, one or more of the user interface objects are provided within a predetermined zone within the three-dimensional environment, and user interface objects located within the predetermined zone follow the user within the three-dimensional environment, while user interface objects located outside the predetermined zone do not follow the user within the three-dimensional environment (e.g., user interface objects located outside the predetermined zone are anchored to the three-dimensional environment). The behaviors described in FIGS. 7A-7J (and FIGS. 8-9) with respect to user interface objects in some examples are applicable to user interface objects in other examples, according to various embodiments, unless otherwise specified in the description.

[0107] 7A-7J illustrate an exemplary computer system (e.g., device 101 or another computer system) in communication with a first display generating component (e.g., display generating component 7100 or another display generating component). In some embodiments, the first display generating component is a head-up display. In some embodiments, the first display generating component is a head-mounted display (HMD). In some embodiments, the first display generating component is a standalone display, a touchscreen, a projector, or another type of display. In some embodiments, the computer system communicates with one or more input devices, including cameras or other sensors and input devices, that detect the movement of a user's hands, the movement of the user's entire body, and / or the movement of the user's head in the physical environment. In some embodiments, the one or more input devices detect the user's movements and the current posture, orientation, and position of the user's hands, face, and entire body. In some embodiments, the one or more input devices include an eye-tracking component that detects the location and movement of the user's gaze. In some embodiments, the first display generating component, and optionally the one or more input devices and the computer system, are part of a head-mounted device (e.g., an HMD or pair of goggles) that moves and rotates with the user's head in the physical environment, changing the user's viewpoint into the three-dimensional environment provided via the first display generating component. In some embodiments, the first display generating component is a head-up display that does not move or rotate with the user's head or the user's entire body, but optionally changes the user's viewpoint into the three-dimensional environment in accordance with movement of the user's head or body relative to the first display generating component. In some embodiments, the first display generating component is optionally moved and rotated by the user's hands, relative to the physical environment, or relative to the user's head, changing the user's viewpoint into the three-dimensional environment in accordance with movement of the first display generating component relative to the user's head or face or relative to the physical environment.

[0108] 7A-7E are block diagrams illustrating the display of user interface object 7104 (e.g., user interface objects 7104-1 to 7104-3 are instances of user interface object 7104) at respective positions in a three-dimensional environment corresponding to a location relative to user 7002 (e.g., the user's viewpoint) within physical environment 7000.

[0109] 7A shows a physical environment 7000 including a user 7002 interacting with a display generation component 7100. In the examples described below, the user 7002 provides input or commands to a computer system using one or both of two hands, namely, hand 7020 and hand 7022. In some of the examples described below, the computer system also uses the position or movement of the user's arm, such as the user's left arm 7028 connected to the user's left hand 7020, as part of the input provided to the computer system by the user. The physical environment 7000 includes a physical object 7014 and physical walls 7004 and 7006. The physical environment 7000 further includes a physical floor 7008.

[0110] As shown in FIG. 7B , a computer system (e.g., display generation component 7100) displays a view of a three-dimensional environment (e.g., environment 7000′, a virtual three-dimensional environment, an augmented reality environment, a pass-through view of the physical environment, or a camera view of the physical environment). In some embodiments, the three-dimensional environment is a virtual three-dimensional environment without a representation of physical environment 7000. In some embodiments, the three-dimensional environment is a mixed reality environment, which is a virtual environment augmented with sensor data corresponding to physical environment 7000. In some embodiments, the three-dimensional environment is an augmented reality environment that includes one or more virtual objects (e.g., user interface object 7104) and a representation of at least a portion of the physical environment surrounding display generation component 7100 (e.g., representations of walls 7004′, 7006′, a representation of a floor 7008′, and a representation of physical objects 7014′). For example, in some embodiments, the representation of the physical environment includes a camera view of the physical environment. In some embodiments, the representation of the physical environment includes a view of the physical environment through a transparent or semi-transparent portion of the display generation component. In some embodiments, the representation 7014' of the physical object is locked (e.g., fixed) to the three-dimensional environment such that the representation 7014' is maintained in its position within the three-dimensional environment as the user moves within the physical environment (e.g., is displayed only when the user's current view includes the portion of the three-dimensional environment to which the representation 7014' of the physical object is fixed).

[0111] 7C-7E illustrate examples of a user's attention to various objects (e.g., physical objects in a physical environment and / or virtual objects) in a three-dimensional environment 7000′ displayed using the display generation component 7100. For example, FIG. 7C illustrates a first view from a user's perspective while the user is paying attention to (e.g., gazing at) user interface object 7104-1. For example, the user's attention is represented by a dashed line from the user's eyes. In some embodiments, the computer system determines that the user is paying attention to a respective portion (e.g., object) of the three-dimensional environment based on sensor data that determines the user's line of sight and / or the user's head position. It will be appreciated that the computer system can use various sensor data to determine the portion of the three-dimensional environment to which the user is currently paying attention.

[0112] In some embodiments, user interface object 7104-1 includes a panel that includes multiple selectable user interface options (e.g., buttons) that are selectable by a user via the user's gaze and / or gestures (e.g., air gestures) with one or more of the user's hands. In some embodiments, the user controls (e.g., modifies) which selectable user interface options are included in the panel (e.g., user interface object 7104-1). For example, the user selects particular application icons, settings, controls, and / or other options to be displayed in the panel such that selected application icons, settings, controls, and / or other options included in the panel are easily accessible by the user (e.g., the panel follows the user as the user moves through the physical environment, allowing the user to interact with the panel as the user moves through the physical environment, as described in more detail below).

[0113] 7D illustrates a user focusing on object 7014′ (e.g., a representation of physical object 7014 in physical environment 7000). In response to the user not focusing on user interface object 7104-1 (e.g., as shown in FIG. 7C), user interface object 7104-1 is updated to user interface object 7104-2, which is displayed as a visually less emphasized version of user interface object 7104-1 (e.g., as indicated by a shaded fill). In some embodiments, user interface object 7104-2 is displayed with faded visual characteristics relative to the visual characteristics of user interface object 7104-1, and user interface object 7104-1 is displayed with unchanged (e.g., unfaded) visual characteristics while the user focuses on user interface object 7104-1. In some embodiments, user interface object 7104-2 is displayed with a faded visual characteristic relative to other objects (e.g., virtual objects and / or physical objects) displayed in three-dimensional environment 7000′. For example, representation 7014′ of physical object is not visually suppressed (e.g., is not modified), but user interface object 7104-2 is visually suppressed. In some embodiments, the user interface object is visually suppressed by blurring user interface object 7104-2, reducing the size of user interface object 7104-2, reducing the opacity of user interface object 7104-2, increasing the translucency of user interface object 7104-2, ceasing to display user interface object 7104-2 entirely, or a combination of visual effects (e.g., fading and blurring simultaneously) that cause user interface object 7104-2 to be suppressed.

[0114] In some embodiments, as shown in FIGS. 7E-7G, while a user is not focusing on user interface objects 7104 (e.g., user interface object 7104-3, user interface object 7104-4, and user interface object 7104-5), user interface objects 7104 continue to be displayed with visual de-emphasis (e.g., as indicated by the shaded fill in FIGS. 7E-7F). In some embodiments, the visual de-emphasis of user interface objects 7104 increases while the user is not focusing on the user interface objects (e.g., as the amount of time the user is not focusing on the user interface objects increases). For example, in response to a user initially shifting the user's attention away from user interface object 7104-1, user interface object 7104-2 is displayed faded by a first amount (e.g., the opacity of the user interface object is decreased by a first amount and / or the translucency of the user interface object is increased by a first amount). In some embodiments, after a predetermined amount of time (e.g., 0.1, 0.2, 0.5, 1, 2, or 5 seconds), user interface object 7104-3 (FIG. 7E) is displayed faded by a second amount greater than the first amount (e.g., user interface object 7104-3 is displayed with a greater amount of visual de-emphasis than user interface object 7104-2).

[0115] In some embodiments, the amount of visual de-emphasis is determined based at least in part on the speed and / or amount (e.g., amount of change in angle and / or amount of distance) at which the user moves the user's attention away from the object. For example, in response to the user moving their eyes away from user interface object 7104-1 quickly (e.g., at a first speed), user interface object 7104-2 is visually de-emphasized by a first amount. In response to the user moving their eyes away (and / or changing their orientation) from user interface object 7104-1 more slowly (e.g., at a second speed slower than the first speed), user interface object 7104-2 is de-emphasized by a second amount less than the first amount. In some embodiments, the amount of visual de-emphasis is based on the amount of change (e.g., change in distance and / or change in angle) between user interface object 7104-1 and the user's current location of attention within the three-dimensional environment (e.g., in addition to or instead of being based on the speed of the user's movement / attention change). For example, if the user moves their attention to an area close to user interface object 7104-1 (e.g., within 5 cm, or within 10 cm, or meeting a predetermined proximity criterion), the user interface object is visually de-emphasized to a lesser extent than if the user moves their attention to an area farther away from user interface object 7104-1 (e.g., more than 5 cm, or more than 10 cm). Thus, as the user moves their attention from user interface object 7104-1, the display of user interface object 7104-2 is updated according to one or more characteristics of the user's movement and / or change in the user's attention.

[0116] 7E-7H , as the user moves within the three-dimensional environment 7000′ (e.g., corresponding to the user moving around the physical environment 7000), the user interface object 7104 continues to be displayed with visual de-emphasis (or is not displayed) while the user is not focusing on the user interface object 7104. In some embodiments, the user interface object 7104 is displayed at various positions within the three-dimensional environment as the user moves within the physical environment (e.g., the user interface object 7104 follows the user), as described in more detail below.

[0117] In some embodiments, user interface object 7104 continues to be displayed with visual de-emphasis until the computer system detects that the user is focusing on user interface object 7104, as shown in Figure 7H. For example, in response to detecting that the user is focusing on user interface object 7104-6, the user interface object is displayed without visual de-emphasis (e.g., user interface object 7104-6 is displayed with the same visual characteristics as user interface object 7104-1 of Figure 7C). In some embodiments, in response to detecting that the user is focusing on user interface object 7104-6, user interface object 7104-6 is displayed in a position within the three-dimensional environment such that user interface object 7104-6 has the same relative position to the user as user interface object 7104-1's (e.g., previous) relative position to the user (e.g., the initial position of the user interface object relative to the user before the user moved within the physical environment).

[0118] 7E-7H are block diagrams illustrating user interface objects 7104 that are displayed in various positions within the three-dimensional environment as user 7002 moves through physical environment 7000. It will be appreciated that changing the position of user interface objects 7104 within the three-dimensional environment can be performed in conjunction with (e.g., simultaneously with) the visual de-emphasis of user interface objects 7104 described above.

[0119] In some embodiments, as shown in Figure 7E, the user (and the user's current viewpoint) moves within the physical environment (e.g., the user moves a first amount of distance to the right), and as the user moves within the physical environment, the view displayed on the display generation component 7100 is updated (e.g., in real time) to include a current view of the three-dimensional environment that reflects the user's movement within the physical environment. For example, as the user moves to the right in Figure 7E (e.g., relative to the view in Figure 7D), the representation of the physical object 7014' appears to be more centered in the user's current view in Figure 7E (compared to the representation of the physical object 7014' displayed at the far right of the user's view in Figure 7D).

[0120] In some embodiments, while the user moves within the physical environment, user interface object 7104-3 is initially maintained in the same position within the three-dimensional environment (e.g., relative to other displayed objects within the three-dimensional environment). For example, in FIG. 7D , user interface object 7104-2 is displayed with its right edge aligned (e.g., vertically) with the left edge of the representation of object 7014′. In response to the user moving a first amount in FIG. 7E , user interface object 7104-3 continues to be displayed in the same position within the three-dimensional environment relative to the representation of object 7014′. For example, user interface object 7104-3 initially appears fixed in the three-dimensional environment. In some embodiments, user interface object 7104-2 (e.g., and user interface object 7104-3) is maintained in the same position within the three-dimensional environment relative to other objects within the three-dimensional environment in response to the user moving less than a threshold amount (e.g., of a change in distance, orientation, and / or position) within the physical environment (e.g., the first amount of movement by the user is less than the threshold amount). In some embodiments, user interface object 7104-2 (e.g., and user interface object 7104-3) is maintained in the same position within the three-dimensional environment relative to other objects within the three-dimensional environment during a first predetermined period of user movement. For example, user interface object 7104-2 is displayed in the same position within the three-dimensional environment for the first 2 seconds (e.g., 0.5 seconds, or 4 seconds) of the user moving within the physical environment.

[0121] In some embodiments, after the user moves more than a threshold amount (e.g., more than a threshold distance, more than a threshold amount of change in orientation and / or position, and / or for a period of time longer than a first predetermined period), user interface object 7104-4 is updated to appear in a position in the three-dimensional environment that is different from its initial position (e.g., before the user began moving). For example, user interface object 7104-3 disappears from the user's current view in FIG. 7F without being updated. Thus, user interface object 7104-4 is moved relative to other objects displayed in the three-dimensional environment to remain within the user's current view (e.g., user interface object 7104-4 remains displayed in its entirety as the user moves in the physical environment). Thus, user interface object 7104-4 is not anchored to the three-dimensional environment, but instead is anchored to the user's current viewpoint.

[0122] In some embodiments, display generation component 7100 displays user interface object 7104-3 with animated movement (e.g., gradual and continuous movement) toward the position of user interface object 7104-4 shown in FIG. 7F. In some embodiments, while user interface object 7104-3 is being moved, user interface object 7104-3 is visually emphasized and suppressed as described above. In some embodiments, user interface objects 7104-3 through 7104-4 are displayed as if the user interface objects are following the user as they move through the physical environment (e.g., such that user interface object 7104 remains entirely visible in each individual user's current view). In some embodiments, when user interface object 7104-3 updates the position of user interface object 7104-4, the rate at which user interface object 7104-3 moves toward the position of user interface object 7104-4 is displayed as moving at a rate slower than the rate at which the user is moving through the physical environment. For example, user interface objects may be delayed in following the user (e.g., only beginning to follow the user after two seconds) and may appear to move slower in the three-dimensional environment than the rate of the user's movement in the physical environment (e.g., the rate of change relative to the user's current viewpoint). Thus, the user interface objects appear to lag behind the user as they move in the physical environment.

[0123] FIG. 7G shows the user continuing to move within the physical environment (e.g., relative to FIGS. 7D-7F). FIG. 7G shows additional lateral movement (e.g., left-right movement) of the user (e.g., and the display generation component 7100) within the physical environment as the user continues to move to the right (e.g., in the same direction as described above) within the physical environment. FIG. 7G also shows movement of the user's posture (e.g., orientation) in the vertical direction (e.g., as indicated by the downward arrow in FIG. 7G). For example, the user moves to the right within the physical environment while (e.g., simultaneously) moving the user's current viewpoint downward (e.g., to include more of the representation of floor 7008′ in FIG. 7G). In some embodiments, user interface object 7104-5 is updated to move as the user moves within the physical environment (e.g., at a slower rate than the user). For example, if the user moves more to the right in Figure 7G relative to Figure 7E, user interface object 7104 will also be displayed as moving to the right (e.g., along with the user) between Figures 7E-7G at a rate slower than the rate of the user's movement. For example, instead of user interface object 7104 remaining displayed in the top-center portion of the user's current viewpoint in Figures 7E-7G (e.g., showing user interface object 7104 moving at the same rate as the user), user interface object 7104 will appear to lag behind while the user is moving.

[0124] In some embodiments, after the user moves beyond a threshold amount of movement (e.g., the user interface object 7104 moves from an initial position to an updated position and remains within the user's current view), the user interface object 7104 continues to follow the user in the three-dimensional environment as the user continues to move within the physical environment. In some embodiments, the user interface object 7104 is moved to a different position in the three-dimensional environment (e.g., as the user moves within the physical environment) to maintain the same spatial relationship to the user (e.g., relative to a portion of the user's body and / or relative to the user's current viewpoint). For example, the user interface object 7104 continues to follow the user to remain within a predetermined portion of the user's current view (e.g., the upper left corner of the user's current view) and / or a predetermined distance away from the user's current view (e.g., within an arm's length of the user).

[0125] The user's viewpoint is optionally updated by any combination of moving the display generating component 7100 laterally within the physical environment, changing the relative angle (e.g., pose) of the display generating component 7100, and / or changing the pose (e.g., orientation) of the user's head (e.g., when the user is looking down toward the floor 7008, such as when the display generating component is an HMD worn by the user). The examples described herein of the user's movement within the physical environment in a particular direction and / or orientation (e.g., rightward and / or downward) are non-limiting examples of the user's movement within the physical environment. For example, other movements of the user (e.g., leftward, upward, and / or combinations of different directions and / or poses) may cause the user interface objects to be displayed with similar behavior (e.g., the user interface objects may move within the user's current viewpoint of the three-dimensional environment to follow the user's movement (optionally with a delay and / or lag)).

[0126] In some embodiments, as shown in Figure 7H, after the user moves more than a threshold amount (e.g., of distance, pose, and / or orientation) within the physical environment, user interface object 7104-6 is re-displayed at a position within the user's current view of the three-dimensional environment that is defined relative to the user (e.g., the user's body and / or the user's viewpoint). For example, in Figure 7C, user interface object 7104-1 is initially displayed at a position within the three-dimensional environment that is defined relative to the user's current viewpoint. For example, user interface object 7104-1 is displayed at a predetermined distance (e.g., perceived depth) from the user and at a height relative to the user (e.g., above the user's current viewpoint or at a predetermined angle (e.g., 45 degrees) above the user's viewpoint when the user is looking straight ahead). In some embodiments, while the user is moving within the physical environment, before the user moves a threshold amount, the user interface object is moved within the user's current view and appears with the lazy follow behavior described with reference to Figures 7E-7G, and after the user moves at least the threshold amount (e.g., as shown in Figure 7H), the user interface object 7104-6 reappears in the same position defined relative to the user's current viewpoint as described in Figure 7C. In some embodiments, the same position defined relative to the user's current viewpoint corresponds to a predetermined zone that is within the user's comfortable viewing distance.

[0127] In some embodiments, the delay and offset behavior of the user interface object 7104 described above (e.g., also referred to herein as delayed follow behavior) is implemented pursuant to the addition of the user interface object 7104 to one of a plurality of predetermined zones. For example, the initial position of the user interface object 7104-1 is set within a first of the plurality of predetermined zones, and any user interface objects placed (e.g., anchored) within one of the plurality of predetermined zones are updated according to the delayed follow behavior described herein. In some embodiments, the user 7002 can also move user interface objects into and out of various zones (e.g., such that while an individual user interface object is not placed within one of the predetermined zones, the delayed follow behavior no longer applies). In some embodiments, while the user selects the user interface object, the plurality of predetermined zones are highlighted (e.g., with an outline of each individual zone) to indicate to the user where the user can place the user interface object in order for the user interface object to have the delayed follow behavior.

[0128] In some embodiments, the predetermined zone covers a predefined portion (e.g., a predefined shape) of the three-dimensional environment. For example, the predetermined zone occupies a position within the three-dimensional environment defined by its length, width, depth, and / or shape (e.g., boundary). For example, a first predetermined zone is located at (e.g., occupies) a first depth (e.g., or range of depths) and has a first width, length, and / or height. In some embodiments, the first predefined zone occupies a portion of the three-dimensional environment that corresponds to a three-dimensional shape, or optionally a two-dimensional shape (e.g., a two-dimensional window or dock). For example, the first predefined zone is a cube at a predefined position within the three-dimensional environment (e.g., moving a user interface object into the cube at a predefined position is moving the user interface object into the first predefined zone).

[0129] In some embodiments, the user interface object 7104 disappears while the user moves the user's head without moving the user's body in the physical environment. In some embodiments, the user interface object remains displayed (e.g., has a visually de-emphasized characteristic) while the user moves the user's body (e.g., torso and head). For example, if the user rotates the user's head (e.g., updates the user's current view of the three-dimensional environment) without changing the user's location (e.g., moving from a first location to a second location in the physical environment) and / or without the user moving the user's torso (e.g., changing the user's body orientation), the user interface object 7104 is not animated to move from a first position to a second position in the three-dimensional environment. Instead, the user interface object 7104 is not displayed during the user's movement and reappears in response to the user's head movement ceasing (e.g., remaining in the user's new head position for a predetermined period of time) (e.g., when the user remains stationary in the second position for a predetermined period of time).

[0130] In some embodiments, the user is further enabled to interact with user interface object 7104-4, as shown in Figure 7I. For example, user interface object 7104-7 is a panel including multiple selectable objects (e.g., application icons, controls in a control center, settings, and / or buttons). In some embodiments, in response to detecting user input (e.g., the user's gaze and / or air gesture) directed at a first selectable object of the multiple selectable objects, the first selectable object is highlighted (e.g., highlighted, outlined, magnified, or distinguished relative to other selectable objects).

[0131] In some embodiments, the plurality of selectable objects include one or more controls for an immersive experience in the three-dimensional environment. For example, user interface object 7104-7 includes a play and / or pause control to immerse the user in the three-dimensional environment in a complete virtual experience and provide the user with options to change the level of immersion (e.g., display more or less pass-through content from the physical environment in the three-dimensional environment). For example, controls for playing and / or pausing the immersive experience in the three-dimensional environment are displayed. In some embodiments, higher levels of immersion in the three-dimensional environment include additional virtual features, such as displaying virtual objects, displaying virtual wallpaper, displaying virtual lighting, etc. Thus, the user can control how much of the physical environment is displayed in the three-dimensional environment as pass-through content relative to the amount of virtual content displayed in the three-dimensional environment.

[0132] For example, as shown in FIG. 7J , in response to user input (e.g., a hand gesture using the user's hand 7020 or a combined hand and gaze gesture), the user may move a user interface object to a different position in the three-dimensional environment relative to the user's current view (e.g., user interface object 7104-8 is displayed at the bottom left of the user's current view in FIG. 7J ). In some embodiments, the new position of user interface object 7104-8 is within a predetermined zone of a plurality of predetermined zones (e.g., user interface object 7104-8 continues to have delayed-follow behavior as the user moves within the physical environment). For example, the user relocates user interface object 7104-8 from a first predetermined zone to a second predetermined zone. In some embodiments, after the user interface object is moved to the second predetermined zone, the user interface object is moved within the three-dimensional environment such that it remains displayed relative to the user in the second predetermined zone within the user's current view after the user moves within the physical environment (e.g., beyond a threshold amount of movement).

[0133] In some embodiments, in response to a user positioning a user interface object near (e.g., within a threshold distance from) a predetermined zone (e.g., while the zone is highlighted as the user selects and moves the user interface object within the three-dimensional environment), the user interface object 7104 snaps to the predetermined zone (e.g., as the user confirms placing the user interface object within the predetermined zone). For example, in response to a user repositioning the user interface object sufficiently close to the predetermined zone, the computer system automatically displays the user interface object snapped to the predetermined zone (e.g., the user releases the pinch and / or drag gesture, causing the user interface object to drop (e.g., and snap) into place, without requiring the user to perfectly align the user interface object with the predetermined zone). In some embodiments, in response to a user interface object snapping into place within the predetermined zone, the computer system outputs an audio and / or haptic indication.

[0134] In some embodiments or situations, the new position of user interface object 7104-8 is not within a predetermined zone of a plurality of predetermined zones. In some embodiments, if user interface object 7104-8 is not located within a predetermined zone (e.g., the user repositions the user interface object to a position in the three-dimensional environment that does not correspond to a predetermined zone), user interface object 7104-8 does not continue to have delayed-follow behavior as the user moves within the physical environment (e.g., user interface object 7104-8 is fixed within the three-dimensional environment so as to be world-locked as the user moves, instead of changing position to remain within the user's current view).

[0135] In some embodiments, the user is only allowed to reposition the user interface object within a predetermined distance from the user. For example, the user interface object is placed into a position within arm's reach of the user. In some embodiments, the user interface object cannot be placed into a position outside the predetermined distance from the user (e.g., beyond arm's reach from the user). For example, in response to the user repositioning the user interface object into a position within the three-dimensional environment that is farther from the user than a predetermined distance from the user, the computer system provides an error warning to the user (e.g., does not allow the user to place the user interface object into a position farther from the user than a predetermined distance from the user). In some embodiments, in response to the user repositioning the user interface object into a position within the three-dimensional environment that is farther from the user than a predetermined distance from the user, the computer system allows the user to place the object into that position, but provides a warning (e.g., a text display) that the user interface object will not follow the user in the three-dimensional environment when placed into that position (e.g., placing the object into a position farther from the user than a predetermined distance fixes the object in the three-dimensional environment so that the user interface object does not move to maintain the same relative spatial relationship with the user as the user moves within the physical environment).

[0136] 7J , the user can also resize user interface object 7104-8. For example, user input (e.g., a pinch gesture with a first hand) is directed at user interface object 7104-8 (e.g., on the resize affordance of user interface object 7104-8), and user interface object 7104-8 increases and / or decreases in size as the user drags the resize affordance outward from the user interface object (e.g., to enlarge the user interface object) or drags the resize affordance inward toward the center of the user interface object (e.g., to decrease the size of the user interface object). In some embodiments, the user is enabled to perform two-handed gestures (e.g., using both hands to perform the gesture). For example, after selecting a user interface object with a user's first hand (e.g., with a pinch gesture), the user can move the user's other hand toward and / or away from the user's first hand (e.g., pinching the user interface object) to decrease and / or increase the size of the user interface object, respectively. In some embodiments, the input gestures used in various examples and embodiments described herein (e.g., with respect to FIGS. 7A-7J and 8-9 ) optionally include discrete, small movement gestures performed by moving a user's finger(s) relative to other finger(s) or portion(s) of the user's hand, optionally without requiring the user's entire hand or arm to move significantly away from their natural location(s) and posture(s) to perform an action (just before or during the gesture) to interact with a virtual or mixed reality environment, according to some embodiments.

[0137] In some embodiments, the input gesture is detected by analyzing data and signals captured by a sensor system (e.g., sensor 190 of FIG. 1 , image sensor 314 of FIG. 3 ). In some embodiments, the sensor system includes one or more imaging sensors (e.g., one or more cameras: a motion RGB camera, an infrared camera, and / or a depth camera). For example, the one or more imaging sensors are components of, or provide data to, a computer system (e.g., computer system 101 of FIG. 1 (e.g., a portable electronic device or HMD)) that includes a display generation component (e.g., display generation component 120 or 7100 of FIGS. 1 , 3, and 4 (e.g., a touchscreen display, a stereoscopic display, and / or a display with a pass-through portion that functions as both a display and a touch-sensitive surface)). In some embodiments, the one or more imaging sensors include one or more rear-facing cameras on a side of the device opposite the device's display. In some embodiments, the input gesture is detected by a sensor system of a head-mounted system (e.g., a VR headset including a stereoscopic display that provides a left image for the user's left eye and a right image for the user's right eye). For example, one or more cameras that are components of the head-mounted system are mounted on the front and / or bottom of the head-mounted system. In some embodiments, one or more imaging sensors are positioned in the space in which the head-mounted system is used (e.g., arrayed around the head-mounted system at various locations in a room) such that the imaging sensors capture images of the head-mounted system and / or a user of the head-mounted system. In some embodiments, the input gesture is detected by a sensor system of a head-up device (e.g., a head-up display, an automobile windshield capable of displaying graphics, a window capable of displaying graphics, a lens capable of displaying graphics). For example, the one or more imaging sensors are mounted on an interior surface of an automobile. In some embodiments, the sensor system includes one or more depth sensors (e.g., a sensor array).For example, the one or more depth sensors include one or more light-based (e.g., infrared) sensors and / or one or more acoustic-based (e.g., ultrasonic) sensors. In some embodiments, the sensor system includes one or more signal emitters, such as light emitters (e.g., infrared emitters) and / or sound emitters (e.g., ultrasonic emitters). For example, light (e.g., light from an infrared light emitter array having a predetermined pattern) is projected onto a hand (e.g., hand 7102) while images of the hand under the illumination of the light are captured by one or more cameras, and the captured images are analyzed to determine the position and / or configuration of the hand. By determining input gestures using signals from image sensors directed at the hand, as opposed to using signals of a touch-sensitive surface or other direct-contact or proximity-based mechanism, a user can freely choose to perform large movements or remain relatively still when providing input gestures with their hand without experiencing constraints imposed by a particular input device or input area.

[0138] In some embodiments, the tap input optionally indicates a thumb tap input on the index finger of the user's hand (e.g., on the side of the index finger adjacent to the thumb). In some embodiments, the tap input is detected without having to lift the thumb from the side of the index finger. In some embodiments, the tap input is detected according to a determination that a downward movement of the thumb is followed by an upward movement of the thumb and the thumb is in contact with the side of the index finger for less than a threshold time. In some embodiments, the tap hold input is detected according to a determination that the thumb moves from an up position to a touch down position and remains in the touch down position for at least a first threshold time (e.g., a tap time threshold or another time threshold longer than the tap time threshold). In some embodiments, the computer system requires that the entire hand remain substantially stationary in a location for at least a first threshold time to detect a thumb tap hold input with the thumb on the index finger. In some embodiments, the touch hold input is detected without requiring the hand to remain substantially stationary (e.g., the entire hand can move while the thumb is resting on the side of the index finger). In some embodiments, a taphole drag input is detected when the thumb touches the side of the index finger and the whole hand moves while the thumb remains stationary on the side of the index finger.

[0139] In some embodiments, the flick gesture optionally indicates a push or flick input of the thumb moving across the index finger (e.g., from the palm side of the index finger to the back side). In some embodiments, the extension movement of the thumb is accompanied by an upward movement away from the side of the index finger, e.g., as in an upward flick input by the thumb. In some embodiments, the index finger moves in a direction opposite to that of the thumb while the thumb moves forward and upward. In some embodiments, a reverse flick input is performed by the thumb moving from an extended position to a retracted position. In some embodiments, the index finger moves in a direction opposite to that of the thumb while the thumb moves backward and downward.

[0140] In some embodiments, the swipe gesture is a swipe input, optionally by movement of the thumb along the index finger (e.g., along the side of the index finger adjacent to the thumb or along the side of the palm). In some embodiments, the index finger is optionally in an extended state (e.g., substantially straight) or a bent state. In some embodiments, the index finger moves between an extended state and a bent state during movement of the thumb in the swipe input gesture.

[0141] In some embodiments, different phalanges of various fingers correspond to different inputs. Thumb tap inputs across various phalanges of various fingers (e.g., index, middle, ring, and optionally pinky) are optionally mapped to different actions. Similarly, in some embodiments, different push or click inputs can be performed by the thumb across different fingers and / or different portions of the fingers to trigger different actions on individual user interface contacts. Similarly, in some embodiments, different swipe inputs performed by the thumb along different fingers and / or in different directions (e.g., towards the distal or proximal end of the finger) trigger different actions in the respective user interface contexts.

[0142] In some embodiments, the computer system processes tap inputs, flick inputs, and swipe inputs as different types of inputs based on the type of thumb movement. In some embodiments, the computer system processes inputs having different finger locations tapped, touched, or swiped by the thumb as different sub-input types (e.g., proximal, intermediate, distal subtypes, or index, middle, ring, or pinky subtypes) of a given input type (e.g., tap input type, flick input type, and / or swipe input type). In some embodiments, the amount of movement performed by the moving finger (e.g., thumb) and / or other movement measures associated with the finger movement (e.g., velocity, initial velocity, ending velocity, duration, direction, and / or movement pattern) are used to quantitatively affect the action triggered by the finger input.

[0143] In some embodiments, the computer system recognizes combination input types that combine a series of thumb movements, such as a tap-swipe input (e.g., the thumb touching down on another finger and then swiping along the side of the finger), a tap-flick input (e.g., the thumb touching down on another finger and then flicking across the finger from the side of the palm to the back of the finger), and a double-tap input (e.g., two consecutive taps on the side of the finger in approximately the same location).

[0144] In some embodiments, the gesture input is performed with the index finger instead of the thumb (e.g., the index finger performs a tap or swipe on the thumb, or the thumb and index finger move toward each other to perform a pinch gesture). In some embodiments, a wrist movement (e.g., a horizontal or vertical wrist flick) is performed immediately before, immediately after (e.g., within a threshold time), or simultaneously with the finger movement input to trigger an additional, different, or modified action in the current user interface context compared to a finger movement input without the modified wrist movement input. In some embodiments, a finger input gesture performed with a user's palm facing the user's face is treated as a different type of gesture than a finger input gesture performed with a user's palm facing away from the user's face. For example, a tap gesture performed with a user's palm facing the user performs an action with added (or reduced) privacy protection compared to an action (e.g., the same action) performed in response to a tap gesture performed with a user's palm facing away from the user's face.

[0145] While one type of finger input may be used to trigger an action type in the examples provided in this disclosure, in other embodiments, other types of finger input are optionally used to trigger the same type of action.

[0146] Additional explanation regarding FIGS. 7A-7J is provided below with reference to methods 800 and 900 described with respect to FIGS. 8-9 below.

[0147] FIG. 8 is a flowchart of a method 800 for visually de-emphasizing a user interface element in a three-dimensional environment while the user is not attending to the user interface element, according to some embodiments.

[0148] In some embodiments, method 800 is performed on a computer system (e.g., computer system 101 of FIG. 1 ) that includes a first display generation component (e.g., display generation component 120 of FIGS. 1 , 3, and 4 ) (e.g., a head-up display, a display, a touchscreen, and / or a projector) and one or more input devices, such as one or more sensors (e.g., a camera (e.g., color sensors, infrared sensors, and other depth-sensing cameras) placed on a user's hand and facing downward or a camera facing forward from the user's head). In some embodiments, method 800 is governed by instructions stored on a non-transitory computer-readable storage medium and executed by one or more processors of the computer system, such as one or more processors 202 of computer system 101 (e.g., control unit 110 of FIG. 1A ). Some operations of method 800 are, optionally, combined and / or the order of some operations is, optionally, changed.

[0149] In some embodiments, the computer system is in communication with a first display generating component (e.g., the first display generating component is a head-up display, a head-mounted display (HMD), a display, a touchscreen, and / or a projector) and one or more input devices (e.g., a camera, a controller, a touch-sensitive surface, a joystick, a button, a glove, a watch, a motion sensor, and / or an orientation sensor). In some embodiments, the first display generating component is first display generating component 7100 described with respect to FIGS. 7A-7J . In some embodiments, the computer system is an integrated device having one or more processors and memory enclosed in the same housing as the first display generating component and at least some of the one or more input devices. In some embodiments, the computer system includes a computing component (e.g., a server, a mobile electronic device such as a smartphone or tablet device, a wearable device such as a watch, a wristband, or earphones, a desktop computer, and / or a laptop computer) including one or more processors and memory that are separate from the first display generating component and / or the one or more input devices. In some embodiments, the first display generating component and the one or more input devices are integrated within and enclosed in the same housing. Many of the features of method 800, according to some embodiments, are described with respect to Figures 7A-7J.

[0150] Method 800 relates to displaying visually de-emphasized user interface elements when a user is not focusing on the user interface elements. The user interface elements remain de-emphasized as the user moves around in a physical environment, and when the user focuses on the user interface elements, the user interface elements are no longer de-emphasized and are displayed for the user in a position within a three-dimensional environment based on the user's current view of the three-dimensional environment. Automatically de-emphasizing and changing the display location of user interface objects based on whether the user is focusing on the user interface objects and based on the user's current viewpoint provides real-time visual feedback as the user shifts attention to different parts of the three-dimensional environment. Providing improved visual feedback to the user improves system usability and makes the user-system interface more efficient (e.g., by assisting the user in providing appropriate inputs when operating / interacting with the system and reducing user errors), as well as reducing power usage and improving system battery life by allowing the user to use the system more quickly and efficiently.

[0151] The computer system displays a first user interface object in a first view of the three-dimensional environment via a first display generation component (802). In some embodiments, the first user interface object includes one or more user interface objects in a predetermined layout (e.g., user interface object 7104-1 includes one or more user interface objects displayed within user interface object 7104-1).

[0152] While displaying the first user interface object, the computer system detects, via one or more input devices, whether the user meets attention criteria for the first user interface object (e.g., whether the user is looking at the first user interface object, such as by determining whether the user meets gaze detection criteria and / or head position criteria) (804). For example, as described above with reference to Figures 7C and 7D, in some embodiments, the computer system detects whether the user is looking at user interface object 7104-1 (e.g., as indicated by the dashed line from the user's eyes to user interface object 7104-1) or whether the user is not looking at user interface object 7104-2 (e.g., as indicated by the dashed line from the user's eyes to physical object representation 7014').

[0153] In response to detecting that the user does not meet an attention criterion for the first user interface object (e.g., the user is not attending to the first user interface object), the computer system displays (806) the first user interface object with a modified appearance, where displaying the first user interface object with a modified appearance includes de-emphasizing the first user interface object relative to one or more other objects (e.g., real objects or virtual objects) in the three-dimensional environment. For example, as described above with reference to FIG. 7D , the computer system visually de-emphasizes (e.g., reduces opacity and / or increases blur) the first user interface object 7104-2 while the user is not attending to the first user interface object.

[0154] While displaying the first user interface object with the modified appearance, the computer system detects (808) a first movement of the user's viewpoint relative to the physical environment via one or more input devices. For example, as described with reference to FIGS. 7D-7H , the user moves their location within the physical environment. In some embodiments, the user's location (and current viewpoint) optionally includes both the user's position in the physical environment (e.g., in three-dimensional space) and the user's posture / or orientation within the physical environment. In some embodiments, the physical environment corresponds to a three-dimensional environment (e.g., at least a portion of the physical environment is displayed as pass-through content), and changes in the user's orientation and / or position in the physical environment update the user's current view of the three-dimensional environment. In some embodiments, the first movement must satisfy a movement criterion (e.g., the user must move at least a threshold amount from a previous position in the physical environment and / or the user must move the user's torso (e.g., not just the user's head) within the physical environment) before the computer system determines (e.g., makes a new determination) whether the user meets the attention criterion. Optionally, changes in the user's posture and / or orientation in the physical environment may satisfy the movement criteria.

[0155] After detecting a first movement of the user's gaze relative to the physical environment (e.g., following or in response to detecting the first movement of the gaze), the computer system detects (810) that the user meets an attention criterion for the first user interface object (e.g., the user is looking at the first user interface object), as described with reference to FIG. 7H.

[0156] In response to detecting that the user meets the attention criterion, the computer system displays (812) the first user interface object in a second view of the three-dimensional environment that is different from the first view of the three-dimensional environment, where displaying the first user interface object in the second view of the three-dimensional environment includes displaying the first user interface object in an appearance that emphasizes the first user interface object relative to one or more other objects (e.g., real or virtual objects) in the three-dimensional environment more than when the first user interface object was displayed with the modified appearance. For example, as described with reference to FIG. 7H , the first user interface object is displayed in a different position in the three-dimensional environment (e.g., compared to its position in the three-dimensional environment of FIG. 7D ) but continues to have a first spatial relationship to a first anchor position that corresponds to the user's current viewpoint (e.g., location and / or position) in the physical environment. Further, as shown in FIG. 7H , user interface object 7104-6 is no longer visually de-emphasized (as while the user is not attending to the user interface object) in response to the user's viewing of user interface object 7104-6. Thus, in some embodiments, the first user interface object is displayed to follow the user as the user moves through the physical environment.

[0157] In some embodiments, the first user interface object has a first spatial relationship to a first anchor position in the three-dimensional environment that corresponds to the location of the user's body in the physical environment. For example, the first user interface object is maintained (e.g., or locked) in the same general location relative to the user's torso, hand, head, or other part of the user's body. In some embodiments, the first spatial relationship is maintained across movements of the user's viewpoint. For example, as described with reference to FIGS. 7C and 7H , the first spatial relationship between the user's viewpoint and instances of user interface objects 7104-1 and 7104-6 is maintained across movements of the user within the physical environment. Automatically displaying a particular user interface object in a position that is maintained (e.g., or locked) in the same general location relative to a part of the user's body, even as the user's viewpoint changes (e.g., by changing the user's current viewpoint as the user moves around the physical environment), provides real-time visual feedback as the user moves around the physical environment, thereby providing the user with improved visual feedback.

[0158] In some embodiments, after detecting a first movement of the user's viewpoint relative to the physical environment, the computer system maintains the display of the first user interface object at the same anchor position within the three-dimensional environment (e.g., until a time threshold is met). In some embodiments, the first user interface object is initially maintained at the same location within the three-dimensional environment while the user's viewpoint moves. In some embodiments, once the user's viewpoint moves beyond a threshold distance and / or once the user's viewpoint moves for a threshold time (e.g., the user moves and stops but does not return to the initial viewpoint), the user interface object moves within the three-dimensional environment, as described with reference to FIGS. 7E-7H . In some embodiments, the rate of movement of the user interface object is slower than the rate of movement of the user's viewpoint. Automatically displaying a particular user interface object in a position that is maintained (e.g., or locked) in the same general location relative to the three-dimensional environment as the user's viewpoint changes (e.g., by changing the user's current viewpoint as the user moves around the physical environment) provides real-time visual feedback as the user moves around the physical environment, thereby providing the user with improved visual feedback.

[0159] In some embodiments, after detecting a first movement of a user's viewpoint relative to the physical environment, pursuant to a determination that the first movement satisfies a time threshold (e.g., the user moves for at least the time threshold, or the user moves and remains at a second viewpoint for at least the time threshold), the computer system moves a first user interface object within the three-dimensional environment to the same position relative to the user's viewpoint (e.g., the same position as before the movement of the user's viewpoint), as described above with reference to FIG. 7H. Automatically displaying a particular user interface object as moving within the three-dimensional environment after the user has moved within the physical environment for longer than a predetermined amount of time (e.g., and / or remained at a different position within the three-dimensional environment for a predetermined amount of time) provides real-time visual feedback as the user moves within the physical environment, thereby providing improved visual feedback to the user.

[0160] In some embodiments, the computer system receives a user input to reposition (e.g., fix) the first user interface object in the three-dimensional environment. In some embodiments, in response to receiving the input to reposition the first user interface object in the three-dimensional environment, the computer system repositions the first user interface object to a respective position in the three-dimensional environment in accordance with the input, for example, as described above with reference to FIG. In some embodiments, after relocating the first user interface object to the respective position within the three-dimensional environment in accordance with the input, the computer system detects an input to change the user's viewpoint, and in response to detecting the input to change the user's viewpoint, the computer system changes the user's viewpoint in accordance with the input to change the user's viewpoint, and displaying the first user interface object in the three-dimensional environment from the user's current viewpoint includes: in accordance with a determination that the first user interface object is located within a first predetermined zone, displaying the first user interface object at a respective position having a first spatial relationship to a first anchor position in the three-dimensional environment that corresponds to a location of the user's viewpoint in the physical environment (e.g., displaying the first user interface object at the respective position after detecting movement of the user's viewpoint relative to the physical environment while the first user interface object is located within the first predetermined zone); and in accordance with a determination that the first user interface object is not located within the first predetermined zone (e.g., any predetermined zone of a plurality of predetermined zones), maintaining the display of the first user interface object at the same anchor position in the three-dimensional environment that does not correspond to a location of the user's body in the physical environment. For example, as described above with reference to Figures 7C-7H, the first predetermined zone is the zone that follows the user's point of view as the user moves through the physical environment.In some embodiments, the first predefined zone follows the user's viewpoint with a delay (e.g., the first predefined zone moves at a speed slower than the user's movement speed). For example, the first predefined zone does not initially move with the user's viewpoint until the user's viewpoint has moved a threshold amount (e.g., at least the threshold amount) and / or until the user's viewpoint has moved for a threshold amount of time (e.g., at least the threshold amount of time). In some embodiments, the first predefined zone is referred to herein as a delayed-following zone. The first predefined zone allows a user to anchor specific user interface objects to the zone, and user interface objects placed within the zone automatically follow the user in the three-dimensional environment even as the user moves around the physical environment, by distinguishing user interface objects placed within the zone from user interface objects placed outside the zone that would otherwise be anchored in the three-dimensional environment so as not to automatically follow the user in the three-dimensional environment, thereby providing improved visual feedback to the user.

[0161] In some embodiments, a first predetermined zone is selected from a plurality of predetermined zones in the three-dimensional environment, the first predetermined zone having a first spatial relationship to a first anchor position in the three-dimensional environment corresponding to the location of the user's viewpoint, and a second predetermined zone of the plurality of predetermined zones having a second spatial relationship (e.g., different from the first spatial relationship of the first predetermined zone) to a second anchor position in the three-dimensional environment corresponding to the location of the user's viewpoint. For example, as described with reference to FIG. 7J , a user can reposition user interface elements into any of a plurality of predetermined zones, each predetermined zone having a different relative spatial arrangement with respect to the user's current viewpoint. In some embodiments, user interface objects positioned within a predetermined zone follow the movement of the user's viewpoint (e.g., with delayed follow behavior). In some embodiments, the first predetermined zone is displayed at a first position within the user's current view of the three-dimensional environment relative to the user, the first position being maintained before and after the user's movement, and the second predetermined zone is displayed at a second position within the user's current view of the three-dimensional environment relative to the user, the second position being maintained before and after the user's movement. Providing the user with the option to change the anchor of a particular user interface object located within any of multiple zones within the three-dimensional environment, each zone having a different relative position to the user's current view, facilitates placing the particular user interface object in a position that is most comfortable or convenient for the user to view, and provides real-time visual feedback to the user as the user selects where to place the user interface object and as the user moves within the physical environment, thereby providing improved visual feedback to the user.

[0162] In some embodiments, in response to detecting that a user has initiated user input to reposition the first user interface object, the computer system displays a first predetermined zone and maintains displaying the visual indication for the first predetermined zone while detecting the user input to reposition the first user interface object in the first predetermined zone. In some embodiments, the computer system optionally stops displaying the visual indication for the first predetermined zone in response to no longer detecting the user input to reposition the first user interface object (e.g., in response to termination of the user input). In some embodiments, the computer system displays the visual indication for each of the predetermined zones among the plurality of predetermined zones. In some embodiments, the computer system provides an outline of the delayed following zone(s) so that the user knows where to place (e.g., drag and drop or prop) the first user interface object such that the first user interface object has delayed following behavior (e.g., the first user interface object follows the user's viewpoint (e.g., with a delay) while the first user interface object is within the delayed following zone). For example, as described with reference to Figure 7J, the computer system optionally provides contours of predetermined zones. Automatically displaying contours of multiple zones in the three-dimensional environment, each zone having a different relative position to the user's current view, makes it easier for the user to select a zone to anchor a particular user interface object, and when the user interface object is placed in a zone, the user interface object follows the user as the user moves in the physical environment, allowing the user to more easily determine where in the three-dimensional environment the user interface object should be placed that is most comfortable or convenient for the user to view, thereby providing the user with improved visual feedback.

[0163] In some embodiments, the attention criteria for the first user interface object include gaze criteria. For example, the computer system uses one or more cameras and / or other sensors to determine whether the user is gazing at (e.g., looking at and / or looking at) the first user interface object, as described above with reference to FIG. 7C. Automatically determining whether a user is gazing at a particular user interface object by detecting whether the user is gazing at the user interface object, and automatically updating the display of the user interface object without requiring additional input from the user when the user is gazing at the user interface object, provides the user with additional control without requiring the user to navigate complex menu hierarchies, thereby providing the user with improved visual feedback without requiring additional user input.

[0164] In some embodiments, the attention criteria with respect to the first user interface object include criteria with respect to the position of the user's head in the physical environment. For example, as the user's head moves (e.g., in a pose, orientation, and / or position) within the physical environment, the computer system determines whether the user's head is in a particular pose, orientation, and / or position relative to the first user interface object, as described above with reference to FIG. 7C . Automatically determining whether the user is attending to a user interface object by detecting whether the user's head is in a particular position without requiring additional input from the user, and automatically updating the display of the user interface object when the user's head is in the particular position, provides the user with additional control without requiring the user to navigate complex menu hierarchies, thereby providing the user with improved visual feedback without requiring additional user input.

[0165] In some embodiments, the computer system detects (e.g., receives) a pinch input (e.g., a pinch input including a movement of two or more fingers of a hand to contact or separate from each other) directed toward a first affordance displayed on at least a portion of the first user interface object, followed by a hand movement that performed the pinch input, and, in response to the hand movement, resizes the first user interface object in accordance with the hand movement. For example, a user may resize the first user interface object as described with reference to FIG. 7J . In some embodiments, the hand movement that performed the pinch input is a drag gesture (e.g., a pinch-and-drag gesture is performed with one hand). For example, the pinch gesture includes a movement of two or more fingers of a hand that contact or separate from each other in conjunction with (e.g., followed by) a drag input that changes the position of the user's hand from a first position (e.g., a start position of the drag) to a second position (e.g., an end position of the drag). In some embodiments, the user maintains the pinch gesture while performing the drag input and releases the pinch gesture (e.g., spreading two or more fingers apart) to end the drag gesture (e.g., at a second position). For example, the pinch input selects a first affordance, and once selected, the user can drag the affordance away from a central portion of the first user interface object (e.g., to increase the size of the first user interface object) and / or drag the affordance toward a central portion of the first user interface object (e.g., to decrease the size of the first user interface object). In some embodiments, the first affordance is displayed at a corner of the first user interface object (e.g., the first affordance is a resize affordance).In some embodiments, a user pinches a first affordance (e.g., to select the affordance) and then, with the same hand that is pinching the first affordance, drags a corner of a first user interface object from a first position to a second position within the three-dimensional environment. For example, the user drags the affordance outward to enlarge the first user interface object or drags a corner of the first user interface object inward to decrease the size of the first user interface object. In some embodiments, the hand movement includes the user changing the distance between two or more fingers performing a pinch input. For example, the user moves the user's thumb and index finger closer together to decrease the size of the user interface object, and opens the pinch gesture (e.g., increases the distance between two fingers (e.g., the user's thumb and index finger)) to increase the size of the user interface object. In another example, the user's whole hand movement increases or decreases the size of the user interface object depending on the direction of the user's hand movement. In some embodiments, the pinch input is directly directed toward the first affordance (e.g., the user performs the pinch input at a position corresponding to the first affordance), or the pinch input is indirectly directed toward the first affordance (e.g., the user performs the pinch input while gazing at the first affordance, and the position of the user's hands while performing the pinch input is not a position corresponding to the first affordance). For example, the user can direct the user's input toward the first affordance by initiating a gesture at or near the first affordance (e.g., within a distance of 0.5 cm, 1 cm, 5 cm, or 0-5 cm, measured from the outer edge of the first affordance or the central portion of the first affordance).In some embodiments, the user can further direct the user's input toward the first affordance by paying attention to the first affordance (e.g., gazing at the first affordance), and while focusing on the first affordance, the user initiates the gesture (e.g., at any position detectable by the computer system). For example, if the user is focusing on the first affordance, the gesture does not need to be initiated at or near the first affordance. Enabling a user to perform a pinch input, such as by pinching a corner of a user interface object or performing a pinch input while gazing at a user interface object, and automatically updating the size of a user interface object by dragging the user's hand with the same hand against the user interface object (e.g., while continuing to perform the pinch input) to increase or decrease the size of the object provides the user with additional control without requiring the user to navigate complex menu hierarchies, and allows the user to intuitively resize the user interface object by selecting the user interface object (e.g., using a pinch input) and dragging the user's hand to different positions in the three-dimensional environment to change the size accordingly, thereby providing the user with improved visual feedback.

[0166] In some embodiments, the computer system receives (e.g., detects) a pinch input by a first hand, a pinch input by a second hand directed toward a first user interface object, and a subsequent change in distance between the first hand and the second hand. In some embodiments, in response to the change in distance between the first hand and the second hand, the computer system changes the size of the first user interface object according to the change in distance between the first hand and the second hand. For example, a user performs an input that is a two-handed gesture (e.g., a pinch-and-drag input). The user pinches a user interface object with the user's first hand, performs a pinch input with the user's second hand (e.g., while maintaining the pinch input with the user's first hand), moves the user's second hand closer to the user's first hand (e.g., using a drag input) to decrease the size of the user interface object, and moves the user's second hand away from the user's first hand to increase the size of the user interface object. In some embodiments, the pinch-and-drag input also moves the first user interface object within the three-dimensional environment (e.g., when the first user interface object is dragged to different predetermined zones, as described with reference to FIG. 7J). In some embodiments, the pinch input performed with the first hand is directly directed at the first user interface object (e.g., the user performs the pinch input at a position corresponding to the first user interface object), or the pinch input is indirectly directed at the first user interface object (e.g., the user performs the pinch input while gazing at the first user interface object, and the position of the user's hand while performing the pinch input is not a position corresponding to the first user interface object).For example, a user can direct the user's input to a first user interface object by initiating a gesture at or near the first user interface object (e.g., within a distance of 0.5 cm, 1 cm, 5 cm, or 0-5 cm, measured from the outer edge of the first user interface object or the central portion of the first user interface object). In some embodiments, the user can further direct the user's input to the first user interface object by paying attention to the first user interface object (e.g., gazing at the first user interface object), and while focusing on the first user interface object, the user initiates the gesture (e.g., in any position detectable by the computer system). For example, if the user is focusing on the first user interface object, the gesture need not be initiated at or near the first user interface object. In some embodiments, a pinch input by a second hand can also be detected directly or indirectly (e.g., the pinch input can be initiated at a position on or near the first user interface object, or at any position while the user focuses on the first user interface object). In some embodiments, after a pinch input by a first hand is detected, a pinch input by a second hand is detected at any position while the first pinch input is maintained. For example, while the first user interface object is selected (e.g., using a pinch input by the first hand), a pinch input by the second hand is detected at any position (e.g., regardless of whether the user focuses on the first user interface object). In some embodiments, the user simultaneously resizes and repositions (e.g., moves) the first user interface object in a three-dimensional environment (e.g., using a combination of gestures).For example, a user may provide the pinch input described above to resize a first user interface object while providing a drag input (e.g., with the user's second hand) to reposition the user interface object (e.g., by dragging the user interface object to another position within the three-dimensional environment). Enabling a user to use both hands, with each hand automatically updating the size of the user, each hand selecting (e.g., pinching) a portion such as a corner of the user interface object, and resizing the user interface object based on changes in the distance between the user's hands while the portion of the user interface object is selected, provides the user with additional control, such that the user can intuitively resize the user interface object to increase its size by increasing the distance between the user's hands or decrease its size by decreasing the distance between the user's hands, thereby providing the user with improved visual feedback without requiring the user to navigate complex menu hierarchies.

[0167] In some embodiments, suppressing the first user interface object relative to one or more other objects in the three-dimensional environment includes suppressing the first user interface object relative to one or more other virtual objects in the three-dimensional environment. For example, as described with reference to FIG. 7D , the first user interface object is visually suppressed while one or more other virtual objects, including an application (e.g., an application window, an application object), a user interface object (e.g., an affordance and / or control), a virtual environment (e.g., an immersive experience), etc., are not visually suppressed (e.g., remain unchanged). Automatically updating the display of the particular user interface object by visually suppressing the particular user interface object relative to other displayed virtual content when the user is not focusing on the particular user interface object provides real-time visual feedback when the user is focusing on different virtual content in the three-dimensional environment, thereby providing the user with improved visual feedback.

[0168] In some embodiments, suppressing the first user interface object relative to one or more other objects in the three-dimensional environment includes suppressing the first user interface object relative to representations of one or more physical objects in the physical environment. For example, as described with reference to FIG. 7D , one or more physical objects in the physical environment are displayed in the three-dimensional environment as pass-through content (e.g., representations 7014′ of the physical objects) without visual suppression, while the first user interface object 7104-2 is visually suppressed. Automatically updating the display of the particular user interface object by visually suppressing the particular user interface object relative to other real-world content from the physical environment displayed in the three-dimensional environment when the user is not focusing on the particular user interface object provides real-time visual feedback when the user is focusing on real and / or virtual content displayed in the three-dimensional environment, thereby providing the user with improved visual feedback.

[0169] In some embodiments, the first user interface object includes a plurality of selectable user interface objects. For example, as described with reference to FIG. 7I, a user can interact with one or more selectable user interface objects displayed within user interface object 7104-7. In some embodiments, the selectable user interface object is an affordance selectable using a gaze and / or air gesture. In some embodiments, the first user interface object includes a panel (e.g., a menu) having a plurality of selectable objects. In some embodiments, the first selectable user interface object from the plurality of selectable user interface objects is an application icon, and in response to a user selecting the application icon, the computer system opens (e.g., launches) an application window for the application corresponding to the application icon. In some embodiments, the first selectable user interface object from the plurality of selectable user interface objects is a control for adjusting a setting (e.g., volume level, brightness level, and / or immersion level) of the three-dimensional environment, and in response to a user selecting the control for adjusting a setting of the three-dimensional environment, the computer system adjusts the setting according to the user selection. Automatically displaying multiple controls that the user can select by gazing at and / or performing gestures directed at the multiple controls provides the user with additional controls that are easily accessed by the user within the displayed user interface object (e.g., that follow the user as they move within their physical environment) without requiring the user to navigate complex menu hierarchies, thereby providing the user with improved visual feedback without requiring additional user input.

[0170] In some embodiments, the computer system displays one or more user interface objects for controlling the immersion level of the three-dimensional environment, and, in response to a user input directed to a first user interface object of the one or more user interface objects for increasing the immersion level of the three-dimensional environment, displays additional virtual content in the three-dimensional environment (e.g., optionally stops displaying pass-through content). In some embodiments, in response to detecting a user input directed to a second user interface object of the one or more user interface objects for decreasing the immersion level of the three-dimensional environment, the computer system displays additional content corresponding to the physical environment (e.g., optionally stops displaying virtual content (e.g., one or more virtual objects)). For example, a user can control how much of the physical environment is displayed as pass-through content in the three-dimensional environment (e.g., none of the physical environment is displayed (e.g., represented in the three-dimensional environment) during a fully immersive experience). In some embodiments, one or more user interface objects described with reference to FIG. 7I include controls for playing and / or pausing the immersive experience in the three-dimensional environment. In some embodiments, one or more user interface objects are displayed within the first user interface object (e.g., within a delayed following zone) such that the one or more user interface objects follow the user as the user moves around the physical environment. Automatically displaying multiple controls that allow the user to control the immersive experience of the three-dimensional environment relative to the physical environment provides additional control for the user without requiring the user to navigate complex menu hierarchies, thereby allowing the user to easily control how much content from the physical environment is displayed in the three-dimensional environment, and provides the user with real-time visual feedback when the user requests to change the level of immersion in the three-dimensional environment, thereby providing the user with improved visual feedback without requiring additional user input.

[0171] In some embodiments, the amount of de-emphasis of the first user interface object relative to one or more other objects in the three-dimensional environment is based (e.g., at least in part) on the angle between the user's detected line of sight and the first user interface object. For example, the angle is defined as "0" when the user's line of sight is directly in front of the user, and as the user's line of sight moves (left, right, up, or down) relative to the first user interface object, the angle increases as the user's line of sight moves away from the first user interface object. For example, the first user interface object becomes more faded / blurred as the user's line of sight is tracked away from the first user interface object. In some embodiments, the amount of de-emphasis of the first user interface object is proportional (e.g., linearly or otherwise) to the amount of change in the angle between the user's line of sight and the first user interface object (e.g., de-emphasis increases as the user's line of sight increases in angle). Automatically updating the display of a particular user interface object by visually de-emphasizing the particular user interface object by varying amounts based on the perceived angle between the user's current view and the user interface object, such that the user interface object appears more out of focus as the user's current view angle moves farther away from the user interface object, provides real-time visual feedback as the user's current view of the three-dimensional environment changes, and provides the user with greater awareness of the user's movement relative to the user interface object, thereby providing the user with improved visual feedback.

[0172] In some embodiments, the amount of de-emphasis of the first user interface object relative to one or more other objects in the three-dimensional environment is based (e.g., at least in part) on the speed of the first movement of the user's gaze. For example, the faster the user's gaze moves (e.g., the faster the head is turned), the greater the de-emphasis of the first user interface object. In some embodiments, the amount of de-emphasis of the first user interface object is proportional (e.g., linear or non-linear) to the speed and / or direction of the movement of the user's gaze (e.g., faster movement causes more de-emphasis, and slower movement causes less de-emphasis, as described above with reference to FIGS. 7D-7G). Automatically updating the display of a particular user interface object by visually de-emphasizing the particular user interface object based on how fast the user is moving in the physical environment, such that the faster the user moves in the physical environment, the more faded the object appears, provides real-time visual feedback as the user moves at different speeds in the three-dimensional environment, thereby providing improved visual feedback to the user.

[0173] In some embodiments, the first user interface object moves within the three-dimensional environment according to the user's movements. For example, the first user interface object is fixed in a position relative to the user's viewpoint so that it appears in the same position relative to the user's viewpoint as the user moves, as described with reference to Figures 7C and 7H. Automatically moving a particular user interface object within the three-dimensional environment to follow the user's current viewpoint while the user moves within the physical environment, while maintaining the same relative spatial relationship between the physical environment and the user's viewpoint, provides real-time visual feedback to the user as they move within the physical environment, and displays the user interface object in a convenient position so that the user can see and interact with the user interface object even as they move within the physical environment, thereby providing improved visual feedback to the user.

[0174] In some embodiments, immediately before and after a first movement of the user's viewpoint relative to the physical environment, a distinct, distinctive position of a first user interface object in the three-dimensional environment has a first spatial relationship with respect to a first anchor position in the three-dimensional environment that corresponds to the location of the user's viewpoint in the physical environment. For example, as described with reference to FIG. 7C , user interface object 7104-1 is displayed (in FIGS. 7E-7G ) at a distinct, distinctive position (e.g., relative to the user's viewpoint) before the user moves, and is redisplayed in the same distinct, distinctive position as user interface object 7104-6 in FIG. 7H (e.g., after the user stops moving in the physical environment). Automatically maintaining a particular user interface object in the three-dimensional environment at the same position relative to the user's current viewpoint as the user moves and changes the user's viewpoint in the physical environment provides real-time visual feedback to the user as the user moves within the physical environment, such that the user can see and interact with the user interface object as the user moves within the physical environment, thereby providing improved visual feedback to the user.

[0175] It should be understood that the particular order described of the operations in FIG. 8 is merely an example, and that the described order is not intended to indicate the only order in which the operations may be performed. Those skilled in the art will recognize various ways to reorder the operations described herein. It should also be noted that other process details described herein with respect to other methods described herein (e.g., method 900) are also applicable in a similar manner to method 800 described above with respect to FIG. 8. For example, the gestures, inputs, physical objects, user interface objects, movements, references, three-dimensional environments, display generating components, representations of physical objects, virtual objects, and / or animations described above with reference to method 900 optionally have one or more of the characteristics of the gestures, inputs, physical objects, user interface objects, movements, references, three-dimensional environments, display generating components, representations of physical objects, virtual objects, and / or animations described herein with reference to other methods described herein (e.g., method 800). For the sake of brevity, those details will not be repeated here.

[0176] FIG. 9 is a flowchart of a method 900 for updating the display of user interface elements in a three-dimensional environment to follow a user as the user changes their current view of the three-dimensional environment, according to some embodiments.

[0177] In some embodiments, method 900 is performed on a computer system (e.g., computer system 101 of FIG. 1 ) that includes a first display generation component (e.g., display generation component 120 of FIGS. 1 , 3, and 4 ) (e.g., a head-up display, a display, a touchscreen, and / or a projector) and one or more input devices, such as one or more sensors (e.g., a camera (e.g., color sensors, infrared sensors, and other depth-sensing cameras) placed on a user's hand and facing downward or a camera facing forward from the user's head). In some embodiments, method 900 is governed by instructions stored on a non-transitory computer-readable storage medium and executed by one or more processors of the computer system, such as one or more processors 202 of computer system 101 (e.g., control unit 110 of FIG. 1A ). Some operations of method 900 are, optionally, combined and / or the order of some operations is, optionally, changed.

[0178] In some embodiments, the computer system is in communication with a first display generating component (e.g., the first display generating component is a head-up display, a head-mounted display (HMD), a display, a touchscreen, and / or a projector) and one or more input devices (e.g., a camera, a controller, a touch-sensitive surface, a joystick, a button, a glove, a watch, a motion sensor, and / or an orientation sensor). In some embodiments, the first display generating component is first display generating component 7100 described with respect to FIGS. 7A-7J . In some embodiments, the computer system is an integrated device having one or more processors and memory enclosed in the same housing as the first display generating component and at least some of the one or more input devices. In some embodiments, the computer system includes a computing component (e.g., a server, a mobile electronic device such as a smartphone or tablet device, a wearable device such as a watch, a wristband, or earphones, a desktop computer, and / or a laptop computer) including one or more processors and memory that are separate from the first display generating component and / or the one or more input devices. In some embodiments, the first display generating component and the one or more input devices are integrated within and enclosed in the same housing. Many of the features of method 900, according to some embodiments, are described with respect to Figures 7A-7J.

[0179] Method 900 relates to displaying user interface elements within a three-dimensional environment, where the display of the user interface elements is updated to follow the user as the user changes their current view of the three-dimensional environment (e.g., by moving around the physical environment). The user interface elements initially do not move as the user's view changes until the user's view changes by more than a threshold amount. After the user's view changes by more than the threshold amount, the user interface elements follow the user (e.g., delaying to follow the user at a slower rate of movement than the user's movement). Automatically changing the display location of user interface objects to follow the user as their current viewpoint changes from the user moving within the physical environment provides real-time visual feedback as the user moves within the physical environment. Providing improved visual feedback to the user improves system usability and makes the user-system interface more efficient (e.g., by assisting the user in providing appropriate inputs when operating / interacting with the system and reducing user errors), as well as reducing power usage and improving the system's battery life by allowing the user to use the system more quickly and efficiently.

[0180] The computer system, via the first display generation component, displays a first user interface object in a first view of the three-dimensional environment at a first position in the three-dimensional environment and at a first spatial arrangement relative to a distinct portion of the user (e.g., relative to the user's current viewpoint of the three-dimensional environment) (902). For example, as described with reference to FIG. 7C, user interface object 7104-1 is initially displayed at a first spatial arrangement relative to the user's current position in the physical environment.

[0181] While displaying the first user interface object, the computer system detects (904) movement of the user's viewpoint from a first location to a second location in the physical environment via one or more input devices. For example, the user's location optionally includes information related to the user's three-dimensional position (e.g., coordinates) in the physical environment, as well as information related to the user's posture and / or orientation in the physical environment. In some embodiments, detecting movement of the user's viewpoint includes detecting torso movement within the physical environment (e.g., the first user interface object is maintained in the same general position relative to the user's body). For example, as described with reference to Figures 7E-7G, the user moves within the physical environment, which changes the user's current view of the three-dimensional environment.

[0182] In response to detecting movement of the user's viewpoint from a first location to a second location (906), following a determination that the movement of the user's viewpoint from the first location to the second location does not satisfy a threshold amount of movement (e.g., a threshold amount of change in the user's angle (orientation) and / or a threshold amount of distance), the computer system maintains the display of the first user interface object in a first position within the three-dimensional environment (908) (e.g., even though the first user interface no longer has the first spatial arrangement relative to individual portions of the user). For example, as shown in FIG. 7E, when the user initially moves within the physical environment (e.g., does not satisfy the threshold amount of movement), user interface object 7104-3 remains fixed (e.g., locked) in the same position within the three-dimensional environment.

[0183] In response to detecting movement of the user's viewpoint from a first location to a second location (906) and determining that the movement of the user's viewpoint from the first location to the second location satisfies a threshold amount of movement (910), the computer system stops displaying the first user interface object at a first position within the three-dimensional environment (912) and displays the first user interface object at a second position within the three-dimensional environment, the second position within the three-dimensional environment having a first spatial arrangement relative to a distinct portion of the user (914). For example, the position of the user interface object changes within the three-dimensional environment (e.g., the user interface object is not fixed to a position within the three-dimensional environment) but remains in the same relative position relative to the user after the user has moved at least a threshold amount. For example, as described with reference to Figures 7C and 7H, the default position of the first user interface object (e.g., shown in Figure 7C) relative to the user is restored after the threshold amount of movement is satisfied in Figure 7H.

[0184] In some embodiments, while maintaining the display of a first user interface object at a first position within the three-dimensional environment (e.g., the first position is a fixed position within the three-dimensional environment), the computer system detects that movement of the user's viewpoint satisfies a threshold amount of movement. In some embodiments, in response to the movement of the user's viewpoint satisfying the threshold amount of movement, the computer system moves the first user interface object from the first position to a second position in the three-dimensional environment (e.g., animates the movement). For example, as described with reference to FIG. 7D , initially, the user moves less than the threshold amount of movement, and the user interface object is maintained in the same fixed first position in the physical environment before the user satisfied the threshold amount of movement (e.g., as shown in FIG. 7H ). Then (e.g., pursuant to a determination that the amount of movement satisfies the threshold amount of movement), the computer system moves the user interface object to the second position within the three-dimensional environment. Automatically displaying a user interface object in an initially locked position relative to the three-dimensional environment as the user moves around the physical environment before displaying a user interface object that moves with the user after the user has moved more than a threshold amount provides real-time visual feedback such that the user interface object remains less than a threshold distance from the user as the user moves around the physical environment (e.g., because the user interface object follows the user after the user has moved at least a threshold amount away from the user interface object's initial position and / or after the user's viewpoint has changed a threshold amount), thereby providing improved visual feedback to the user.

[0185] In some embodiments, while displaying the first user interface object at a second position within the three-dimensional environment, the computer system detects, via one or more input devices, movement of the user's viewpoint from a second location to a third location within the physical environment. In some embodiments, in response to detecting movement of the user's viewpoint from the second location to the third location, and in accordance with a determination that the movement of the user's viewpoint from the second location to the third location does not satisfy a second threshold amount of movement (e.g., the same and / or different threshold amount of movement as movement from the first location to the second location), the computer system maintains the display of the first user interface object at the second position within the three-dimensional environment (e.g., even though the first user interface object no longer has the first spatial arrangement relative to distinct portions of the user). In some embodiments, following a determination that movement of the user's viewpoint from the second location to the third location satisfies the second movement threshold amount, the computer system stops displaying the first user interface object at the second position within the three-dimensional environment and displays the first user interface object at a third position within the three-dimensional environment, the third position within the three-dimensional environment having the first spatial configuration relative to a discrete portion (e.g., third location) of the user (e.g., third location). For example, the user interface object continues to have the delayed-follow behavior described with reference to FIGS. 7D-7H even if the user subsequently moves (e.g., continues to move continuously or sporadically) within the physical environment. Automatically changing the display location of the user interface object to follow the user as their current viewpoint changes in accordance with the user moving within the physical environment provides real-time visual feedback as the user moves within the physical environment.

[0186] In some embodiments, while moving a first user interface object from a first position to a second position within the three-dimensional environment, the computer system visually suppresses the first user interface object relative to one or more other objects within the three-dimensional environment. For example, as shown in FIGS. 7E-7G, user interface objects 7104-3 through 7104-5 are visually suppressed in the three-dimensional environment while the user is moving. Automatically updating the display of a particular user interface object by visually suppressing the particular user interface object relative to other displayed content while the user is moving around the physical environment provides real-time visual feedback as the user is moving around within the three-dimensional environment, and by de-emphasizing the user interface object while the user is moving (and not interacting with the user interface object), reduces the amount of visual load on the user (e.g., the user's level of distraction), thereby providing improved visual feedback to the user.

[0187] In some embodiments, visually de-emphasizing the first user interface object includes displaying the first user interface object with reduced opacity, as described above with reference to Figures 7D-7E. For example, while the first user interface object is visually de-emphasized (e.g., while the user is moving), the first user interface object appears more translucent (e.g., fades) relative to other objects displayed in the three-dimensional environment. Automatically updating the display of a particular user interface object by reducing the opacity of the particular user interface object relative to other displayed content while the user is moving around the physical environment provides real-time visual feedback as the user is moving around in the three-dimensional environment, reducing the amount of visual load on the user (e.g., the user's level of distraction) by unobtrusively displaying the user interface object while the user is moving (e.g., not interacting with the user interface object), thereby providing improved visual feedback to the user.

[0188] In some embodiments, visually de-emphasizing the first user interface object includes displaying the first user interface object with a blurred visual effect, as described above with reference to Figures 7D-7E. For example, the first user interface object appears blurred relative to other objects displayed in the three-dimensional environment (e.g., while the user is moving). Automatically updating the display of a particular user interface object by blurring the particular user interface object relative to other displayed content while the user is moving around the physical environment provides real-time visual feedback as the user is moving around within the three-dimensional environment, and by displaying the de-emphasized user interface object while the user is moving (e.g., not interacting with the user interface object), reduces the amount of visual load on the user (e.g., the user's level of distraction), thereby providing improved visual feedback to the user.

[0189] In some embodiments, pursuant to a determination that the user does not satisfy the attention criterion for the first user interface object, the computer system visually de-emphasizes the first user interface object relative to one or more other objects within the three-dimensional environment, as described with reference to Figure 7D. Automatically updating the display of the particular user interface object by visually de-emphasizing the particular user interface object relative to other displayed content when the user is not attending to the particular user interface object provides real-time visual feedback when the user is attending to different virtual content within the three-dimensional environment, and reduces the amount of visual load (e.g., or level of distraction) for the user by displaying the user interface object in a less obtrusive manner relative to other content while the user is not attending to the user interface object, thereby providing improved visual feedback to the user.

[0190] In some embodiments, the attention criterion for the first user interface object includes a gaze criterion, as described above with reference to FIG. 7D . For example, the user satisfies the gaze criterion according to a determination that the user has gazed at (e.g., looked at) the first user interface object for at least a threshold amount of time. Automatically determining whether a user is gazed at a particular user interface object by detecting whether the user is gazed at the user interface object, and automatically updating the display of the user interface object without requiring additional input from the user when the user is gazed at the user interface object, provides the user with additional control without requiring the user to navigate complex menu hierarchies, thereby providing the user with improved visual feedback without requiring additional user input.

[0191] In some embodiments, the attention criterion for the first user interface object includes a criterion related to the position of the user's head in the physical environment, as described above with reference to FIG. 7D . For example, the criterion related to the user's head position is satisfied according to a determination that the user's current head position matches a predetermined head position and / or a determination that the user's head position remains in the predetermined head position for at least a threshold amount of time. Automatically determining whether a user is attending to a user interface object by detecting whether the user's head is in a particular position without requiring additional input from the user, and automatically updating the display of the user interface object when the user's head is in the particular position, provides the user with additional control without requiring the user to navigate complex menu hierarchies, thereby providing the user with improved visual feedback without requiring additional user input.

[0192] In some embodiments, the amount of visual de-emphasis of the first user interface object is based (e.g., at least in part) on the speed of movement of the user's gaze. For example, the first user interface object appears to fade (e.g., be displayed with reduced opacity) and / or appear blurrier as the user moves faster in the physical environment. In some embodiments, the visual de-emphasis is gradual (e.g., the first user interface object appears more faded over a period of time), such that the amount of visual de-emphasis increases as the user moves over a longer period of time. In some embodiments, the rate of the amount of visual de-emphasis during the gradual de-emphasis is based on the speed of movement of the user's gaze (e.g., the rate of the amount of visual de-emphasis is proportional to the speed of movement of the user). For example, an increase in the speed of movement of the user's gaze results in an increase in the amount of fading and / or blurring. Automatically updating the display of a particular user interface object by visually de-emphasizing the particular user interface object relative to other displayed content, with an amount of visual de-emphasis determined based on the speed of the user's movement, such that the faster the user moves while the user moves through the physical environment, the more de-emphasized the particular user interface object appears to be, provides real-time visual feedback as the user moves through the three-dimensional environment at different speeds, thereby providing improved visual feedback to the user.

[0193] In some embodiments, while displaying the first user interface object, the computer system displays a second user interface object (e.g., a virtual object, an application, or a representation of a physical object, such as representation 7014′ of the physical object) at a fourth position within the first view of the three-dimensional environment, the fourth position having a second spatial arrangement relative to a location within the three-dimensional environment. For example, the second user interface object is fixed to an object or portion of the three-dimensional environment such that the second user interface object does not move (e.g., is maintained at the same position within the three-dimensional environment) as the user's viewpoint moves. Automatically displaying one or more user interface objects at a locked position relative to the three-dimensional environment provides real-time visual feedback to the user as the user moves around the physical environment without locking specific other user interface objects that follow the user as the user moves, such that the user knows where the locked user interface elements are located in the three-dimensional environment relative to the three-dimensional environment, and the user can view or interact with the locked user interface elements by returning to their fixed location within the three-dimensional environment, thereby providing the user with improved visual feedback.

[0194] In some embodiments, while displaying the first user interface object, the computer system displays a third user interface object at a fifth position within the first view of the three-dimensional environment, the third user interface object having a third spatial configuration relative to a discrete part of the user (e.g., relative to a part of the user's body (e.g., head, torso) or relative to the user's viewpoint). For example, the third user interface object is anchored to the user's hand, and as the user moves within the physical environment, the third user interface object is displayed in a position anchored to the user's hand (e.g., displayed in response to the user raising the user's hand so that it is within the user's current view of the three-dimensional environment). Automatically displaying one or more user interface objects in positions that maintain (e.g., are locked to) the same general location relative to the user's body part as the user moves around the three-dimensional environment, without locking specific other user interface objects to the user's body part as the user moves, provides real-time visual feedback as the user moves around the physical environment, thereby providing the user with improved visual feedback.

[0195] In some embodiments, in response to detecting a user input repositioning the first user interface object to a sixth position within the three-dimensional environment, the computer system moves the first user interface object to the sixth position within the three-dimensional environment, the sixth position having a fourth spatial arrangement relative to a discrete portion of the user, and updates, for the first user interface object, the first spatial arrangement of the first user interface object relative to the discrete portion of the user to the fourth spatial arrangement relative to the discrete portion of the user. For example, as described with reference to FIG. 7J , the first user interface object 7104-8 is positioned (e.g., repositioned and fixed) in a different zone (e.g., directly in front of the user) having a different spatial arrangement relative to the user (e.g., the user's torso, the user's head, the user's viewpoint), with the new zone in the left corner defined at a different angle and / or distance from a portion of the user (e.g., the fourth spatial arrangement). In some embodiments, after updating the first user interface object to have the fourth spatial configuration, in response to movement of the user's viewpoint that meets a threshold amount of movement, the first user interface object is maintained in the fourth spatial configuration relative to a distinct portion of the user (e.g., continues to have the delayed-follow behavior described above). Enabling a user to change the anchor position of a particular user interface object to have a different spatial relationship to the user's current view so that the particular user interface object remains in the same position relative to the user's current view even as the user's current view changes provides real-time visual feedback as the user moves through their physical environment and facilitates the user placing the particular user interface object in a position that is most comfortable or convenient for the user to view, thereby providing improved visual feedback to the user.

[0196] In some embodiments, the sixth position within the three-dimensional environment is within a predetermined distance from the user, and in response to detecting a user input to reposition the first user interface object to a seventh position within the three-dimensional environment, the computer system moves the first user interface object to the seventh position, where the seventh position within the three-dimensional environment is more than the predetermined distance from the user, and at the seventh position, the first user interface object is fixed to a portion of the three-dimensional environment. For example, the user may place the first user interface object at a position in three-dimensional space that is outside any of the predetermined zones. In some embodiments, following a determination that the user has placed the first user interface object at a position within the three-dimensional environment (e.g., the seventh position) such that the first user interface object is fixed to the three-dimensional environment (e.g., its position is independent of the user's current viewpoint and the first user interface object does not maintain an individual spatial relationship to an individual portion of the user while the user moves the user's current viewpoint of the three-dimensional environment), the first user interface object no longer has delayed-follow behavior. In some embodiments, in response to the user placing the first user interface object at a seventh position that is outside a predetermined distance from the user (e.g., outside the user's arm's reach), a textual display is displayed indicating that the first user interface object will not have delayed follow behavior while the object is placed at the seventh position. In some embodiments, the sixth position cannot be more than a predetermined distance away. For example, the sixth position must be within arm's reach of the user (e.g., a predetermined zone is within the user's arm's reach) in order for the user interface object to continue to have delayed follow behavior while the user interface object is placed at the sixth position.By automatically providing the user with the option to change the anchor of certain user interface objects to have a different spatial relationship to the user's current view that are within a predetermined distance (e.g., within arm's reach) of the user, while allowing the user to place (e.g., relocate) other user interface objects to fixed positions in the three-dimensional environment (e.g., positions beyond the predetermined distance), the system allows the user to place objects in areas that are most comfortable or convenient for the user and provides real-time visual feedback as the user moves through the physical environment, thereby providing the user with improved visual feedback.

[0197] In some embodiments, while detecting user input to relocate the first user interface object to an eighth position within the three-dimensional environment, the computer system displays visual indications of one or more predetermined zones. In some embodiments, in response to detecting movement of the user's viewpoint in accordance with a determination that the eighth position is within a predetermined zone of the one or more predetermined zones, the eighth position having a fifth spatial configuration relative to a distinct portion of the user, the computer system displays the first user interface object at a ninth position within the three-dimensional environment having the fifth spatial configuration relative to a distinct portion of the user. In some embodiments, in response to detecting movement of the user's viewpoint in accordance with a determination that the eighth position is not within a predetermined zone of the one or more predetermined zones, the computer system maintains the display of the first user interface object at the eighth position within the three-dimensional environment. For example, as described with reference to FIG. 7J , if a user repositions a first user interface object outside of a predetermined zone, the first user interface object no longer has a delayed follow behavior (e.g., the user interface object is fixed to the three-dimensional environment and no longer maintains the first user interface object as having a first spatial arrangement relative to a distinct portion of the user). For example, a user interface object positioned outside any of the predetermined zones does not follow the user's viewpoint as the user's viewpoint moves (e.g., the user interface object positioned outside the predetermined zone is fixed to the three-dimensional environment). In some embodiments, the computer system displays an outline (e.g., or other visual indication) of a predetermined zone (e.g., a delayed follow zone) as described with reference to FIG. 7J . In some embodiments, a user interface object positioned in any of the predetermined zones follows the user's viewpoint as the user's viewpoint moves.In some embodiments, the visual indication of the predetermined zone is displayed while the user is moving the first user interface object (e.g., in response to the user initiating a gesture to reposition the first user interface object). Automatically displaying the outlines of multiple zones in the three-dimensional environment, each zone having a different relative position to the user's current view, makes it easier for the user to select a zone in which to anchor a particular user interface object, and when the user interface object is placed within a zone, the user interface object follows the user as the user moves within the physical environment, allowing the user to more easily determine where in the three-dimensional environment the user interface object should be placed that is most comfortable or convenient for the user to view, thereby providing the user with improved visual feedback.

[0198] In some embodiments, in response to detecting user input to reposition the first user interface object to the sixth position, the first user interface object snaps to a sixth position within the three-dimensional environment. For example, as described with reference to FIG. 7J , snapping the user interface object to the sixth position includes automatically moving the user interface object to the sixth position (e.g., without the user continuing to drag the first user interface object toward the sixth position) pursuant to determining that the first user interface object has moved within a predetermined threshold distance from the sixth position while detecting user input to reposition the first user interface object to the sixth position. In some embodiments, in response to user input moving the first user interface object out of the sixth position (e.g., out of a predetermined zone), the first user interface object remains displayed in the sixth position until the user input moves the first user interface object beyond a threshold amount of movement from the sixth position. In some embodiments, in response to the first user interface object snapping to the sixth position, a haptic and / or audio indication is provided (e.g., simultaneously while displaying the first user interface object in the sixth position). For example, the user input is a pinch gesture (e.g., or pinch-and-drag gesture) that includes bringing two or more fingers of a hand into contact with or out of contact with one another (e.g., optionally in conjunction with (e.g., followed by) a drag input that changes the position of the user's hand from a first position (e.g., a start position for a drag) to a second position (e.g., an end position for a drag). In some embodiments, the user maintains the pinch gesture while performing the drag input and releases the pinch gesture (e.g., spreading the two or more fingers apart) to end the drag gesture (e.g., at the second position).For example, a pinch input may select a first user interface object, and once selected, the user may drag the first user interface object to reposition the first user interface object to the sixth position (e.g., near the sixth position before the user interface object was snapped to the sixth position). In some embodiments, the pinch input may be directly directed at the first user interface object (e.g., the user performs the pinch input at a position corresponding to the first user interface object), or the pinch input may be indirectly directed at the first user interface object (e.g., the user performs the pinch input while gazing at the first affordance, and the position of the user's hands while performing the pinch input is not at a position corresponding to the first user interface object). For example, the user may direct the user's input to the first user interface object by initiating a gesture at or near the first user interface object (e.g., within a distance of 0.5 cm, 1 cm, 5 cm, or 0-5 cm, measured from the outer edge of the first user interface object or the central portion of the first user interface object). In some embodiments, the user may further direct the user's input to the first user interface object by paying attention to the first user interface object (e.g., gazing at the first user interface object), and the user initiates the gesture while gazing at the first user interface object (e.g., at any position detectable by the computer system). For example, if the user is gazing at the first user interface object, the gesture need not be initiated at or near the first user interface object (e.g., the user performs a drag gesture while gazing at the first user interface object).The present invention automatically snaps a user interface object in response to a user repositioning the user interface object within a predetermined distance of the snap position, moves the user interface object to the predetermined snap position without requiring the user to precisely align the user interface object over the snap position, provides real-time visual feedback to the user as the user repositions the user interface object, and provides a visual indication confirming that the user interface object has been successfully repositioned at the snap position, thereby providing improved visual feedback to the user.

[0199] In some embodiments, the computer system displays one or more user interface objects for controlling the immersion level of the three-dimensional environment. In some embodiments, in response to user input directed to a first user interface object of the one or more user interface objects for increasing the immersion level of the three-dimensional environment, the computer system displays additional virtual content within the three-dimensional environment (e.g., and optionally stops or reduces the display of pass-through content). In some embodiments, in response to detecting user input directed to a second user interface object of the one or more user interface objects for decreasing the immersion level of the three-dimensional environment, the computer system displays additional content corresponding to the physical environment (e.g., displays the additional pass-through content and optionally stops or reduces the display of virtual content (e.g., one or more virtual objects)). In some embodiments, the one or more user interface objects include controls for playing and / or pausing the immersive experience in the three-dimensional environment. For example, a user can control how much of the physical environment is displayed as pass-through content in the three-dimensional environment (e.g., none of the physical environment is displayed (e.g., otherwise represented in the three-dimensional environment) during a fully immersive experience). For example, the user input is a tap input (e.g., an air gesture or a pinch gesture), which is optionally a tap input of the thumb on top of the index finger of the user's hand (e.g., on the side of the index finger adjacent to the thumb). In some embodiments, the tap input is detected without having to lift the thumb from the side of the index finger. For example, the user performs a tap input directed at a first user interface object to increase the level of immersion in the three-dimensional environment.In some embodiments, the user input is directed directly to the first user interface object (e.g., the user performs a tap input at a position corresponding to the first user interface object), or the user input is directed indirectly to the first user interface object (e.g., the user performs a tap input while gazing at the first affordance, and the user's hand position while performing the tap input is not at a position corresponding to the first user interface object). For example, the user can direct the user input to the first user interface object by initiating a gesture at or near the first user interface object (e.g., within a distance of 0.5 cm, 1 cm, 5 cm, or 0-5 cm, measured from the outer edge of the first user interface object or the central portion of the first user interface object). In some embodiments, the user can further direct the user's input toward the first user interface object by paying attention to the first user interface object (e.g., gazing at the first user interface object), and while focusing on the first user interface object, the user initiates the gesture (e.g., at any position detectable by the computer system). For example, if the user focuses on the first user interface object, the gesture does not need to be initiated at or near the first user interface object. Automatically displaying multiple controls that allow the user to control the immersive experience of the three-dimensional environment relative to the physical environment provides additional control for the user without requiring the user to navigate complex menu hierarchies, thereby allowing the user to easily control how much content from the physical environment is displayed in the three-dimensional environment, and provides the user with real-time visual feedback when the user requests to change the immersion level in the three-dimensional environment, thereby providing the user with improved visual feedback without requiring additional user input.

[0200] In some embodiments, the computer system displays one or more user interface objects for controlling the experience of the three-dimensional environment. In some embodiments, in response to user input directed to a first user interface object of the one or more user interface objects, the computer system performs a first action that modifies content in the three-dimensional environment (e.g., playing or pausing the first content). In some embodiments, in response to detecting user input directed to a second user interface object of the one or more user interface objects, the computer system performs a second action in the three-dimensional environment that is different from the first action (e.g., pausing or playing the second content). In some embodiments, the one or more user interface objects include controls for playing and / or pausing an immersive experience in the three-dimensional environment. For example, in response to a user selecting a first user interface object, the computer system displays a first type of virtual wallpaper (corresponding to a first type of virtual experience) in the three-dimensional environment. In response to a user selecting a second user interface object, the computer system displays the three-dimensional environment with virtual lighting effects. For example, the user input is a tap input (e.g., an air gesture or a pinch gesture), which is optionally a tap input of the thumb on top of the index finger of the user's hand (e.g., on the side of the index finger adjacent to the thumb). In some embodiments, the tap input is detected without having to lift the thumb from the side of the index finger. For example, a user performs a tap input directed at a first user interface object to play (e.g., or pause) virtual content displayed in the three-dimensional environment.In some embodiments, the user input is directed directly to the first user interface object (e.g., the user performs a tap input at a position corresponding to the first user interface object), or the user input is indirectly directed to the first user interface object (e.g., the user performs a tap gesture while gazing at the first affordance, and the position of the user's hand while performing the tap input is not a position corresponding to the first user interface object). For example, the user can direct the user input to the first user interface object by initiating a gesture at or near the first user interface object (e.g., within a distance of 0.5 cm, 1 cm, 5 cm, or 0-5 cm, measured from the outer edge of the first user interface object or the central portion of the first user interface object). In some embodiments, the user can further direct the user's input to the first user interface object by paying attention to the first user interface object (e.g., gazing at the first user interface object), and while focusing on the first user interface object, the user initiates the gesture (e.g., at any position detectable by the computer system). For example, if the user focuses on the first user interface object, the gesture does not need to be initiated at or near the first user interface object. Thus, the user interface objects correspond to controls for modifying (e.g., displaying or stopping displaying) various virtual content for the virtual experience in the three-dimensional environment. In some embodiments, in response to selecting a play control user interface object, the virtual content for the virtual experience is displayed, and in response to selecting a pause control user interface object, the display of the virtual content for the virtual experience is stopped.Automatically displaying multiple controls that allow a user to control the immersive experience of the three-dimensional environment relative to the physical environment provides additional control for the user without requiring the user to navigate complex menu hierarchies, thereby allowing the user to easily control how much content from the physical environment is displayed in the three-dimensional environment, and provides the user with real-time visual feedback when the user requests to change the level of immersion in the three-dimensional environment, thereby providing the user with improved visual feedback without requiring additional user input.

[0201] In some embodiments, in response to a movement of the user's viewpoint satisfying a threshold amount of movement, in accordance with a determination that the user's body (e.g., torso, head, or hand) moves within the physical environment, the computer system moves the first user interface object from a first position to a second position within the three-dimensional environment. In some embodiments, in accordance with a determination that the user's head moves from a first head position to a second head position within the physical environment (e.g., without detecting movement of the user's body (e.g., detecting only head rotation)), the computer system stops displaying the first user interface object at the first position within the three-dimensional environment and displays the first user interface object at a second position within the three-dimensional environment (e.g., without animating the movement) (e.g., after stopping displaying the first user interface object at the first position). In some embodiments, the user's head must also remain at the second head position for at least a threshold amount of time (e.g., 1 second, 2 seconds, 10 seconds, or a time amount of 0.5-10 seconds) before the first user interface object is displayed at the second position within the three-dimensional environment. In some embodiments, in response to a user moving their body (e.g., torso, head, or hands) within the physical environment, the first user interface object continues to be displayed and is animated to move with the user's body movements, but in response to detecting a rotation of the user's head (e.g., without moving the user's body) that changes the user's viewpoint, the user interface object disappears and reappears in the new viewpoint (while the head remains pointed toward the new viewpoint).In some embodiments, as described above with reference to method 800 and FIGS. 7C-7H , determining that the user's head has moved from a first head position to a second head position includes detecting that the user does not satisfy an attention criterion for the first user interface object, and in response to detecting that the user does not satisfy the attention criterion for the first user interface object, the first user interface object is displayed with a modified appearance (e.g., faded, no longer displayed, or otherwise visually de-emphasized). In some embodiments, displaying the first user interface object at a second position within the three-dimensional environment is performed in response to detecting that the user satisfies the attention criterion for the first user interface object (e.g., as described with reference to FIG. 7H ). Automatically displaying a particular user interface object to move in the three-dimensional environment while the user's torso moves in the physical environment, without displaying the particular user interface object in the three-dimensional environment as the user's head moves without moving the user's torso, provides real-time visual feedback as the user moves in the physical environment, thereby providing improved visual feedback to the user.

[0202] In some embodiments, after determining that the movement satisfies the threshold amount, the first user interface object is displayed at a plurality of respective positions within the three-dimensional environment as the user's viewpoint moves relative to the physical environment, where at a first of the plurality of respective positions, the first user interface object has (e.g., continues to have) the first spatial arrangement relative to a distinct portion of the user, and at a second of the plurality of respective positions, the first user interface object has (e.g., continues to have) the first spatial arrangement relative to a distinct portion of the user. For example, the plurality of distinct positions in the three-dimensional environment may apply to any number of locations within the three-dimensional environment (e.g., the first user interface object is enabled to be displayed in any portion of the three-dimensional environment within the user's current viewpoint to maintain the first spatial arrangement relative to a distinct portion of the user). Thus, as the user moves within the physical environment, the first user interface object may be displayed at various positions within the three-dimensional environment such that the first user interface object appears to be continuously moving within the three-dimensional environment as the user moves. 7E-7G, the user interface objects are displayed in additional positions as the user continues to move within the physical environment. Automatically moving certain user interface objects within the three-dimensional environment to follow the user's current viewpoint while the user moves within the physical environment, while maintaining the same spatial relationship between the user interface objects and the user's viewpoint, provides real-time visual feedback as the user moves within the physical environment, thereby providing improved visual feedback to the user.

[0203] It should be understood that the particular order described of the operations in FIG. 9 is merely an example, and that the described order is not intended to indicate the only order in which the operations may be performed. Those skilled in the art will recognize various ways to reorder the operations described herein. It should also be noted that other process details described herein with respect to other methods described herein (e.g., method 800) are also applicable in a similar manner to method 900 described above with respect to FIG. 9. For example, the gestures, inputs, physical objects, user interface objects, movements, references, three-dimensional environments, display generating components, representations of physical objects, virtual objects, and / or animations described above with reference to method 900 optionally have one or more of the characteristics of the gestures, inputs, physical objects, user interface objects, movements, references, three-dimensional environments, display generating components, representations of physical objects, virtual objects, and / or animations described herein with reference to other methods described herein (e.g., method 800). For the sake of brevity, those details will not be repeated here.

[0204] The operations described above with reference to Figures 8 and 9 are optionally performed by the components shown in Figures 1-6. In some embodiments, aspects and / or operations of methods 800 and 900 may be interchanged, substituted, and / or added between these methods, and for the sake of brevity, the details of which will not be repeated here.

[0205] Furthermore, for methods described herein in which one or more steps are conditioned on one or more conditions being satisfied, it should be understood that the described method can be repeated in multiple iterations, such that over the course of the iterations, all of the conditions on which the method steps are conditioned are satisfied in different iterations of the method. For example, if a method requires performing a first step if a condition is satisfied and a second step if the condition is not satisfied, one skilled in the art will understand that the steps recited in the claim are repeated in a particular order until the conditions are satisfied and then no longer satisfied. Thus, a method described with one or more steps that depend on one or more conditions being satisfied can be rewritten as a method that is repeated until each condition recited in the method is satisfied. However, this is not required for system or computer-readable medium claims in which the system or computer-readable medium includes instructions for performing a conditional action based on the satisfaction of the corresponding one or more conditions, and thus can determine whether a contingency is met without explicitly repeating the method steps until all conditions on which the method steps are conditioned are satisfied. Those skilled in the art will also understand that, as with methods having conditional steps, the system or computer-readable storage medium may repeat the steps of the method as many times as necessary to ensure that all of the conditional steps have been performed.

[0206] The foregoing has been described with reference to specific embodiments for purposes of explanation. However, the exemplary discussion above is not intended to be exhaustive or to limit the invention to the precise form disclosed. Many modifications and variations are possible in light of the above teachings. These embodiments were chosen and described in order to best explain the principles of the invention and its practical application, and thereby enable others skilled in the art to best utilize the invention and the various described embodiments with various modifications suited to the particular uses contemplated.

Claims

1. A first computer system in communication with a first display generation component and one or more first input devices, displaying a first user interface object in a first view of a three-dimensional environment via the first display generation component; Detecting, via the one or more input devices, whether the user satisfies an attention criterion for the first user interface object while displaying the first user interface object; displaying the first user interface object with a modified appearance in response to detecting that the user does not satisfy the attention criterion for the first user interface object, wherein displaying the first user interface object with the modified appearance includes de-emphasizing the first user interface object relative to one or more other objects in the three-dimensional environment; detecting, via the one or more input devices, a first movement of the user's viewpoint relative to a physical environment while displaying the first user interface object with the modified appearance; detecting that the user satisfies the attention criterion with respect to the first user interface object after detecting the first movement of the user's viewpoint relative to the physical environment; displaying the first user interface object in a second view of the three-dimensional environment that is different from the first view of the three-dimensional environment in response to detecting that the user satisfies the attention criterion, wherein displaying the first user interface object in the second view of the three-dimensional environment includes displaying the first user interface object with an appearance that highlights the first user interface object relative to one or more other objects in the three-dimensional environment more than when the first user interface object is displayed with the modified appearance; A method comprising:

2. The method of claim 1 , wherein the first user interface object has a first spatial relationship to a first anchor position in the three-dimensional environment that corresponds to a location of the user's body in the physical environment.

3. 10. The method of claim 1, further comprising: after detecting the first movement of the user's viewpoint relative to the physical environment, maintaining the display of the first user interface object at the same anchor position within the three-dimensional environment.

4. after detecting the first movement of the gaze point of the user relative to the physical environment; moving the first user interface object within the three-dimensional environment to the same position relative to the viewpoint of the user in accordance with a determination that the first movement satisfies a time threshold; The method of claim 1 or 2, further comprising:

5. receiving a user input to reposition the first user interface object in the three-dimensional environment; in response to receiving the input for repositioning the first user interface object in the three-dimensional environment, repositioning the first user interface object to a respective position in the three-dimensional environment according to the input; detecting an input to change the user's viewpoint after repositioning the first user interface object to the respective position within the three-dimensional environment in accordance with the input; and responsive to detecting the input for changing the viewpoint of the user, changing the viewpoint of the user in accordance with the input for changing the viewpoint of the user and displaying the first user interface object in the three-dimensional environment from the current viewpoint of the user; displaying the first user interface object at a discrete position having a first spatial relationship to a first anchor position in the three-dimensional environment corresponding to a location of the user's viewpoint in the physical environment in accordance with determining that the first user interface object is located within a first predetermined zone; maintaining the display of the first user interface object at a same anchor position in the three-dimensional environment that does not correspond to a location of the user's body in the physical environment in accordance with a determination that the first user interface object is not located within the first predetermined zone; The method of claim 1 , further comprising:

6. the first predetermined zone is selected from a plurality of predetermined zones within the three-dimensional environment; the first predetermined zone has a first spatial relationship to a first anchor position within the three-dimensional environment corresponding to the location of the user's viewpoint; 6. The method of claim 5, wherein a second predetermined zone of the plurality of predetermined zones has a second spatial relationship to a second anchor position within the three-dimensional environment that corresponds to the location of the viewpoint of the user.

7. displaying a visual indication of the first predetermined zone in response to detecting that the user has initiated the user input to reposition the first user interface object; maintaining display of the visual indication relative to the first predetermined zone while detecting the user input to relocate the first user interface object to the first predetermined zone; The method of claim 5 or 6, further comprising:

8. The method of claim 1 , wherein the attention criteria for the first user interface object include a gaze criteria.

9. The method of claim 1 , wherein the attention criteria for the first user interface object include criteria for the position of the user's head in the physical environment.

10. Detecting a pinch input directed toward a first affordance displayed on at least a portion of the first user interface object, followed by a hand movement that performed the pinch input; In response to the movement of the hand, resizing the first user interface object according to the movement of the hand; 10. The method of claim 1, further comprising:

11. receiving a pinch input by a first hand and a pinch input by a second hand directed at the first user interface object, followed by a change in distance between the first hand and the second hand; In response to the change in distance between the first hand and the second hand, changing a size of the first user interface object according to the change in distance between the first hand and the second hand; The method of claim 1 , further comprising:

12. 12. The method of claim 1, wherein highlighting and suppressing the first user interface object relative to the one or more other objects in the three-dimensional environment comprises highlighting and suppressing the first user interface object relative to one or more other virtual objects in the three-dimensional environment.

13. 13. The method of claim 1, wherein suppressing the first user interface object relative to the one or more other objects in the three-dimensional environment comprises suppressing the first user interface object relative to representations of one or more physical objects in the physical environment.

14. The method of claim 1 , wherein the first user interface object comprises a plurality of selectable user interface objects.

15. displaying one or more user interface objects for controlling a level of immersion of the three-dimensional environment; and displaying additional virtual content in the three-dimensional environment in response to user input directed at a first user interface object of the one or more user interface objects to increase a level of immersion in the three-dimensional environment; displaying additional content corresponding to the physical environment in response to detecting a user input directed at a second user interface object of the one or more user interface objects to reduce an immersion level of the three-dimensional environment; and 15. The method of any one of claims 1 to 14, further comprising:

16. 16. The method of claim 1, wherein an amount of de-emphasis of the first user interface object relative to the one or more other objects in the three-dimensional environment is based on an angle between the detected line of sight of the user and the first user interface object.

17. 17. The method of claim 1, wherein an amount of de-emphasis of the first user interface object relative to the one or more other objects in the three-dimensional environment is based on a rate of the first movement of the viewpoint of the user.

18. The method of claim 4 , wherein the first user interface object moves within the three-dimensional environment according to the user's movements.

19. 2. The method of claim 1 , wherein immediately before and after the first movement of the user's viewpoint relative to the physical environment, a distinct characteristic position of the first user interface object in the three-dimensional environment has a first spatial relationship to a first anchor position in the three-dimensional environment that corresponds to a location of the user's viewpoint in the physical environment.

20. a first display generation component; one or more input devices; one or more processors; a memory storing one or more programs; A computer system comprising:

20. A computer system, wherein the one or more programs are configured to be executed by the one or more processors, the one or more programs including instructions to perform the method of any one of claims 1 to 19.

21. 20. A computer-readable storage medium having stored thereon one or more programs, the one or more programs including instructions that, when executed by a computer system including a first display generation component and one or more input devices, cause the computer system to perform the method of any one of claims 1 to 19.

22. 20. A graphical user interface on a computer system including a first display generation component, one or more input devices, a memory, and one or more processors executing one or more programs stored in the memory, the graphical user interface comprising a user interface displayed according to the method of any one of claims 1 to 19.

23. a first display generation component; one or more input devices; means for carrying out the method according to any one of claims 1 to 19; A computer system comprising:

24. 1. An information processing device for use in a computer system including a first display generation component and one or more input devices, the information processing device comprising: An information processing device comprising means for carrying out the method according to any one of claims 1 to 19.

25. A first computer system in communication with a first display generation component and one or more first input devices, displaying, via the first display generation component, a first user interface object in a first view of a three-dimensional environment at a first position within the three-dimensional environment and at a first spatial arrangement relative to a discrete portion of the user; detecting, via the one or more input devices, movement of the user's viewpoint from a first location to a second location within a physical environment while displaying the first user interface object; In response to detecting the movement of the user's viewpoint from the first location to the second location, maintaining a display of the first user interface object at the first position within the three-dimensional environment in accordance with a determination that the movement of the user's viewpoint from the first location to the second location does not satisfy a threshold amount of movement; in response to determining that the movement of the user's viewpoint from the first location to the second location satisfies the threshold amount of movement; ceasing to display the first user interface object at the first position within the three-dimensional environment; and displaying the first user interface object at a second position within the three-dimensional environment, the second position within the three-dimensional environment having the first spatial arrangement with respect to the distinct portion of the user; A method comprising:

26. detecting that the movement of the user's viewpoint satisfies the threshold amount of movement while maintaining display of the first user interface object at the first position within the three-dimensional environment; moving the first user interface object from the first position to the second position within the three-dimensional environment in response to the movement of the user's viewpoint satisfying the threshold amount of movement; 26. The method of claim 25, further comprising:

27. detecting, via the one or more input devices, movement of the user's viewpoint from the second location to a third location within the physical environment while displaying the first user interface object at the second position within the three-dimensional environment; In response to detecting the movement of the user's viewpoint from the second location to the third location, maintaining the display of the first user interface object at the second position within the three-dimensional environment in accordance with a determination that the movement of the user's viewpoint from the second location to the third location does not satisfy a second threshold amount of movement; and in response to determining that the movement of the user's viewpoint from the second location to the third location satisfies a second threshold amount of movement; ceasing to display the first user interface object at the second position within the three-dimensional environment; and displaying the first user interface object at a third position within the three-dimensional environment, the third position within the three-dimensional environment having the first spatial configuration with respect to the distinct portion of the user; 27. The method of claim 25 or 26, further comprising:

28. 28. The method of claim 26 or 27, further comprising visually de-emphasizing the first user interface object relative to one or more other objects in the three-dimensional environment while moving the first user interface object from the first position to the second position in the three-dimensional environment.

29. 30. The method of claim 28, wherein visually de-emphasizing the first user interface object comprises displaying the first user interface object with reduced opacity.

30. 30. The method of claim 28 or 29, wherein visually de-emphasizing the first user interface object comprises displaying the first user interface object with a blurred visual effect.

31. 31. The method of claim 25, further comprising: visually de-emphasizing the first user interface object relative to one or more other objects in the three-dimensional environment in accordance with a determination that the user does not satisfy an attention criterion with respect to the first user interface object.

32. The method of claim 31 , wherein the attention criteria for the first user interface object include a gaze criteria.

33. The method of claim 31 , wherein the attention criteria for the first user interface object include criteria for the position of the user's head in the physical environment.

34. 34. The method of claim 28, wherein an amount of visual de-emphasis of the first user interface object is based on a rate of the movement of the eyepoint of the user.

35. 35. The method of claim 25, further comprising, while displaying the first user interface object, displaying a second user interface object at a fourth position within the first view of the three-dimensional environment, the fourth position having a second spatial orientation relative to a location within the three-dimensional environment.

36. 36. The method of claim 25, further comprising, while displaying the first user interface object, displaying a third user interface object at a fifth position within the first view of the three-dimensional environment, the third user interface object having a third spatial arrangement relative to a distinct portion of the user.

37. in response to detecting a user input to reposition the first user interface object to a sixth position within the three-dimensional environment; moving the first user interface object to a sixth position within the three-dimensional environment, the sixth position having a fourth spatial configuration with respect to the discrete portion of the user; For the first user interface object, updating the first spatial arrangement of the first user interface object relative to the respective portion of the user to the fourth spatial arrangement relative to the respective portion of the user; 37. The method of any one of claims 25 to 36, further comprising:

38. the sixth position within the three-dimensional environment is within a predetermined distance from the user; 38. The method of claim 37, further comprising: in response to detecting user input that relocates the first user interface object to a seventh position within the three-dimensional environment, moving the first user interface object to the seventh position, the seventh position within the three-dimensional environment being greater than the predetermined distance from the user, and wherein at the seventh position the first user interface object is fixed to a portion of the three-dimensional environment.

39. displaying visual indications of one or more predetermined zones while detecting a user input to relocate the first user interface object to an eighth position within the three-dimensional environment; and In response to detecting movement of the user's viewpoint according to a determination that the eighth position is within a predetermined zone of the one or more predetermined zones, the eighth position having a fifth spatial configuration relative to the respective portion of the user, displaying the first user interface object at a ninth position within the three-dimensional environment having the fifth spatial configuration relative to the respective portion of the user; 39. The method of claim 25, further comprising: maintaining the display of the first user interface object at the eighth position within the three-dimensional environment in response to detecting movement of the user's viewpoint in accordance with a determination that the eighth position is not within a predetermined zone of the one or more predetermined zones.

40. 39. The method of claim 37 or 38, wherein in response to detecting the user input to reposition the first user interface object to the sixth position, the first user interface object snaps to the sixth position within the three-dimensional environment.

41. displaying one or more user interface objects for controlling a level of immersion of the three-dimensional environment; and displaying additional virtual content in the three-dimensional environment in response to user input directed at a first user interface object of the one or more user interface objects to increase a level of immersion in the three-dimensional environment; displaying additional content corresponding to the physical environment in response to detecting a user input directed at a second user interface object of the one or more user interface objects to reduce an immersion level of the three-dimensional environment; and 41. The method of any one of claims 25 to 40, further comprising:

42. displaying one or more user interface objects for controlling an experience in the three-dimensional environment; performing a first operation to modify content within the three-dimensional environment in response to user input directed at a first user interface object of the one or more user interface objects; performing a second action in the three-dimensional environment, different from the first action, in response to detecting user input directed at a second user interface object of the one or more user interface objects; 42. The method of any one of claims 25 to 41, further comprising:

43. In response to the movement of the user's viewpoint satisfying the threshold amount of movement, moving the first user interface object from the first position to the second position in the three-dimensional environment in accordance with a determination that the user's body moves in the physical environment; in response to determining that the user's head moves from a first head position to a second head position within the physical environment; ceasing to display the first user interface object at the first position within the three-dimensional environment; and displaying the first user interface object at the second position within the three-dimensional environment; and 43. The method of any one of claims 25 to 42, further comprising:

44. 44. The method of claim 25, wherein after the determination that the movement satisfies the threshold amount, the first user interface object is displayed at a plurality of respective positions within the three-dimensional environment as the user's viewpoint moves relative to the physical environment, wherein at a first position of the plurality of respective positions, the first user interface object has the first spatial arrangement relative to the distinct portion of the user, and at a second position of the plurality of distinct positions, the first user interface object has the first spatial arrangement relative to the distinct portion of the user.

45. a first display generation component; one or more input devices; one or more processors; a memory storing one or more programs; A computer system comprising:

45. A computer system, wherein the one or more programs are configured to be executed by the one or more processors, the one or more programs comprising instructions for performing the method of any one of claims 25 to 44.

46. 45. A computer-readable storage medium having stored thereon one or more programs, the one or more programs including instructions that, when executed by a computer system including a first display generation component and one or more input devices, cause the computer system to perform the method of any one of claims 25 to 44.

47. 45. A graphical user interface on a computer system including a first display generation component, one or more input devices, memory, and one or more processors executing one or more programs stored in the memory, the graphical user interface comprising a user interface displayed according to the method of any one of claims 25 to 44.

48. a first display generation component; one or more input devices; means for carrying out the method of any one of claims 25 to 44; A computer system comprising:

49. 1. An information processing device for use in a computer system including a first display generation component and one or more input devices, the information processing device comprising:

45. An information processing device comprising means for carrying out the method of any one of claims 25 to 44.

Citation Information

Patent Citations

  • Information processing device and image forming method

    JP2019220185A

  • Information processing device, information processing method, and program

    WO2015064165A1