Device, method, and graphical user interface for presenting virtual objects within a virtual environment

The computer system enhances interaction in augmented and virtual reality environments by using touch-sensitive displays, eye-tracking, and haptic feedback to streamline user inputs, addressing inefficiencies and conserving battery life.

JP7702573B2Active Publication Date: 2025-07-03APPLE INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024518499
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-09-25
Filing Date
2022-09-16
Publication Date
2025-07-03
Estimated Expiration
2042-09-16

AI Technical Summary

Technical Problem

Existing methods for interacting with augmented and virtual reality environments are cumbersome, inefficient, and place a significant cognitive burden on users due to complex manipulation of virtual objects and insufficient feedback, leading to energy wastage, particularly in battery-operated devices.

Method used

A computer system equipped with touch-sensitive displays, eye-tracking, hand-tracking, and haptic output generators, along with a graphical user interface, reduces the number and complexity of user inputs by providing intuitive feedback and spatial arrangement updates of virtual objects, enhancing interaction efficiency and conserving battery life.

Benefits of technology

The system improves user interaction efficiency, reduces cognitive burden, and extends battery life by minimizing unnecessary inputs and optimizing power consumption in augmented and virtual reality environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007702573000001
    Figure 0007702573000001
  • Figure 0007702573000002
    Figure 0007702573000002
  • Figure 0007702573000003
    Figure 0007702573000003
Patent Text Reader

Abstract

In some embodiments, the electronic device updates the spatial arrangement of one or more virtual objects in the three dimensional environment. In some embodiments, the electronic device updates the positions of multiple virtual objects together. In some embodiments, the electronic device displays the objects in the three dimensional environment based on the estimated location of a floor in the three dimensional environment. In some embodiments, the electronic device moves (e.g., repositions) the objects in the three dimensional environment.
Need to check novelty before this filing date? Find Prior Art

Description

Cross - reference to related applications

[0001] Cross - reference between related applications This application claims the benefit of U.S. Provisional Patent Application No. 63 / 261,667, filed on September 25, 2021, the content of which is hereby incorporated by reference in its entirety for all purposes.

Technical Field

[0002] This relates generally to a computer system having a display generation component and one or more input devices that present a graphical user interface in a virtual environment, including but not limited to, via the display generation component.

Background Art

[0003] The development of computer systems for augmented reality has advanced significantly in recent years. Exemplary augmented reality environments include at least some virtual elements that replace or enhance the physical world. Input devices such as cameras, controllers, joysticks, touch - sensitive surfaces, and touch - screen displays for computer systems and other electronic computing devices are used to interact with virtual / augmented reality environments. Exemplary virtual elements include virtual objects containing digital images, videos, text, icons, and control elements such as buttons and other graphics.

Summary of the Invention

[0004] Some methods and interfaces for interacting with an environment (such as an application, an augmented reality environment, a mixed reality environment, and a virtual reality environment) that includes at least some virtual elements are cumbersome, inefficient, and limited. For example, systems that provide insufficient feedback for performing actions associated with virtual objects, systems that require a series of inputs to achieve a desired result in an augmented reality environment, and systems where the manipulation of virtual objects is complex and error-prone impose a large cognitive burden on the user and degrade the experience in the virtual / augmented reality environment. In addition, those methods take more time than necessary, thereby wasting energy. This latter consideration is particularly important in battery-operated devices.

[0005] Accordingly, there is a need for a computer system having improved methods and interfaces for providing a computer-generated experience that makes interacting with the computer system more efficient and intuitive for the user. Such methods and interfaces can optionally complement or replace conventional methods of providing an augmented reality experience to the user. Such methods and interfaces reduce the number, degree, and / or type of inputs from the user by assisting the user in understanding the connection between the input provided and the device response to that input, thereby creating a more efficient human-machine interface.

[0006] The above-mentioned drawbacks and other problems associated with the user interface of a computer system are reduced or eliminated by the disclosed system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a notebook computer, a tablet computer, or a handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device such as a wristwatch or a head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also known as a "touch screen" or "touch screen display"). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, the computer system has one or more output devices in addition to a display generation component, and the output devices include one or more haptic output generators and / or one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, a memory, and one or more modules, programs, or sets of instructions stored in the memory for performing a plurality of functions. In some embodiments, the user interacts with the GUI through a stylus and / or finger contact and gestures on a touch-sensitive surface, the movement of the user's eyes and hands in space relative to the GUI (and / or the computer system) when captured by a camera and other motion sensors, and voice input when captured by one or more audio input devices.In some embodiments, functions performed through the interaction optionally include image editing, drawing, presenting, word processing, spreadsheet creation, game play, making a phone call, video conferencing, sending an email, instant messaging, training support, digital photo shooting, digital video shooting, web browsing, digital music playback, note taking, and / or digital video playback. The executable instructions for performing those functions are optionally included in a transient computer-readable storage medium and / or a non-transient computer-readable storage medium, or other computer program products configured to be executed by one or more processors.

[0007] There is a need for an improved method and interface for navigating a user interface and an electronic device having such an interface. Such a method and interface can complement or replace conventional methods for interacting with a graphical user interface. Such a method and interface reduce the number, degree, and / or type of input from a user and create a more efficient human-machine interface. In the case of battery-operated computing devices, such a method and interface conserve power and lengthen the battery charging intervals.

[0008] In some embodiments, the electronic device updates the spatial arrangement of one or more virtual objects in a three-dimensional environment. In some embodiments, the electronic device updates the positions of a plurality of virtual objects together. In some embodiments, the electronic device displays an object in a three-dimensional environment based on the estimated location of a floor within the three-dimensional environment. In some embodiments, the electronic device moves (e.g., repositions) an object within the three-dimensional environment.

[0009] Note that the various embodiments described above can be combined with any other embodiment described herein. The features and advantages described herein are not exhaustive, and in particular, many additional features and advantages will become apparent to those skilled in the art upon review of the drawings, specification, and claims. Further, note that the language used herein has been selected solely for readability and for the purpose of explanation, and not for the purpose of defining or limiting the subject matter of the invention.

Brief Description of the Drawings

[0010] To better understand the various embodiments described, the following "Modes for Carrying Out the Invention" should be referred to in conjunction with the following drawings, where like reference numerals refer to corresponding parts throughout the following figures.

[0011]

Figure 1

[0012]

Figure 2

[0013]

Figure 3

[0014]

Figure 4

[0015]

Figure 5

[0016]

Figure 6A

[0017]

Figure 6B

[0018]

Figure 7A

Figure 7B

Figure 7C

Figure 7D

Figure 7E

Figure 7F

Figure 7G

[0019]

Figure 8A

Figure 8B

Figure 8C

Figure 8D

Figure 8E

Figure 8F

Figure 8G

Figure 8H

Figure 8I

Figure 8J

Figure 8K

[0020]

Figure 9A

Figure 9B

Figure 9C

Figure 9D

Figure 9E

Figure 9F

Figure 9G

[0021]

Figure 10A

Figure 10B

Figure 10C

Figure 10D

Figure 10E

Figure 10F

Figure 10G

Figure 10H

Figure 10I

Figure 10J

Figure 10K

[0022]

Figure 11A

Figure 11B

Figure 11C

Figure 11D

Figure 11E

[0023]

Figure 12A

Figure 12B

Figure 12C

Figure 12D

Figure 12E

Figure 12F

Figure 12G

[0024]

Figure 13A

Figure 13B

Figure 13C

Figure 13D

[0025]

Figure 14A

Figure 14B

Figure 14C

Figure 14D

Figure 14E

Figure 14F

Figure 14G

DETAILED DESCRIPTION OF THE INVENTION

[0026] The present disclosure relates to a user interface for providing a computer-generated reality (XR) experience to a user according to some embodiments.

[0027] The systems, methods, and GUIs described herein provide an improved way for an electronic device to present content corresponding to physical locations indicated within a navigation user interface element.

[0028] In some embodiments, a computer system displays one or more virtual objects (e.g., a user interface of an application, representations of other users, content items, etc.) within a three-dimensional environment. In some embodiments, the computer system evaluates the spatial arrangement of the virtual objects relative to the user's perspective within the three-dimensional environment according to one or more spatial criteria described in more detail below. In some embodiments, the computer system detects user input corresponding to a request to update the position and / or orientation of the virtual objects to satisfy one or more spatial criteria. In some embodiments, in response to the input, the computer system updates the position and / or orientation of the virtual objects to satisfy one or more spatial criteria. Updating the position and / or orientation of the virtual objects in this way provides an efficient way to enable the user to access, view, and / or interact with the virtual objects, which in turn reduces power consumption and improves the battery life of the computer system by enabling the user to use the computer system more quickly and efficiently.

[0029] In some embodiments, a computer system updates the position and / or orientation of a plurality of virtual objects relative to the perspective of a user of the computer system in a three-dimensional environment. In some embodiments, the computer system receives input corresponding to a request to update the position and / or orientation of a plurality of virtual objects relative to the user's perspective. In some embodiments, in response to the input, the computer system updates the position and / or orientation of the plurality of virtual objects while maintaining the spatial relationships between the plurality of virtual objects relative to the user's perspective. Updating the position and / or orientation of the plurality of virtual objects in this way provides an efficient way to update the view of the three-dimensional environment from the user's perspective, which in turn reduces power consumption and improves the battery life of the computer system by enabling the user to use the computer system more quickly and efficiently.

[0030] In some embodiments, the computer system presents objects within a three-dimensional environment based on the estimated location of a physical floor within the physical environment of the computer system. In some embodiments, while the computer system is presenting an object within the three-dimensional environment based on the estimated location of a physical floor within the physical environment of the computer system, the computer system determines a new estimated location of the physical floor within the physical environment of the computer system. In some embodiments, the computer system continues to present an object within the three-dimensional environment based on a previous estimated location of a physical floor within the physical environment of the computer system until certain conditions / criteria are met. Displaying an object within the three-dimensional environment based on a previous estimated location of a physical floor until certain conditions / criteria are met provides an efficient way to update the location of the object within the three-dimensional environment when those conditions / criteria are met, rather than before those conditions / criteria are met, which further reduces power usage and improves the battery life of the computer system by enabling the user to use the computer system more quickly and efficiently.

[0031] In some embodiments, a computer system presents one or more user interface objects in a three-dimensional environment. In some embodiments, the computer system receives a request to move (e.g., reposition) an individual user interface object within the three-dimensional environment to a new location within the three-dimensional environment. In some embodiments, while moving (e.g., repositioning) an individual user interface object to a new location within the three-dimensional environment, the computer system visually de-emphasizes one or more portions of the individual user interface object. Visually de-emphasizing portions of an individual user interface object while moving the individual user interface object within the three-dimensional environment reduces a potential loss of sense of direction that can lead to dizziness or motion sickness symptoms, and thus provides a mechanism by which a user can safely interact with the three-dimensional environment, thereby reducing the cognitive burden on the user when interacting with the three-dimensional environment.

[0032] Figures 1-6 provide an illustration of an exemplary computer system for providing an XR experience (as described below with respect to methods 800, 1000, 1200, and 1400) to a user. Figures 7A-7G show exemplary techniques for updating the spatial arrangement of one or more virtual objects in a three-dimensional environment according to some embodiments. Figures 8A-8K are flowcharts showing a method for updating the spatial arrangement of one or more virtual objects in a three-dimensional environment according to some embodiments. Figures 9A-9G show exemplary techniques for updating the positions of a plurality of virtual objects together according to some embodiments. Figures 10A-10K are flowcharts showing a method for updating the positions of a plurality of virtual objects together according to some embodiments. Figures 11A-11E show exemplary techniques for displaying an object within a three-dimensional environment based on an estimated location of a floor of the three-dimensional environment according to some embodiments. Figures 12A-12G are flowcharts showing a method for displaying an object in a three-dimensional environment based on an estimated location of a floor of the three-dimensional environment according to some embodiments. Figures 13A-13D show exemplary techniques for moving an object within a three-dimensional environment according to some embodiments of the present disclosure. Figures 14A-14G are flowcharts showing a method for moving an object within a three-dimensional environment according to some embodiments.

[0033] The processes described below enhance the operability of a device and streamline the user interface with the device by various techniques, including other techniques, such as assisting the user in making appropriate inputs when operating / interacting with the device and reducing user errors, to provide improved visual feedback to the user, reduce the number of inputs required to perform an operation, provide additional control options without cluttering the user interface with additional controls being displayed, perform an operation without requiring further user input when a set of conditions is met, improve privacy and / or security, and / or otherwise. These techniques also reduce power consumption and improve the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0034] Furthermore, in the method described herein, conditioned on one or more conditions being met by one or more steps, it should be understood that the described method can be repeated in multiple iterations such that, over the course of the repetitions, all of the conditions conditioned on by the steps of the method are met in different repetitions of the method. For example, if a method requires performing a first step when a condition is met and a second step when the condition is not met, one of ordinary skill in the art will understand that the steps recited in the claims will be repeated in any order until the condition is met and then ceases to be met. Thus, a method described in terms of one or more steps that depend on one or more met conditions can be rewritten as a method that is repeated until each condition described in the method is met. However, this is not required in claims for a system or computer-readable medium that includes instructions to perform conditional operations based on the fulfillment of the corresponding one or more conditions, and thus can determine whether an event has been fulfilled without explicitly repeating the steps of the method until all of the conditions for the steps of the method being conditional are met. One of ordinary skill in the art will also understand that, similar to a method with conditional steps, a system or computer-readable storage medium can repeat the steps of the method the number of times necessary to ensure that all of the conditional steps have been performed.

[0035] In some embodiments, as shown in FIG. 1, an XR experience is provided to a user via an operating environment 100 that includes a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted device (HMD), a display, a projector, a touch screen, etc.), one or more input devices 125 (e.g., an eye tracking device 130, a hand tracking device 140, other input devices 150), one or more output devices 155 (e.g., speakers 160, a haptic output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a tactile sensor, an orientation sensor, a proximity sensor, a temperature sensor, a location sensor, a motion sensor, a speed sensor, etc.), and optionally one or more peripheral devices 195 (e.g., home appliances, wearable devices, etc.). In some embodiments, one or more of the input device 125, the output device 155, the sensor 190, and the peripheral device 195 are integrated with the display generation component 120 (e.g., within a head-mounted device or a handheld device).

[0036] When describing an XR experience, various related but distinct environments are individually referred to using various terms for the sake of the user to perceive and / or interact with (e.g., using inputs detected by the computer system 101 to generate audio, visual, and / or tactile feedback corresponding to the various inputs provided to the computer system 101 that generates the XR experience) the user. The following is a subset of these terms.

[0037] Physical Environment: The physical environment refers to the physical world that people can perceive and / or interact with without the assistance of an electronic system. Physical environments such as a physical park include physical objects such as physical trees, physical buildings, and physical people. People can directly perceive and / or interact with the physical environment through senses such as vision, touch, hearing, taste, and smell.

[0038] Extended Reality: In contrast, an extended reality (XR) environment refers to an environment that is wholly or partially simulated and with which people can perceive and / or interact through an electronic system. In XR, a subset of a person's body movements or their representations are tracked, and in response, one or more characteristics of one or more virtual objects simulated within the XR environment are adjusted to behave according to at least one physical law. For example, an XR system can detect the rotation of a person's head and, in response, adjust the graphic content and sound field presented to the person in a similar way to how such views and sounds would change in the physical environment. Depending on the situation (e.g., for accessibility reasons), the adjustment of the characteristics (s) of the virtual object(s) in the XR environment may be made in response to a representation of body movement (e.g., a voice command). People may use any one of these senses including vision, hearing, touch, taste, and smell to perceive and / or interact with XR objects. For example, a person can perceive and / or interact with an audio object that creates a 3D or spatial audio environment that provides the perception of a point audio source within a 3D space. In another example, an audio object can enable audio transparency that selectively incorporates ambient sound from the physical environment, with or without including computer-generated audio. In some XR environments, people may perceive and / or interact with only audio objects.

[0039] Examples of XR include virtual reality and mixed reality.

[0040] Virtual Reality: A virtual reality (VR) environment refers to an imitation environment designed to be based entirely on computer-generated sensory inputs for one or more senses. A VR environment includes multiple virtual objects that a person can perceive and / or interact with. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can perceive and / or interact with virtual objects in a VR environment through a simulation of the person's presence within the computer-generated environment and / or through a simulation of a subset of the person's physical movements within the computer-generated environment.

[0041] Mixed Reality: In contrast to a VR environment designed to be based entirely on computer-generated sensory inputs, a mixed reality (MR) environment refers to an imitation environment designed to incorporate sensory inputs or their representations from the physical environment in addition to including computer-generated sensory inputs (e.g., virtual objects). On the virtual continuum, an MR environment is anywhere between, but not including, the complete physical environment at one end and the virtual reality environment at the other end. In some MR environments, the computer-generated sensory inputs can respond to changes in the sensory inputs from the physical environment. Also, some electronic systems for presenting an MR environment may track the location and / or orientation with respect to the physical environment to enable virtual objects to interact with real objects (i.e., physical articles or their representations from the physical environment). For example, the system may take movement into account so that a virtual tree appears stationary with respect to the physical ground.

[0042] Examples of mixed reality include augmented reality and augmented virtuality.

[0043] Augmented Reality: An augmented reality (AR) environment refers to an emulated environment in which one or more virtual objects are superimposed on a physical environment or a representation thereof. For example, an electronic system for presenting an AR environment may have a transparent or translucent display through which a person can directly view the physical environment. The system may be configured to present virtual objects on the transparent or translucent display, whereby a person can use the system to perceive virtual objects superimposed on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture an image or video of the physical environment, which is a representation of the physical environment. The system synthesizes the image or video with virtual objects and presents the composite on the opaque display. A person uses this system to indirectly view the physical environment through the image or video of the physical environment and perceive virtual objects superimposed on the physical environment. As used herein, the video of the physical environment shown on the opaque display is referred to as a "pass-through video," meaning that the system uses one or more image sensors (singular or plural) to capture an image of the physical environment and uses those images when presenting the AR environment on the opaque display. Further alternatively, the system may have a projection system that projects virtual objects into the physical environment or onto a physical surface, for example, as holograms, whereby a person can use the system to perceive virtual objects superimposed on the physical environment. The augmented reality environment also refers to an emulated environment in which a representation of the physical environment is transformed by computer-generated sensory information. For example, when providing a pass-through video, the system may transform one or more sensor images to map to a selected perspective (e.g., viewpoint) different from the perspective captured by the imaging sensor. As another example, the representation of the physical environment may be transformed by graphically modifying (e.g., magnifying) a portion thereof, thereby making the modified portion a modified version that represents the original captured image but is non-photorealistic. As a further example, the representation of the physical environment may be transformed by graphically removing or obscuring a portion thereof.

[0044] Augmented Virtual: An augmented virtual (AV) environment refers to an imitative environment in which a virtual environment or a computer-generated environment incorporates one or more sensory inputs from a physical environment. The sensory inputs can be representations of one or more characteristics of the physical environment. For example, an AV park may have virtual trees and virtual buildings, but people with faces are realistically reproduced from images of physical people. As another example, virtual objects may adopt the shape or color of physical items imaged by one or more imaging sensors. As a further example, virtual objects can adopt shadows that match the position of the sun in the physical environment.

[0045] Viewpoint-locked virtual object: When a computer system displays a virtual object at the same location and / or position within the user's viewpoint even when the user's viewpoint shifts (e.g., changes), the virtual object is viewpoint-locked. In embodiments where the computer system is a head-mounted device, the user's viewpoint is locked in the forward direction of the user's head (e.g., the user's viewpoint is at least a portion of the user's field of view when the user is looking straight ahead). Thus, the user's viewpoint remains fixed even when the user's line of sight moves without the user moving their head. In embodiments where the computer system has a display generation component (e.g., a display screen) that can be repositioned relative to the user's head, the user's viewpoint is the augmented reality view presented to the user on the computer system's display generation component. For example, a viewpoint-locked virtual object that is displayed in the upper left corner of the user's viewpoint when the user's viewpoint is in a first orientation (e.g., the user's head is facing north) will continue to be displayed in the upper left corner of the user's viewpoint even when the user's viewpoint changes to a second orientation (e.g., the user's head is facing west). In other words, the location and / or position at which the viewpoint-locked virtual object is displayed within the user's viewpoint is independent of the user's position and / or orientation in the physical environment. In embodiments where the computer system is a head-mounted device, the user's viewpoint is locked to the orientation of the user's head such that the virtual object is also referred to as a "head-locked virtual object".

[0046] Environment-Locked Virtual Object: A virtual object is environment-locked (or "world-locked") when a computer system displays the virtual object at a location and / or position within the user's view that is based on (e.g., selected with reference to and / or fixed to) a location and / or object within a three-dimensional environment (e.g., a physical environment or a virtual environment). When the user's view shifts, the location and / or object within the environment relative to the user's view changes, and as a result, the environment-locked virtual object is displayed at a different location and / or position within the user's view. For example, an environment-locked virtual object locked to a tree directly in front of the user is displayed at the center of the user's view. If the user's view shifts to the right (e.g., the user's head is turned to the right) such that the tree becomes leftward in the user's view (e.g., the position of the tree in the user's view shifts), the environment-locked virtual object locked to the tree is displayed leftward in the user's view. In other words, the location and / or position at which the environment-locked virtual object is displayed in the user's view depends on the location and / or position and / or orientation of the location and / or object in the environment to which the virtual object is locked. In some embodiments, the computer system uses a stationary reference frame (e.g., a coordinate system fixed to a fixed location and / or object in the physical environment) to determine the position at which to display the environment-locked virtual object in the user's view. The environment-locked virtual object can be locked to a stationary part of the environment (e.g., the floor, a wall, a table, or other stationary object), or to a movable part of the environment (e.g., a vehicle, an animal, a person, or a representation of a part of the user's body such as the user's hand, wrist, arm, foot, etc. that moves independently of the user's view), such that the virtual object moves as the view or the part of the environment moves in order to maintain a fixed relationship between the virtual object and the part of the environment.

[0047] In some embodiments, an environment-locked or viewpoint-locked virtual object exhibits a delayed following behavior that reduces or delays the movement of the environment-locked or viewpoint-locked virtual object relative to the movement of a reference point that the virtual object is following. In some embodiments, when exhibiting the delayed following behavior, the computer system intentionally delays the movement of the virtual object when detecting movement of a reference point (e.g., a part of the environment, a viewpoint, or a point fixed relative to the viewpoint such as a point between 5 and 300 cm from the viewpoint) that the virtual object is following. For example, when the reference point (e.g., a part of the environment or a viewpoint) moves at a first speed, the virtual object is moved by the device so as to remain locked to the reference point but moves at a second speed that is slower than the first speed (e.g., until the reference point stops or decelerates its movement, at which time the virtual object begins to catch up to the reference point). In some embodiments, when the virtual object exhibits the delayed following behavior, the device ignores a small amount of movement of the reference point (e.g., movement of the reference point that is less than a threshold amount of movement such as movement of 0 to 5 degrees or 0 to 50 cm). For example, when the reference point (e.g., the part of the environment or the viewpoint to which the virtual object is locked) moves by a first amount, the distance between the reference point and the virtual object increases (e.g., because the virtual object is displayed so as to maintain a fixed or substantially fixed position relative to a viewpoint or a part of the environment that is different from the reference point to which the virtual object is locked), and when the reference point (e.g., the part of the environment or the viewpoint to which the virtual object is locked) moves by a second amount that is greater than the first amount, the distance between the reference point and the virtual object first increases (e.g., because the virtual object is displayed so as to maintain a fixed or substantially fixed position relative to a viewpoint or a part of the environment that is different from the reference point to which the virtual object is locked), and then the virtual object is moved by the computer system so as to maintain a fixed or substantially fixed position relative to the reference point, so that it decreases as the amount of movement of the reference point increases beyond a threshold (e.g., a "delayed following" threshold).In some embodiments, a virtual object that maintains a position substantially fixed relative to a reference point includes the virtual object being displayed within a threshold distance (e.g., 1, 2, 3, 5, 15, 20, 50 cm) of the reference point in one or more dimensions (e.g., up / down, left / right, and / or forward / backward relative to the position of the reference point).

[0048] Hardware: There are many different types of electronic systems that enable a person to perceive and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, heads-up displays (HUDs), vehicle windshields with integrated display functionality, windows with integrated display functionality, displays formed as lenses designed to be placed on a person's eye (similar to contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable controllers or handheld controllers with or without tactile feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system may have one or more speakers (singular or plural) and an integrated opaque display. Alternatively, a head-mounted system may be configured to receive an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing an image or video of the physical environment and / or one or more microphones for capturing the audio of the physical environment. A head-mounted system may have a transparent or translucent display instead of an opaque display. The transparent or translucent display may have a medium through which light representing an image is directed towards the person's eye. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scan light sources, or any combination of these technologies. The medium may be an optical waveguide, hologram medium, optical coupler, optical reflector, or any combination thereof. In one embodiment, the transparent or translucent display may be configured to selectively become opaque. A projection-based system can employ retinal projection technology that projects a graphical image onto a person's retina. The projection system may also be configured to project virtual objects into the physical environment, for example, as holograms or as physical surfaces.In some embodiments, the controller 110 is configured to manage and adjust the user's XR experience. In some embodiments, the controller 110 includes a suitable combination of software, firmware, and / or hardware. The controller 110 will be described in more detail below with reference to FIG. 2. In some embodiments, the controller 110 is a computing device that is local or remote to the scene 105 (e.g., a physical environment). For example, the controller 110 is a local server located within the scene 105. In another example, the controller 110 is a remote server (e.g., a cloud server, a central server, etc.) located outside the scene 105. In some embodiments, the controller 110 is communicatively coupled to a display generation component 120 (e.g., an HMD, a display, a projector, a touch screen, etc.) via one or more wired or wireless communication channels 144 (e.g., BLUETOOTH, IEEE802.11x, IEEE802.16x, IEEE802.3x, etc.). In another example, the controller 110 is included within a housing (e.g., a physical housing) of one or more of the display generation component 120 (e.g., an HMD, or a portable electronic device including a display and one or more processors, etc.), one or more of the input devices 125, one or more of the output devices 155, one or more of the sensors 190, and / or one or more of the peripheral devices 195, or shares the same physical housing or support structure as one or more of the above.

[0049] In some embodiments, the display generation component 120 is configured to provide the user with an XR experience (e.g., at least the visual component of the XR experience). In some embodiments, the display generation component 120 includes a suitable combination of software, firmware, and / or hardware. The display generation component 120 will be described in more detail below with reference to FIG. 3. In some embodiments, the functions of the controller 110 are provided by and / or combined with the display generation component 120.

[0050] According to some embodiments, the display generation component 120 provides an XR experience to the user while the user is virtually and / or physically present within the scene 105.

[0051] In some embodiments, the display generation component is worn on a part of the user's body (e.g., the user's own head or hand, etc.). Thus, the display generation component 120 includes one or more XR displays provided for displaying XR content. For example, in various embodiments, the display generation component 120 surrounds the user's field of view. In some embodiments, the display generation component 120 is a handheld device (such as a smartphone or a tablet) configured to present XR content, and the user holds a device having a display directed towards the user's field of view and a camera directed towards the scene 105. In some embodiments, the handheld device is optionally disposed within a housing worn on the user's head. In some embodiments, the handheld device is optionally disposed on a support in front of the user (e.g., a tripod). In some embodiments, the display generation component 120 is an XR chamber, housing, or room configured to present XR content while the user is not wearing or holding the display generation component 120. Many user interfaces described with reference to one type of hardware for displaying XR content (e.g., a device on a handheld or tripod) can be implemented on another type of hardware for displaying XR content (e.g., an HMD or other wearable computing device). For example, a user interface showing an interaction with XR content triggered based on an interaction occurring within the space in front of a handheld or tripod-mounted device can be implemented in the same way as an HMD where the interaction occurs within the space in front of the HMD and the response of the XR content is displayed via the HMD. Similarly, a user interface showing an interaction with CRG content triggered based on the movement of a handheld or tripod-mounted device relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eye(s), head, or hand)) can be implemented in the same way as an HMD where the movement is caused by the movement of the HMD relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eye(s), head, or hand)).

[0052] Although the relevant features of the operating environment 100 are shown in FIG. 1, those skilled in the art will understand from this disclosure that various other features are not shown for the sake of brevity so as not to obscure more appropriate aspects of the exemplary embodiments disclosed herein.

[0053] FIG. 2 is a block diagram of an example of the controller 110 according to some embodiments. Although certain features are shown, those skilled in the art will understand from this disclosure that various other features are not shown for the sake of brevity so as not to obscure more appropriate aspects of the embodiments disclosed herein. Thus, by way of non-limiting example, in some embodiments, the controller 110 includes one or more processing units 202 (e.g., microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), graphics processing units (GPUs), central processing units (CPUs), processing cores, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FIREWIRE, THUNDERBOLT, IEEE802.3x, IEEE802.11x, IEEE802.16x, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), BLUETOOTH, ZIGBEE, or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these and various other components.

[0054] In some embodiments, one or more communication buses 204 include circuitry for interconnecting and controlling communication between system components. In some embodiments, one or more I / O devices 206 include at least one of a keyboard, a mouse, a touchpad, a joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, and the like.

[0055] Memory 220 includes high-speed random access memory such as dynamic random-access memory (DRAM), static random-access memory (SRAM), double-data-rate random-access memory (DDRRAM), or other random access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. Memory 220 includes a non-transitory computer-readable storage medium. In some embodiments, memory 220, or the non-transitory computer-readable storage medium of memory 220, stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 230 and XR experience module 240.

[0056] Operating system 230 includes instructions for processing various basic system services and instructions for performing hardware-dependent tasks. In some embodiments, XR experience module 240 is configured to manage and coordinate one or more XR experiences for one or more users (e.g., a single XR experience for one or more users, or multiple XR experiences for each group of one or more users). To that end, in various embodiments, XR experience module 240 includes a data acquisition unit 241, a tracking unit 242, an adjustment unit 246, and a data transmission unit 248.

[0057] In some embodiments, the data acquisition unit 241 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the display generation component 120 of FIG. 1, and optionally one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. For that purpose, in various embodiments, the data acquisition unit 241 includes its instructions and / or logic, and the heuristics and metadata therefor.

[0058] In some embodiments, the tracking unit 242 is configured to map the scene 105 and track at least the position / location of the display generation component 120 with respect to the scene 105 of FIG. 1 and optionally with respect to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. For that purpose, in various embodiments, the tracking unit 242 includes its instructions and / or logic, and the heuristics and metadata therefor. In some embodiments, the tracking unit 242 includes a hand tracking unit 244 and / or an eye tracking unit 243. In some embodiments, the hand tracking unit 244 is configured to track the position / location of one or more parts of the user's hand and / or the movement of one or more parts of the user's hand with respect to the scene 105 of FIG. 1, with respect to the display generation component 120, and / or with respect to a coordinate system defined for the user's hand. The hand tracking unit 244 will be described in more detail below with respect to FIG. 4. In some embodiments, the eye tracking unit 243 is configured to track the position and movement of the user's line of sight (or more generally the user's eyes, face, or head) with respect to the scene 105 (e.g., the physical environment and / or the user (e.g., the user's hand)) or with respect to the XR content displayed via the display generation component 120. The eye tracking unit 243 will be described in more detail below with respect to FIG. 5.

[0059] In some embodiments, the adjustment unit 246 is configured to manage and adjust the XR experience presented to the user by the display generation component 120 and, optionally, by one or more of the output device 155 and / or the peripheral device 195. For that purpose, in various embodiments, the adjustment unit 246 includes its instructions and / or logic, as well as the heuristics and metadata therefor.

[0060] In some embodiments, the data transmission unit 248 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the display generation component 120 and, optionally, to one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. For that purpose, in various embodiments, the data transmission unit 248 includes its instructions and / or logic, as well as the heuristics and metadata therefor.

[0061] Although the data acquisition unit 241, the tracking unit 242 (including, for example, the eye tracking unit 243 and the hand tracking unit 244), the adjustment unit 246, and the data transmission unit 248 are shown as being present on a single device (e.g., the controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 241, the tracking unit 242 (including, for example, the eye tracking unit 243 and the hand tracking unit 244), the adjustment unit 246, and the data transmission unit 248 may be disposed within separate computing devices.

[0062] Furthermore, FIG. 2 is more intended to illustrate the functions of various features that may exist in a particular embodiment, as contrasted with the structural overview of the embodiments described herein. As will be recognized by those skilled in the art, the separately shown matters can be combined, and some matters can be separated. For example, some of the functional modules separately shown in FIG. 2 can be implemented within a single module, and the various functions of a single functional block can be executed by one or more functional blocks in various embodiments. The actual number of modules, as well as the specific division of particular functions and how functions are assigned therebetween, vary depending on the implementation form and, in some embodiments, depend in part on a particular combination of hardware, software, and / or firmware selected for a particular implementation form.

[0063] FIG. 3 is a block diagram of an example of a display generation component 120 according to some embodiments. Although certain features are shown, those skilled in the art will understand from this disclosure that various other features are not shown for the sake of brevity so as not to obscure more suitable aspects of the embodiments disclosed herein. For that purpose, by way of non-limiting example, in some embodiments, the display generation component 120 (e.g., an HMD) includes one or more processing units 302 (e.g., a microprocessor, ASIC, FPGA, GPU, CPU, processing core, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE802.3x, IEEE802.11x, IEEE802.16x, GSM, CDMA, TDMA, GPS, infrared, BLUETOOTH, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more XR displays 312, one or more optional inward and / or outward image sensors 314, a memory 320, and one or more communication buses 304 for interconnecting these and various other components.

[0064] In some embodiments, one or more communication buses 304 include circuitry that interconnects and controls communication between system components. In some embodiments, one or more I / O devices and sensors 306 include at least one of an inertial measurement unit (IMU), accelerometer, gyroscope, thermometer, one or more physiological sensors (e.g., blood pressure monitor, heart rate monitor, blood oxygen sensor, blood glucose sensor, etc.), one or more microphones, one or more speakers, a tactile engine, one or more depth sensors (e.g., structured light, time-of-flight, etc.).

[0065] In some embodiments, one or more XR displays 312 are configured to provide a user with an XR experience. In some embodiments, one or more XR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light emitting field effect transistor (OLET), organic light emitting diode (OLED), surface conduction electron emission device display (SED), field emission display (FED), quantum dot light emitting diode (QD-LED), MEMS, and / or similar display types. In some embodiments, one or more XR displays 312 correspond to waveguide displays, such as diffractive, reflective, polarized, holographic, etc. For example, the display generation component 120 (e.g., an HMD) includes a single XR display. In another example, the display generation component 120 includes an XR display for each eye of the user. In some embodiments, one or more XR displays 312 can present MR or VR content. In some embodiments, one or more XR displays 312 can present MR or VR content.

[0066] In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of a user's face, including the user's eyes (and may be referred to as an eye tracking camera). In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hand(s) and optionally the user's arm(s) (and may be referred to as a hand tracking camera). In some embodiments, one or more image sensors 314 are configured to face forward to acquire image data corresponding to a scene that the user views when the display generation component 120 (e.g., an HMD) is not present (and may be referred to as a scene camera). One or more optional image sensors 314 can include one or more RGB cameras, one or more infrared (IR) cameras, one or more event-based cameras, and / or the like (e.g., including a complementary metal-oxide semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor).

[0067] Memory 320 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid state storage devices. Memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. Memory 320 includes a non-transitory computer-readable storage medium. In some embodiments, memory 320, or the non-transitory computer-readable storage medium of memory 320, stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 330 and an XR presentation module 340.

[0068] The operating system 330 includes instructions for processing various basic system services and instructions for performing hardware-dependent tasks. In some embodiments, the XR presentation module 340 is configured to present XR content to a user via one or more XR displays 312. For that purpose, in various embodiments, the XR presentation module 340 includes a data acquisition unit 342, an XR presentation unit 344, an XR map generation unit 346, and a data transmission unit 348.

[0069] In some embodiments, the data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the controller 110 of FIG. 1. For that purpose, in various embodiments, the data acquisition unit 342 includes its instructions and / or logic, as well as the heuristics and metadata therefor.

[0070] In some embodiments, the XR presentation unit 344 is configured to present XR content via one or more XR displays 312. For that purpose, in various embodiments, the XR presentation unit 344 includes its instructions and / or logic, as well as the heuristics and metadata therefor.

[0071] In some embodiments, the XR map generation unit 346 is configured to generate an XR map (e.g., a 3D map of a composite reality scene or a map of a physical environment in which computer-generated objects can be placed to generate augmented reality) based on media content data. For that purpose, in various embodiments, the XR map generation unit 346 includes its instructions and / or logic, as well as the heuristics and metadata therefor.

[0072] In some embodiments, the data transmission unit 348 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the controller 110 and optionally one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. For that purpose, in various embodiments, the data transmission unit 348 includes its instructions and / or logic, as well as the heuristics and metadata therefor.

[0073] The data acquisition unit 342, the XR presentation unit 344, the XR map generation unit 346, and the data transmission unit 348 are shown as being present on a single device (e.g., the display generation component 120 of FIG. 1), but in other embodiments, it should be understood that any combination of the data acquisition unit 342, the XR presentation unit 344, the XR map generation unit 346, and the data transmission unit 348 may be arranged within separate computing devices.

[0074] Furthermore, FIG. 3 is more intended to illustrate the functions of various features that may exist in a particular implementation, as opposed to the structural overview of the embodiments described herein. As will be recognized by those skilled in the art, the separately shown matters can be combined, and some matters can be separated. For example, several functional modules separately shown in FIG. 3 can be implemented within a single module, and the various functions of a single functional block can be executed by one or more functional blocks in various embodiments. The actual number of modules, as well as the specific division of a particular function and how the functions are assigned among them, vary depending on the implementation, and in some embodiments, they depend in part on the particular combination of hardware, software, and / or firmware selected for a particular implementation.

[0075] FIG. 4 is a schematic diagram of an exemplary embodiment of the hand tracking device 140. In some embodiments, the hand tracking device 140 (FIG. 1) is for the scene 105 of FIG. 1 (e.g., for a part of the physical environment surrounding the user, for the display generation component 120, or for a part of the user (e.g., the user's face, eyes, or head), and / or for a coordinate system defined with respect to the user's hand), and / or the location / position of one or more parts of the user's hand, and / or the movement of one or more parts of the user's hand is tracked by the hand tracking unit 244 (FIG. 2). In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).

[0076] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras, etc.) that captures three-dimensional scene information including at least the hand 406 of a human user. The image sensor 404 captures hand images at a resolution sufficient to distinguish fingers and their respective positions. The image sensor 404 typically captures images of other parts of the user's body, or all of the body, and can have either a zoom function or a dedicated sensor with high magnification to capture hand images at the desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with other image sensors that capture the physical environment of the scene 105, or functions as an image sensor that captures the physical environment of the scene 105. In some embodiments, the image sensor 404 is positioned relative to the user or the user's environment such that the field of view of the image sensor or a portion thereof is used to define an interaction space in which hand movements captured by the image sensor are processed as input to the controller 110.

[0077] In some embodiments, the image sensor 404 outputs a sequence of frames including 3D map data (and optionally also color image data) to the controller 110, thereby extracting high-level information from the map data. This high-level information is typically provided to an application running on the controller via an application programming interface (API), and in response drives the display generation component 120. For example, the user can interact with software running on the controller 110 by moving their hand 406 and changing the pose of their hand.

[0078] In some embodiments, the image sensor 404 projects a spot pattern onto a scene that includes the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral shift of the spots of the pattern. This approach is advantageous in that the user does not need to hold or wear any kind of beacon, sensor, or other marker. This gives the depth coordinates of points in the scene relative to a predetermined reference plane at a specific distance from the image sensor 404. In the present disclosure, it is assumed that the image sensor 404 defines a series of orthogonal x, y, and z axes such that the depth coordinates of points in the scene correspond to the z component measured by the image sensor. Alternatively, the image sensor 404 (e.g., a hand tracking device) can use other 3D mapping methods such as stereoscopic imaging or time-of-flight measurement based on single or multiple cameras or other types of sensors.

[0079] In some embodiments, the hand tracking device 140 captures and processes a temporal sequence of depth maps that include the user's hand while the user moves their hand (e.g., the entire hand or one or more fingers). Software operating on a processor within the image sensor 404 and / or the controller 110 processes the 3D map data to extract hand patch descriptors within these depth maps. The software compares these descriptors to patch descriptors stored in the database 408 based on a previous learning process to estimate the hand pose in each frame. The pose typically includes the 3D locations of the user's hand joints and fingertips.

[0080] The software can also analyze the trajectories of the hand and / or fingers over multiple frames within a sequence to identify gestures. The pose estimation function described herein may be interleaved with the motion tracking function, such that patch-based pose estimation is only performed once per two (or more) frames, while tracking is used to detect changes in pose occurring over the remaining frames. Pose, motion, and gesture information is provided to an application program running on the controller 110 via the API described above. This program can, for example, move and modify an image presented on the display generation component 120 or perform other functions in response to pose and / or gesture information.

[0081] In some embodiments, gestures include air gestures. An air gesture is a gesture that is detected without (or independently of) the user touching an input element that is part of a device (e.g., computer system 101, one or more input devices 125, and / or hand tracking device 140), and is based on detected movement of a part of the user's body (e.g., head, one or more arms, one or more hands, one or more fingers, and / or one or more legs) in the air, including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground, or the distance of the user's hand relative to the ground), movement of the user's body relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of the user's other hand relative to the user's one hand, and / or movement of the user's finger relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tap gesture including movement of the hand in a predetermined pose by a predetermined amount and / or speed, or a shake gesture including rotation of a part of the user's body by a predetermined speed or amount).

[0082] In some embodiments, the input gestures used in the various examples and embodiments described herein are air gestures performed by movement of a user's finger(s) relative to other finger(s) or portion(s) of the user's hand for interacting with an XR environment (e.g., a virtual or mixed reality environment) according to some embodiments. In some embodiments, an air gesture is a gesture that is detected without the user touching (or independently of) an input element that is part of the device, and is based on detected movement of a part of the user's body, such as movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground, or the distance of the user's hand relative to the ground), movement of the user's body relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of the user's other hand relative to the user's one hand, and / or movement of the user's finger relative to another finger or portion of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tap gesture including movement of the hand in a predetermined pose with a predetermined amount and / or speed, or a shake gesture including rotation of a part of the user's body at a predetermined speed or amount).

[0083] In some embodiments where the input gesture is an air gesture (i.e., there is no physical contact with an input device that provides the computer system with information regarding which user interface element is the target of the user input, such as contact with a user interface element displayed on a touch screen, or contact with a mouse or trackpad to move a cursor to a user interface element), the gesture takes into account the user's attention (e.g., line of sight) to determine the target of the user input (e.g., in the case of direct input, as described below). Thus, in implementations that include air gestures, the input gesture is detected attention (e.g., line of sight) to a user interface element in combination with (e.g., simultaneously with) movement of the user's finger(s) and / or hand to perform a pinch and / or tap input, as described in more detail below, for example.

[0084] In some embodiments, an input gesture directed to a user interface object is executed directly or indirectly with reference to the user interface object. For example, a user input is executed directly on a user interface object in response to the user performing an input gesture with their hand at a position corresponding to the position of the user interface object in a three-dimensional environment (e.g., as determined based on the user's current perspective). In some embodiments, an input gesture is executed indirectly on a user interface object in accordance with a user who performs the input gesture while the position of the user's hand is not at a position corresponding to the position of the user interface object in the three-dimensional environment while detecting the user's attention (e.g., line of sight) to the user interface object. For example, in the case of a direct input gesture, the user can direct the user's input to the user interface object by starting a gesture at or near a position corresponding to the display position of the user interface object (e.g., within a distance of 0.5 cm, 1 cm, 5 cm, or 0 - 5 cm measured from the outer edge of the option or the central part of the option). In the case of an indirect input gesture, the user can direct the user's input to the user interface object by paying attention to the user interface object (e.g., by gazing at the user interface object), and while paying attention to the option, the user starts the input gesture (e.g., at any position detectable by the computer system) (e.g., at a position not corresponding to the display position of the user interface object).

[0085] In some embodiments, the input gestures (e.g., air gestures) used in the various examples and embodiments described herein include pinch inputs and tap inputs for interacting with a virtual or augmented reality environment according to some embodiments. For example, the pinch inputs and tap inputs described below are performed as air gestures.

[0086] In some embodiments, the pinch input is part of an air gesture that includes one or more of a pinch gesture, a long pinch gesture, a pinch and drag gesture, or a double pinch gesture. For example, the pinch gesture, which is an air gesture, involves moving two or more fingers of a hand so as to contact each other, i.e., optionally including an interruption immediately after contacting each other (e.g., within 0 to 1 second). The long pinch gesture, which is an air gesture, involves moving two or more fingers of a hand so as to contact each other for at least a threshold amount of time (e.g., at least 1 second) before detecting an interruption in the contact with each other. For example, the long pinch gesture includes the user holding a pinch gesture (e.g., when two or more fingers are in contact), and the long pinch gesture continues until an interruption in the contact between two or more fingers is detected. In some embodiments, the double pinch gesture, which is an air gesture, includes two (e.g., or more) pinch inputs (e.g., performed with the same hand) that are detected continuously and directly with each other (e.g., within a predetermined period). For example, the user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., breaks the contact between two or more fingers), and then performs a second pinch input within a predetermined period (e.g., within 1 second or within 2 seconds) after releasing the first pinch input.

[0087] In some embodiments, a pinch-and-drag gesture, which is an air gesture, includes a pinch gesture (e.g., a pinch gesture or a long pinch gesture) that is performed in relation to (e.g., after) a drag input that changes the position of the user's hand from a first position (e.g., the start position of the drag) to a second position (e.g., the end position of the resistance). In some embodiments, the user maintains the pinch gesture while performing the drag input and releases the pinch gesture (e.g., opens two or more fingers) to end the drag gesture (e.g., at the second position). In some embodiments, the pinch input and the drag input are performed by the same hand (e.g., the user pinches two or more fingers together and touches each other, and moves the same hand to a second position in the air with a drag gesture). In some embodiments, the pinch input is performed by the user's first hand and the drag input is performed by the user's second hand (e.g., the user's second hand moves from the first position to the second position in the air while the user continues the pinch input with the user's first hand). In some embodiments, an input gesture, which is an air gesture, includes an input (e.g., a pinch input and / or a tap input) that is performed using both of the user's hands. For example, the input gesture includes two (e.g., or more) pinch inputs that are performed in relation to each other (e.g., simultaneously with a predetermined period or within a predetermined period). For example, a first pinch gesture (e.g., a pinch input, a long pinch input, or a pinch-and-drag input) performed using the user's first hand, and in relation to performing the pinch input using the first hand, a second pinch input is performed using the other hand (e.g., the second hand of the user's two hands). In some embodiments, movement between the user's two hands (e.g., to increase and / or decrease the distance or relative orientation between the user's two hands).

[0088] In some embodiments, a tap input that is performed as an air gesture (e.g., directed at a user interface element) includes a movement (s) of the user's finger (s) toward the user interface element, optionally a movement of the user's hand toward the user interface element with the user's finger (s) extended toward the user interface element, a downward movement of the user's finger (e.g., mimicking a mouse click operation or a tap on a touch screen), or other predefined movement of the user's hand. In some embodiments, a tap input that is performed as an air gesture is detected based on movement characteristics of a finger or hand that performs a tap gesture movement away from the user's perspective and / or toward an object that is the target of the tap input where the end of the movement follows. In some embodiments, an end of the movement is detected based on a change in movement characteristics of a finger or hand that performs a tap gesture (e.g., movement away from the user's perspective and / or toward the object that is the target of the tap input, reversal of the direction of movement of the finger or hand, and / or reversal of the direction of acceleration of the movement of the finger or hand).

[0089] In some embodiments, a user's attention is determined to be directed to a portion of a three-dimensional environment based on detection of a line of sight directed at the portion of the three-dimensional environment (optionally, without requiring other conditions). In some embodiments, for a device to determine that a user's attention is directed to a portion of a three-dimensional environment, while the user's viewpoint is within a distance threshold from the portion of the three-dimensional environment, at least a threshold duration (e.g., dwell time), detection of a line of sight directed at the portion of the three-dimensional environment that requires, and / or detection of a line of sight directed at the portion of the three-dimensional environment that requires one or more additional conditions such as the line of sight being directed at the portion of the three-dimensional environment. Based on this, it is determined that the user's attention is directed to the portion of the three-dimensional environment. If one of the additional conditions is not met, the device determines that the attention is not directed to the portion of the three-dimensional environment where the line of sight is directed (e.g., until one or more additional conditions are met).

[0090] In some embodiments, the detection of the prepared state configuration of the user or a part of the user is detected by the computer system. The detection of the prepared state configuration of the hand is used by the computer system as an indication that the user is likely preparing to interact with the computer system using one or more air gesture inputs (e.g., pinch, tap, pinch and drag, double pinch, long pinch, or other air gestures described herein) performed by the hand. For example, the prepared state of the hand is determined based on whether the hand has a predetermined hand shape (e.g., a pre-pinch shape where the thumb and one or more fingers are extended and spaced to perform a pinch or grab gesture, or a pre-tap where one or more fingers are extended and the palm is facing away from the user), whether the hand is in a predetermined position relative to the user's perspective (e.g., under the user's head, above the user's waist, extended at least 15 cm, 20 cm, 25 cm, 30 cm, or 50 cm from the body), and / or whether the hand has moved in a particular way (e.g., moved towards the area in front of the user above the user's waist and under the user's head, or away from the user's body or legs). In some embodiments, the prepared state is used to determine whether the interaction element of the user interface responds to attention (e.g., line of sight) input.

[0091] In some embodiments, the software may be downloaded in electronic form to the controller 110, for example, over a network, or alternatively, may be provided on a tangible non-transitory medium such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is similarly stored in a memory associated with the controller 110. Alternatively or additionally, some or all of the described functions of the computer may be performed by dedicated hardware such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP). The controller 110 is shown in FIG. 4 as a separate unit from the image sensor 404 by way of example, but some or all of the processing functions of the controller may be associated with the image sensor 404 by a suitable microprocessor and software, or by dedicated circuitry within the housing of the image sensor 404 (e.g., a hand tracking device), or in other ways. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with the display generation component 120 (e.g., in a television set, a handheld device, or a head-mounted device), or using any other suitable computerized device such as a game console or a media player. The sensing function of the image sensor 404 can similarly be integrated with a computer or other computerized device controlled by the sensor output.

[0092] Figure 4 further includes a schematic diagram of a depth map 410 captured by an image sensor 404 according to some embodiments. The depth map includes a matrix of pixels having respective depth values, as described above. The pixels 412 corresponding to the hand 406 are segmented in this map from the background and the wrist. The luminance of each pixel in the depth map 410 is inversely proportional to the depth value, i.e., the measured z - distance from the image sensor 404, and the tone becomes darker as the depth increases. The controller 110 processes these depth values to identify and segment components (i.e., groups of adjacent pixels) of an image having the characteristics of a human hand. These characteristics can include, for example, the overall size, shape, and movement from frame to frame of a sequence of depth maps.

[0093] Figure 4 also schematically shows a hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406 according to some embodiments. In Figure 4, the hand skeleton 414 is superimposed on the background 416 of the hand segmented from the original depth map. In some embodiments, the hand (e.g., knuckles, fingertips, the center of the palm, the end of the hand connected to the wrist, etc.), and optionally major feature points on the wrist or arm connected to the hand are identified and placed on the hand skeleton 414. In some embodiments, the locations and movements of these major feature points over multiple image frames are used by the controller 110 to determine, according to some embodiments, the hand gesture or the current state of the hand being performed by the hand.

[0094] FIG. 5 shows an exemplary embodiment of an eye tracking device 130 (FIG. 1). In some embodiments, the eye tracking device 130 is controlled by an eye tracking unit 245 (FIG. 2) to track the position and movement of the user's line of sight with respect to scene 105 or XR content displayed via display generation component 120. In some embodiments, the eye tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, when the display generation component 120 is a head-mounted device such as a headset, helmet, goggles, or glasses, or a handheld device disposed on a wearable frame, the head-mounted device includes both a component for generating XR content for viewing by the user and a component for tracking the user's line of sight with respect to the XR content. In some embodiments, the eye tracking device 130 is separate from the display generation component 120. For example, when the display generation component is a handheld device or an XR chamber, the eye tracking device 130 is optionally a device separate from the handheld device or the XR chamber. In some embodiments, the eye tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, the head-mounted eye tracking device 130 is optionally used with a display generation component worn on the head or a display generation component not worn on the head. In some embodiments, the eye tracking device 130 is not a head-mounted device and is optionally used in combination with a head-mounted display generation component. In some embodiments, the eye tracking device 130 is not a head-mounted device and is optionally part of a non-head-mounted display generation component.

[0095] In some embodiments, the display generation component 120 uses a display mechanism (e.g., left and right near-eye display panels) that presents a frame including left and right images in front of the user's eyes to provide the user with a 3D virtual view. For example, the head-mounted display generation component may include left and right optical lenses (referred to herein as eyepieces) positioned between the display and the user's eyes. In some embodiments, the display generation component may include or be coupled to one or more external video cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display on which the user can directly view the physical environment and on which virtual objects can be displayed. In some embodiments, the display generation component projects virtual objects onto the physical environment. The virtual objects are projected, for example, onto a physical surface or as a hologram, whereby an individual can use the system to observe virtual objects superimposed on the physical environment. In such cases, separate display panels and image frames for the left and right eyes may not be required.

[0096] As shown in FIG. 5, in some embodiments, the gaze tracking device 130 includes at least one eye tracking camera (e.g., an infrared (IR) or near-IR (NIR) camera), and an illumination source that emits light (e.g., IR or NIR light) toward the user's eyes (e.g., an IR or NIR light source such as an array or ring of LEDs). The eye tracking camera may be directed at the user's eyes to directly receive reflected IR or NIR light from the light source, or alternatively, may be directed at a “hot” mirror disposed between the user's eyes and a display panel that reflects IR or NIR light from the eyes while allowing visual light to pass through to the eye tracking camera. The gaze tracking device 130 optionally captures an image of the user's eyes (e.g., as a video stream captured at 60-120 frames per second (fps)), analyzes the image to generate gaze tracking information, and communicates the gaze tracking information to the controller 110. In some embodiments, both of the user's eyes are tracked separately by respective eye tracking cameras and illumination sources. In some embodiments, only one of the user's eyes is tracked by an individual eye tracking camera and illumination source.

[0097] In some embodiments, the eye tracking device 130 is calibrated using a device-specific calibration process to determine the parameters of the eye tracking device for a particular operating environment 100, such as the 3D geometric relationships and parameters of the LED, camera, hot mirror (if present), eyepiece, and display screen. The device-specific calibration process may be performed at a factory or other facility prior to delivery of the AR / VR device to the end user. The device-specific calibration process may be an automatic calibration process or a manual calibration process. The user-specific calibration process may include an estimation of the eye parameters of a particular user, such as pupil location, foveal location, optical axis, visual axis, interpupillary distance, etc. According to some embodiments, once the device-specific and user-specific parameters for the eye tracking device 130 are determined, the images captured by the eye tracking camera are processed using the glint assist method to determine the user's current visual axis and viewpoint with respect to the display.

[0098] As shown in FIG. 5, the eye tracking device 130 (e.g., 130A or 130B) includes a line-of-sight tracking system including one or more eyepieces 520, at least one eye tracking camera 540 (e.g., an infrared (IR) or near-IR (NIR) camera) positioned on the side of the user's face where eye tracking is to be performed, and an illumination source 530 (e.g., an IR or NIR light source such as an array or ring of NIR light-emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eye(s) 592. The eye tracking camera 540 is positioned between the user's eye(s) 592 and a display 510 (e.g., a display panel on the left or right side of a head-mounted display, or a display of a handheld device, a projector, etc.), and may be directed toward a mirror 550 that reflects IR or NIR light from the eye(s) 592 while transmitting visible light (as shown, for example, at the top of FIG. 5), or may be directed toward the user's eye(s) 592 to receive the reflected IR or NIR light from the eye(s) 592 (as shown, for example, at the bottom of FIG. 5).

[0099] In some embodiments, the controller 110 renders an AR or VR frame 562 (e.g., the left and right frames of the left and right display panels) and provides the frame 562 to the display 510. The controller 110 uses the eye tracking input 542 from the eye tracking camera 540 for various purposes, such as when processing the frame 562 for display. The controller 110 optionally estimates the user's viewpoint on the display 510 based on the eye tracking input 542 obtained from the eye tracking camera 540 using a glint assist method or other suitable method. The viewpoint estimated from the eye tracking input 542 is optionally used to determine the direction in which the user is currently looking.

[0100] Examples of several possible use cases of the user's current line of sight direction are described below, but this is not intended to be limiting. As an exemplary use case, the controller 110 can render virtual content differently based on the determined line of sight direction of the user. For example, the controller 110 may generate virtual content at a higher resolution in the central visual region determined from the user's current line of sight direction than in the peripheral region. As another example, the controller may position or move virtual content within the view based at least in part on the user's current line of sight direction. As another example, the controller may display specific virtual content within the view based at least in part on the user's current line of sight direction. As another exemplary use case in an AR application, the controller 110 can capture the physical environment of the XR experience and direct an external camera to focus in the determined direction. The autofocus mechanism of the external camera can then focus on an object or surface within the environment that the user is currently looking at on the display 510. As another exemplary use case, the eyepiece 520 may be a focusable lens, and the line of sight tracking information is used by the controller to adjust the focus of the eyepiece 520 so that the virtual object the user is currently looking at has appropriate binocular convergence to match the convergence of the user's eyes 592. The controller 110 can utilize the line of sight tracking information to direct and adjust the focus of the eyepiece 520 so that the nearby object the user is looking at appears at the correct distance.

[0101] In some embodiments, the eye-tracking device is part of a head-mounted device that includes a display (e.g., display 510), two eyepieces (e.g., eyepiece(s) 520), an eye-tracking camera (e.g., eye-tracking camera(s) 540), and a light source (e.g., light source 530 (e.g., IR LED or NIR LED)) attached to a wearable housing. The light source emits light (e.g., IR light or NIR light) towards the user's eye(s) 592. In some embodiments, the light source may be arranged in a ring or circularly around each lens, as shown in FIG. 5. In some embodiments, eight light sources 530 (e.g., LEDs) are arranged around each lens 520 as an example. However, more or fewer light sources 530 may be used, and other arrangements and locations of the light sources 530 may be employed.

[0102] In some embodiments, the display 510 emits light within the visible light range and does not emit light within the IR or NIR range, so as not to introduce noise into the eye-tracking system. Note that the location and angle of the eye-tracking camera(s) 540 are given as an example and are not intended to be limiting. In some embodiments, a single eye-tracking camera 540 is located on each side of the user's face. In some embodiments, two or more NIR cameras 540 can be used on each side of the user's face. In some embodiments, a camera 540 with a wider field of view (FOV) and a camera 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, a camera 540 operating at one wavelength (e.g., 850 nm) and a camera 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.

[0103] An embodiment of a gaze tracking system as shown in FIG. 5 can be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide a user with an experience of computer-generated reality, virtual reality, augmented reality, and / or augmented virtuality.

[0104] FIG. 6A shows a glint-assisted gaze tracking pipeline according to some embodiments. In some embodiments, the gaze tracking pipeline is implemented by a glint-assisted gaze tracking system (e.g., an eye tracking device 130 as shown in FIGS. 1 and 5). The glint-assisted gaze tracking system can maintain a tracking state. Initially, the tracking state is off or "no". When in the tracking state, the glint-assisted gaze tracking system uses prior information from the previous frame when analyzing the current frame to track the pupil contour and glint within the current frame. When not in the tracking state, the glint-assisted gaze tracking system attempts to detect the pupil and glint within the current frame, and if successful, initializes the tracking state to "yes" and continues to the next frame in the tracking state.

[0105] As shown in FIG. 6A, a gaze tracking camera can capture left and right images of the user's left and right eyes. The captured images are then input into the gaze tracking pipeline for processing starting at 610. As indicated by the arrow returning to element 600, the gaze tracking system can continue to capture images of the user's eyes, for example, at a rate of 60 to 120 frames per second. In some embodiments, each set of captured images may be input into the pipeline for processing. However, in some embodiments, or under some conditions, not all captured frames are processed by the pipeline.

[0106] At 610, for the currently captured image, if the tracking state is yes, this method proceeds to element 640. At 610, if the tracking state is no, as shown at 620, the image is analyzed to detect the user's pupils and glints within the image. At 630, if the pupils and glints are detected normally, the method proceeds to element 640. If not detected normally, the method returns to element 610 to process the next image of the user's eyes.

[0107] At 640, when proceeding from element 610, the current frame is analyzed to track the pupils and glints based in part on previous information from the previous frame. At 640, when proceeding from element 630, the tracking state is initialized based on the detected pupils and glints within the current frame. The result of the processing at element 640 is checked to confirm that the result of the tracking or detection is reliable. For example, the result can be checked to determine whether a sufficient number of glints for performing pupil and gaze estimation are tracked or detected normally in the current frame. At 650, if the result is not reliable, the tracking state is set to no at element 660 and the method returns to element 610 to process the next image of the user's eyes. At 650, if the result is reliable, the method proceeds to element 670. At 670, the tracking state is set to yes (if not already yes), and the pupil and glint information is passed to element 680 to estimate the user's viewpoint.

[0108] FIG. 6A is intended to function as an example of an eye-tracking technique that can be used in a particular implementation. As will be recognized by those skilled in the art, other eye-tracking techniques that currently exist or are developed in the future can be used in the computer system 101 to provide an XR experience to the user according to various embodiments, instead of, or in combination with, the glint-assisted eye-tracking technique described herein.

[0109] In some embodiments, the captured portion of the real-world environment 602 is used to provide a user with an XR experience, e.g., a mixed reality environment in which one or more virtual objects are superimposed over a representation of the real-world environment 602.

[0110] FIG. 6B shows an exemplary environment of an electronic device 101 for providing an XR experience, according to some embodiments. In FIG. 6B, the real-world environment 602 includes the electronic device 101, the user 608, and real-world objects (e.g., the table 604). As shown in FIG. 6B, the electronic device 101 is optionally mounted on a tripod or otherwise fixed to the real-world environment 602 such that one or more hands of the user 608 are free (e.g., the user 608 is not optionally holding the device 101 with one or more hands). As described above, the device 101 optionally has one or more groups of sensors positioned on different sides of the device 101. For example, the device 101 optionally includes a sensor group 612-1 and a sensor group 612-2 located on the "back" side and the "front" side of the device 101, respectively (e.g., capable of capturing information from respective faces of the device 101). As used herein, the front side of the device 101 is the side facing the user 608, and the back side of the device 101 is the side facing away from the user 608.

[0111] In some embodiments, the sensor group 612-2 includes an eye-tracking unit (e.g., the eye-tracking unit 245 described above with reference to FIG. 2) that includes one or more sensors for tracking the user's eyes and / or line of sight, and the eye-tracking unit can "look at" the user 608 and track the user's eye(s) in the manner described above. In some embodiments, the eye-tracking unit of the device 101 can capture the movement, orientation, and / or line of sight of the user 608's eyes and process the movement, orientation, and / or line of sight as an input.

[0112] In some embodiments, the sensor group 612-1 includes a hand tracking unit (e.g., the hand tracking unit 243 described above with reference to FIG. 2) that can track one or more hands of the user 608 held on the "rear" side of the device 101, as shown in FIG. 6B. In some embodiments, the hand tracking unit is optionally included in the sensor group 612-2 such that the user 608 can additionally or alternatively hold one or more hands on the "front" side of the device 101 while the device 101 tracks the position of one or more hands. As described above, the hand tracking unit of the device 101 can capture the movement, position, and / or gesture of one or more hands of the user 608 and process the movement, position, and / or gesture as an input.

[0113] In some embodiments, the sensor group 612-1 optionally includes one or more sensors (e.g., the image sensor 404 described above with reference to FIG. 4) configured to capture an image of the real-world environment 602 including the table 604. As described above, the device 101 can capture an image of a portion (e.g., part or all) of the real-world environment 602 and present the captured portion of the real-world environment 602 to the user via one or more display generation components of the device 101 (e.g., a display of the device 101 optionally located on the side of the device 101 facing the user, opposite the side of the device 101 facing the captured portion of the real-world environment 602).

[0114] In some embodiments, the captured portion of the real-world environment 602 is used to provide the user with an XR experience, e.g., a composite reality environment in which one or more virtual objects are superimposed on a representation of the real-world environment 602.

[0115] Accordingly, the description herein describes some embodiments of a three-dimensional environment (e.g., an XR environment) that includes representations of real-world objects and representations of virtual objects. For example, the three-dimensional environment optionally includes a representation of a table that exists within a physical environment and that is captured and displayed within the three-dimensional environment (e.g., actively via a camera and display of a computer system or passively via a transparent or translucent display of the computer system). As described above, the three-dimensional environment is optionally a mixed reality system based on a physical environment that is captured by one or more sensors of a computer system and displayed via a display generation component. As a mixed reality system, the computer system can optionally selectively display portions and / or objects of the physical environment such that they appear to exist within the three-dimensional environment displayed by the computer system. Similarly, in the real world, the computer system can optionally display virtual objects within the three-dimensional environment such that they appear to exist within the real world (e.g., the physical environment) by placing virtual objects at respective locations within the three-dimensional environment that have corresponding locations in the real world. For example, the computer system can optionally display a vase such that it appears as if a real vase is placed on a table within the physical environment. In some embodiments, individual locations within the three-dimensional environment have corresponding locations within the physical environment.Thus, when a computer system is described as displaying a virtual object at an individual location with respect to a physical object (e.g., the location of a user's hand, or near it, or on a physical table, or near it, etc.), the computer system displays the virtual object at a specific location within a three-dimensional environment such that the virtual object appears as if it is at or near the physical object within the physical world. (For example, the virtual object is displayed at a location within the three-dimensional environment corresponding to the location within the physical environment where the virtual object would be displayed if it were a real object at that specific location.)

[0116] In some embodiments, real-world objects that exist within the physical environment displayed within the three-dimensional environment (e.g., and / or real-world objects viewable through a display generation component) can interact with virtual objects that exist only within the three-dimensional environment. For example, the three-dimensional environment can include a table and a vase placed on the table, where the table is a view (or representation) of the physical table within the physical environment and the vase is a virtual object.

[0117] Similarly, the user can optionally use one or more hands to interact with virtual objects within the three-dimensional environment as if the virtual objects were real objects within the physical environment. For example, as described above, one or more sensors of the computer system can optionally capture one or more of the user's hands and display a representation of the user's hand within the three-dimensional environment (e.g., in a manner similar to that described above for displaying real-world objects within the three-dimensional environment), or, in some embodiments, due to the transparency / translucency of a portion of the display generation component that displays the user interface, or a projection of the user interface onto a transparent / translucent surface, or a projection of the user interface onto the user's eye or the user's field of view, the user's hand can be seen through the display generation component by the user's ability to see the physical environment through the user interface. Thus, in some embodiments, the user's hands are displayed at individual locations within the three-dimensional environment and are treated as if they were objects within the three-dimensional environment that can interact with virtual objects within the three-dimensional environment as if they were actual physical objects within the physical environment. In some embodiments, the computer system can update the display of the representation of the user's hand within the three-dimensional environment in conjunction with the movement of the user's hand within the physical environment.

[0118] In some of the embodiments described below, for example, for the purpose of determining whether a physical object is directly interacting with a virtual object (e.g., whether a hand is touching, grasping, holding a virtual object, etc., or is within a threshold distance from the virtual object), the computer system can optionally determine the "effective" distance between a physical object in the physical world and a virtual object in the three-dimensional environment. For example, a hand that directly interacts with a virtual object can optionally include one or more of a finger of a hand pressing a virtual button, a user's hand grasping a virtual vase, two fingers of a user's hand pinching / holding the user interface of an application together, and other types of interactions described herein. For example, when determining whether a user is interacting with a virtual object and / or how the user is interacting with the virtual object, the computer system can optionally determine the distance between the user's hand and the virtual object. In some embodiments, the computer system determines the distance between the user's hand and the virtual object by determining the distance between the location of the hand in the three-dimensional environment and the location of the target virtual object in the three-dimensional environment. For example, one or more of the user's hands are located at a specific position in the physical world, which the computer system can optionally capture and display at a specific corresponding position in the three-dimensional environment (e.g., the position in the three-dimensional environment where the hand is displayed if the hand is a virtual hand rather than a physical hand). The position of the hand in the three-dimensional environment is optionally compared with the position of the target virtual object in the three-dimensional environment to determine the distance between one or more of the user's hands and the virtual object. In some embodiments, the computer system optionally determines the distance between the physical object and the virtual object by comparing positions in the physical world (as opposed to, for example, comparing positions in the three-dimensional environment).For example, when determining the distance between one or more hands of a user and a virtual object, the computer system optionally determines the corresponding location in the physical world of the virtual object (e.g., the position where the virtual object would be located in the physical world if the virtual object were a physical object rather than a virtual object), and then determines the distance between the corresponding physical position and one or more hands of the user. In some embodiments, the same technique is optionally used to determine the distance between any physical object and any virtual object. Thus, as described herein, when determining whether a physical object is in contact with a virtual object or whether a physical object is within a threshold distance of a virtual object, the computer system optionally performs any of the techniques described above to map the location of the physical object to the three-dimensional environment and / or to map the location of the virtual object to the physical environment.

[0119] In some embodiments, the same or similar techniques are used to determine where and what the user's line of sight is directed at and / or where and what a physical stylus held by the user is directed at. For example, if the user's line of sight is directed at a particular position in the physical environment, the computer system optionally determines the corresponding position (e.g., the virtual position of the line of sight) in the three-dimensional environment, and if a virtual object is located at that corresponding virtual position, the computer system optionally determines that the user's line of sight is directed at that virtual object. Similarly, the computer system can optionally determine where the stylus is pointing in the physical environment based on the orientation of the physical stylus. In some embodiments, based on this determination, the computer system determines the corresponding virtual position in the three-dimensional environment corresponding to the location in the physical environment that the stylus is pointing at, and optionally determines that the stylus is pointing at the corresponding virtual position in the three-dimensional environment.

[0120] Similarly, the embodiments described herein may refer to the location of a user (e.g., a user of a computer system), and / or the location of a computer system within a three-dimensional environment. In some embodiments, a user of a computer system is holding, wearing, or otherwise located in or near a computer system. Thus, in some embodiments, the location of the computer system is used as a proxy for the location of the user. In some embodiments, the location of a computer system and / or a user within a physical environment corresponds to an individual location within a three-dimensional environment. For example, if a user stands at a location facing an individual portion of a physical environment displayed by a display generation component, the location of the computer system is the location within the physical environment (and its corresponding location within the three-dimensional environment) where the user would view an object within the physical environment in the same position, orientation, and / or size (e.g., absolutely and / or relative to each other) as the object was displayed by the display generation component of the computer system within the three-dimensional environment. Similarly, if a virtual object displayed within a three-dimensional environment were a physical object within a physical environment (e.g., the virtual object were located at the same location within the physical environment as within the three-dimensional environment and had the same size and orientation within the physical environment as within the three-dimensional environment), the location of the computer system and / or the user is the position where the user would view the virtual object within the physical environment in the same position, orientation, and / or size (e.g., absolutely, and / or relative to each other, and in the context of real-world objects) as the virtual object was displayed by the display generation component of the computer system within the three-dimensional environment.

[0121] In the present disclosure, various input methods are described with respect to interaction with a computer system. If one example is provided using one input device or input method and another example is provided using another input device or input method, it should be understood that each example may be compatible with and optionally utilize the input device or input method described in another example. Similarly, various output methods are described with respect to interaction with a computer system. If one example is provided using one output device or output method and another example is provided using another output device or output method, it should be understood that each example may be compatible with and optionally utilize the output device or output method described in another example. Similarly, various methods are described with respect to interaction with a virtual environment or a mixed reality environment via a computer system. If one example is provided using interaction with a virtual environment and another example is provided using a mixed reality environment, it should be understood that each example may be compatible with and optionally utilize the method described in another example. Accordingly, the present disclosure discloses embodiments that are combinations of features of multiple examples without exhaustively listing all features of the embodiments in the description of each exemplary embodiment.

[0122] User Interface and Related Processes Here, attention is paid to embodiments of a user interface ("UI") and related processes that may be executed in a computer system such as a portable multifunctional device or a head-mounted device, including a display generation component, one or more input devices, and (optionally) one or more cameras.

[0123] Figures 7A-7G show examples of how an electronic device updates the spatial arrangement of one or more virtual objects in a three-dimensional environment according to some embodiments.

[0124] FIG. 7A shows an electronic device 101 that displays a three-dimensional environment 702 via a display generation component 120. In some embodiments, it should be understood that the electronic device 101 may utilize one or more of the techniques described with reference to FIGS. 7A-7G within a two-dimensional environment without departing from the scope of the present disclosure. As described above with reference to FIGS. 1-6, the electronic device 101 optionally includes a display generation component 120 (e.g., a touch screen) and a plurality of image sensors 314. The image sensors optionally include one or more of a visible light camera, an infrared camera, a depth sensor, or any other sensor, and the electronic device 101 can be used to capture one or more images of the user or a part of the user while the user is interacting with the electronic device 101. In some embodiments, the display generation component 120 is a touch screen that can detect gestures and movements of the user's hand. In some embodiments, the user interface shown below can be implemented on a head-mounted display that includes a display generation component that displays the user interface to the user, sensors that detect the physical environment and / or the movement of the user's hand (e.g., external sensors facing outward from the user), and / or the user's line of sight (e.g., internal sensors facing inward toward the user's face).

[0125] In FIG. 7A, the electronic device 101 displays a three-dimensional environment 702 that includes, for example, a first user interface 704 of a first application, a second user interface 706 of a second application, a representation 708 of a second user of a second electronic device having access to the three-dimensional environment 702, and a representation 710 of the second electronic device. In some embodiments, the electronic device 101 displays the three-dimensional environment 702 from the perspective of the user of the electronic device 101. In some embodiments, the perspective of the user of the electronic device 101 is located at a location within the three-dimensional environment 702 that corresponds to the physical location of the electronic device 101 within the physical environment of the electronic device 101.

[0126] In some embodiments, the first application is private to the electronic device 101 and is not shared with the second electronic device. Thus, the electronic device 101 displays the user interface 704 of the first application, while the second electronic device does not display the user interface 704 of the first application. In some embodiments, the second application is shared between the electronic device 101 and the second electronic device. Accordingly, the first electronic device 101 and the second electronic device display the user interface 706 of the second application.

[0127] In some embodiments, the representation 708 of the second user and / or the representation 710 of the second electronic device are views of the second user and the second electronic device through the transparent portion of the display generation component 120 (e.g., a true or actual passthrough). For example, the second electronic device is located in the physical environment of the electronic device 101 having the same spatial relationship to the electronic device 120 as the spatial relationship between the representation 710 of the second electronic device and the user's perspective in the three-dimensional environment 702. In some embodiments, the representation 708 of the second user and / or the representation 710 of the second electronic device are representations (e.g., virtual or video passthrough) displayed via the display generation component 120. For example, the second electronic device is located remotely from the electronic device 101 or has a spatial relationship to the electronic device 101 that is different from the spatial relationship between the representation 710 of the second electronic device and the user's perspective within the three-dimensional environment. In some embodiments, the electronic device 101 displays the representations 708 and 710 via the display generation component, and the second electronic device is located in the physical environment of the electronic device 101 having the same spatial relationship to the electronic device 120 as the spatial relationship between the representation 710 of the second electronic device and the user's perspective within the three-dimensional environment 702. In some embodiments, the electronic device 101 displays the representation 708 of the second user without displaying the representation 710 of the second electronic device.

[0128] In some embodiments, the spatial arrangement of user interfaces 704 and 706 and representations 708 and 710 with respect to the user's perspective of the electronic device 101 meets one or more criteria that specify a range of positions and / or orientations of (e.g., virtual) objects with respect to the user's perspective within the three-dimensional environment 702. For example, the one or more criteria define the spatial arrangement of user interfaces 704 and 706 and representations 708 and 710 with respect to the user's perspective of the electronic device 101 that enables the user of the electronic device 101 (and optionally a second user of a second electronic device) to view and interact with the user interfaces 704 and 706 and representations 708 and 710. For example, the spatial arrangement of user interfaces 704 and 706 and representations 708 and 710 with respect to the user's perspective shown in FIG. 7A enables the user of the electronic device 101 to view and / or interact with the user interfaces 704 and 706.

[0129] In some embodiments, as shown in FIG. 7A, the electronic device 101 updates the user's perspective from which the electronic device 101 displays the three-dimensional environment 702 in response to detecting movement of the electronic device 101 (e.g., and / or the display generation component 120) in the physical environment of the electronic device 101 (e.g., and / or the display generation component 120). In some embodiments, the electronic device 101 updates the user's perspective in response to detecting movement of the user of the electronic device 101 in the physical environment of the user and / or the electronic device 101 and / or the display generation component 120 (e.g., via the input device 314).

[0130] Figure 7B shows an example of an electronic device 101 that updates the display of a three-dimensional environment 702 in response to detection of movement of the electronic device 101 shown in Figure 7A. In some embodiments, the electronic device 101 updates the three-dimensional environment 702 as well, in response to input corresponding to a request to update the position of virtual objects (e.g., user interfaces 704 and 706 and representations 708 and 710) together or separately, according to one or more steps of method 1000. As shown in Figure 7B, the position of the electronic device 101 in the physical environment of the electronic device 101 is updated in response to the movement shown in Figure 7A, and the user's perspective in the three-dimensional environment 702 in which the electronic device 101 displays the three-dimensional environment 702 is updated accordingly. For example, user interfaces 704 and 706 and representations 708 and 710 are farther away from the user's perspective in Figure 7B than in Figure 7A.

[0131] In some embodiments, the spatial arrangement of user interfaces 704 and 706 and representations 708 and 710 with respect to the user's perspective of the electronic device 101 does not meet one or more criteria. For example, the criteria are not met because user interface 704 and / or user interface 706 is greater than a threshold distance from the user's perspective in the three-dimensional environment (e.g., a threshold distance for readability and / or ability of the user to provide input to user interface 704 and / or 706). Referring to method 800, additional criteria are provided below. In some embodiments, because the spatial arrangement does not meet one or more criteria, when selected, the electronic device 101 causes the three-dimensional environment 702 to be updated on the electronic device 101 such that the spatial arrangement of user interfaces 704 and 706 and representations 708 and 710 with respect to the user's perspective meets one or more criteria (e.g., "recents" the three-dimensional environment 702), by displaying a selectable option 712.

[0132] In some embodiments, the electronic device 101 detects an input corresponding to a request to recenter the three-dimensional environment 702, such as detecting a selection of the selectable recentering option 712 or detecting an input directed to a mechanical input device communicating with the electronic device 101, such as the button 703.

[0133] In some embodiments, the electronic device 101 detects a selection of one of the user interface elements, such as the selectable option 712, by detecting an indirect selection input, a direct selection input, an air gesture selection input, or an input device selection input. In some embodiments, detecting a selection of a user interface element includes detecting that the user's hand 713a has performed an individual gesture (e.g., "hand state B"). In some embodiments, detecting an indirect selection input includes detecting that the user's hand 713a is performing a selection gesture (e.g., "hand state B"), such as a pinch hand gesture in which the user touches their thumb to another finger of their hand, while detecting the user's line of sight directed at an individual user interface element via the input device 314. In some embodiments, detecting a direct selection input includes, via the input device 314, detecting that the user's hand 713a is performing a selection gesture (e.g., "hand state B"), such as a pinch gesture within a predetermined threshold distance (e.g., 1, 2, 3, 5, 10, 15, or 30 centimeters) of the location of an individual user interface element, or a "press" gesture in which the user's hand or finger "presses" the location of an individual user interface element while in a pointing hand shape. In some embodiments, detecting an air gesture input includes detecting the user's line of sight directed at an individual user interface element while detecting a press gesture at the location of an air gesture user interface element displayed in the three-dimensional environment 702 via the display generation component 120. In some embodiments, detecting an input device selection includes detecting an operation of a mechanical input device (e.g., a stylus, a mouse, a keyboard, a trackpad, etc.) in a predetermined manner corresponding to a selection of a user interface element while the cursor controlled by the input device is associated with the location of an individual user interface element and / or while the user's line of sight is directed at the individual user interface element.

[0134] In some embodiments, button 703 is a multi-functional button. For example, in response to detecting that the user has pressed button 703a for less than a threshold period (e.g., 0.1, 0.2, 0.3, 0.5, 1, or 2 seconds), electronic device 101 displays the home user interface (e.g., the user interface of the operating system of electronic device 101), and in response to detecting that the user has pressed button 703 for more than the threshold period, electronic device 101 recenters the three-dimensional environment 702. In some embodiments, electronic device 101 communicates with a faucet or dial configured to detect when the user presses or turns the dial. In some embodiments, in response to detecting that the user has turned the dial, the electronic device updates the level of visual emphasis of virtual objects (e.g., user interfaces 704 and 706, representations 708 and 710) relative to other portions of the three-dimensional environment 702 according to the direction and magnitude in which the dial was turned. In some embodiments, in response to detecting that the user has pressed the dial for less than the threshold period, electronic device 101 displays the home user interface, and in response to detecting that the user has pressed the dial for more than the threshold period, electronic device 101 recenters the three-dimensional environment 702.

[0135] In some embodiments, the input directed to the recentering option 712 of FIG. 7B and the input directed to button 703 correspond to a request to recenter the three-dimensional environment 702. Although FIG. 7B shows both inputs, it should be understood that in some embodiments, electronic device 101 does not detect the inputs simultaneously, but rather detects one of these inputs at a time. In some embodiments, in response to one (or more) of the inputs shown in FIG. 7B, electronic device 101 updates the three-dimensional environment 702 as shown in FIG. 7C. In some embodiments, electronic device 101 does not recenter the three-dimensional environment 702 unless an input corresponding to a request to do so is received (e.g., even if the spatial arrangement of virtual objects relative to the user's perspective does not meet one or more criteria).

[0136] FIG. 7C shows an electronic device 101 that displays a three-dimensional environment 702 updated according to one or more of the inputs shown in FIG. 7B. In some embodiments, the user's perspective of the electronic device 101 is away from the user interfaces 704 and 706 and the representations 708 and 710 as shown in FIG. 7A before the electronic device 101 receives the recentering input shown in FIG. 7B. Thus, in response to the recentering input, the electronic device 101 updates the positions of the user interfaces 704 and 706 and the representations 708 and 710 in the three-dimensional environment 702 such that the spatial arrangement of the user interfaces 704 and 706 and the representations 708 and 710 with respect to the user's perspective satisfies one or more of the above criteria. For example, the electronic device 101 updates the position and / or orientation of the user interfaces 704 and 706 and the representations 708 and 710 based on the updated perspective of the user in the three-dimensional environment 702 such that the spatial arrangement of the user interfaces 704 and 706 and the representations 708 and 710 with respect to the updated perspective of the user satisfies one or more of the criteria. In some embodiments, when the spatial arrangement of the user interfaces 704 and 706 and the representations 708 and 710 with respect to the updated perspective of the user satisfies one or more of the criteria, the user interfaces 704 and 706 are at a distance from and / or an orientation with respect to the user's perspective of the electronic device 101 that facilitates user interaction with the user interfaces.

[0137] In some embodiments, the electronic device 101 uses a spatial template associated with one or more virtual objects to recenter the three-dimensional environment 702. For example, the user interface 706 of the second application is shared content (e.g., video content) consumed by the user of the electronic device 101 and the second user of the second electronic device, and is associated with a shared content spatial template. In some embodiments, the shared content spatial template positions the user interface 706 and the representations 708 and 710 with respect to the perspective of the user of the electronic device 101 such that the perspectives of the user of the electronic device 101 and the user of the second electronic device (e.g., represented by the representations 708 and 710) are on the same side of the user interface 706 at a distance at which the content included in the user interface 706 is visible to the user.

[0138] Accordingly, as shown in FIGS. 7B - 7C, in some embodiments, in response to detecting a request to recenter the three - dimensional environment 702 after the user's perspective has been updated (e.g., in response to the movement of the electronic device 101 shown in FIG. 7A), the electronic device 101 recenters the three - dimensional environment 702 by updating the positions and / or orientations of the user interfaces 704 and 706 and the representations 708 and 710. In some embodiments, the electronic device 101 detects an input corresponding to a request to move the user interfaces 704 and 706 and the representations 708 and 710 away from the user's perspective before detecting the recentering input. In some embodiments, in response to detecting the recentering input after detecting the request(s) to move the user interfaces 704 and 706 and the representations 708 and 710, the electronic device 101 updates the user's perspective within the three - dimensional environment 702 (e.g., as opposed to updating the positions and / or orientations of the user interfaces 704 and 706 and / or the representations 708 and 710). For example, the electronic device 101 moves the user's perspective to a new location within the three - dimensional environment (e.g., independent of the location of the electronic device 101 within the physical environment) and updates the spatial arrangement of the user interfaces 704 and 706 and the representations 708 and 710 with respect to the user's perspective to meet one or more of the above criteria.

[0139] FIG. 7C also shows an example of an electronic device 101 that detects an input corresponding to a request to update the position of a user interface 704 of a first application within a three-dimensional environment 702. For example, hand 713a provides a selection input directed to the user interface 704 of the first application, as described above. In some embodiments, after detecting a portion of the selection input, e.g., the thumb is touching another finger in a pinch hand shape, but before separating the thumb and finger to complete the pinch gesture, the electronic device 101 detects the movement of hand 713a. In some embodiments, the electronic device 101 updates the position of the user interface 704 of the first application according to the detected movement of hand 713a during the selection input, as shown in FIG. 7D.

[0140] FIG. 7D shows an example of an electronic device 101 that displays the user interface 704 of the first application at an updated location within the three-dimensional environment 702 in response to the input shown in FIG. 7C. In some embodiments, updating the location of the user interface 704 of the first application causes the spatial arrangement of the user interfaces 704 and 706 and the representations 708 and 710 with respect to the user's perspective to not meet one or more criteria. For example, the user interface 704 of the first application is not entirely within the field of view of the electronic device 101. As another example, the spacing between the user interface 704 of the first application and the user interface 706 of the second application exceeds a predetermined threshold included in one or more criteria described below with reference to method 800. In some embodiments, in response to detecting that one or more criteria are not met, the electronic device 101 displays a re-centering option 712.

[0141] In some embodiments, the electronic device 101 detects an input corresponding to a request to recenter the three-dimensional environment 702. For example, the electronic device 101 detects a selection of the recentering option 712 by the hand 713a or detects an input via the button 703 in FIG. 7D, as described in more detail above. In some embodiments, in response to detecting an input corresponding to a request to recenter the three-dimensional environment 101 after the position of the user interface 704 has been updated, the electronic device 101 updates the position of the user interface 704 to meet one or more criteria, as shown in FIG. 7E.

[0142] FIG. 7E shows an electronic device 101 that displays the user interface 704 of a first application at an updated location (e.g., and / or orientation) within the three-dimensional environment 702 in response to the recentering input shown in FIG. 7D. In some embodiments, the position (e.g., and / or orientation) of the user interface 704 of the first application has been updated before receiving an input corresponding to a request to recenter the three-dimensional environment 702 (e.g., as opposed to a movement of the user's perspective that causes one or more criteria not to be met), so the electronic device 101 updates the position (e.g., and / or orientation) of the user interface 704 of the first application in the three-dimensional environment 702 in response to the recentering input (e.g., without updating the user's perspective). For example, the electronic device 101 updates the position and / or orientation of the user interface 704 of the first application to be within a threshold distance (e.g., included in one or more criteria) of the user interface 706 of the second application, ensuring that the user interface 704 of the first application is within the field of view of the electronic device 101. In some embodiments, if the electronic device 101 detects the movement of a different virtual object, such as the user interface 706 of the second application, before detecting the recentering input, the electronic device 101 instead updates the position (e.g., and / or orientation) of that virtual object in response to the recentering input.

[0143] In some embodiments, before the electronic device 101 detects a recentering input, regardless of which electronic device provided a previous request to reposition and / or reorient the user interface 704, the electronic device 101 updates the position of the user interface 704 of the first application in response to the recentering input. For example, if a second user of a second electronic device updates the position of the user interface 706 of a second application (e.g., an application that both electronic devices have access to), the second electronic device and the electronic device 101 display the user interface 706 at the updated position and / or orientation within the three-dimensional environment 702 according to the input provided by the second user of the second electronic device. If the electronic device 101 detects a recentering input while the user interface 706 is being displayed at the position and orientation according to the input detected by the second electronic device, the electronic device 101 updates the position and / or orientation of the user interface 706 in the three-dimensional environment 702 to meet one or more of the criteria described above. In some embodiments, updating the position of the user interface 706 within the three-dimensional environment 702 causes the electronic device 101 and the second electronic device to display the user interface 706 at the updated position and / or location.

[0144] Figure 7F shows an example of an electronic device 101 that displays a three-dimensional environment 702 including a user interface 714 of a third application. As shown in Figure 7F, a part of the user interface 714 of the third application is within the field of view of the electronic device 101, and a part of the user interface 714 of the third application is outside the field of view of the electronic device 101. The second user 708 is included in the three-dimensional environment 702, but since the second user 708 is not within the field of view of the electronic device 101, the electronic device 101 does not display a representation of the second user or a representation of the second electronic device. In some embodiments, the user's perspective, the user interface 714 of the third application, and the spatial arrangement of the second user 708 do not meet one or more of the above criteria. For example, one or more criteria are not met because at least the second user 708 is not within the field of view of the electronic device 101 and a part of the user interface 714 of the third application is not within the field of view of the electronic device 101. Since one or more criteria are not met, the electronic device 101 displays a recentering option 712 within the three-dimensional environment 702.

[0145] In some embodiments, as shown in Figure 7F, the electronic device 101 detects an input corresponding to a request to recenter the three-dimensional environment 702. In some embodiments, the input includes detecting an input directed to the button 703. In some embodiments, the input includes detecting a selection of the recentering option 712 by the hand 713a. Although Figure 7F shows both inputs, it should be understood that in some embodiments, the inputs are not detected simultaneously. In some embodiments, in response to the recentering input, the electronic device 101 updates the three-dimensional environment 702 as shown in Figure 7G.

[0146] FIG. 7G shows an example of an electronic device 101 that displays a three-dimensional environment 702 in response to the recentering input described above with reference to FIG. 7F. In some embodiments, in response to the recentering input, the electronic device 101 updates the user's viewpoint within the three-dimensional environment 702 such that the field of view of the electronic device 101 includes the representation 708 of a second user, the representation 710 of a second electronic device, and the user interface 714 of a third application.

[0147] In some embodiments, the electronic device 101 updates the three-dimensional environment 702 according to a shared activity space template associated with the user interface 714 of the third application. For example, the user interface 714 of the third application includes a virtual board game or other content intended to be viewed by the user from different sides of the user interface 714 of the third application. Thus, in some embodiments, updating the three-dimensional environment 702 in response to a recentering request includes updating the spatial arrangement of the user's viewpoint, the user interface 714 of the third application, the representation 708 of the second user, and the representation 710 of the second electronic device to position the user interface 714 of the third application between the user's viewpoint and the representation 708 of the second user.

[0148] In some embodiments, the electronic device 101 updates the three-dimensional environment 702 according to another spatial template depending on which spatial template is applied to the three-dimensional environment 702. For example, in some embodiments, the three-dimensional environment 702 is associated with a group activity space template that may not be associated with a specific user interface within the three-dimensional environment 702. For example, the group activity space template is used for (e.g., virtual) meetings between users. In some embodiments, the electronic device 101 applies the group activity space template by recentering the user's viewpoints such that the user's viewpoints face each other and multiple users are in positions within the field of view of the user of the electronic device 101.

[0149] In some embodiments, in FIG. 7G, the three-dimensional environment 702 is updated according to the position and / or movement of the representation 708 of the second user. For example, recentering the three-dimensional environment 702 includes updating the three-dimensional environment 702 so that the users face each other with respect to the user interface 714 of the third application. In some embodiments, when a recentering input is received, if the second user 708 has a position different from the position shown in FIG. 7F within the three-dimensional environment 702, the electronic device 101, in response to the recentering input shown in FIG. 7F, centers the three-dimensional environment 702 in a manner different from the method shown in FIG. 7G (e.g., according to the position of the second user 708 within the three-dimensional environment 702).

[0150] Additional or alternative details regarding the embodiments shown in FIGS. 7A-7G are provided below in the description of method 800 described with reference to FIGS. 8A-8K.

[0151] FIGS. 8A-8K are flowcharts showing a method for updating the spatial arrangement of one or more virtual objects in a three-dimensional environment according to some embodiments. In some embodiments, method 800 is executed in a computer system (e.g., computer system 101 of FIG. 1) that includes a display generation component (e.g., display generation component 120 of FIGS. 1, 3, and 4) (e.g., a head-up display, a display, a touch screen, a projector, etc.) and one or more cameras (e.g., a camera that the user holds non-manually and faces downward (e.g., a color sensor, an infrared sensor, and other depth sensing cameras) or a camera that faces forward from the user's head). In some embodiments, method 800 is stored on a non-transitory computer-readable storage medium and is controlled by instructions executed by one or more processors of the computer system, such as one or more processors 202 of computer system 101 (e.g., control unit 110 of FIG. 1A). Some operations of method 800 are optionally combined and / or the order of some operations is optionally changed.

[0152] In some embodiments, method 800 is executed on an electronic device (e.g., 101) that communicates with a display generation component and one or more input devices (e.g., a mobile device (e.g., a tablet, smartphone, media player, or wearable device), or a computer). In some embodiments, the display generation component is a display integrated with the electronic device (optionally a touch screen display), an external display such as a monitor, projector, television, or a hardware component (optionally embedded or external) that projects a user interface and makes the user interface visible to one or more users. In some embodiments, the one or more input devices include an electronic device or component that can receive user input (e.g., capture user input, detect user input, etc.) and transmit information related to the user input to the electronic device. Examples of input devices include a touch screen, a mouse (e.g., external), a trackpad (optionally integrated or external), a touchpad (optionally integrated or external), a remote control device (e.g., external), another mobile device (e.g., separate from the electronic device), a handheld device (e.g., external), a controller (e.g., external), a camera, a depth sensor, an eye tracking device, and / or a motion sensor (e.g., a hand tracking device, a hand motion sensor), etc. In some embodiments, the electronic device communicates with a hand tracking device (e.g., one or more cameras, depth sensors, proximity sensors, touch sensors (e.g., touch screen, trackpad)). In some embodiments, the hand tracking device is a wearable device such as a smart glove. In some embodiments, the hand tracking device is a handheld input device such as a remote control or a stylus.

[0153] In some embodiments, such as FIG. 7A, while an electronic device (e.g., 101) is displaying a three-dimensional environment (e.g., 702) that includes a plurality of virtual objects (e.g., 704, 706) having a first spatial arrangement relative to a current viewpoint of a user of the electronic device (e.g., 101) via a display generation component (e.g., 120), the electronic device detects (802a) a movement of the user's current viewpoint from a first viewpoint to a second viewpoint in the three-dimensional environment (e.g., 702) via one or more input devices (e.g., 314). In some embodiments, the three-dimensional environment includes virtual objects such as application windows, operating system elements, representations of other users, and / or representations of content items and / or physical objects within the physical environment of the electronic device. In some embodiments, the representation of the physical object is displayed in the three-dimensional environment via a display generation component (e.g., a virtual passthrough or a video passthrough). In some embodiments, the representation of the physical object is a view of the physical object within the physical environment of the electronic device that is visible through a transparent portion of the display generation component (e.g., a true passthrough or an actual passthrough). In some embodiments, the electronic device displays the three-dimensional environment from the user's viewpoint at a location in the three-dimensional environment corresponding to the physical location of the electronic device within the physical environment of the electronic device. In some embodiments, the three-dimensional environment is a computer-generated reality (XR) environment (e.g., a virtual reality (VR) environment, a mixed reality (MR) environment, or an augmented reality (AR) environment) that is generated, displayed, or otherwise made viewable by the device. In some embodiments, detecting a movement of the user's viewpoint includes detecting a movement of at least a portion of the user (e.g., the user's head, torso, hand, etc.). In some embodiments, detecting a movement of the user's viewpoint includes detecting a movement of the electronic device or the display generation component. In some embodiments, displaying the three-dimensional environment from the user's viewpoint includes displaying the three-dimensional environment from a viewing direction associated with the location of the user's viewpoint in the three-dimensional environment.In some embodiments, updating the user's perspective causes the electronic device to display a plurality of virtual objects from a view associated with the location of the user's updated perspective. For example, if the user's perspective moves to the left, the electronic device updates the positions of the plurality of virtual objects displayed via the display generation component to move to the right.

[0154] In some embodiments, such as in FIG. 7B, in response to detecting a movement corresponding to the movement of the user's current perspective from a first perspective to a second perspective, the electronic device (e.g., 101) displays (e.g., 802b), via the display generation component (e.g., 120), a three-dimensional environment (e.g., 702) from a second perspective that includes a plurality of virtual objects having a second spatial arrangement different from the first spatial arrangement with respect to the user's current perspective. In some embodiments, the locations of the plurality of virtual objects in the three-dimensional environment remain the same, and the location of the user's perspective in the three-dimensional environment changes, thereby changing the spatial arrangement of the plurality of objects with respect to the user's perspective. For example, if the user's perspective moves away from a plurality of virtual objects, the electronic device displays the virtual objects at the same location within the three-dimensional environment prior to the movement of the user's perspective and increases the amount of space between the user's perspective and the virtual objects (e.g., by displaying the virtual objects in a smaller size, by increasing the stereoscopic depth of the objects, etc.).

[0155] In some embodiments, such as FIG. 7B, while displaying a three-dimensional environment (e.g., 702) from a second perspective that includes a plurality of virtual objects (e.g., 704, 706) having a second spatial arrangement relative to the user's current perspective, an electronic device (e.g., 101) receives input (802c) corresponding to a request to update the spatial arrangement of the plurality of virtual objects relative to the user's current perspective so as to meet one or more criteria that specify a range of distances or orientations of the virtual objects relative to the user's current perspective (e.g., not based on the user's previous perspective) via one or more input devices (e.g., 314). In some embodiments, the input and / or one or more inputs described with reference to method 800 are gesture inputs. In some embodiments, a gesture is a gesture that is detected without the user touching (or independently of) an input element that is part of the device, and is based on detected movement of a part of the user's body, including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground, or the distance of the user's hand relative to the ground), movement of the user's body relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of the user's other hand relative to one of the user's hands, and / or movement of the user's finger relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tap gesture including movement of the hand in a predetermined posture with a predetermined amount and / or speed, or a shake gesture including rotation of a part of the user's body at a predetermined speed or amount).

[0156] As described in more detail below, in some embodiments, the input corresponding to a request to update the spatial arrangement of a plurality of objects with respect to a user's perspective to meet one or more criteria is an input directed to a hardware button, switch, etc. that is in communication with (e.g., incorporated in) the electronic device. As described in more detail below, in some embodiments, the input corresponding to a request to update a three-dimensional environment to meet one or more criteria is an input directed to selectable options presented via a display generation component. In some embodiments, the user's perspective is a second perspective, and the plurality of virtual objects are presented in a second spatial arrangement with respect to the user's perspective, but the spatial orientation of the plurality of virtual objects is not based on the user's current perspective. For example, the spatial arrangement of the plurality of objects is based on a different perspective of the user than the user's first or second perspective. In some embodiments, the electronic device first presents a virtual object at a location within the three-dimensional environment based on the perspective of the user when the virtual object was first presented. In some embodiments, the virtual object is first positioned according to one or more criteria including criteria that are met when an interactive portion of the virtual object is directed toward the user's perspective, the virtual object does not obstruct the view of other virtual objects from the user's perspective, the virtual object is within a threshold distance (e.g., 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 1000, or 2000 centimeters) of the user's perspective, the virtual objects are within a threshold distance (e.g., 1, 5, 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 1000, or 2000 centimeters) of each other, etc. In some embodiments, the input is different from an input that requests updating the position of one or more objects within a three-dimensional environment (e.g., with respect to the user's perspective), such as the input described herein with reference to method 1000.

[0157] In some embodiments, such as FIG. 7C, in response to an input corresponding to a request to update a three-dimensional environment (e.g., 702), an electronic device (e.g., 101) displays a three-dimensional environment (e.g., 702) from a second perspective (e.g., 802b) that includes displaying a plurality of virtual objects (e.g., 704, 706) in a third spatial arrangement that is different from a second spatial arrangement with respect to the user's perspective via a display generation component (e.g., 120), and the third spatial arrangement of the plurality of virtual objects satisfies one or more criteria. In some embodiments, the one or more criteria are satisfied when the interactive portion of the virtual object is oriented towards the user's perspective, the virtual object does not obstruct the view of other virtual objects from the user's perspective, the virtual object is within a threshold distance (e.g., 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 1000, or 2000 centimeters) of the user's perspective, the virtual objects are within a threshold distance (e.g., 1, 5, 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 1000, or 2000 centimeters) of each other, etc. In some embodiments, the third spatial arrangement is the same as the first spatial arrangement (e.g., the first spatial arrangement was based on the user's first perspective). In some embodiments, the third spatial arrangement is different from the first spatial arrangement (e.g., the first spatial arrangement was not based on the user's first perspective, and real objects within the three-dimensional environment prevent the virtual objects in the first spatial arrangement while the user's perspective is the second perspective). In some embodiments, displaying the plurality of virtual objects in the third spatial arrangement includes updating the location (e.g., and / or orientation) of one or more of the virtual objects while maintaining the user's second perspective at a fixed location within the three-dimensional environment. In some embodiments, in response to the input, the electronic device updates the position of the virtual object from a location that is not necessarily oriented around the user's perspective to a location that is oriented around the user's perspective.

[0158] Updating the three-dimensional environment to include displaying a plurality of virtual objects in a third spatial arrangement that meets one or more criteria in response to an input provides an efficient way to display virtual objects in a spatial arrangement based on the user's updated perspective, thereby improving the user interaction with the electronic device and thereby improving the placement of virtual objects in the three-dimensional environment and enabling the user to use the electronic device quickly and efficiently.

[0159] In some embodiments, such as FIG. 7B, receiving an input corresponding to a request to update the spatial arrangement of a plurality of virtual objects includes receiving the input via a hardware input device (e.g., 703) of one or more input devices (804). In some embodiments, the hardware input device detects the input by detecting a physical operation of a mechanical input device by the user. In some embodiments, the mechanical input device is a button, switch, dial, etc.

[0160] Updating the spatial arrangement of a plurality of virtual objects in response to an input received via a mechanical input device provides an efficient and consistent way to update the spatial arrangement of a plurality of virtual objects in a three-dimensional environment, thereby enabling the user to use the electronic device quickly and efficiently using an enhanced input mechanism.

[0161] In some embodiments, such as FIG. 7B, the input corresponding to the request to update the spatial arrangement of a plurality of virtual objects satisfies one or more first input criteria (806a). In some embodiments, the one or more first input criteria are associated with the input corresponding to the request to update the spatial arrangement of a plurality of virtual objects. In some embodiments, the electronic device is configured to receive a plurality of inputs corresponding to different respective operations via a mechanical input device that satisfies a different set of criteria in order to determine the corresponding operation for which the received input is to be performed. For example, the mechanical input device is a button or a dial configured to be pressed like a button, and one or more first criteria are satisfied when the electronic device detects that the mechanical input device has been pressed during a time period that satisfies a threshold time period (e.g., 0.1, 0.2, 0.3, 0.5, 1, 2, 3, 4, or 5 seconds) (e.g., longer or shorter).

[0162] In some embodiments, such as FIG. 7B, the electronic device (e.g., 101) receives a second input (806b) via a hardware input device (e.g., 703). In some embodiments, the second input does not satisfy one or more first input criteria.

[0163] In some embodiments, in response to receiving a second input, the electronic device (e.g., 101) performs an individual operation corresponding to the second input (806c) without updating the spatial arrangement of the plurality of virtual objects (e.g., displaying a home user interface via a display generation component) according to a determination that the second input meets one or more second input criteria that are different from one or more first input criteria. For example, one or more first input criteria are met when the input includes detecting that a mechanical input device has been pressed for a threshold time period, and one or more second input criteria are met when the input includes detecting that a mechanical input device has been pressed for less than a threshold time period. As another example, one or more first input criteria are met when the input includes detecting that a mechanical input device has been pressed for less than a threshold time period, and one or more second input criteria are met when the input includes detecting that a mechanical input device has been pressed for more than a threshold time period. In some embodiments, according to a determination that the second input meets one or more third input criteria that are different from one or more first input criteria and one or more second input criteria, the electronic device modifies the amount of visual emphasis with which the electronic device displays one or more representations of physical objects within a three-dimensional environment via a display generation component. In some embodiments, one or more third input criteria include criteria that are met when the input includes a directional aspect such as rotation of a dial, selection of a direction button, or movement of a contact on a touch-sensitive surface. In response to the third input criteria being met, the electronic device optionally modifies the relative visual emphasis of the representation of the physical environment displayed within the three-dimensional environment with respect to other portions of the three-dimensional environment (e.g., without updating the spatial arrangement of the plurality of virtual objects and / or without performing an individual operation). In some embodiments, the electronic device modifies the amount of relative visual emphasis according to a movement metric of the input that meets the third input criteria (e.g., the direction, speed, duration, distance, etc. of the movement or directional component of the input).

[0164] Determining whether an input received via a mechanical input device corresponds to a request to update the spatial arrangement of a plurality of virtual objects based on one or more input criteria provides an efficient way to perform a plurality of operations in response to an input received via one mechanical input device, thereby improving user interaction with an electronic device by providing a streamlined human-machine interface.

[0165] In some embodiments, such as FIG. 7B, receiving an input corresponding to a request to update the spatial arrangement of a plurality of virtual objects (e.g., 704, 706) includes detecting a selection of a user interface element (e.g., 712) displayed in a three-dimensional environment (e.g., 702) via a display generation component (e.g., 120) (808). In some embodiments, the selection of a user interface element is detected via an input device (e.g., keyboard, trackpad, stylus, hand tracking device, eye tracking device) that manipulates a cursor and detects the selection of the user interface element at the location where the cursor is displayed. In some embodiments, the selection of a user interface element is detected via an input device (e.g., touch sensor surface, hand tracking device) that tracks the location corresponding to an individual part of the user (e.g., hand) and detects the selection when a selection gesture (e.g., shape of an individual hand, three-dimensional gesture performed by the hand, touch gesture detected via a touch sensor surface) is performed at a location corresponding to a selectable user interface element. In some embodiments, detecting the selection of a user interface element includes detecting the attention (e.g., line of sight) of the user directed towards the user interface element while detecting a predetermined gesture of an individual part of the user (e.g., hand). For example, the gesture is a pinch gesture where the thumb touches another finger of the hand. As another example, the gesture is a pressing gesture where an extended finger of the hand presses a user interface element within a three-dimensional environment (e.g., while one or more other fingers are bent towards the palm of the hand).

[0166] Receiving an input corresponding to a request to update a spatial arrangement of a plurality of virtual objects by detecting a selection of a user interface element provides an efficient way to show the user how to update the spatial arrangement of the plurality of virtual objects, thereby enhancing the user's interaction with the electronic device using improved visual feedback and enabling the user to use the device more quickly and efficiently.

[0167] In some embodiments, such as in FIG. 7B, displaying a plurality of virtual objects (e.g., 704, 706) in a second spatial arrangement includes (810a) displaying the plurality of virtual objects (e.g., 704, 706) at a first position within a three-dimensional environment (e.g., 702) (e.g., relative to an individual reference point within the three-dimensional environment) via a display generation component (e.g., 120). In some embodiments, the electronic device displays a plurality of virtual objects having a second spatial arrangement in response to detecting a movement of the user's viewpoint.

[0168] In some embodiments, such as in FIG. 7C, in response to receiving an input, the electronic device (e.g., 101) moves (810a)(810b) a plurality of virtual objects from a first position to a second position within the three-dimensional environment (e.g., relative to an individual reference point within the three-dimensional environment). In some embodiments, the electronic device updates the positions of a plurality of virtual objects within the three-dimensional environment (e.g., without updating the position of the user's viewpoint within the three-dimensional environment) to meet one or more criteria in response to receiving an input. In some embodiments, the electronic device updates the positions of a plurality of virtual objects to meet one or more criteria and updates the position of the user's viewpoint in response to receiving an input.

[0169] Updating the positions of a plurality of virtual objects in response to receiving an input after detecting a movement of a user's perspective to meet one or more criteria provides an efficient way to present the plurality of virtual objects in a comfortable manner for the user, thereby improving user interaction with the electronic device by reducing the cognitive burden, time, and input required to interact with the plurality of virtual objects.

[0170] In some embodiments, such as FIG. 7C, while displaying a three-dimensional environment (e.g., 702) including an individual virtual object (e.g., 704) of a plurality of virtual objects (e.g., 704, 706) at a first position within the three-dimensional environment via a display generation component (e.g., 120), the spatial arrangement of the plurality of virtual objects relative to the user's current perspective is a first individual spatial arrangement that meets one or more criteria, and the electronic device (e.g., 101) receives (812a) an input (e.g., similar to the object movement input described with reference to methods 1000 and / or 1400) corresponding to a request to update the position of an individual virtual object (e.g., 704) within the three-dimensional environment via one or more input devices. In some embodiments, the input is directed to a user interface element displayed proximate to the individual virtual object, and the user interface element, when selected, causes the electronic device to initiate a process of moving the individual virtual object within the three-dimensional environment.

[0171] In some embodiments, such as FIG. 7D, in response to an input corresponding to a request to update the position of an individual virtual object (e.g., 704) within a three-dimensional environment (e.g., 702), an electronic device (e.g., 101) displays a plurality of virtual objects (e.g., 704, 706) in a second individual spatial arrangement that does not meet one or more criteria, including displaying the individual virtual object (e.g., 704) at a second position different from the first position within the three-dimensional environment (e.g., 702) via a display generation component (e.g., 120) (812b). In some embodiments, the electronic device does not detect additional inputs corresponding to requests to update the positions of other virtual objects within the three-dimensional environment. In some embodiments, the electronic device receives one or more additional inputs corresponding to requests to update the positions of one or more other virtual objects in a manner that continues to meet one or more criteria.

[0172] In some embodiments, such as FIG. 7D, while displaying a three-dimensional environment (e.g., 702) that includes displaying an individual virtual object (e.g., 706) at a second position within the three-dimensional environment, an electronic device (e.g., 101) receives a second input corresponding to a request to update the spatial arrangement of a plurality of virtual objects (e.g., 704, 704) relative to the user's current perspective via one or more input devices so as to meet one or more criteria (812c).

[0173] In some embodiments, such as FIG. 7E, in response to receiving a second input, the electronic device (e.g., 101) updates the position of an individual virtual object (e.g., 704) to meet one or more criteria without updating the position of one or more other virtual objects (e.g., 706) within the plurality of virtual objects (e.g., without updating the position and / or orientation of the user's viewpoint within the three-dimensional environment) (812b). In some embodiments, if the movement of a second virtual object in the three-dimensional environment fails to meet one or more criteria, in response to the second input, the electronic device updates the position of the second virtual object to meet one or more criteria without updating the position of one or more other virtual objects.

[0174] Updating the position of an individual virtual object in response to a second input without updating the position of one or more other virtual objects provides enhanced access to the individual virtual object while maintaining the display of the other virtual objects at their respective positions, thereby reducing the cognitive burden, time, and input required to interact with the plurality of virtual objects.

[0175] In some embodiments, such as FIG. 7C, displaying a plurality of virtual objects (e.g., 704, 706) in a third spatial arrangement includes, according to a determination that the three-dimensional environment (e.g., 702) is associated with a first spatial template, displaying, via a display generation component (e.g., 120), the plurality of virtual objects (e.g., 704, 706) in the third spatial arrangement, and displaying, via the display generation component (e.g., 120), an individual virtual object (e.g., 706) of the plurality of virtual objects in an orientation relative to a current viewpoint of a user that meets one or more criteria associated with the first spatial template (814). In some embodiments, the spatial template identifies one or more criteria used by an electronic device to select positions and orientations of virtual objects and / or viewpoints of one or more users within the three-dimensional environment. In some embodiments, the spatial template is set based on virtual objects included in the three-dimensional environment (e.g., spatial templates associated with respective virtual objects) and / or based on user-defined settings. Exemplary spatial templates according to some embodiments are described in more detail below. For example, if the spatial template is a shared content spatial template and the individual virtual objects are content items consumed by a plurality of users within the three-dimensional environment, the electronic device displays the individual virtual objects in positions and orientations that direct individual faces of the content item including the content toward the viewpoints of the users within the three-dimensional environment.

[0176] In some embodiments, such as FIG. 7G, displaying a plurality of virtual objects (e.g., 714) in a third spatial arrangement includes, according to a determination that the three-dimensional environment (e.g., 702) is associated with a second spatial template, displaying, via a display generation component (e.g., 120), the plurality of virtual objects in the third spatial arrangement by displaying individual objects of the plurality of virtual objects (e.g., 714) in an orientation that meets one or more criteria associated with the second spatial template (814c) relative to the user's current viewpoint (814). For example, if the spatial template is a shared activity space template and the individual virtual object is a virtual board game being played by a plurality of users within the three-dimensional environment, the electronic device displays the individual virtual object between the user's viewpoints within the three-dimensional environment such that the user's viewpoints face different sides of the individual virtual object.

[0177] In response to an input corresponding to a request to update the spatial arrangement of the plurality of virtual objects, displaying the individual virtual objects in different orientations to meet different criteria associated with different spatial templates provides an efficient and versatile way of presenting the virtual objects in a three-dimensional environment at locations and orientations that facilitate user interaction with the virtual objects for various functions, thereby enabling rapid and efficient user-to-user electronic devices.

[0178] In some embodiments, such as FIG. 7C, displaying a plurality of virtual objects (e.g., 704, 706) in a third spatial arrangement includes, in accordance with a determination that the three-dimensional environment (e.g., 702) is associated with a shared content space template, displaying, via a display generation component (e.g., 120), an individual object (e.g., 706) of the plurality of virtual objects in a posture that directs an individual face of the individual object (e.g., 706) toward the user's perspective and a second perspective of a second user (or more users) within the three-dimensional environment (e.g., 702) (816). In some embodiments, the shared content space template is associated with virtual objects that include a user interface for a content or content (e.g., delivery, playback, streaming, etc.) application. In some embodiments, the individual face of an individual virtual object includes visual content of a content item (e.g., displaying content such as a movie, a television program, etc.). In some embodiments, the shared content space template enables a plurality of users to consume a content item together from the same side of the content item by directing the content toward the perspectives of both (or more) users.

[0179] Positioning an individual face of an individual virtual object toward the user's perspective provides an efficient way to facilitate shared consumption of a content item associated with the individual virtual object, thereby enabling the user to quickly and efficiently configure the three-dimensional environment for shared consumption of the content item.

[0180] In some embodiments, such as FIG. 7G, displaying a plurality of virtual objects (e.g., 714, 708) in a third spatial arrangement includes, in accordance with a determination that the three-dimensional environment (e.g., 702) is associated with a shared activity space template, via a display generation component (e.g., 120), displaying an individual object (e.g., 714) among the plurality of virtual objects with a first side of the individual object (e.g., 714) facing a user's perspective and a second side different from the first side of the individual object facing a second perspective of a second user (818). In some embodiments, the shared activity space template is associated with an individual virtual object configured for interaction by a plurality of users from a plurality of sides of the virtual object. For example, the individual virtual object is a virtual board game for a plurality of users to play or a virtual table with the perspectives of a plurality of users arranged around it. In some embodiments, the shared activity space template enables a user to view an individual virtual object simultaneously with a representation of another user within the three-dimensional environment that is displayed at the location of the perspective of the other user within the three-dimensional environment.

[0181] Positioning different sides of an individual virtual object towards different users within a three-dimensional environment provides an efficient way to facilitate interaction between users and / or with the individual virtual object, thereby enabling users to quickly and efficiently configure the three-dimensional environment for a shared activity associated with the individual virtual object.

[0182] In some embodiments, such as FIG. 7G, displaying a plurality of virtual objects (e.g., 708, 710) in a third spatial arrangement includes, in accordance with a determination that the three-dimensional environment (e.g., 702) is associated with a group activity space template, displaying a representation of a second user (e.g., 708) in a pose directed toward the current viewpoint of a user of the electronic device (820) via a display generation component (e.g., 120). In some embodiments, the group activity space template is optionally associated with user-defined settings for interacting with one or more other users within the three-dimensional environment in a manner that is not related to individual virtual objects other than the representations of one or more other users (e.g., virtual meetings between users). In some embodiments, the group activity space template causes the electronic device to position the viewpoint of the user (and, e.g., the corresponding representation of the user within the three-dimensional environment) at a location such that the user's viewpoint is directed toward the representations and / or viewpoints of one or more other users within the three-dimensional environment (e.g., without updating the positions of the representations and / or viewpoints of the other users).

[0183] Directing the user's viewpoint toward the representation of a second user within the three-dimensional environment provides an efficient way to facilitate interaction between users, thereby enabling the user to quickly and efficiently configure the three-dimensional environment for group activities.

[0184] In some embodiments, such as FIG. 7B, while displaying a three-dimensional environment (e.g., 702) that includes a second user associated with a second perspective within the three-dimensional environment in which the spatial arrangement of a plurality of virtual objects (e.g., 704, 706) with respect to the user's current perspective (and the second perspective of the second user) is a first individual spatial arrangement that meets one or more criteria, the electronic device (e.g., 101) detects an indication of movement of the second perspective of the second user from a first individual perspective to a second individual perspective within the three-dimensional environment (822a). In some embodiments, the movement of the second perspective of the second user causes one or more criteria not to be met in the electronic device. In some embodiments, the movement of the second perspective of the second user causes one or more criteria not to be met in a second electronic device being used by the user. In some embodiments, the electronic device detects an indication of movement of the second perspective of the second user without receiving one or more inputs corresponding to a request to update the user's perspective and / or the pose of the plurality of virtual objects. In some embodiments, the electronic device further detects movement of the user's perspective and / or movement of one or more of the virtual objects in a manner that continues to meet one or more criteria in the electronic device and / or a second electronic device being used by the second user.

[0185] In some embodiments, such as FIG. 7B, while displaying a three-dimensional environment having the second perspective of the second user at the second individual perspective, the electronic device (e.g., 101) receives a second input corresponding to a request to update the spatial arrangement of a plurality of virtual objects (e.g., 704, 706) with respect to the user's current perspective via one or more input devices (e.g., 314) so as to meet one or more criteria (822b).

[0186] In some embodiments, such as FIG. 7C, an electronic device (e.g., 101) updates the spatial arrangement of a plurality of virtual objects (e.g., 704, 706) to a second individual spatial arrangement that meets one or more criteria according to a second input and according to a second individual perspective of a second user (822c). In some embodiments, the electronic device updates the spatial arrangement of the plurality of virtual objects such that one or more criteria are met with respect to the perspective of the user and the second perspective of the second user. For example, when the second perspective of the second user moves to the left of the perspective of the user, updating the spatial arrangement of the plurality of virtual objects includes moving or orienting one or more of the virtual objects toward the left with respect to the perspective of the user (e.g., in a direction based on the perspective of the user and the new perspective of the second user). As another example, when the second perspective of the second user moves to the right of the perspective of the user, updating the spatial arrangement of the plurality of virtual objects includes moving or orienting one or more of the virtual objects toward the right with respect to the perspective of the user.

[0187] Updating the spatial arrangement according to the second individual perspective of the second user provides an efficient way to meet one or more criteria in the electronic device and the second electronic device of the second user in response to one input, thereby reducing the number, time, and cognitive burden of the inputs required to configure the three-dimensional environment in a comfortable way for the user and the second user.

[0188] In some embodiments, such as FIG. 7C, while displaying, via a display generation component (e.g., 120), a three-dimensional environment (e.g., 702) in which the spatial arrangement of a plurality of virtual objects (e.g., 704, 706) relative to the user's current viewpoint is a first individual spatial arrangement that meets one or more criteria, the electronic device (e.g., 101) receives (824a), via one or more input devices, a sequence of one or more inputs corresponding to a request to update one or more positions of a plurality of virtual objects (e.g., 704) in the three-dimensional environment (e.g., 702). In some embodiments, the sequence of inputs is an input directed to a user interface element for repositioning each virtual object within the three-dimensional environment, as described with reference to method 1400. In some embodiments, the sequence of inputs is an input corresponding to a request to reposition a plurality of virtual objects together according to one or more steps of method 1000, which is described in more detail below.

[0189] In some embodiments, such as FIG. 7D, in response to receiving a sequence of one or more inputs, the electronic device (e.g., 101) displays, via the display generation component (e.g., 120), each position within the three-dimensional environment (e.g., 702) of a plurality of virtual objects (e.g., 704, 706) according to the sequence of one or more inputs in a second individual spatial arrangement that does not meet one or more criteria (824b). In some embodiments, the electronic device does not receive an input corresponding to a request to update the user's viewpoint. In some embodiments, the electronic device receives an input corresponding to a request to update the user's viewpoint such that the one or more criteria are not met.

[0190] In some embodiments, such as FIG. 7D, while displaying a three-dimensional environment (e.g., 702) that includes displaying a plurality of virtual objects (e.g., 704, 706) at respective positions within the three-dimensional environment (e.g., 702), an electronic device (e.g., 101) receives a second input (824c) via one or more input devices (e.g., 314) in response to a request to update the spatial arrangement of the plurality of virtual objects (e.g., 704, 706) relative to the user's current viewpoint to satisfy one or more criteria.

[0191] In some embodiments, such as FIG. 7E, in response to receiving the second input, an electronic device (e.g., 101) updates the positions of the plurality of virtual objects (e.g., 704, 706) to a third individual spatial arrangement that satisfies one or more criteria according to the respective positions of the plurality of virtual objects (e.g., 704, 706) within the three-dimensional environment (e.g., 702) (824d). In some embodiments, the third individual spatial arrangement is the same as the first individual spatial arrangement. In some embodiments, the third individual spatial arrangement is different from the first individual spatial arrangement. In some embodiments, the third individual spatial arrangement involves a minimum movement and / or reorientation of the plurality of virtual objects from (e.g., the respective positions and / or orientations at which the virtual objects are displayed in response to a sequence of inputs corresponding to requests to update one or more positions of the plurality of virtual objects in the three-dimensional environment) the positions and / or orientations in the three-dimensional environment required to satisfy one or more criteria.

[0192] Updating the positions of the plurality of virtual objects according to the respective positions of the plurality of virtual objects in the three-dimensional environment after the plurality of virtual objects have been repositioned to satisfy one or more criteria provides an efficient way to facilitate user interaction with the plurality of virtual objects, thereby reducing the cognitive burden, time, and input required for convenient user interaction with the plurality of virtual objects.

[0193] In some embodiments, while a display generation component (e.g., 120) is displaying a three-dimensional environment (e.g., 120) in which the spatial arrangement of a plurality of virtual objects (e.g., 704, 706) with respect to the user's current viewpoint is a first individual spatial arrangement that meets one or more criteria such as in FIG. 7A, an electronic device (e.g., 101) detects (826a) one or more indications of a request by a second user within the three-dimensional environment (e.g., 702) to update one or more positions of the plurality of virtual objects (e.g., 704, 706) within the three-dimensional environment (e.g., 702). In some embodiments, a request by a second user in the three-dimensional environment includes one or more requests to update the position of one or more virtual objects (e.g., individually). In some embodiments, a request by a second user in the three-dimensional environment does not include a request by the second user to update a second viewpoint of the second user in the three-dimensional environment.

[0194] In some embodiments, in response to detecting one or more indications, the electronic device (e.g., 101) displays (826b), via the display generation component (e.g., 120), the plurality of virtual objects (e.g., 704, 706) at respective positions within the three-dimensional environment (e.g., 702) according to a sequence of inputs in a second individual spatial arrangement that does not meet one or more criteria such as in FIG. 7B. In some embodiments, the electronic device does not receive an input corresponding to a request to update the pose of a virtual object or the user's viewpoint and does not detect an indication of a request by the second user to update the position of the second viewpoint of the second user. In some embodiments, the electronic device receives an input corresponding to a request to update the pose of a virtual object and / or the user's viewpoint and / or detects an indication of a request by the second user to update the position of the second viewpoint of the user, but does not cause these updates to not meet one or more criteria.

[0195] In some embodiments, while displaying a three-dimensional environment (e.g., 702) that includes displaying a plurality of virtual objects (e.g., 704, 706) at respective positions within the three-dimensional environment (e.g., 702), an electronic device (e.g., 101) receives a second input (826c) via one or more input devices (e.g., 314) in response to a request to update the spatial arrangement of the plurality of virtual objects (e.g., 704, 706) relative to the user's current viewpoint to meet one or more criteria such as those in FIG. 7B.

[0196] In some embodiments, in response to receiving the second input, an electronic device (e.g., 101) updates the positions of the plurality of virtual objects (e.g., 704, 706) to a third individual spatial arrangement that meets one or more criteria, such as those in FIG. 7C, according to the respective positions of the plurality of virtual objects (e.g., 704, 706) within the three-dimensional environment (e.g., 702) (826d). In some embodiments, the third individual spatial arrangement is the same as the first individual spatial arrangement. In some embodiments, the third individual spatial arrangement is different from the first individual spatial arrangement. In some embodiments, the third individual spatial arrangement involves a minimum movement and / or reorientation of the plurality of virtual objects within the three-dimensional environment (e.g., from the respective positions and / or orientations at which the virtual objects are displayed in response to an indication corresponding to a request to update one or more positions of the plurality of virtual objects within the three-dimensional environment) required to meet one or more criteria.

[0197] Updating the positions of the plurality of virtual objects according to the respective positions of the plurality of virtual objects in the three-dimensional environment to meet one or more criteria after the plurality of virtual objects have been repositioned by a second user provides an efficient way to facilitate user interaction with the plurality of virtual objects, thereby reducing the cognitive burden, time, and input required for convenient user interaction with the plurality of virtual objects.

[0198] In some embodiments, while the spatial arrangement of a plurality of virtual objects (e.g., 704, 706) does not meet one or more criteria such as in FIG. 7B, in accordance with a determination that no input corresponding to a request to update the spatial arrangement of the plurality of objects has been received, the electronic device (e.g., 101) maintains (828) the spatial arrangement of the plurality of virtual objects (e.g., 704, 706) until an input corresponding to a request to update the spatial arrangement of the plurality of objects is received. In some embodiments, the electronic device does not update the spatial arrangement of the plurality of objects to meet one or more criteria unless and until an input corresponding to a request to update the spatial arrangement of the plurality of objects is received. In some embodiments, the electronic device does not update the spatial arrangement of one or more virtual objects to automatically meet one or more criteria.

[0199] Maintaining the spatial arrangement of a plurality of virtual objects until an input is received provides an efficient way to display the virtual objects in a familiar location and / or orientation within the three-dimensional environment, thereby enhancing user interaction with the electronic device by enabling the user to quickly and efficiently position the virtual objects within the three-dimensional environment for interaction.

[0200] In some embodiments, an electronic device (e.g., 101) displays (830a) a plurality of virtual objects (e.g., 704, 706) in a three-dimensional environment (e.g., 702), including displaying a first virtual object of the plurality of virtual objects (e.g., 704, 706) at a location that exceeds a predetermined threshold distance (e.g., 1, 2, 3, 4, 5, 10, 15, 30, or 50 meters) from the user's current perspective, such as in FIG. 7B. In some embodiments, a second virtual object is positioned at a location that exceeds a predetermined threshold distance from the user's perspective. In some embodiments, a plurality of (e.g., all) virtual objects are positioned at locations that exceed a predetermined threshold distance from the user's perspective. In some embodiments, one or more criteria include criteria that are satisfied when a plurality of (e.g., all) virtual objects are positioned within a predetermined threshold distance from the user's current perspective.

[0201] In some embodiments, such as in FIG. 7B, while a plurality of virtual objects (e.g., 704, 706) are being displayed in a three-dimensional environment (e.g., 702), including displaying a first virtual object of the plurality of virtual objects (e.g., 704, 706) at a location that exceeds a predetermined threshold distance from the user's current perspective, the electronic device (e.g., 101) receives a second input (830b) via one or more input devices in response to a request to update the spatial arrangement of the plurality of virtual objects (e.g., 704, 706) with respect to the user's current perspective so as to satisfy one or more criteria.

[0202] In some embodiments, such as FIG. 7C, in response to receiving a second input, an electronic device (e.g., 101) updates the user's perspective to an individual perspective within a predefined threshold distance of a first virtual object (e.g., 704, 706) (830c), and the spatial arrangement of a plurality of virtual objects (e.g., 704, 706) with respect to the individual perspective meets one or more criteria. In some embodiments, the electronic device updates the user's perspective without updating the location and / or orientation of the virtual object. In some embodiments, the electronic device updates the user's perspective and updates the location and / or orientation of the virtual object. In some embodiments, if it is possible to meet one or more criteria by updating the user's perspective without updating the position and / or orientation of the virtual object, the electronic device updates the user's perspective without updating the position and / or orientation of the virtual object in response to the second input.

[0203] Updating the user's perspective to meet one or more criteria in response to a second perspective when a first virtual object is located farther from the user's perspective than a predefined threshold distance provides an efficient way to access the virtual object for interaction, thereby reducing the cognitive burden, time, and input required to interact with the virtual object.

[0204] In some embodiments, such as FIG. 7D, an electronic device (e.g., 101) displays (832a) a plurality of virtual objects (e.g., 704, 706) having a first interval between a first virtual object (e.g., 704) of the plurality of virtual objects and a second virtual object (e.g., 704) of the plurality of virtual objects via a display generation component (e.g., 120), and the first interval does not meet one or more interval criteria of one or more criteria. In some embodiments, the one or more interval criteria include criteria that are met when the distance between each (e.g., pair of) virtual objects is between a first predetermined threshold distance (e.g., 1, 2, 3, 5, 10, 15, 30, 50, or 100 centimeters) and a second predetermined threshold distance (e.g., 1, 2, 3, 5, 10, 15, 30, or 50 meters).

[0205] In some embodiments, such as FIG. 7D, while displaying a plurality of virtual objects (e.g., 704, 706) having a first interval between a first virtual object (e.g., 704) and a second virtual object (e.g., 706), an electronic device (e.g., 101) receives (832b) a second input corresponding to a request to update the spatial arrangement of the plurality of virtual objects (e.g., 704, 706) with respect to the current viewpoint of the user via one or more input devices (e.g., 314) so as to meet one or more criteria (e.g., including one or more interval criteria).

[0206] In some embodiments, such as FIG. 7E, in response to receiving a second input, an electronic device (e.g., 101) displays (832c) a plurality of virtual objects (e.g., 704, 706) having a second spacing between a first virtual object (e.g., 704) and a second virtual object (e.g., 706) via a display generation component (e.g., 120), and the second spacing meets one or more spacing criteria. In some embodiments, displaying a plurality of virtual objects having a second spacing between a first virtual object and a second virtual object includes displaying a plurality of virtual objects having a spatial arrangement that meets one or more criteria. In some embodiments, the electronic device updates the position and / or orientation of one or more virtual objects other than the first and second virtual objects as necessary to meet one or more criteria. In some embodiments, the second spacing is smaller than the first spacing. In some embodiments, displaying a plurality of virtual objects having a second spacing between a first virtual object and a second virtual object includes changing the position and / or orientation of the first and / or second object, such as updating the position (e.g., without updating the orientation) or updating both the position and the orientation.

[0207] Displaying a plurality of virtual objects having a second spacing between a first virtual object and a second virtual object in response to a second input provides an efficient way for a user to view and / or interact with the first and second virtual objects simultaneously, thereby reducing the time, input, and cognitive burden required to interact with the first and second virtual objects.

[0208] In some embodiments, such as FIG. 7A, detecting a movement of a user's current viewpoint in a three-dimensional environment from a first viewpoint to a second viewpoint includes detecting a movement of an electronic device (e.g., 101) in the physical environment of the electronic device (e.g., 101) or a movement of a display generation component (e.g., 120) in the physical environment of the display generation component (e.g., 120) via one or more input devices (e.g., 314) (834). In some embodiments, detecting a movement of an electronic device or a display generation component includes detecting a movement of the user (e.g., where the electronic device or the display generation component is a wearable device).

[0209] Updating the user's viewpoint in response to detecting a movement of an electronic device or a display generation component provides an efficient and intuitive way to traverse the three-dimensional environment, thereby improving the user interaction with the three-dimensional environment.

[0210] FIGS. 9A-9G illustrate examples of a method by which an electronic device updates the positions of a plurality of virtual objects together, according to some embodiments.

[0211] FIG. 9A shows an electronic device 101a that displays a three-dimensional environment 902a via a display generation component 120a. In some embodiments, it should be understood that the electronic device 101a may utilize one or more of the techniques described with reference to FIGS. 9A-9G within a two-dimensional environment without departing from the scope of the present disclosure. As described above with reference to FIGS. 1-6, the electronic device 101a optionally includes a display generation component 120a (e.g., a touch screen) and a plurality of image sensors 314a. The image sensors optionally include one or more of a visible light camera, an infrared camera, a depth sensor, or any other sensor, and the electronic device 101a can be used to capture one or more images of the user or a part of the user while the user is interacting with the electronic device 101a. In some embodiments, the display generation component 120a is a touch screen capable of detecting gestures and movements of the user's hand. In some embodiments, the user interface shown below can also be implemented in a head-mounted display that includes a display generation component that displays the user interface to the user, and sensors that detect the physical environment and / or the movement of the user's hand (e.g., external sensors facing outward from the user), and / or the user's line of sight (e.g., internal sensors facing inward toward the user's face).

[0212] In FIG. 9A, the first electronic device 101a and the second electronic device 101b have access to a three-dimensional environment 902 that includes a representation of the first electronic device 101a, a representation of the second electronic device 101b, a first user 916a of the first electronic device 101a, a second user 916b of the second electronic device 101b, a user interface 904 of a first application accessible to both electronic devices 101a and 101b, a user interface 906 of a second application accessible to the first electronic device 101a but not to the second electronic device 101b, and a user interface 908 of a third application accessible to both electronic devices 101a and 101b. In some embodiments, the first electronic device 101a and the second electronic device 101b are disposed within a shared physical environment, and the spatial arrangement of the first electronic device 101a and the second electronic device 101b within the physical environment corresponds to (e.g., is the same as) the spatial arrangement of the first electronic device 101a and the second electronic device 101b within the three-dimensional environment 902. In some embodiments, the first electronic device 101a and the second electronic device 101b are located in different physical environments (e.g., different rooms, different buildings, different cities, etc.), and the spatial arrangement of the first electronic device 101a and the second electronic device 101b in the physical world is different from the spatial arrangement of the first electronic device 101a and the second electronic device 101b in the three-dimensional environment 902. In some embodiments, the electronic device 101b has one or more of the characteristics of the electronic device 101a.

[0213] Figures 9A-9G include a top-down view of a three-dimensional environment 902 presented by electronic devices 101a and 101b. The first electronic device 101a presents the three-dimensional environment 902a from the perspective of a first user 916a within the three-dimensional environment 902, including, for example, displaying a user interface 904a of a first application, a user interface 906a of a second application, a user interface 908a of a third application, a representation 916b of a second user, and a representation of the second electronic device 101b. As another example, the second electronic device 101b presents the three-dimensional environment 902b from the perspective of a second user 916b in the three-dimensional environment 902, including displaying a user interface 908b of a first application, a user interface 904b of a third application, a representation 916a of the first user, and a representation of the first electronic device 101a. In some embodiments, the second electronic device 101b does not have access to the user interface 906 of the second application, and thus the second electronic device 101b does not display the user interface 906 of the second application. However, FIG. 9A shows a location 906b within the three-dimensional environment 902b where the user interface 906 of the second application is displayed when the second electronic device 101b has access to the user interface 906 of the second application.

[0214] In some embodiments, the first electronic device 101a is associated with a digital origin 912 within the three-dimensional environment 902. In some embodiments, the second electronic device 101b is associated with a different digital origin (not shown) within the three-dimensional environment 902. In some embodiments, the digital origin 912 is a location within the three-dimensional environment 902 that the electronic device 101a uses to select the location and / or orientation of virtual objects (e.g., user interfaces 904, 906, 908) and / or the perspective of the first user 101a in response to a request to re-center the three-dimensional environment 902a according to one or more steps of method 800. For example, in response to a request to re-center the three-dimensional environment 902a, the electronic device 101a evaluates one or more criteria regarding the spatial arrangement of virtual objects and / or the perspective of user 916a relative to the digital origin 912 and updates the position and / or orientation of one or more virtual objects and / or the perspective of user 916a to satisfy the one or more criteria. In some embodiments, the electronic device 101a updates the perspective of the first user 916a to be at the location of the digital origin 912 in response to a request to re-center. In some embodiments, the electronic device 101a updates the position and / or orientation of user interfaces 904a, 904b, and 904c relative to the digital origin 912 to select a position and / or orientation that facilitates user interaction with user interfaces 904a, 904b, and 904c. For example, the one or more criteria include criteria that specify a range of distances from the digital origin 912 and / or a range of orientations relative to the digital origin 912, as described with reference to methods 800 and / or 1000.

[0215] In some embodiments, the electronic device 101a selects a digital origin 912 at the start of an AR / VR session. For example, the digital origin 912 is selected based on the position and / or orientation of the electronic device 101a in the physical environment and the position and / or orientation of the electronic device 101a in the three-dimensional environment 902. In some embodiments, the electronic device 101a sets or resets the digital origin 912 in response to detecting a transition of the display generation component 120a from a state of not being in a predetermined posture to a predetermined posture for a first user over a threshold amount of time (e.g., 10, 20, 30, or 45 seconds, or 1, 2, 3, 5, 10, or 15 minutes). In some embodiments, the predetermined posture is a range of distance and / or orientation with respect to the user. For example, when the display generation component 120a is a wearable device, the display generation component 120a is in a predetermined posture when the user is wearing the display generation component 120a.

[0216] As will be described in more detail below with respect to method 1000 in FIGS. 9B-9F, in some embodiments, there are several situations in which the digital origin is disabled. In some embodiments, the digital origin is disabled when one or more disabling criteria are met. For example, the disabling criteria relate to the number of objects (e.g., user interfaces 904a, 904b, 904c, and the viewpoints of users 916a) having a spatial relationship to the digital origin 912 that do not meet one or more of the spatial criteria described with reference to methods 1000 and / or 800. As another example, the disabling criteria relate to the extent to which the objects do not meet one or more of the spatial criteria and / or the spatial relationships between the objects themselves. For example, the digital origin 912 is invalid when one or more of the spatial criteria are not met, and updating the location of the digital origin 912 satisfies one or more of the spatial criteria with fewer or minimal updates to the positions and / or orientations of the virtual objects than would be required if the digital origin remained at its current location. In some embodiments, the electronic device 101a immediately updates the location of the digital origin 912 when the location of the digital origin 912 becomes invalid. In some embodiments, the electronic device 101a does not update the location of the digital origin 912 until a request to re-center the three-dimensional environment 902a is received and until such request is received.

[0217] In some embodiments, the electronic device 101a updates the positions of the user interfaces 904a, 906a, and 908a in response to various user inputs. In some embodiments, the user interfaces 904a, 906a, and 908a, as described above with reference to method 800, when selected, are each associated with a selectable option that causes the electronic device 101a to initiate a process of updating the position and / or orientation of the individual user interface 904a, 906a, or 908a. Additionally, in some embodiments, in response to a request to update the position and / or orientation of virtual objects (e.g., user interfaces 904a, 906a, and 908a, representation of the second user 916b, and representation of the second electronic device 101b) in the three-dimensional environment 902a together, the electronic device 101a updates the position and / or orientation of the virtual objects with respect to the user 916a's perspective without updating the spatial arrangement of the virtual objects with respect to other virtual objects. For example, the electronic device 101a moves the virtual objects together as a group. In some embodiments, when selected, the electronic device 101a displays a selectable option that causes the electronic device 101a to initiate a process of updating the position and / or orientation of the virtual objects as a group. As another example, as shown in FIG. 9A, in response to detecting a selection input by both hands (e.g., by hands 913a and 913b), the electronic device 101a initiates a process of updating the position and / or orientation of the virtual objects together as a group. For example, as shown in FIG. 9A, the electronic device 101a detects that the user has made a pinch hand shape (e.g., a hand shape where the thumb touches another finger of the same hand) with both hands 913a and 913b for at least a threshold time (e.g., 0.1, 0.2, 0.5, 1, 2, or 3 seconds). In some embodiments, in response to detecting a pinch hand shape by both hands 913a and 913b, the electronic device 101a initiates a process of updating the position and / or orientation of the virtual objects according to the movement of hands 913a and 913b while the pinch hand shape is maintained.

[0218] Figure 9B shows an example where the electronic device 101a detects the start of movement of hands 913a and 913b while the hands are in a pinch hand shape. In some embodiments, in response to detecting the start of movement of hands 913a and 913b in the pinch hand shape (or, in some embodiments, in response to detecting the pinch hand shape of hands 913a and 913b without yet detecting movement of hands 913a and 913b), the electronic device 101a updates the display of virtual objects (e.g., user interfaces 904a, 906a, and 908a, the representation of the second user 916b, and the representation of the electronic device 101b) to change the visual characteristics of the virtual objects, indicating that further input will cause the electronic device 101a to update the position and / or orientation of the virtual objects in the three-dimensional environment 902a. For example, the electronic device 101a increases the amount of translucency and / or dimness of the user interfaces 904a, 906a, and 908a and the representations of the second user 916b and the second electronic device 101b in response to detecting the pinching hand shape and / or the movement of hands 913a and 913b in the pinching hand shape. In some embodiments, the electronic device 101a increases the amount of visual emphasis of the user interfaces 904a, 906a, and 908a and the representations of the second user 916b and the second electronic device 101b while the input of Figure 9B is being received, such as by blurring and / or darkening and / or dimming areas of the three-dimensional environment 902a other than the user interfaces 904a, 906a, and 908a and the representations of the second user 916b and the second electronic device 101b.In some embodiments, while the electronic device 101a updates the positions of the user interfaces 904a, 906a, and 908a and the representations of the second user 916b and the second electronic device 101b in accordance with the movement of the pinch-hand shaped hands 913a and 913b, the electronic device 101a maintains an increased translucency and / or dimness of the representations of the user interfaces 904a, 906a, and 908a and the second user 916b and the second electronic device 101b. In some embodiments, while the first electronic device 101a is detecting a two-handed pinch hand shape, the second electronic device 101b displays the representation of the first user 916a and the representation of the first electronic device 101a with an increased translucency and / or dimness relative to the amount of translucency and / or dimness with which the second electronic device 101b displayed the representation of the first user 916a and the representation of the first electronic device 101a in FIG. 9A.

[0219] As shown in FIG. 9B, the first electronic device 101a detects movement of the hands 913a and 913b in a direction away from the body of the first user 916a corresponding to movement away from the perspective of the first user 916a in the three-dimensional environment 902a. In response to the input shown in FIG. 9B, the electronic device 101a updates the position and / or orientation of the user interfaces 904a, 906a, and 908a with respect to the perspective of the first user 916a and the representations of the second user 916b and the second electronic device 101b without changing the spatial relationship between two or more of the user interfaces 904a, 906a, and 908a and the representations of the second user 916b and the second electronic device 101b, as shown in FIG. 9C.

[0220] FIG. 9C shows how the three-dimensional environment 902a displayed by the first electronic device 101a is updated as the hands 913a and 913b move while the first electronic device 101a is in the pinch-hand shape, moving the user interfaces 904a, 906a, and 908a and the representations of the second user 916b and the second electronic device 101b. For example, the first electronic device 101a updates the positions of the user interfaces 904a, 906a, and 908a and the representation of the second user 916b and the second electronic device 101b to move further away from the user's perspective while maintaining the spatial relationship between the user interfaces 904a, 906a, and 908a and the representation of the second user 916b and the second electronic device 101b. In some embodiments, the amount by which the user interfaces 904a, 906a, and 908a and the representation of the second user 916b and the second electronic device 101b move with respect to the perspective of the first user 916a corresponds to the amount of movement of the hands 913a and 913b. The user interfaces 904a, 906a, and 908a and the representation of the second user 916b and the second electronic device 101b appear to move away from the perspective of the first user 916a as shown in the top-down view of the three-dimensional environment 906, but the perspective of the user 916a moves away from the user interfaces 904, 906, and 908, and the representation of the second user 916b and the second electronic device 101b, and the user interfaces 904, 906, and 908 and the representation of the second user 916b and the second electronic device 101b remain at their respective locations within the three-dimensional environment 906 in response to the input shown in FIG. 9B. In some embodiments, in response to the first electronic device 101a receiving the input shown in FIG. 9B, the second electronic device 101b continues to display the user interfaces 904b and 908b at the same locations within the three-dimensional environment 902b from the perspective of the second user 916b in FIG. 9C, similar to FIG. 9B.

[0221] In some embodiments, in response to an input that updates the position and / or orientation of user interfaces 904a, 906a, and 908a, and the representation of second user 916b and second electronic device 101b relative to the perspective of first user 916a, first electronic device 101a updates the location of digital origin 912. In some embodiments, first electronic device 101a updates the location of digital origin 912 in response to the input shown in FIG. 9B according to a determination that one or more spatial criteria are met at the location of the updated digital origin. For example, as shown in FIG. 9C, first electronic device 101a updates digital origin 912 to be at the location of the perspective of first user 916a in the three-dimensional environment 902. In some embodiments, the updated location of digital origin 912 meets one or more spatial criteria regarding user interfaces 904a, 906a, and 908a, and the representation of second user 916b and second electronic device 101b.

[0222] While the first electronic device 101a detects two-handed input corresponding to a request to update the positions and / or orientations of user interfaces 904a, 906a, and 908a, and the representation of a second user 916b and a second electronic device 101b together, the second electronic device 101b continues to display the representation of the first user 916a and the first electronic device 101a with increased translucency and / or dimness, for example. In some embodiments, the second electronic device 101b does not update the location where the second electronic device 101b displays the representation of the first user 916a and the first electronic device 101a until the first electronic device 101a detects the end of the input, for the second electronic device 101b to update the locations of user interfaces 904a, 906a, and 908a and the representation of the second user 916b and the second electronic device 101b together. For example, detecting the end of the input includes detecting that the user has stopped making a pinch-hand shape with one or more of hands 913a and 913b. As shown in FIG. 9C, the first electronic device 101a detects the continuation of two-handed pinch input provided by hands 913a and 913b. In response to the continuation of the input shown in FIG. 9C, the electronic device 101a updates the spatial relationship among the user interfaces 904a, 906a, and 908a, the representation of the second user 916b and the second electronic device 101b, and the perspective of the first user 916a according to the continued movement of the pinch-hand shaped hands 913a and 913b.

[0223] FIG. 9D shows an example of how the first electronic device 101a updates the three-dimensional environment 902a in accordance with the continuation of the input shown in FIG. 9C. In some embodiments, the electronic device 101a updates the positions of the user interfaces 904a, 906a, and 908a, and the representations of the second user 916b and the second electronic device 101b relative to the viewpoint of the user 916a within the three-dimensional environment 902 according to the direction and amount of movement of the hands 913a and 913b in FIG. 9C. For example, the electronic device 101a updates the three-dimensional environment 902a to display the user interfaces 904a, 906a, and 908a and the representations of the second user 916b and the second electronic device 101b further away from the user's viewpoint. As shown in the top-down view of the three-dimensional environment 902, the electronic device 101a updates the viewpoint of the first user 916a within the three-dimensional environment 902 without updating the positions and / or orientations of the user interfaces 904, 906, and 908, and the representations of the second user 916b and the second electronic device 101b in the three-dimensional environment 902. Thus, the second electronic device 101b optionally continues to display the user interfaces 904b and 908b at the locations where the user interfaces 904b and 908b were displayed in FIG. 9C.

[0224] In some embodiments, in response to the first electronic device 101a detecting the end of an input, the second electronic device 101b updates the positions of the representation of the first user 916a and the representation of the first electronic device 101a, and updates the positions and / or orientations of the user interfaces 904a, 906a, and 908a, and the representations of the second user 916b and the second electronic device 101b with respect to the perspective of the first user 101a. As shown in FIG. 9D, the second electronic device 101b displays the representation of the first user 916a and the representation of the first electronic device 101a at the updated location, for example, with a reduced translucency and / or dimness relative to the translucency and / or dimness at which the representations of the first electronic device 101a and the first user 916a were displayed in FIG. 9C. In some embodiments, in response to the first electronic device 101a detecting the end of the input shown in FIG. 9C, the second electronic device 101b displays an animation that updates the positions of the representations of the first user 916a and the first electronic device 101a. In some embodiments, in response to the first electronic device 101a detecting the end of the input shown in FIG. 9C, the second electronic device 101b stops displaying the representations of the first user 916a and the first electronic device 101a at the location shown in FIG. 9C and starts displaying the representations of the first user 916a and the first electronic device 101a at the location of FIG. 9D (e.g., via a "teleporting" effect).

[0225] In some embodiments, the digital origin 912 is invalidated when the first user 916a is separated from one or more of the user interfaces 904a, 906a, and 908a, and the representations of the second user 916b and the second electronic device 101b by a threshold distance (e.g., 1, 2, 3, 5, 10, or 15 meters). In FIG. 9D, for example, the digital origin 912 is shown by a dashed line to illustrate that it is invalidated due to the first user 916a updating the positions of the user interfaces 904a, 906a, and 908a and the representations of the second user 916b and the second electronic device 101b relative to the perspective of the first user 916a such that one or more of the virtual objects are greater than the threshold distance from the perspective of the first user 916a. In some embodiments, the digital origin 912 becomes invalid when one or more of the user interfaces 904a, 906a, and 908a, and the representations of the second user 916b and the second electronic device 101b are moved away from the perspective of the first user 916a by a threshold distance in response to the second user 916b moving the representations of one or more of the user interfaces 904a and 908a, and the second user 916b and the second electronic device 101b (e.g., by individually moving one or more applications) away from the perspective of the user within the three-dimensional environment 902.

[0226] In some embodiments, as shown in FIG. 9D, when the user interfaces 904a, 906a, and 908a, and the representations of the second user 916b and the second electronic device 101b are moved away from the perspective of the first user 101a, one or more of the spatial references described above with reference to method 800 are no longer satisfied. In some embodiments, in response to one or more of the spatial references no longer being satisfied, the first electronic device 101a, when selected, displays a selectable option 914a to recenter the three-dimensional environment 902a on the first electronic device 101a according to one or more steps of method 800. In some embodiments, since the digital origin 912 is invalid, in response to detecting the selection of option 914a, the electronic device 101a updates the digital origin 912 and recenters the three-dimensional environment 902a based on the updated digital origin 912. In some embodiments, the first electronic device 101a selects a new location for the digital origin 912 based on the position and / or orientation of the objects within the three-dimensional environment 902 and one or more of the spatial references described with reference to methods 800 and / or 1000.

[0227] In some embodiments, the digital origin 912 is invalidated in response to one or more users moving beyond a threshold amount (e.g., 1, 2, 3, 5, 10, 15, or 25 meters) within the three-dimensional environment 902. For example, in FIG. 9D, the second user 916b moves himself and / or the electronic device 101b, as a result of which the location of the second user 916b in the three-dimensional environment 902 is updated as shown in FIG. 9E. In response to the second user 916b thus updating his position within the three-dimensional environment 902, the digital origin 912 is invalidated. FIG. 9E shows, in some embodiments, a method by which the electronic device 101a automatically selects a new location for the digital origin 912 without receiving a user input (e.g., selection of option 914a) that requests the electronic device 101a to update the digital origin 912 in response to the digital origin 912 being invalidated. For example, the updated digital origin 912 is located such that when the first user 916a re-centers the three-dimensional environment 902a to update the perspective of the digital origin 912, both users 916a and 916b are on the same side of the user interface 904 of the first application (e.g., the shared content application described above with reference to method 800) and on the opposite side of the user interface 908 of the third application (e.g., the shared activity application described above with reference to method 800).

[0228] As shown in FIG. 9E, in response to the second user 916b updating his perspective in the three-dimensional environment 902b, the second electronic device 101b updates the position and / or orientation of the user interfaces 904b and 908b with respect to the perspective of the second user 916b, and the first electronic device 101a stops displaying the representation of the second user 916b and the second electronic device 101b because the second user 916b is no longer within the field of view of the first electronic device 101a.

[0229] In some embodiments, as described below with reference to FIGS. 9F-9G, the first electronic device 101a invalidates the digital origin 912 in response to a request to share an application with the second electronic device 101b. In FIG. 9F, the first electronic device 101a displays a three-dimensional environment 902a that includes a user interface 918a of a fourth application that is accessible to the first electronic device 101a but not to the second electronic device 101b. Since the user interface 918 of the fourth application is not accessible to the second electronic device 101b, the second electronic device 101b displays a representation of the first user 916a and the first electronic device 101a without displaying the user interface 918 of the fourth application. However, FIG. 9F shows a location 918b where the user interface 918 of the fourth application would be displayed if the fourth application were accessible to the second electronic device 101b. FIG. 9F also shows the digital origin 912 associated with the first electronic device 101a. In some embodiments, the location of the digital origin 912 in FIG. 9F is valid because the user interface 918a of the fourth application and the representation of the second user 916b are in locations and orientations that satisfy one or more spatial criteria. In some embodiments, the spatial criteria do not require the second user 916b to have an unobstructed view of the user interface 918 of the fourth application because the second electronic device 101b does not have access to the fourth application.

[0230] As shown in FIG. 9F, the first electronic device 101a detects a selection (e.g., via hand 913a) of a selectable option 920 that, when selected, causes the first electronic device 101a to share a fourth application with the second electronic device 101b. In some embodiments, the selection input is one of a direct selection input, an indirect selection input, an air gesture selection input, or a selection input using an input device, as described above with reference to method 800. In response to the input, the first electronic device 101a makes the user interface 918 of the fourth application accessible to the second electronic device 101b, as shown in FIG. 9G.

[0231] FIG. 9G shows the second electronic device 101b displaying the user interface 918b of the fourth application in response to an input for sharing the fourth application detected by the first electronic device 101a in FIG. 9F. In response to the input shown in FIG. 9F, the first electronic device 101a disables the digital origin 912. FIG. 9G shows the updated location of the digital origin 912 where the first electronic device 101a moves the perspective of the first user 101a when the first electronic device 101a attempts to recenter the three-dimensional environment 902a. In some embodiments, the digital origin 912 associated with the first electronic device 101a is disabled in response to a request to share the user interface 918 of the fourth application because the representation 916a of the first user hides the user interface 918b of the fourth application from the view of the second user 916b.

[0232] Additional or alternative details regarding the embodiments shown in FIGS. 9A - 9G are provided below in the description of method 1000, which is described with reference to FIGS. 10A - 10K.

[0233] Figures 10A - 10K are flowcharts showing a method for updating the positions of a plurality of virtual objects together according to some embodiments. In some embodiments, method 1000 is executed in a computer system (e.g., computer system 101 of FIG. 1) that includes a display generation component (e.g., display generation component 120 of FIGS. 1, 3, and 4) (e.g., a head-up display, a display, a touch screen, a projector, etc.) and one or more cameras (e.g., a camera placed in the user's hand and facing downward (e.g., a color sensor, an infrared sensor, and other depth sensing cameras) or a camera facing forward from the user's head). In some embodiments, method 1000 is stored in a non-transitory computer-readable storage medium and is controlled by instructions executed by one or more processors of the computer system, such as one or more processors 202 of computer system 101 (e.g., control unit 110 of FIG. 1A). Some operations of method 1000 are optionally combined and / or the order of some operations is optionally changed.

[0234] In some embodiments, method 1000 is executed in an electronic device that communicates with a display generation component and one or more input devices (e.g., a mobile device (e.g., a tablet, smartphone, media player, or wearable device), or a computer). In some embodiments, the display generation component is a display integrated with the electronic device (optionally a touch screen display), an external display such as a monitor, projector, television, or a hardware component (optionally embedded or external) that projects a user interface and makes the user interface visible to one or more users. In some embodiments, the one or more input devices include an electronic device or component that can receive user input (e.g., capture user input, detect user input, etc.) and transmit information related to the user input to the electronic device. Examples of input devices include a touch screen, a mouse (e.g., external), a trackpad (optionally integrated or external), a touchpad (optionally integrated or external), a remote control device (e.g., external), another mobile device (e.g., separate from the electronic device), a handheld device (e.g., external), a controller (e.g., external), a camera, a depth sensor, an eye tracking device, and / or a motion sensor (e.g., a hand tracking device, a hand motion sensor), etc. In some embodiments, the electronic device communicates with a hand tracking device (e.g., one or more cameras, depth sensors, proximity sensors, touch sensors (e.g., touch screen, trackpad)). In some embodiments, the hand tracking device is a wearable device such as a smart glove. In some embodiments, the hand tracking device is a handheld input device such as a remote control or a stylus.

[0235] In some embodiments, such as FIG. 9B, while displaying a three-dimensional environment (e.g., 902a) from the perspective of a user (e.g., 916a) of an electronic device (e.g., 101a) via a display generation component (e.g., 120a), the three-dimensional environment (e.g., 902a) includes a plurality of virtual objects (e.g., 904a, 906a) at a first location within a first spatial arrangement relative to the user's perspective within the three-dimensional environment (e.g., 902a), and the electronic device (e.g., 101) receives an input (1002a) corresponding to a request to move one or more of the plurality of virtual objects (e.g., 904a, 906a) via one or more input devices. In some embodiments, the input and / or one or more inputs described with reference to method 1000 are air gesture inputs as described with reference to method 800. In some embodiments, the three-dimensional environment includes virtual objects such as application windows, operating system elements, representations of other users, and / or representations of content items and / or physical objects within the physical environment of the electronic device. In some embodiments, the representation of the physical object is displayed in the three-dimensional environment via a display generation component (e.g., a virtual passthrough or a video passthrough). In some embodiments, the representation of the physical object is a view of the physical object within the physical environment of the electronic device that is visible through a transparent portion of the display generation component (e.g., a true passthrough or an actual passthrough). In some embodiments, the electronic device displays a three-dimensional environment from the perspective of the user at a location within the three-dimensional environment corresponding to the physical location of the electronic device within the physical environment of the electronic device. In some embodiments, the three-dimensional environment is generated, displayed, or otherwise made visually perceivable by the device (e.g., a computer-generated reality (XR) environment such as a virtual reality (VR) environment, a mixed reality (MR) environment, or an augmented reality (AR) environment). In some embodiments, receiving the input includes detecting a predefined gesture performed by a predefined portion of the user (e.g., the hand(s)).In some embodiments, the default gesture is a pinch gesture in which the user touches a finger with the thumb and moves the finger and thumb apart, performed by one or more of the user's hands. For example, the input includes the user forming a pinch hand shape with both hands and moving both hands to cause the electronic device to update the positions of a plurality of virtual objects according to the movement of the user's hands. In some embodiments, the default gesture is a press gesture in which the user moves their hand(s) to a default location (e.g., a location corresponding to one or more user interface elements being moved, a location corresponding to an air gesture user interface element different from one or more user interface elements to which the input is directed) while forming a pointing hand shape in which one or more fingers are extended and one or more fingers are curled against the palm. In some embodiments, the input for moving one or more of the plurality of virtual objects has one or more of the characteristics of object movement input described with reference to method 1400.

[0236] In some embodiments, in response to receiving an input (1002b), such as in FIG. 9C, according to a determination that the input meets one or more first criteria (e.g., the input is directed to an individual object of a plurality of virtual objects rather than to the rest of the plurality of virtual objects) (1002c), an electronic device (e.g., 101) updates a three-dimensional environment (e.g., 902a) (1002d) to move an individual object of a plurality of virtual objects (e.g., 902a) in the three-dimensional environment (e.g., 904a) from a first location to a second location different from the first location according to the input, and the individual object at the second location (e.g., 904a) has a second spatial arrangement with respect to the user's (e.g., 916a) perspective that is different from the first spatial arrangement with respect to the user's (e.g., 916a) perspective. In some embodiments, the input changes the location of an individual object within the three-dimensional environment. In some embodiments, the input changes the location of an individual object with respect to other virtual objects within the three-dimensional environment. In some embodiments, the input changes the location of an individual object within the three-dimensional environment with respect to the user's perspective. In some embodiments, updating the location of the individual object includes changing the distance of the virtual object from the user's perspective. In some embodiments, the second individual location is based on the movement (e.g., speed, duration, distance, etc.) of a predefined part (e.g., a hand (singular or plural)) of the user providing the input. In some embodiments, the one or more first criteria include criteria that are met when the electronic device detects a predefined gesture performed with one hand rather than both hands of the user. In some embodiments, the one or more first criteria include criteria that are met when the input is directed to a user interface element, and the user interface element causes the electronic device to initiate a process of moving an individual virtual object without moving other virtual objects within the three-dimensional environment in response to the input directed to the user interface element.

[0237] In some embodiments, in response to receiving an input (1002b), according to a determination that the input meets one or more first criteria (e.g., the input is directed to an individual object among a plurality of virtual objects without being directed to the remainder of the one or more virtual objects) (1002c), an electronic device (e.g., 101) maintains one or more second objects (e.g., 906a in FIG. 9C) among the plurality of virtual objects at a first location in a three-dimensional environment in a first spatial arrangement with respect to the user's perspective (1002e). In some embodiments, the input causes the electronic device to maintain the locations of other virtual objects in the three-dimensional environment. In some embodiments, the input causes the electronic device to maintain the locations of other virtual objects with respect to the user's perspective. In some embodiments, the input causes the electronic device to maintain the locations of each second object with respect to other second objects.

[0238] In some embodiments, such as FIG. 9C, in response to receiving an input (1002b), according to a determination that the input meets one or more second criteria (e.g., the input is directed to a plurality of virtual objects), an electronic device (e.g., 101) updates a three-dimensional environment (e.g., 902a) to move a plurality of virtual objects (e.g., 904a, 906a) within the three-dimensional environment according to the input (1002f). After the movement, the plurality of virtual objects (e.g., 904a, 904b) have a third spatial arrangement with respect to the user's (e.g., 916a) perspective that is different from the first spatial arrangement with respect to the user's (e.g., 916a) perspective. In some embodiments, the one or more second criteria include criteria that are met when the electronic device detects a predefined gesture performed with both of the user's hands. In some embodiments, the one or more second criteria include criteria that are met when the input is directed to a user interface element, and the user interface element causes the electronic device to initiate a process of moving a plurality of virtual objects within the three-dimensional environment in response to the input directed to the user interface element. In some embodiments, the input causes the electronic device to update the locations of a plurality of virtual objects within the three-dimensional environment. In some embodiments, the input causes the electronic device to update the location of a virtual object with respect to the user's perspective. In some embodiments, during and / or in response to the input (e.g., while and / or after a plurality of elements are moving together based on the input), one or more (e.g., all) of the virtual objects maintain a spatial relationship with respect to each other and / or with respect to other virtual objects. In some embodiments, when the electronic device is communicating with a second electronic device having and / or associated with a second user, the electronic device maintains the locations of the virtual objects within the three-dimensional environment and updates the location of the user's perspective within the three-dimensional environment, so that the spatial orientation of the virtual objects with respect to the user's perspective is updated without changing the location of the virtual objects with respect to the perspective of the second user.

[0239] Updating a three-dimensional environment to move a plurality of objects according to a determination that an input meets one or more second criteria improves user interaction with an electronic device by reducing the input required to move a plurality of virtual objects in response to a single input, thereby enabling a user to use the electronic device quickly and efficiently.

[0240] In some embodiments, such as FIG. 9C, in response to receiving an input, updating a three-dimensional environment (e.g., 902a) to move a plurality of virtual objects (e.g., 904a, 904b) within the three-dimensional environment according to a determination that the input meets one or more second criteria includes moving a first virtual object (e.g., 904a) of the plurality of virtual objects in an individual (e.g., lateral, angular) direction by an individual (e.g., lateral, angular) amount and moving a second virtual object (e.g., 906a) of the plurality of virtual objects in an individual direction by an individual amount based on the direction and magnitude of the input (1004). In some embodiments, when moving a plurality of virtual objects, the virtual objects maintain an individual spatial orientation relative to other virtual objects and are displayed in different spatial arrangements relative to the user's perspective. For example, in response to an input to move a plurality of virtual objects in a first direction by a first amount, the electronic device moves the first virtual object and the second virtual object in the first direction by the first amount within the three-dimensional environment. As another example, in response to an input to rotate a plurality of objects in a second direction by a second amount, the electronic device moves the first virtual object and the second virtual to rotate the first virtual object and the second virtual object in the second direction by the second amount (e.g., around an individual reference point within the three-dimensional environment).

[0241] Responsive to receiving an input, moving the first and second virtual objects in individual directions by individual amounts provides an efficient way to maintain the individual spatial arrangements of the plurality of virtual objects while adjusting the positions and / or orientations of the plurality of virtual objects (e.g., collectively), thereby enabling a user to use the device quickly and efficiently with fewer inputs.

[0242] In some embodiments, such as FIG. 9C, one or more second criteria are satisfied when receiving an input includes detecting a first portion (e.g., a first hand) of a user (e.g., 913a) in a posture that satisfies one or more posture criteria and detecting a second portion (e.g., a second hand) of a user (e.g., 913b) in a posture that satisfies one or more posture criteria (1006). In some embodiments, one or more second criteria include criteria that are satisfied when the electronic device detects that the user has made a pinch hand shape with both hands (e.g., touching the thumb to another finger of the same hand) via one or more input devices (e.g., a hand tracking device). In some embodiments, responsive to detecting that the user has made a pinch hand shape with both hands, the electronic device initiates a process of moving a plurality of virtual objects, and responsive to detecting hand movement while the pinch hand shape is maintained, the electronic device moves the plurality of objects according to one or more movements (e.g., direction, speed, duration, distance, etc.) of the hand. In some embodiments, responsive to detecting an input directed to a user interface element for moving an individual virtual object, including detecting a pinch gesture with one hand (e.g., not both hands), the electronic device initiates a process of moving an individual virtual object (e.g., without moving a plurality of virtual objects).

[0243] Detecting the first and second portions of the user in a posture that satisfies one or more posture criteria and moving a plurality of objects accordingly provides an efficient way to update the positions of the plurality of virtual objects at once, thereby enabling the user to use the electronic device quickly and efficiently.

[0244] In some embodiments, such as FIG. 9A, while the electronic device (e.g., 101a) is not receiving input, the electronic device (e.g., 101) displays (1008a) a plurality of virtual objects (e.g., 904a, 906a) within a three-dimensional environment (e.g., 101a) via a display generation component (e.g., 120a) with a first amount of visual emphasis (e.g., opacity, vividness, color contrast, size, etc.). In some embodiments, while the electronic device is not receiving input, the electronic device displays the plurality of virtual objects in full color, opacity, and vividness. In some embodiments, while the electronic device is not receiving input, the electronic device displays a plurality of virtual objects having a first amount of visual emphasis relative to the rest of the three-dimensional environment.

[0245] In some embodiments, such as FIG. 9B, in accordance with a determination that the input satisfies one or more second criteria, while receiving the input, the electronic device (e.g., 101) displays (1008db) a plurality of virtual objects (e.g., 904a, 906a) within a three-dimensional environment (e.g., 902a) via a display generation component (e.g., 120a) with a second amount of visual emphasis that is less than the first amount of visual emphasis (e.g., opacity, vividness, color contrast, size, etc.). In some embodiments, while the electronic device is receiving input, the electronic device displays a plurality of virtual objects with reduced opacity and / or vividness, and / or displays a plurality of virtual objects in a modified color such as a lighter color, a darker color, or a color with lower chroma / contrast. In some embodiments, while the electronic device is receiving input, the electronic device displays a plurality of virtual objects with a second amount of visual emphasis that is less than the first amount of visual emphasis relative to the rest of the three-dimensional environment.

[0246] Reducing the visual emphasis of multiple objects while an input is being received provides an efficient way of presenting a portion of a three-dimensional environment that is proximate to and / or overlapped by multiple virtual objects while receiving input for updating the positions of the multiple virtual objects, which enhances user interaction with the electronic device by providing enhanced visual feedback while an input for updating the positions of the multiple virtual objects is being received, thereby enabling a user to quickly and efficiently interact with the electronic device using the enhanced visual feedback.

[0247] In some embodiments, such as FIG. 9B, an electronic device (e.g., 101b) displays (1010a) a representation of a second user (e.g., 916a) of a second electronic device (e.g., 101a) within a three-dimensional environment (e.g., 902b) via a display generation component (e.g., 120b). In some embodiments, the electronic device and the second electronic device are communicating with each other (e.g., via a network connection). In some embodiments, the electronic device and the second electronic device have access to the three-dimensional environment. In some embodiments, the electronic device displays the representation of the second user at a location within the three-dimensional environment corresponding to the perspective of the second user. In some embodiments, the second electronic device displays the representation of the user of the electronic device at a location within the three-dimensional environment corresponding to the perspective of the user of the electronic device.

[0248] In some embodiments, such as FIG. 9B, while a second electronic device (e.g., 101a) detects an individual input that meets one or more second criteria (e.g., an input for updating the positions of a plurality of virtual objects relative to the view of a second user's perspective), the electronic device (e.g., 101) displays (1010b) the representation of the second user (e.g., 916a) with a first amount of visual emphasis (e.g., opacity, vividness, color contrast, size, visual emphasis relative to the rest of the three-dimensional environment, etc.) via a display generation component (e.g., 120a). In some embodiments, while the second electronic device is detecting an individual input, the electronic device displays the representation of the second user with a reduced visual emphasis relative to the visual emphasis with which the representation of the second user was displayed while the second electronic device was not detecting an individual input. For example, while the second electronic device is detecting an individual input, the electronic device displays a representation of the second user with reduced opacity and / or vividness and / or a lighter, thinner, darker color and / or reduced contrast and / or saturation compared to how the electronic device displayed the representation of the second user while the second electronic device was not detecting an individual input. In some embodiments, the electronic device maintains the display of the representation of the second user at an individual location within the three-dimensional environment while the second electronic device is detecting an individual input. In some embodiments, the electronic device displays the representation of the second user with a first amount of visual emphasis while moving the representation of the second user according to the individual input detected by the second electronic device. For example, while the second electronic device detects an input corresponding to a request to move a plurality of virtual objects closer to the perspective of the second user, the electronic device displays the representation of the second user in a state where the amount of the first visual emphasis moves toward the plurality of virtual objects while the virtual objects move relative to the perspective of the user without displaying the movement of the virtual objects relative to the perspective of the user. In some embodiments, the electronic device receives an indication of an individual input from the second electronic device.

[0249] In some embodiments, such as FIG. 9A, while a second electronic device (e.g., 101a) does not detect an individual input that meets one or more second criteria, an electronic device (e.g., 101b) displays (1010c) a representation of a second user (e.g., 916a) with a second amount of visual emphasis (e.g., visual emphasis relative to the rest of the three-dimensional environment) that is greater than the first amount of visual emphasis via a display generation component (e.g., 120b). In some embodiments, while the second electronic device is not detecting an individual input, the electronic device displays the representation of the second user with a greater (or full) opacity and / or vividness and / or increased degree of saturation.

[0250] Displaying a representation of a second user while the second electronic device is detecting an individual input provides an efficient way to indicate to the user that the second user is involved in providing an individual input, thereby improving communication between users and enabling the user to use the electronic device quickly and efficiently.

[0251] In some embodiments, such as FIG. 9A, an electronic device (e.g., 101b) displays (1012a) a representation of a second user (e.g., 916a) of a second electronic device located at a third location within a three-dimensional environment (e.g., 902b) via a display generation component (e.g., 120b). In some embodiments, the electronic device and the second electronic device communicate with each other (e.g., via a network connection). In some embodiments, the electronic device and the second electronic device have access to the three-dimensional environment. In some embodiments, the third location is the location in the three-dimensional environment of the perspective of the second user. In some embodiments, the second electronic device displays a representation of the user of the electronic device at a location within the three-dimensional environment of the perspective of the user of the electronic device. In some embodiments, the electronic device displays additional representations of other users having access to the three-dimensional environment at locations within the three-dimensional environment of the perspectives of the other users.

[0252] In some embodiments, such as FIG. 9B, while displaying the representation of a second user (e.g., 916a) at a third location within a three-dimensional environment (e.g., 902b), an electronic device (e.g., 101b) receives an indication (1012b) that a second electronic device (e.g., 101a) has received an individual input that meets one or more second criteria. In some embodiments, the individual input is an input received at the second electronic device that corresponds to a request to update the position of a virtual object relative to the perspective of the second user in the three-dimensional environment.

[0253] In some embodiments, such as FIG. 9D, in response to this indication, the electronic device (e.g., 101b) displays the representation (e.g., 916a) of the second user at a fourth location within the three-dimensional environment (e.g., 902b) according to the individual input (1012c), without displaying an animation of the representation of the second user moving from the third location to the fourth location via a display generation component (e.g., 120b). In some embodiments, the electronic device displays a "jump" or "teleport" representation of the second user from the third location to the fourth location. In some embodiments, the fourth location is based on the individual input received by the second electronic device. For example, in response to an individual input that rotates a plurality of virtual objects clockwise relative to a reference point (e.g., inside the boundary around the plurality of virtual objects), in response to the indication of the individual input, the electronic device updates the location of the representation of the second user from the third location to a fourth location counterclockwise around the plurality of virtual objects. In some embodiments, the electronic device displays the representation of the second user at the third location within the three-dimensional environment until it detects the end of an input received by the second electronic device that meets one or more second criteria.

[0254] Displaying the representation of the second user at the fourth location without displaying the animation of the representation of the second user moving from the third location to the fourth location provides an efficient way to update the representation of the second user according to the perspective of the second user with reduced distraction, thereby improving user interaction with the electronic device and enabling the user to use the electronic device quickly and efficiently.

[0255] In some embodiments, one or more second criteria are satisfied when the input is directed to a user interface element that is displayed via a display generation component (e.g., 120a), and include a criterion for moving a plurality of virtual objects (e.g., 904a, 906a of FIG. 9A) according to the input on the electronic device (e.g., 101a) when the input is directed to the user interface element (1014). In some embodiments, the electronic device also displays, for each of the plurality of virtual objects, a respective user interface element associated with that virtual object that causes the electronic device to move the individual virtual object (e.g., without moving the plurality of virtual objects) when an input that satisfies one or more first criteria is directed to the individual user interface element. In some embodiments, in response to detecting a selection of a user interface element for moving a plurality of virtual objects on the electronic device, the electronic device begins a process of moving the plurality of virtual objects and, in response to further input including movement (e.g., another directional input such as interaction of a user's individual part with a directional button or key, operation of a joystick, etc.), the electronic device moves the plurality of objects according to the movement. In some embodiments, detecting a selection of a user interface element includes detecting the user's attention (e.g., line of sight) directed to the user interface element and detecting that the user performs a predefined gesture with the user's individual part (e.g., the user's hand), such as making the pinching hand shape, via (e.g., a hand tracking device). In some embodiments, after detecting a selection of a user interface element, the electronic device moves a plurality of virtual objects according to the movement of the user's individual part (e.g., the user's hand), while the user's individual part maintains the pinching hand shape. In some embodiments, detecting an input directed to a user interface element includes detecting an input provided by one hand (not both hands) of the user.

[0256] Moving a plurality of virtual objects in response to input directed to individual user interface elements provides an efficient way to teach the user how to move the plurality of virtual objects, thereby enabling the user to use the electronic device quickly and efficiently.

[0257] In some embodiments, in accordance with a determination that the input meets one or more second criteria (1016a), and in accordance with a determination that the input includes a first magnitude of movement, such as in FIG. 9B, an electronic device (e.g., 101) moves a plurality of virtual objects (e.g., 904a, 906a) in a three-dimensional environment (e.g., 902a) by a second amount (1016b). In some embodiments, the second amount of movement (e.g., speed, duration, distance) corresponds to the first magnitude of movement (e.g., speed, duration, distance).

[0258] In some embodiments, in accordance with a determination that the input meets one or more second criteria (1016a), and in accordance with a determination that the input includes a third magnitude of movement, such as in FIG. 9C, that is different from the first magnitude of movement, the electronic device (e.g., 101) moves a plurality of virtual objects (e.g., 904a, 906a) in a three-dimensional environment (e.g., 902a) by a fourth amount that is different from the second amount (1016c). In some embodiments, the fourth amount of movement (e.g., speed, duration, distance) corresponds to the third magnitude of movement (e.g., speed, duration, distance). In some embodiments, if the first magnitude of movement is greater than the third magnitude of movement, the second amount is greater than the fourth amount. In some embodiments, if the first magnitude of movement is less than the third magnitude of movement, the second amount is less than the fourth amount. For example, if the first magnitude of movement (e.g., of the user's hand(s)) includes a movement that is faster or has a longer distance and / or duration than the third magnitude of movement (e.g., of the user's hand(s)), the second amount is greater than the fourth amount. As another example, if the first magnitude of movement (e.g., of the user's hand(s)) includes a movement that is slower or has a shorter distance and / or duration than the third magnitude of movement (e.g., of the user's hand(s)), the second amount is less than the fourth amount.

[0259] Moving a plurality of virtual objects by an amount corresponding to the magnitude of movement of the input provides an efficient way for the user to control the amount of movement of the plurality of objects with enhanced control, thereby enabling the user to use the electronic device quickly and efficiently.

[0260] In some embodiments, before receiving an input corresponding to a request to move one or more of a plurality of virtual objects (e.g., 904a, 906a in FIG. 9A), an electronic device (e.g., 101a) displays, via a display generation component (e.g., 120a), a plurality of user interface elements associated with the plurality of virtual objects (e.g., 904a, 906a) (1018), and one or more first criteria are satisfied when the input is directed to an individual user interface element associated with an individual object among the plurality of user interface elements, and when the input is directed to the individual user interface element, include criteria for causing the electronic device to initiate a process of moving an individual object among the plurality of virtual objects in a three-dimensional environment (e.g., 902a). In some embodiments, the plurality of virtual objects includes a first virtual object and a second virtual object. In some embodiments, when selected, the electronic device displays a first selectable element proximate to the first virtual object that causes the electronic device to initiate a process of moving the first virtual object within the three-dimensional environment. In some embodiments, when selected, the electronic device displays a second selectable element proximate to the second virtual object that causes the electronic device to initiate a process of moving the second virtual object within the three-dimensional environment. In some embodiments, in response to detecting the user's attention (e.g., line of sight) directed to the first virtual object, optionally, while detecting an individual part of the user (e.g., hand) in a predetermined shape, such as a pre-pinch hand shape where the thumb is within a predetermined threshold distance (e.g., 0.5, 1, 2, 3, 4, or 5 centimeters) of another finger of the hand, or a pointing hand shape where one or more fingers are extended and one or more fingers are bent towards the palm of the hand, optionally, while the hand is within a predetermined threshold distance (e.g., 1, 2, 3, 5, 10, 15, 30, 50, or 100 centimeters) of the first object, the first selectable element is displayed.In some embodiments, in response to detecting the user's attention (e.g., line of sight) directed to a second virtual object, the electronic device optionally detects an individual part of the user (e.g., a hand) in a predefined shape, such as the shape of the hand before pinching or the pointing hand shape, while optionally displaying a second selectable element while the hand is within a predefined threshold distance (e.g., 1, 2, 3, 5, 10, 15, 30, 50, or 100 centimeters) of the second object.

[0261] In response to an input directed to an individual user interface element that initiates a process of moving an individual virtual object to the electronic device, moving an individual virtual object (e.g., without moving a plurality of virtual objects) provides an efficient way to select which virtual object to move, thereby improving the user interaction with the electronic device and enabling the user to use the electronic device quickly and efficiently.

[0262] In some embodiments, while displaying a three-dimensional environment (e.g., 902a), an electronic device (e.g., 101a) receives (1020a) an input (e.g., an input described with reference to method 800, etc.) corresponding to a request to update the three-dimensional environment (e.g., 902a) via one or more input devices to meet one or more spatial references to a digital origin (e.g., 912 in FIG. 9D). In some embodiments, the request corresponds to (e.g., is, includes) a request to update the spatial arrangement of a plurality of virtual objects with respect to the user's current viewpoint so as to meet one or more criteria that specify a range of distances and / or orientations of the virtual objects with respect to the user's current viewpoint and / or the digital origin according to one or more steps of method 800. In some embodiments, the digital origin is located at a location within the three-dimensional environment that the electronic device uses to evaluate one or more location and / or orientation criteria with respect to a plurality of virtual objects and / or the user's viewpoint when updating the three-dimensional environment in response to a request to update the spatial arrangement of a plurality of virtual objects with respect to the user's current viewpoint so as to meet one or more spatial criteria that specify a range of distances and / or orientations of the virtual objects with respect to the user's current viewpoint and / or the digital origin according to one or more steps of method 800. In some embodiments, the digital origin and the user's viewpoint are located at the same location within the three-dimensional environment. In some embodiments, the digital origin and the user's viewpoint are located at different locations within the three-dimensional environment.

[0263] In some embodiments, in response to an input, an electronic device (e.g., 101a) updates (1020b) a three-dimensional environment (e.g., 902a) to satisfy one or more spatial references relative to a digital origin. In some embodiments, updating the three-dimensional environment to satisfy one or more spatial references relative to a digital origin corresponds to (e.g., consists of, includes) updating the spatial arrangement of one or more virtual objects and / or updating the user's viewpoint according to one or more steps of method 800. In some embodiments, if a virtual object satisfies one or more references relative to the digital origin but the user's viewpoint does not satisfy one or more references relative to the digital origin, the electronic device updates the user's viewpoint in response to the input. In some embodiments, if the user's viewpoint satisfies one or more references relative to the digital origin but one or more virtual objects do not satisfy one or more references relative to the digital origin, the electronic device updates the location and / or orientation of one or more virtual objects in response to the input.

[0264] Updating the three-dimensional environment to satisfy one or more spatial references relative to a digital origin improves the user interaction with the electronic device by providing a consistent experience when the user makes a request to update the three-dimensional environment, thereby enabling the user to use the electronic device quickly and efficiently.

[0265] In some embodiments, such as FIG. 9A, the digital origin (e.g., 912) is determined (1022) when the electronic device (e.g., 101a) starts an augmented reality or virtual reality session that includes displaying a three-dimensional environment (e.g., 902a). In some embodiments, the digital origin is determined when the electronic device detects that the spatial orientation of the electronic device or the display generation component relative to the user has reached a predetermined spatial orientation. For example, the electronic device and / or the display generation component is a wearable device, and the predetermined spatial location is when the user is wearing and / or has the electronic device and / or the display generation component on. In some embodiments, the digital origin is initially at a location within the three-dimensional environment of the user's perspective (e.g., when the AR or VR session starts, or when the spatial orientation of the electronic device or the display generation component relative to the user first reaches a predetermined spatial orientation).

[0266] Determining the digital origin when the AR or VR session starts improves the user interaction with the electronic device by establishing a reference point at the start of the session, thereby providing the user with a consistent experience and enabling the user to use the electronic device quickly and efficiently with reduced user errors.

[0267] In some embodiments, such as FIG. 9A, before receiving an input that meets one or more second criteria, a digital origin (e.g., 912) is placed at a first distinct location within a three-dimensional environment (e.g., 902a) (1024a). In some embodiments, the first distinct location of the digital origin is selected at the start of a VR or AR session. In some embodiments, the first distinct location of the digital origin is a location selected in response to a request to update the three-dimensional environment so that the digital origin meets one or more spatial criteria received while the digital origin is in an invalid state, as described in more detail below. In some embodiments, the first distinct location of the digital origin is a location selected in response to an input prior to meeting one or more second criteria, as described below.

[0268] In some embodiments, such as FIG. 9C, in response to receiving an input that meets one or more second criteria, in accordance with the determination that the input meets one or more second criteria, an electronic device (e.g., 101a) updates the digital origin (e.g., 912) to be placed at a second distinct location that is different from the first distinct location within the three-dimensional environment (e.g., 902a) (1024b). In some embodiments, the electronic device updates the digital origin to maintain the spatial relationship between the digital origin and a plurality of virtual objects that existed when the input was received. In some embodiments, the electronic device updates the digital origin to change the spatial relationship between the digital origin and a plurality of virtual objects that existed when the input was received. In some embodiments, the user's viewpoint does not move in response to receiving an input that meets one or more criteria. In some embodiments, the electronic device updates the digital origin to be located at the location of the user's viewpoint within the three-dimensional environment. In some embodiments, after updating the location of the digital origin, in response to receiving an input corresponding to a request that meets one or more spatial criteria for the digital origin and / or the user's viewpoint, in accordance with one or more steps of method 800, the electronic device updates the user's viewpoint (e.g., to be located at the location of the digital origin).

[0269] Updating the digital origin in response to an input that meets one or more second criteria is used to update the location in a three-dimensional environment to evaluate the spatial orientation of the virtual object and the user's perspective according to the updated spatial arrangement of a plurality of objects requested by the user, thereby improving the user interaction with the electronic device and enabling the user to use the electronic device quickly and efficiently.

[0270] In some embodiments, such as FIG. 9F, while a three-dimensional environment (e.g., 902) is accessible to a second user of a second electronic device (e.g., 101b) (1026a), the electronic device (e.g., 101a) displays, via a display generation component (e.g., 120a), a plurality of virtual objects (e.g., 918a) within the three-dimensional environment (e.g., 101a) that includes an individual virtual object (e.g., 918a) to which the electronic device (e.g., 101a) is accessible while the digital origin (e.g., 912) is located at a first individual location within the three-dimensional environment but is not accessible to the second electronic device (e.g., 101b) (1026b). In some embodiments, the second electronic device does not display the individual virtual object while the individual virtual object is not accessible to the second electronic device. In some embodiments, the spatial relationship between two or more of each virtual object, the user's perspective, and the digital origin meets one or more spatial criteria while the individual virtual object is not accessible to the second electronic device. In some embodiments, if the individual virtual object were accessible to the second electronic device, one or more spatial relationships between two or more of each virtual object, the user's perspective, the digital origin, and the perspective of the second user of the second electronic device would not meet one or more spatial criteria. In some embodiments, the individual virtual object is accessible to the electronic device but not to the second electronic device, and the individual virtual object is a "private" object (e.g., of the user).

[0271] In some embodiments, such as FIG. 9F, while a three-dimensional environment (e.g., 902) is accessible to a second user of a second electronic device (e.g., 101b) (1026a), an individual virtual object (e.g., 918a) is accessible to an electronic device (e.g., 101a) but not to the second electronic device (e.g., 101b), and while a digital origin (e.g., 912) is located at a first individual location within the three-dimensional environment (e.g., 902), the electronic device (e.g., 101) receives an input (1026c) corresponding to a request to make the individual virtual object (e.g., 918a) accessible to the electronic device (e.g., 101a) and the second electronic device (e.g., 101b) via one or more input devices (e.g., 314).

[0272] In some embodiments, such as FIG. 9G, while a three-dimensional environment (e.g., 902) is accessible to a second user of a second electronic device (1026a), in response to an input (1026d) corresponding to a request to make an individual object (e.g., 918a) accessible to an electronic device (e.g., 101a) and the second electronic device (e.g., 101b), the electronic device (e.g., 101a) updates (1026e) the individual virtual object (e.g., 918a) to be accessible to the electronic device (e.g., 101a) and the second electronic device (e.g., 101b). In some embodiments, while the individual virtual object is accessible to the electronic device and the second electronic device, the electronic device displays the virtual object in the three-dimensional environment via a display generation component, and the second electronic device displays the virtual object in the three-dimensional environment via a second display generation component that communicates with the second electronic device. In some embodiments, the individual virtual object is accessible to the electronic device and the second electronic device, and the individual virtual object is a "shared" object.

[0273] In some embodiments, such as FIG. 9G, while a three-dimensional environment (e.g., 902) is accessible to a second user of a second electronic device (e.g., 101b) (1026a), in response to an input corresponding to a request to make an individual object (e.g., 918) accessible to an electronic device (e.g., 101a) and the second electronic device (e.g., 101b) (1026d), the electronic device (e.g., 101) updates a digital origin (e.g., 912) to be located at a second individual location different from a first individual location within the three-dimensional environment (e.g., 902) according to the location of the individual virtual object (e.g., 918a) (and / or the location of a second perspective of the three-dimensional environment associated with the second electronic device). In some embodiments, while the digital origin is located at the second individual location, one or more spatial relationships among two or more of each virtual object, the user's perspective, the digital origin, and the second user's perspective of the second electronic device satisfy one or more spatial criteria. In some embodiments, the digital origin is associated with the perspective of the user of the electronic device, and the second user's perspective is associated with a second digital origin different from the digital origin. In some embodiments, the digital origin is associated with the three-dimensional environment, the perspective of the first user, and the perspective of the second user. In some embodiments, the electronic device does not update the perspective of use in response to an input corresponding to a request to make an individual object accessible to the electronic device and the second electronic device. In some embodiments, while the digital origin is at the second individual location, in response to receiving an input corresponding to a request to update the spatial arrangement of a plurality of virtual objects, the user's perspective, and the digital origin according to one or more steps of method 800, the electronic device updates the user's perspective (e.g., to be located at the second individual location).In some embodiments, when both the first electronic device and the second electronic device receive input corresponding to a request to update the spatial arrangement of a plurality of virtual objects, the user's perspective, and the digital origin according to one or more steps of method 800, the perspectives of both users are updated to satisfy one or more spatial criteria, thereby positioning the perspectives of the user and the individual virtual objects to facilitate interaction with and viewing of the individual virtual objects by both users.

[0274] Updating the digital origin in response to the input to make an individual virtual object accessible on the second electronic device provides an efficient way to establish a reference point within a three-dimensional environment that is consistent with sharing the individual virtual object with the second electronic device, thereby enabling the user to use the electronic device quickly and efficiently.

[0275] In some embodiments, such as FIG. 9D, the user's perspective is a first perspective within a three-dimensional environment (e.g., 902), the digital origin (e.g., 912) is located at a first distinct location within the three-dimensional environment (e.g., 902), and while having an active status (e.g., the digital origin is active) at the first distinct location within the three-dimensional environment (e.g., 902), the electronic device (e.g., 101b) receives an input (1028a) corresponding to a request to update the user's perspective to a second perspective different from the first perspective within the three-dimensional environment via one or more input devices. In some embodiments, the digital origin has an active status when the spatial arrangement(s) among two or more of the digital origin, one or more virtual objects, and the user's perspective satisfy one or more spatial criteria. In some embodiments, the one or more spatial criteria include criteria that are satisfied when a threshold number of spatial relationships (e.g., 1, 2, or 3, and / or 25%, 50%, or 75%) between the digital origin and the virtual object and / or the user's perspective satisfy one or more criteria that specify a range of distances or a range of orientations of the virtual object as described above with reference to method 800. For example, the criteria for determining whether the digital origin is active include criteria that are satisfied when 50% of the spatial relationships satisfy the criteria. In some embodiments, while the digital origin is active, in response to an input corresponding to a request to update the three-dimensional environment so as to satisfy one or more criteria that specify a range of distances or a range of orientations of the virtual object with respect to the user's current perspective as described above with reference to method 800, the electronic device maintains the location of the digital origin and updates the spatial relationship between the virtual object and the user's perspective according to the digital origin.

[0276] In some embodiments, such as FIG. 9E, in response to an input (1028b) corresponding to a request to update the user's perspective to a second perspective within a three-dimensional environment (e.g., 902b), the electronic device (e.g., 101b) displays (1028c) the three-dimensional environment (902b) from the second perspective via a display generation component (e.g., 120b).

[0277] In some embodiments, such as FIG. 9E, in response to an input corresponding to a request (1028b) to update the user's perspective to a second perspective within a three-dimensional environment (e.g., 902b), and according to a determination that the second perspective is within a threshold distance of the first perspective, the electronic device (e.g., 101) maintains (1028d) the active status of the digital origin (e.g., 912) at a first individual location within the three-dimensional environment (e.g., 902b).

[0278] In some embodiments, such as FIG. 9E, in response to an input corresponding to a request to update the user's perspective to a second perspective within the three-dimensional environment (1028b), and in accordance with a determination that the second perspective is farther from the first perspective than a threshold distance, the electronic device (e.g., 101b) updates the status of the digital origin (e.g., 912) to an invalid status (e.g., determines that the digital origin is not valid) (1028e). In some embodiments, in response to updating the valid status of the digital origin to an invalid status, the electronic device selects a new location within the three-dimensional environment for the digital origin (e.g., according to one or more criteria regarding the spatial relationship between the digital origin and the virtual object and / or the user's perspective). In some embodiments, while the digital origin is not valid, the electronic device does not update the location of the digital origin until it receives an input corresponding to a request to update the three-dimensional environment to satisfy one or more criteria that specify a range of distances or a range of orientations of the virtual object with respect to the current perspective of the user as described above with reference to method 800, and then the electronic device updates the location of the digital origin and optionally updates the spatial relationship between the virtual object and the user's perspective according to the updated digital origin. In some embodiments, the criteria for the valid status of the digital origin include criteria that are satisfied when the user's perspective remains within the threshold distance and not satisfied when the user's perspective moves beyond the threshold distance. In some embodiments, the electronic device does not update the user's perspective in accordance with a determination that the second perspective is not valid or when updating the location of the digital origin. In some embodiments, in response to an input corresponding to a request to update the spatial arrangement of the virtual object and the user's perspective according to one or more spatial criteria by one or more steps of method 800, the electronic device updates the user's perspective according to the updated digital origin.

[0279] Invalidating the digital origin in response to movement of the user's perspective beyond a threshold distance provides an efficient way to update the user interface according to the user's updated perspective, which improves user interaction with the electronic device by enabling the user to select a perspective within a three-dimensional environment, thereby enabling the user to use the electronic device quickly and efficiently.

[0280] In some embodiments, a device (e.g., and / or an electronic device) including a display generation component (e.g., 120a) is in a posture that satisfies one or more posture criteria with respect to an individual part of the user (e.g., the user's head) (e.g., the user wears the display generation component on the head in a predetermined manner), while the digital origin (e.g., 912 such as in FIG. 9A) is located at a first individual location within a three-dimensional environment (e.g., 902), the electronic device (e.g., 101a) detects (1030a) movement of the display generation component (e.g., 120a) to a posture that does not satisfy one or more posture criteria with respect to the individual part of the user. In some embodiments, the display generation component (e.g., and / or the electronic device) is a wearable device, and the one or more posture criteria are satisfied when the user wears the display generation component (e.g., and / or the electronic device) on his or her head. In some embodiments, the one or more posture criteria are not satisfied when the user does not wear the display generation component (e.g., and / or the electronic device) on his or her head, or when the user wears the display generation component (e.g., and / or the electronic device) on his or her head but not in front of his or her face in a predetermined posture.

[0281] In some embodiments, while a device including a display generation component (e.g., 120a) is in a posture with respect to an individual part of a user that does not meet one or more posture criteria, the electronic device (e.g., 101) detects (1030b) a movement of the display generation component (e.g., 120a) to a posture with respect to an individual part of the user that meets one or more posture criteria. In some embodiments, after not wearing the display generation component (e.g., and / or the electronic device), the user begins to wear the display generation component (e.g., and / or the electronic device) in a posture that meets one or more posture criteria.

[0282] In some embodiments, in response to a movement of the display generation component (e.g., 120) to a posture with respect to an individual part of the user that meets one or more posture criteria (1030c), the device including the display generation component (e.g., 120a) determines that the electronic device (e.g., 101a) maintains the digital origin (e.g., 912 such as in FIG. 9A) at a first individual location within the three-dimensional environment (e.g., 902) according to the determination that the device was in a posture with respect to an individual part of the user that did not meet one or more posture criteria for a time less than a predetermined time threshold (e.g., 1, 2, 3, 5, 30, or 45 seconds, or 1, 2, 3, 5, or 10 minutes). In some embodiments, if the posture does not meet one or more posture criteria for a time less than a predetermined time threshold, the electronic device does not reset the digital origin and does not update the user's viewpoint.

[0283] In some embodiments, in response to movement of a display generation component (e.g., 120a) to a pose with respect to an individual part of a user that satisfies one or more pose criteria (e.g., 1030c), a device that includes the display generation component (e.g., 120a) determines that the pose of the device has been in a pose with respect to an individual part of the user that does not satisfy one or more pose criteria for a time longer than a predetermined time threshold, and in accordance with this determination, an electronic device (e.g., 101a) updates (1030e) a digital origin (e.g., 912 such as in FIG. 9A) to a second individual location that is different from a first individual location within a three-dimensional environment (e.g., 902). In some embodiments, the second individual location is associated with (e.g., is) the location of the user's perspective within the three-dimensional environment based on the physical location of the user and / or the display generation component (e.g., and / or the electronic device). In some embodiments, after a threshold period has elapsed, when the movement of the display generation component to a pose with respect to an individual part of the user satisfies one or more criteria and the digital origin remains valid at the first individual location, the electronic device maintains the digital origin at the first individual location. In some embodiments, updating the digital origin does not include updating the user's perspective. In some embodiments, in response to an input corresponding to a request to update the spatial arrangement of virtual objects and the user's perspective according to one or more spatial criteria by one or more steps of method 800, the electronic device updates the user's perspective according to the updated digital origin.

[0284] Updating the digital origin in response to movement of the display generation component to a pose that satisfies the pose criteria when the pose criteria have not been satisfied for a predetermined time threshold enables the user to update the three-dimensional environment when the pose of the display generation component (e.g., and / or the electronic device) satisfies one or more pose criteria, thereby improving the user interaction with the electronic device by enabling the user to use the electronic device quickly and efficiently.

[0285] In some embodiments, such as FIG. 9A, while displaying a three-dimensional environment (e.g., 902a) that includes a plurality of virtual objects (e.g., 904a, 906a) in a first individual spatial arrangement relative to a user's perspective, the first individual spatial arrangement includes displaying the plurality of virtual objects (e.g., 904a, 906a) within a predetermined threshold distance of the user's perspective, a digital origin (e.g., 912) is located at a first individual location within the three-dimensional environment (e.g., 902a), and while having an active status at the first individual location within the three-dimensional environment (e.g., 902a), the electronic device (e.g., 101) detects (1032a) an indication of one or more inputs corresponding to a request to update the spatial arrangement of the plurality of virtual objects (e.g., 904a, 906a) relative to the user's perspective. In some embodiments, the one or more inputs meet one or more first criteria, such as the inputs described with reference to method 1400, and do not meet one or more second criteria (e.g., the one or more inputs correspond to a request to individually update the position of a virtual object).

[0286] In some embodiments, such as FIG. 9D, in response to detecting an indication of one or more inputs, the electronic device (e.g., 101) displays (1032b) a plurality of virtual objects (e.g., 904a, 904b) in a second individual spatial arrangement relative to the user's perspective via a display generation component (e.g., 120a).

[0287] In some embodiments, such as FIG. 9C, according to a determination that a second individual spatial arrangement includes displaying at least one of a plurality of virtual objects (e.g., 904a, 906a) within a predefined threshold distance (e.g., 1, 2, 3, 5, 10, 15, or 30 meters) from a user's perspective, an electronic device (e.g., 101) maintains the active status of the digital origin at a first individual location within the three-dimensional environment (1032c). In some embodiments, while the digital origin has an active status, in response to a request to update the spatial arrangement of the virtual objects to meet one or more of the criteria described above with reference to method 800, the electronic device updates the three-dimensional environment according to the first individual location of the digital origin. In some embodiments, when the electronic device maintains the digital origin at the first individual location, it does not update the user's perspective.

[0288] In some embodiments, such as FIG. 9D, according to a determination that a second individual spatial arrangement includes displaying a plurality of virtual objects (e.g., 904a, 906a) beyond a predefined threshold distance from a user's perspective, an electronic device (e.g., 101) updates the status of the digital origin to an inactive status (1032d). In some embodiments, while the digital origin has an inactive status, in response to a request to update the spatial arrangement of the virtual objects to meet one or more of the criteria described above with reference to method 800, the electronic device updates the three-dimensional environment according to the second individual location of the digital origin. In some embodiments, the second individual location is determined in response to updating the active status of the digital origin to an inactive status. In some embodiments, the second individual location is determined in response to a request to update the spatial arrangement of the virtual objects to meet one or more of the criteria described above with reference to method 800.

[0289] Disabling the digital origin when multiple objects are farther from the user than a predefined threshold distance provides an efficient way to update the three-dimensional environment to display virtual objects closer to the user's perspective using the updated digital origin, thereby enhancing the user interaction with the electronic device by enabling the electronic device to be used quickly and efficiently.

[0290] In some embodiments, such as FIG. 9F, while the location of the digital origin (e.g., 912) meets one or more digital origin criteria, the electronic device (e.g., 101) maintains the digital origin (e.g., 912) at a first distinct location within the three-dimensional environment (1034a). In some embodiments, the digital origin criteria include one or more of the above-described criteria, such as the user's perspective that does not move beyond a threshold distance, the user's perspective within a threshold distance of one or more vir...

Claims

Claim 1 A method comprising: in an electronic device that communicates with a display generation component and one or more input devices, while displaying, via the display generation component, a three-dimensional environment including a plurality of virtual objects having a first spatial arrangement with respect to a current viewpoint of a user of the electronic device, detecting, via the one or more input devices, a movement of the current viewpoint of the user in the three-dimensional environment from a first viewpoint to a second viewpoint; in response to detecting the movement of the current viewpoint of the user from the first viewpoint to the second viewpoint, displaying, via the display generation component, the three-dimensional environment from the second viewpoint including the plurality of virtual objects having a second spatial arrangement different from the first spatial arrangement with respect to the current viewpoint of the user; while displaying the three-dimensional environment from the second viewpoint including the plurality of virtual objects having the second spatial arrangement with respect to the current viewpoint of the user, receiving, via the one or more input devices, an input corresponding to a request to update a spatial arrangement of the plurality of virtual objects with respect to the current viewpoint of the user so as to satisfy one or more criteria specifying a range of distances or a range of orientations of the virtual objects with respect to the current viewpoint of the user; in response to the input corresponding to the request to update the three-dimensional environment, displaying, via the display generation component, the three-dimensional environment from the second viewpoint including displaying the plurality of virtual objects having a third spatial arrangement different from the second spatial arrangement with respect to the viewpoint of the user, wherein the third spatial arrangement of the plurality of virtual objects satisfies the one or more criteria. Claim 2 The method of claim 1, wherein receiving the input corresponding to the request to update the spatial arrangement of the plurality of virtual objects comprises receiving the input via a hardware input device of the one or more input devices. Claim 3 The input corresponding to the request to update the spatial arrangement of the plurality of virtual objects satisfies one or more first input criteria, and the method comprises: receiving a second input via the hardware input device; in response to receiving the second input, performing an individual operation corresponding to the second input without updating the spatial arrangement of the plurality of virtual objects according to a determination that the second input meets one or more second input criteria different from the one or more first input criteria; The method according to claim 2, further comprising:

4. Receiving the input corresponding to the request to update the spatial arrangement of the plurality of virtual objects includes detecting a selection of a user interface element displayed in the three-dimensional environment via the display generation component. The method according to claim 1.

5. Displaying the plurality of virtual objects having the second spatial arrangement includes displaying the plurality of virtual objects at a first position within the three-dimensional environment via the display generation component, and the method includes: in response to receiving the input, further comprising moving the plurality of virtual objects from the first position to a second position within the three-dimensional environment. The method according to claim 1.

6. While displaying the three-dimensional environment including an individual virtual object among the plurality of virtual objects, the spatial arrangement of the plurality of virtual objects with respect to the current viewpoint of the user satisfies the one or more criteria. Receiving an input corresponding to a request to update the position of the individual virtual object within the three-dimensional environment via the one or more input devices; in response to receiving the input corresponding to the request to update the position of the individual virtual object in the three-dimensional environment, via the display generation component, having a second individual spatial arrangement that does not meet the one or more criteria. Displaying the plurality of virtual objects, including displaying the individual virtual object at a second position different from the first position in the three-dimensional environment. While displaying the three-dimensional environment including displaying the individual virtual object at the second position within the three-dimensional environment, receiving, via the one or more input devices, a second input corresponding to a request to update a spatial arrangement of the plurality of virtual objects with respect to a current viewpoint of the user so as to satisfy the one or more criteria. Further comprising, in response to receiving the second input, updating the position of the individual virtual object so as to satisfy the one or more criteria without updating the positions of one or more other virtual objects within the plurality of virtual objects. The method according to claim 1. **Claim 7** Displaying the plurality of virtual objects having the third spatial arrangement According to a determination that the three-dimensional environment is associated with a first spatial template, displaying, via the display generation component, the plurality of virtual objects having the third spatial arrangement, including displaying, via the display generation component, an individual virtual object among the plurality of virtual objects in an orientation with respect to the current viewpoint of the user that satisfies one or more criteria associated with the first spatial template. According to a determination that the three-dimensional environment is associated with a second spatial template, displaying, via the display generation component, the plurality of virtual objects having the third spatial arrangement, including displaying, via the display generation component, the individual object among the plurality of virtual objects in an orientation with respect to the current viewpoint of the user that satisfies one or more criteria associated with the second spatial template. The method according to claim 1, including. **Claim 8** Displaying the plurality of virtual objects having the third spatial arrangement According to a determination that the three-dimensional environment is associated with a shared content space template, displaying, via the display generation component, the individual object among the plurality of virtual objects in a posture in which an individual surface of the individual object faces the viewpoint of the user and the second viewpoint of the second user within the three-dimensional environment. The method according to claim 7, including. **Claim 9** Displaying the plurality of virtual objects having the third spatial arrangement comprises: displaying, via the display generation component, the individual object among the plurality of virtual objects, according to a determination that the three-dimensional environment is associated with a shared activity space template, in a posture in which a first side of the individual object faces the user's viewpoint and a second side of the individual object, different from the first side, faces a second viewpoint of a second user, the method according to claim 7.

10. Displaying the plurality of virtual objects having the third spatial arrangement comprises: displaying, via the display generation component, a representation of a second user in a posture directed toward the current viewpoint of the user of the electronic device, according to a determination that the three-dimensional environment is associated with a group activity space template, the method according to claim 7.

11. While displaying the three-dimensional environment including a second user associated with a second viewpoint within the three-dimensional environment via the display generation component, the spatial arrangement of the plurality of virtual objects with respect to the current viewpoint of the user is a first individual spatial arrangement that satisfies the one or more criteria, and detecting an indication of movement of the second user from a first individual viewpoint to a second individual viewpoint of the second viewpoint within the three-dimensional environment; While displaying the three-dimensional environment having the second viewpoint of the second user at the second individual viewpoint, receiving a second input corresponding to a request to update the spatial arrangement of the plurality of virtual objects with respect to the current viewpoint of the user so as to satisfy the one or more criteria via the one or more input devices; and further updating, in response to receiving the second input, the spatial arrangement of the plurality of virtual objects according to the second individual viewpoint of the second user to a second individual spatial arrangement that satisfies the one or more criteria, the method according to claim 1.

12. While displaying the three-dimensional environment, which is a first individual spatial arrangement in which the spatial arrangement of the plurality of virtual objects with respect to the current viewpoint of the user satisfies the one or more criteria, via the display generation component, receiving, via the one or more input devices, one or more sequences of inputs corresponding to a request to update one or more positions of the plurality of virtual objects within the three-dimensional environment In response to receiving the one or more sequences of inputs, via the display generation component, displaying the plurality of virtual objects at respective positions within the three-dimensional environment according to the one or more sequences of inputs in a second individual spatial arrangement that does not satisfy the one or more criteria While displaying the three-dimensional environment including displaying the plurality of virtual objects at the respective positions within the three-dimensional environment, receiving, via the one or more input devices, a second input corresponding to a request to update the spatial arrangement of the plurality of virtual objects with respect to the current viewpoint of the user so as to satisfy the one or more criteria The method according to claim 1, further comprising, in response to receiving the second input, updating the respective positions of the plurality of virtual objects to a third individual spatial arrangement that satisfies the one or more criteria according to the respective positions of the plurality of virtual objects within the three-dimensional environment

13. While displaying the three-dimensional environment, which is a first individual spatial arrangement in which the spatial arrangement of the plurality of virtual objects with respect to the current viewpoint of the user satisfies the one or more criteria, via the display generation component, detecting one or more indications of a request by a second user in the three-dimensional environment to update one or more positions of the plurality of virtual objects within the three-dimensional environment In response to detecting the one or more indications, via the display generation component, displaying the plurality of virtual objects at respective positions within the three-dimensional environment according to the one or more indications in a second individual spatial arrangement that does not satisfy the one or more criteria While displaying the three-dimensional environment including displaying the plurality of virtual objects at the respective positions in the three-dimensional environment, via the one or more input devices, receiving a second input corresponding to a request to update a spatial arrangement of the plurality of virtual objects with respect to a current viewpoint of the user so as to satisfy the one or more criteria. In response to receiving the second input, updating the respective positions of the plurality of virtual objects to a third individual spatial arrangement that satisfies the one or more criteria according to the respective positions of the plurality of virtual objects in the three-dimensional environment. The method according to claim 1, further comprising.

14. While the spatial arrangement of the plurality of virtual objects does not satisfy the one or more criteria. According to a determination that the input corresponding to the request to update the spatial arrangement of the plurality of virtual objects has not been received, maintaining the spatial arrangement of the plurality of virtual objects until the input corresponding to the request to update the spatial arrangement of the plurality of virtual objects is received. The method according to claim 1, further comprising.

15. Displaying the plurality of virtual objects in the three-dimensional environment, including displaying a first virtual object among the plurality of virtual objects at a location beyond a predetermined threshold distance from the current viewpoint of the user. While the plurality of virtual objects are being displayed in the three-dimensional environment, including displaying the first virtual object among the plurality of virtual objects at the location beyond the predetermined threshold distance from the current viewpoint of the user, via the one or more input devices, receiving a second input corresponding to a request to update a spatial arrangement of the plurality of virtual objects with respect to a current viewpoint of the user so as to satisfy the one or more criteria. In response to receiving the second input, updating the viewpoint of the user to an individual viewpoint within the predetermined threshold distance of the first virtual object, wherein the spatial arrangement of the plurality of virtual objects with respect to the individual viewpoint satisfies the one or more criteria. The method according to claim 1, further comprising.

16. Displaying, via the display generation component, a plurality of virtual objects having a first interval between a first virtual object among the plurality of virtual objects and a second virtual object among the plurality of virtual objects, the first interval not satisfying one or more interval criteria among the one or more criteria. While displaying the plurality of virtual objects having the first interval between the first virtual object and the second virtual object, receiving, via the one or more input devices, a second input corresponding to a request to update a spatial arrangement of the plurality of virtual objects with respect to a current viewpoint of the user so as to satisfy the one or more criteria. In response to receiving the second input, displaying, via the display generation component, a plurality of virtual objects having a second interval between the first virtual object and the second virtual object, the second interval satisfying the one or more interval criteria. The method according to claim 1 further includes this step.

17. Detecting the movement of the current viewpoint of the user in the three-dimensional environment from the first viewpoint to the second viewpoint includes detecting the movement of the electronic device within a physical environment of the electronic device or the movement of the display generation component within a physical environment of the display generation component via the one or more input devices. The method according to claim 1 includes this step.

18. An electronic device, One or more processors, A memory, One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for executing the method according to any one of claims 1 to 17. An electronic device comprising these components.

19. One or more programs including instructions, when the instructions are executed by one or more processors of an electronic device, cause the electronic device to execute the method according to any one of claims 1 to 17.

Citation Information

Patent Citations

  • Method and device for controlling virtual object to perform shortcut operation, equipment and medium

    CN110413171A

  • A system for rendering a shared digital interface from each user's perspective.

    JP2014514653A

  • Systems, Methods, and Graphical User Interfaces for Displaying and Manipulating Virtual Objects in Augmented Reality Environments

    US20210295602A1

  • Information processing device, information processing method, and program

    WO2020066682A1

  • Head-mounted information processing device and head-mounted display system

    WO2020179027A1