Interaction within hybrid spatial group in multi-user communication session

The system addresses inconsistent interactions in multi-user communication by moving shared objects based on co-location, ensuring spatial truth for co-located users and adapting for non-co-located participants, thereby enhancing user experience.

JP2025114484APending Publication Date: 2025-08-05APPLE INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024225104
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-12-12
Filing Date
2024-12-20
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

Existing multi-user communication systems struggle to effectively manage interactions in three-dimensional environments based on whether participants are co-located or non-co-located, leading to inconsistent user experiences.

Method used

A system and method that allows a first electronic device to move shared objects within a three-dimensional environment in response to user input, distinguishing between co-located and non-co-located participants by adjusting object movement relative to the viewer's perspective without updating the visual representation of non-co-located users.

Benefits of technology

Enhances interaction consistency and user experience by maintaining spatial truth and authenticity for co-located users while accommodating non-co-located participants in multi-user communication sessions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025114484000001_ABST
    Figure 2025114484000001_ABST
Patent Text Reader

Abstract

To provide systems and methods for facilitating interactions, including movement of a content that is shared in a multi-user communication session on the basis of whether participants in the multi-user communication session are collocated or non-collocated.SOLUTION: In some examples, a first electronic device presents a three-dimensional environment including a first object of a first type and a visual representation of a user of a second electronic device. In some examples the first electronic device receives a request to move the first object. In some examples, in response to receiving the request, in accordance with a determination that the second electronic device is collocated with the first electronic device in a first physical environment, the first electronic device moves the first object of the first type in the three-dimensional environment according to the first input without updating presentation of the visual representation of the user of the second electronic device.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Provisional Patent Application No. 63 / 614,489, filed December 22, 2023, U.S. Provisional Patent Application No. 63 / 614,489, filed September 24, 2024, and U.S. Patent Application No. 18 / --,---, filed December 2024, the contents of which are incorporated by reference herein in their entirety for all purposes.

[0002] (Technical field) The present invention generally relates to systems and methods for establishing a multi-user communication session in which at least a subset of the participants in the multi-user communication session are co-located in a physical environment. [Background technology]

[0003] Some computer graphical environments provide computer-generated two-dimensional and / or three-dimensional environments in which at least some objects displayed for user viewing are virtual. In some examples, the three-dimensional environment is presented by multiple devices communicating in a multi-user communication session. In some examples, an avatar (e.g., a representation) of each non-co-located user participating in the multi-user communication session (e.g., via a computing device) is displayed within the three-dimensional environment of the multi-user communication session. In some examples, content may be shared within the three-dimensional environment for viewing and interaction by multiple users participating in the multi-user communication session. Summary of the Invention

[0004] Some examples of the present disclosure are directed to systems and methods for facilitating interaction, including movement, of content shared in a multi-user communication session based on whether participants in the multi-user communication session are co-located or non-co-located. In some examples, the method is executed on a first electronic device in communication with one or more displays, one or more input devices, and a second electronic device, the first electronic device being in a communication session with the second electronic device. In some examples, the first electronic device presents, via the one or more displays, a three-dimensional environment including a visual representation of a first object of a first type (e.g., a shared virtual object) and a user of the second electronic device. In some examples, while presenting the three-dimensional environment including the first object of the first type and a visual representation of the user of the second electronic device, the first electronic device receives, via the one or more input devices, a first input corresponding to a request to move the first object within the three-dimensional environment. In some examples, in response to receiving the first input, in accordance with a determination that one or more criteria are met, including criteria that are met when the second electronic device is co-located with the first electronic device in the first physical environment, the first electronic device moves a first object of a first type within the three-dimensional environment relative to a point of view of the first electronic device in accordance with the first input without updating a presentation of a visual representation of a user of the second electronic device. In some examples, in accordance with a determination that the one or more criteria are not met because the second electronic device is not co-located with the first electronic device in the first physical environment, the first electronic device moves a first object of the first type and a visual representation of a user of the second electronic device within the three-dimensional environment relative to a point of view of the first electronic device in accordance with the first input.

[0005] A full description of these examples is set forth in the Drawings and Detailed Description, and it should be understood that this Summary is not intended to limit the scope of the present disclosure in any way.

[0006] For a better understanding of the various examples described herein, reference should be made to the following detailed description in conjunction with the following drawings, in which like reference numerals refer to corresponding parts throughout: [Brief explanation of the drawings]

[0007] [Figure 1] 1 illustrates an electronic device presenting an augmented reality environment, according to some examples of the present disclosure.

[0008] [Figure 2] 1 shows a block diagram of an example architecture for a system according to some examples of the present disclosure.

[0009] [Figure 3] 1 illustrates an example of a spatial group in a multi-user communication session including a first electronic device and a second electronic device, according to some examples of the present disclosure.

[0010] [Figure 4A] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 4B] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 4C] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 4D] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 4E] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 4F]1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 4G] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 4H] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 4I] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 4J] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 4K] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 4L] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 4M] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 4N] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 4O] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 4P]1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 4Q] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 4R] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 4S] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 4T] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 4U] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 4V] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 4W] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 4X] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 4Y] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 4Z]1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 4AA] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure.

[0011] [Figure 5A] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 5B] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 5C] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 5D] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. [Figure 5E] 1 illustrates exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure.

[0012] [Figure 6] FIG. 1 shows a flow diagram illustrating an exemplary process for moving objects within a three-dimensional environment within a multi-user communication session based on whether the multi-user communication session includes collocated or non-collocated users, in accordance with some examples of the present disclosure.

[0013] [Figure 7A] 1 illustrates exemplary interactions within a multi-user communication session involving co-located users, according to some examples of the present disclosure. [Figure 7B] 1 illustrates exemplary interactions within a multi-user communication session involving co-located users, according to some examples of the present disclosure. [Figure 7C] 1 illustrates exemplary interactions within a multi-user communication session involving co-located users, according to some examples of the present disclosure. [Figure 7D] 1 illustrates exemplary interactions within a multi-user communication session involving co-located users, according to some examples of the present disclosure. [Figure 7E] 1 illustrates exemplary interactions within a multi-user communication session involving co-located users, according to some examples of the present disclosure. [Figure 7F] 1 illustrates exemplary interactions within a multi-user communication session involving co-located users, according to some examples of the present disclosure. [Figure 7G] 1 illustrates exemplary interactions within a multi-user communication session involving co-located users, according to some examples of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0014] Some examples of the present disclosure are directed to systems and methods for facilitating interaction, including movement, of content shared in a multi-user communication session based on whether participants in the multi-user communication session are co-located or non-co-located. In some examples, the method is executed on a first electronic device in communication with one or more displays, one or more input devices, and a second electronic device, the first electronic device being in a communication session with the second electronic device. In some examples, the first electronic device presents, via the one or more displays, a three-dimensional environment including a visual representation of a first object of a first type (e.g., a shared virtual object) and a user of the second electronic device. In some examples, while presenting the three-dimensional environment including the first object of the first type and a visual representation of the user of the second electronic device, the first electronic device receives, via the one or more input devices, a first input corresponding to a request to move the first object within the three-dimensional environment. In some examples, in response to receiving the first input, in accordance with a determination that one or more criteria are met, including criteria that are met when the second electronic device is co-located with the first electronic device in the first physical environment, the first electronic device moves a first object of a first type within the three-dimensional environment relative to a point of view of the first electronic device in accordance with the first input without updating a presentation of a visual representation of a user of the second electronic device. In some examples, in accordance with a determination that the one or more criteria are not met because the second electronic device is not co-located with the first electronic device in the first physical environment, the first electronic device moves a first object of the first type and a visual representation of a user of the second electronic device within the three-dimensional environment relative to a point of view of the first electronic device in accordance with the first input.

[0015] As used herein, a spatial group corresponds to a group or number of participants (e.g., users) in a multi-user communication session. In some examples, a spatial group within a multi-user communication session has a spatial arrangement that determines the location of users and content located within the spatial group. In some examples, users within the same spatial group within a multi-user communication session experience spatial truth according to the spatial arrangement of the spatial group. In some examples, when a user of a first electronic device is in a first spatial group and a user of a second electronic device is in a second spatial group in a multi-user communication session, the users experience spatial truth localized to their respective spatial groups. In some examples, a user of a first electronic device and a user of a second electronic device are grouped into separate spatial groups within a multi-user communication session, but when the first electronic device and the second electronic device return to the same operating state, the user of the first electronic device and the user of the second electronic device are regrouped into the same spatial group within the multi-user communication session.

[0016] As used herein, a hybrid spatial group corresponds to a group or number of participants (e.g., users) in a multi-user communication session in which at least a subset of the participants are not co-located in a physical environment. For example, as described through one or more examples in this disclosure, a hybrid spatial group includes at least two participants co-located in a first physical environment and at least one participant not co-located with the at least two participants in the first physical environment (e.g., at least one participant is located in a second physical environment different from the first physical environment). In some examples, a hybrid spatial group within a multi-user communication session has a spatial arrangement that determines the location of users and content located within the spatial group. In some examples, users in the same hybrid spatial group within a multi-user communication session experience spatial authenticity according to the spatial arrangement of the spatial group, as similarly described above.

[0017] In some examples, initiating a multi-user communication session may include interaction with one or more user interface elements. In some examples, a user's gaze may be tracked by the electronic device as input for targeting selectable options / affordances within individual user interface elements displayed within the three-dimensional environment. For example, the gaze may be used to identify one or more options / affordances for selection using another selection input. In some examples, individual options / affordances may be selected using hand-tracking input detected via an input device in communication with the electronic device. In some examples, objects displayed within the three-dimensional environment may be moved and / or reoriented within the three-dimensional environment according to movement input detected via the input device.

[0018] FIG. 1 illustrates an electronic device 101 presenting an extended reality (XR) environment (e.g., a computer-generated environment that optionally includes representations of physical and / or virtual objects) according to some examples of the present disclosure. In some examples, as shown in FIG. 1, the electronic device 101 is a head-mounted display or other head-mountable device configured to be worn on the head of a user of the electronic device 101. Examples of the electronic device 101 are described below with reference to the architectural block diagram of FIG. 2. As shown in FIG. 1, the electronic device 101 and a table 106 are located in a physical environment. The physical environment may include physical features such as physical surfaces (e.g., floors, walls) or physical objects (e.g., tables, lamps, etc.). In some examples, the electronic device 101 may be configured to detect and / or capture images of the physical environment, including the table 106 (shown within the field of view of the electronic device 101).

[0019] 1, electronic device 101 includes one or more internal image sensors 114a (e.g., eye-tracking cameras described below with reference to FIG. 2) oriented toward the user's face. In some examples, internal image sensor 114a is used for eye tracking (e.g., detecting the user's gaze). Internal image sensor 114a is optionally positioned on left and right portions of display 120 to enable eye tracking of the user's left and right eyes. In some examples, electronic device 101 also includes external image sensors 114b and 114c facing outward from the user to detect and / or capture the physical environment of electronic device 101 and / or movements of the user's hands or other body parts.

[0020] In some examples, display 120 has a field of view visible to the user (e.g., which may or may not correspond to the field of view of external image sensors 114b and 114c). Because display 120 is optionally part of a head-mounted device, the field of view of display 120 is optionally the same as or similar to the field of view of the user's eyes. In other examples, the field of view of display 120 may be smaller than the field of view of the user's eyes. In some examples, electronic device 101 may be an optical see-through device in which display 120 is a transparent or translucent display through which a portion of the physical environment can be viewed directly. In some examples, display 120 may be contained within a transparent lens or may overlap all or only a portion of a transparent lens. In other examples, electronic device 101 may be a video pass-through device in which display 120 is an opaque display configured to display images of the physical environment captured by external image sensors 114b and 114c. While a single display 120 is shown, it should be understood that display 120 may include a stereo pair of displays.

[0021] 1 , which is not present in the physical environment but is displayed in the XR environment positioned on top of a real-world table 106 (or a representation thereof). Optionally, the virtual object 104 may be displayed on a surface of the table 106 in the XR environment displayed via the display 120 of the electronic device 101 in response to detecting the plane of the table 106 in the physical environment 100.

[0022] It should be understood that virtual object 104 is exemplary of a virtual object, and that one or more different virtual objects (e.g., of various dimensionalities, such as two-dimensional or other three-dimensional virtual objects) can be included and rendered within the three-dimensional XR environment. For example, the virtual object may represent an application or a user interface displayed within the XR environment. In some examples, the virtual object may represent content corresponding to an application and / or displayed via a user interface in the XR environment. In some examples, virtual object 104 is optionally interactive and configured to respond to user input (e.g., air gestures such as an air pinch gesture, an air tap gesture, and / or an air touch gesture) such that a user may virtually touch, tap, move, rotate, or otherwise interact with virtual object 104.

[0023] In some examples, displaying an object within the three-dimensional environment may include interaction with one or more user interface objects within the three-dimensional environment. For example, initiating the display of an object within the three-dimensional environment may include interaction with one or more virtual options / affordances displayed within the three-dimensional environment. In some examples, a user's gaze may be tracked by the electronic device as input to identify one or more virtual options / affordances for selection when initiating the display of the object within the three-dimensional environment. For example, the gaze may be used to identify one or more virtual options / affordances for selection using another selection input. In some examples, the virtual options / affordances may be selected using hand-tracking input detected via an input device in communication with the electronic device. In some examples, an object displayed within the three-dimensional environment may be moved and / or reoriented within the three-dimensional environment according to movement input detected via the input device.

[0024] In the following description, an electronic device is described that communicates with display generating components and one or more input devices. It should be understood that the electronic device optionally communicates with one or more other physical user interface devices, such as a touch-sensitive surface, a physical keyboard, a mouse, a joystick, a hand-tracking device, an eye-tracking device, a stylus, etc. It should also be understood that, as noted above, the described electronic device, display, and touch-sensitive surface are optionally distributed across two or more devices. Thus, as used in this disclosure, information displayed on or by an electronic device may optionally be used to describe information output by the electronic device for display on another display device (touch-sensitive or non-touch-sensitive). Similarly, as used in this disclosure, input received at an electronic device (e.g., touch input received on a touch-sensitive surface of the electronic device or touch input received on the surface of a stylus) is optionally used to describe input received at a separate input device from which the electronic device receives input information.

[0025] The device typically supports a variety of applications, such as one or more of a drawing application, a presentation application, a word processing application, a website creation application, a disc authoring application, a spreadsheet application, a gaming application, a telephony application, a video conferencing application, an email application, an instant messaging application, a training support application, a photo management application, a digital camera application, a digital video camera application, a web browsing application, a digital music playback application, a television channel browsing application, and / or a digital video playback application.

[0026] 2 shows a block diagram of an example architecture of a system 201 according to some examples of the present disclosure. In some examples, the system 201 includes multiple devices. For example, the system 201 includes a first electronic device 260 and a second electronic device 270, where the first electronic device 260 and the second electronic device 270 communicate with each other. In some examples, the first electronic device 260 and the second electronic device 270 are portable devices, such as a mobile phone, a smartphone, a tablet computer, a laptop computer, an auxiliary device that communicates with another device, a head-mounted display, etc. In some examples, the first electronic device 260 and the second electronic device 270 correspond to the electronic device 101 described above with reference to FIG. 1.

[0027] As shown in FIG. 2 , first electronic device 260 optionally includes various sensors (e.g., one or more hand tracking sensors 202A, one or more location sensors 204A, one or more image sensors 206A, one or more touch-sensitive surface 209A, one or more movement and / or orientation sensors 210A, one or more eye tracking sensors 212A, one or more microphones 213A or other audio sensors, one or more body tracking sensors (e.g., torso and / or head tracking sensors), one or more display generating components 214A, one or more speakers 216A, one or more processors 218A, one or more memories 220A, and / or communications circuitry 222A. In some examples, second electronic device 270 includes various sensors (e.g., one or more hand tracking sensors 202B, one or more location sensors 204B, one or more image sensors 206B, one or more touch-sensitive surface 209A, one or more movement and / or orientation sensors 210A, one or more eye tracking sensors 212A, one or more microphones 213A or other audio sensors, one or more body tracking sensors (e.g., torso and / or head tracking sensors), one or more display generating components 214A, one or more speakers 216A, one or more processors 218A, one or more memories 220A, and / or communications circuitry 222A). In some examples, second electronic device 270 includes various sensors (e.g., one or more hand tracking sensors 202B, one or more location sensors 204B, one or more image sensors 206B, one or more touch-sensitive surface 209A, one or more movement and / or orientation sensors 210A, one or more eye tracking sensors 212A, one or more microphones 213A or other audio sensors, one or more body tracking sensors (e.g., torso and / or head tracking sensors), one or 2. The electronic device 260 and the second electronic device 270 may optionally include a surface 209B, one or more movement and / or orientation sensors 210B, one or more eye tracking sensors 212B, one or more microphones 213B or other audio sensors, one or more body tracking sensors (e.g., torso and / or head tracking sensors), one or more display generating components 214B, one or more speakers 216, one or more processors 218B, one or more memories 220B, and / or communications circuitry 222B. In some examples, the one or more display generating components 214A, 214B correspond to the display 120 of FIG. 1. One or more communications buses 208A and 208B are optionally used for communications between the above-mentioned components of the electronic devices 260 and 270, respectively. The first electronic device 260 and the second electronic device 270 optionally communicate via a wired or wireless connection between the two devices (e.g., via communications circuitry 222A, 222B).

[0028] Communications circuitry 222A, 222B optionally includes circuitry for communicating with electronic devices, networks such as the Internet, an intranet, wired and / or wireless networks, cellular networks, and wireless local area networks (LANs). Communications circuitry 222A, 222B optionally includes circuitry for communicating using short-range communications such as near-field communications (NFC) and / or Bluetooth.

[0029] The processor(s) 218A, 218B include one or more general-purpose processors, one or more graphics processors, and / or one or more digital signal processors. In some examples, the memory 220A, 220B is a non-transitory computer-readable storage medium (e.g., flash memory, random access memory, or other volatile or non-volatile memory or storage) that stores computer-readable instructions configured to be executed by the processor(s) 218A, 218B to perform the techniques, processes, and / or methods described below. In some examples, the memory 220A, 220B can include two or more non-transitory computer-readable storage media. A non-transitory computer-readable storage medium can be any medium (other than a signal) that can tangibly store or carry computer-executable instructions used by or in connection with an instruction execution system, apparatus, or device. In some examples, the storage medium is a transient computer-readable storage medium. In some examples, the storage medium is a non-transitory computer-readable storage medium. Non-transitory computer-readable storage media may include, but are not limited to, magnetic, optical, and / or semiconductor storage devices. Examples of such storage include magnetic disks, compact discs (CDs), digital versatile discs (DVDs), or optical discs based on Blu-ray technology, as well as persistent solid-state memory such as flash and solid-state drives.

[0030] In some examples, display generating component(s) 214A, 214B include a single display (e.g., a liquid crystal display (LCD), organic light emitting diode (OLED), or other type of display). In some examples, display generating component(s) 214A, 214B include multiple displays. In some examples, display generating component(s) 214A, 214B can include a display with touch capabilities (e.g., a touchscreen), a projector, a holographic projector, a retina projector, a transparent or translucent display, etc. In some examples, electronic devices 260 and 270 include touch-sensitive surface(s) 209A and 209B, respectively, for receiving user inputs such as tap and swipe inputs or other gestures. In some examples, display generating component(s) 214A, 214B and touch-sensitive surface(s) 209A, 209B form touch-sensitive display(s) (e.g., touchscreens integrated with or external to electronic devices 260 and 270, respectively, in communication with electronic devices 260 and 270, respectively).

[0031] Electronic devices 260 and 270 optionally include image sensor(s) 206A and 206B, respectively. Image sensor(s) 206A / 206B optionally include one or more visible light image sensors, such as charge-coupled device (CCD) sensors and / or complementary metal-oxide semiconductor (CMOS) sensors, operable to acquire images of physical objects from the real-world environment. Image sensor(s) 206A / 206B also optionally include one or more infrared (IR) sensors, such as passive or active IR sensors, for detecting infrared light from the real-world environment. For example, an active IR sensor includes an IR emitter for emitting infrared light into the real-world environment. Image sensor(s) 206A / 206B also optionally include one or more cameras configured to capture movement of physical objects in the real-world environment. Image sensor(s) 206A / 206B also optionally include one or more depth sensors configured to detect the distance of a physical object from electronic device 260 / 270. In some examples, information from the one or more depth sensors can enable the device to identify an object in the real-world environment and distinguish it from other objects in the real-world environment. In some examples, the one or more depth sensors can enable the device to determine the texture and / or topography of an object in the real-world environment.

[0032] In some examples, the electronic devices 260 and 270 use a combination of a CCD sensor, an event camera, and a depth sensor to detect the physical environment around the electronic devices 260 and 270. In some examples, the image sensor(s) 206A / 206B include a first image sensor and a second image sensor. The first image sensor and the second image sensor cooperate and are optionally configured to capture different information of physical objects in the real-world environment. In some examples, the first image sensor is a visible light image sensor and the second image sensor is a depth sensor. In some examples, the electronic devices 260 / 270 use the image sensor(s) 206A / 206B to detect the position and orientation of the electronic devices 260 / 270 and / or the display generating component(s) 214A / 214B within the real-world environment. For example, the electronic device 260 / 270 uses the image sensor(s) 206A / 206B to track the position and orientation of the display generating component(s) 214A / 214B relative to one or more fixed objects in the real-world environment.

[0033] In some examples, electronic device 260 / 270 includes microphone(s) 213A / 213B or other audio sensors. Device 260 / 270 uses microphone(s) 213A / 213B to detect sounds from the user and / or the user's real-world environment. In some examples, microphone(s) 213A / 213B include an array of microphones (multiple microphones) optionally working in concert, such as to identify ambient noise or to localize a sound source within a space in the real-world environment.

[0034] In some examples, the device 260 / 270 includes location sensor(s) 204A / 204B and / or display generation component(s) 214A / 214B for detecting the location of the device 260 / 270. For example, the location sensor(s) 204A / 204B may include a Global Positioning System (GPS) receiver that receives data from one or more satellites, allowing the electronic device 260 / 270 to determine the device's absolute position in the physical world.

[0035] In some examples, the electronic device 260 / 270 includes orientation sensor(s) 210A / 210B for detecting the orientation and / or movement of the electronic device 260 / 270 and / or the display generating component(s) 214A / 214B. For example, the electronic device 260 / 270 uses the orientation sensor(s) 210A / 210B to track changes in the position and / or orientation of the electronic device 260 / 270 and / or the display generating component(s) 214A / 214B relative to physical objects in a real-world environment, etc. The orientation sensor(s) 210A / 210B optionally include one or more gyroscopes and / or one or more accelerometers.

[0036] Electronic device 260 / 270, in some examples, includes hand tracking sensor(s) 202A / 202B and / or eye tracking sensor(s) 212A / 212B (and / or other body tracking sensor(s), such as leg, torso, and / or head tracking sensor(s)). Hand tracking sensor(s) 202A / 202B are configured to track the position / location of one or more parts of a user's hand and / or the movement of one or more parts of a user's hand relative to the extended reality environment, relative to display generating component(s) 214A / 214B, and / or relative to another defined coordinate system. The eye tracking sensor(s) 212A / 212B are configured to track the position and movement of a user's gaze (more generally, eyes, face, or head) relative to the real world or extended reality environment and / or relative to the display generation component(s) 214A / 214B. In some examples, the hand tracking sensor(s) 202A / 202B and / or the eye tracking sensor(s) 212A / 212B are implemented together with the display generation component(s) 214A / 214B. In some examples, the hand tracking sensor(s) 202A / 202B and / or the eye tracking sensor(s) 212A / 212B are implemented separately from the display generation component(s) 214A / 214B.

[0037] In some examples, hand tracking sensor(s) 202A / 202B (and / or other body tracking sensor(s), such as leg, torso, and / or head tracking sensor(s)) can use image sensor(s) 206A / 206B (e.g., one or more IR cameras, 3D cameras, depth cameras, etc.) to capture three-dimensional information from the real world, including one or more body parts (e.g., the hands, legs, or torso of a human user). In some examples, the hand can be resolved with sufficient resolution to distinguish fingers and their respective positions. In some examples, the one or more image sensor(s) 206A / 206B are positioned relative to the user to define the field of view of the image sensor(s) 206A / 206B and an interaction space in which the position, orientation, and / or movement of the fingers / hands captured by the image sensor(s) are used as input (e.g., to distinguish from the user's stationary hand or other hands of other people in the real-world environment). Tracking fingers / hands for input (e.g., gestures, touches, taps, etc.) can be advantageous in that it does not require the user to touch, hold, or wear any kind of beacon, sensor, or other marker.

[0038] In some examples, the eye tracking sensor(s) 212A / 212B include at least one eye tracking camera (e.g., an infrared (IR) camera) and / or an illumination source (e.g., an IR light source such as an LED) that emits light toward the user's eyes. The eye tracking camera may be aimed at the user's eyes to receive reflected IR light from the light source directly or indirectly from the eyes. In some examples, both eyes are tracked separately by respective eye tracking cameras and illumination sources, and focus / gaze can be determined from tracking of both eyes. In some examples, one eye (e.g., a dominant eye) is tracked by one or more respective eye tracking cameras / illumination sources.

[0039] Electronic devices 260 / 270 and system 201 are not limited to the components and configuration of FIG. 2 and may include fewer, other, or additional components in multiple configurations. In some examples, system 201 may be implemented in a single device. One or more people using system 201 are optionally referred to herein as one or more users of the device(s). Attention now turns to exemplary simultaneous displays of three-dimensional environments on a first electronic device (e.g., corresponding to electronic device 260) and a second electronic device (e.g., corresponding to electronic device 270). As described below, the first electronic device can communicate with a second electronic device in a multi-user communication session. In some examples, an avatar (e.g., a representation) of a user of the first electronic device may be displayed within the three-dimensional environment at the second electronic device, and an avatar of a user of the second electronic device may be displayed within the three-dimensional environment at the first electronic device. In some examples, a user of the first electronic device and a user of the second electronic device may be associated with a spatial group in the multi-user communication session. In some examples, interaction with content within the three-dimensional environment while the first electronic device and the second electronic device are within a multi-user communication session can associate the user of the first electronic device and the user of the second electronic device with different spatial groups within the multi-user communication session.

[0040] 3 illustrates an example of a spatial group 340 in a multi-user communication session including a first electronic device 360 and a second electronic device 370, according to some examples of the present disclosure. In some examples, the first electronic device 360 may present a three-dimensional environment 350A, and the second electronic device 370 may present a three-dimensional environment 350B. The first electronic device 360 and the second electronic device 370 may be similar to the electronic devices 101 or 260 / 270 and / or may be head-mountable and / or projection-based systems / devices (including hologram-based systems / devices) configured to generate and present three-dimensional environments, such as, for example, a head-up display (HUD), a head-mounted display (HMD), a window with integrated display capabilities, a display formed as a lens designed to be placed over a person's eye (e.g., similar to a contact lens), etc. In the example of FIG. 3, a first user is optionally wearing a first electronic device 360 and a second user is optionally wearing a second electronic device 370, so that the three-dimensional environment 350A / 350B can be defined by X, Y, and Z axes as seen from the perspective of the electronic devices (e.g., a viewpoint associated with electronic device 360 / 370, which may be a head-mounted display).

[0041] 3 , first electronic device 360 may be in a first physical environment that includes table 306 and window 309. Thus, three-dimensional environment 350A presented using first electronic device 360 optionally includes captured portions of the physical environment surrounding first electronic device 360, such as a representation of table 306′ and a representation of window 309′. Similarly, second electronic device 370 may be in a second physical environment that is different from (e.g., separate from) the first physical environment, including floor lamp 307 and coffee table 308. Thus, three-dimensional environment 350B presented using second electronic device 370 optionally includes captured portions of the physical environment surrounding second electronic device 370, such as a representation of floor lamp 307′ and a representation of coffee table 308′. Additionally, the three-dimensional environments 350A and 350B may include representations of the floor, ceiling, and walls of the rooms in which the first electronic device 360 and the second electronic device 370 are located, respectively.

[0042] As mentioned above, in some examples, the first electronic device 360 is optionally in a multi-user communication session with the second electronic device 370. For example, the first electronic device 360 and the second electronic device 370 are configured (e.g., via the communication circuits 222A / 222B) to present a shared three-dimensional environment 350A / 350B that includes one or more shared virtual objects (e.g., content such as images, video, audio, a representation of an application's user interface, etc.). As used herein, the term "shared three-dimensional environment" refers to a three-dimensional environment independently presented, displayed, and / or visible on two or more electronic devices, in which content, applications, data, etc. may be shared and / or presented to users of the two or more electronic devices. In some examples, while the first electronic device 360 is in a multi-user communication session with the second electronic device 370, an avatar corresponding to a user of one electronic device is optionally displayed within the three-dimensional environment displayed via the other electronic device. 3, at first electronic device 360, avatar 315 corresponding to the user of second electronic device 370 is displayed within three-dimensional environment 350A. Similarly, at second electronic device 370, avatar 317 corresponding to the user of first electronic device 360 is displayed within three-dimensional environment 350B.

[0043] In some examples, the presentation of avatar 315 / 317 as part of the shared three-dimensional environment is optionally accompanied by audio effects corresponding to the voice of the user of electronic device 370 / 360. For example, avatar 315 displayed within three-dimensional environment 350A using first electronic device 360 is optionally accompanied by audio effects corresponding to the voice of the user of second electronic device 370. In some such examples, when the user of second electronic device 370 speaks, the user's voice may be detected by second electronic device 370 (e.g., via microphone(s) 213B) and transmitted to first electronic device 360 (e.g., via communications circuitry 222B / 222A), such that the detected voice of the user of second electronic device 370 may be presented as audio (e.g., using speaker(s) 216A) to the user of first electronic device 360 within three-dimensional environment 350A. In some examples, audio effects corresponding to the voice of the user of the second electronic device 370 may be spatialized to appear to the user of the first electronic device 360 as emanating from the location of the avatar 315 within the shared three-dimensional environment 350A (e.g., despite being output from speakers of the first electronic device 360). Similarly, avatar 317 displayed within the three-dimensional environment 350B using the second electronic device 370 is, optionally, accompanied by audio effects corresponding to the voice of the user of the first electronic device 360. In some such examples, when a user of the first electronic device 360 speaks, the user's voice may be detected by the first electronic device 360 (e.g., via microphone(s) 213A) and transmitted to the second electronic device 370 (e.g., via communication circuitry 222A / 222B), so that the detected voice of the user of the first electronic device 360 may be presented as audio (e.g., using speaker(s) 216B) to the user of the second electronic device 370 within the three-dimensional environment 350B.In some examples, audio effects corresponding to the voice of the user of the first electronic device 360 may be spatialized so that they appear to the user of the second electronic device 370 to emanate from the location of the avatar 317 within the shared three-dimensional environment 350B (e.g., despite being output from the speakers of the first electronic device 360).

[0044] In some examples, during a multi-user communication session, avatars 315 / 317 are displayed within three-dimensional environments 350A / 350B at respective orientations that correspond to and / or are based on the orientation of electronic devices 360 / 370 (and / or the users of electronic devices 360 / 370) within the physical environment surrounding electronic devices 360 / 370. For example, as shown in FIG. 3 , in three-dimensional environment 350A, avatar 315 optionally faces toward the viewpoint of a user of first electronic device 360, and in three-dimensional environment 350B, avatar 317 optionally faces toward the viewpoint of a user of second electronic device 370. As a particular user moves their electronic device (and / or themselves) within the physical environment, the user's viewpoint changes accordingly, and therefore the orientation of the user's avatar within the three-dimensional environment may also change. For example, referring to FIG. 3 , if the user of the first electronic device 360 looks leftward within the three-dimensional environment 350A such that the first electronic device 360 is rotated (e.g., by a corresponding amount) to the left (e.g., counterclockwise), the user of the second electronic device 370 will see the avatar 317 corresponding to the user of the first electronic device 360 rotate to the right (e.g., clockwise) relative to the perspective of the user of the second electronic device 370 as the first electronic device 360 moves.

[0045] Additionally, in some examples, during a multi-user communication session, the viewpoint of three-dimensional environment 350A / 350B and / or the location of the viewpoint of three-dimensional environment 350A / 350B optionally change in accordance with movement of electronic device 360 / 370 (e.g., by the user of electronic device 360 / 370). For example, if during a communication session, first electronic device 360 is moved closer toward table 306′ and / or representation of avatar 315 (e.g., because the user of first electronic device 360 has moved forward in the physical environment surrounding first electronic device 360), the viewpoint of three-dimensional environment 350A changes accordingly, such that representation of table 306′, representation of window 309′, and avatar 315 appear larger in the field of view. In some examples, each user may interact with three-dimensional environment 350A / 350B independently, such that changes in perspective of three-dimensional environment 350A and / or interactions with virtual objects in three-dimensional environment 350A by first electronic device 360, optionally, do not affect what is shown in three-dimensional environment 350B at second electronic device 370, and vice versa.

[0046] In some examples, the avatar 315 / 317 is a representation (e.g., a full-body rendering) of the user of the electronic device 370 / 360. In some examples, the avatar 315 / 317 is a representation of a portion of the user of the electronic device 370 / 360 (e.g., a rendering of the head, face, head and torso, etc.). In some examples, the avatar 315 / 317 is a user-personalized, user-selected, and / or user-created representation displayed within the three-dimensional environment 350A / 350B representing the user of the electronic device 370 / 360. While the avatars 315 / 317 shown in FIG. 3 correspond to full-body representations of the user of the electronic device 370 / 360, respectively, it should be understood that alternative avatars, such as those described above, may be provided.

[0047] As described above, while first electronic device 360 and second electronic device 370 are in a multi-user communication session, three-dimensional environment 350A / 350B may be a shared three-dimensional environment presented using electronic devices 360 / 370. In some examples, content viewed by one user at one electronic device may be shared with another user at another electronic device in the multi-user communication session. In some such examples, content may be experienced (e.g., viewed and / or interacted with) by both users (e.g., via their respective electronic devices) within the shared three-dimensional environment. For example, as shown in FIG. 3 , three-dimensional environment 350A / 350B includes a shared virtual object 310 (e.g., optionally a three-dimensional virtual sculpture) that is viewable by and interactive with both users. As shown in FIG. 3 , shared virtual object 310 may be displayed with a grabber affordance (e.g., handlebars) 335 that can be selected to initiate movement of shared virtual object 310 within three-dimensional environment 350A / 350B.

[0048] In some examples, three-dimensional environment 350A / 350B includes non-shared content that is private to one user in a multi-user communication session. For example, in FIG. 3 , first electronic device 360 displays a private application window 330 within three-dimensional environment 350A, which is, optionally, an object that is not shared between first electronic device 360 and second electronic device 370 in the multi-user communication session. In some examples, private application window 330 may be associated with an individual application (e.g., a media player application, a web browsing application, a messaging application, etc.) running on first electronic device 360. Because private application window 330 is not shared with second electronic device 370, second electronic device 370 optionally displays a representation of private application window 330″ within three-dimensional environment 350B. As shown in FIG. 3 , in some examples, the representation of the private application window 330″ may be a faded, occluded, discolored, and / or translucent representation of the private application window 330 that prevents a user of the second electronic device 370 from viewing the content of the private application window 330.

[0049] As described above, in some examples, the user of the first electronic device 360 and the user of the second electronic device 370 are in a spatial group 340 within a multi-user communication session. In some examples, the spatial group 340 may be a baseline (e.g., first or default) spatial group within the multi-user communication session. For example, when the user of the first electronic device 360 and the user of the second electronic device 370 initially join the multi-user communication session, the user of the first electronic device 360 and the user of the second electronic device 370 are automatically (and initially, as described in more detail below) associated (e.g., grouped) with the spatial group 340 within the multi-user communication session. In some examples, while the users are in the spatial group 340 as shown in FIG. 3 , the user of the first electronic device 360 and the user of the second electronic device 370 have a first spatial arrangement (e.g., a first spatial template) within the shared three-dimensional environment. For example, the user of the first electronic device 360 and the user of the second electronic device 370, including the objects displayed in the shared three-dimensional environment, have spatial authenticity within the spatial group 340. In some examples, spatial authenticity requires consistent spatial placement between the users (or their representations) and the virtual objects. For example, the distance between the viewpoint of the user of the first electronic device 360 and the avatar 315 corresponding to the user of the second electronic device 370 may be the same as the distance between the viewpoint of the user of the second electronic device 370 and the avatar 317 corresponding to the user of the first electronic device 360. As described herein, if the location of the viewpoint of the user of the first electronic device 360 moves, the avatar 317 corresponding to the user of the first electronic device 360 moves within the three-dimensional environment 350B according to the movement of the location of the user's viewpoint relative to the viewpoint of the user of the second electronic device 370.Additionally, when a user of first electronic device 360 performs an interaction on shared virtual object 310 (e.g., moves virtual object 310 within three-dimensional environment 350A), second electronic device 370 changes the display of shared virtual object 310 within three-dimensional environment 350B in accordance with the interaction (e.g., moves virtual object 310 within three-dimensional environment 350B).

[0050] It should be understood that in some examples, more than two electronic devices may be communicatively linked in a multi-user communication session. For example, in a situation where three electronic devices are communicatively linked in a multi-user communication session, a first electronic device will display two avatars, rather than just one avatar, corresponding to the users of the other two electronic devices. Thus, it should be understood that the various processes and example interactions described herein with reference to first electronic device 360 and second electronic device 370 in a multi-user communication session optionally apply to a situation where more than two electronic devices are communicatively linked in a multi-user communication session.

[0051] In some examples, it may be advantageous to provide a mechanism for moving virtual objects shared in a multi-user communication session involving collocated and non-collocated users (e.g., collocated and non-collocated electronic devices associated with the users). For example, it may be desirable to enable a user who is collocated in a first physical environment and participating in a multi-user communication session with one or more users who are not collocated in the first physical environment to collaboratively move and / or reposition virtual content shared and presented within a three-dimensional environment that is optionally viewable by and / or interactive with the collocated and non-collocated users in the multi-user communication session. As used herein, with respect to the first electronic device, a collocated user corresponds to a local user, and a non-collocated user corresponds to a remote user. Also as described above, the three-dimensional environment optionally includes avatars corresponding to remote users of electronic devices that are not collocated in the multi-user communication session. In some examples, as discussed below, rearrangement of virtual objects (e.g., avatars and / or shared virtual content) in a three-dimensional environment within a multi-user communication session is based on whether the multi-user communication session includes co-located users (e.g., with respect to the first electronic device), non-co-located users, or both.

[0052] 4A-4AA illustrate exemplary interactions within a multi-user communication session involving collocated and non-collocated users, according to some examples of the present disclosure.

[0053] 4A-4C illustrate exemplary interactions within a multi-user communication session involving co-located users. In some examples, while a first electronic device 101a is in a multi-user communication session with a second electronic device 101b, a three-dimensional environment 450A is presented using the first electronic device 101a (e.g., via display 120a) and a three-dimensional environment 450B is presented using the second electronic device 101b (e.g., via display 120b). In some examples, electronic devices 101a / 101b optionally correspond to or are similar to electronic devices 360 / 370 described above and / or electronic devices 260 / 270 of FIG. 2. In some examples, as shown in FIG. 4A, first electronic device 101a is being used (e.g., worn on the head) by a first user 402, and second electronic device 101b is being used (e.g., worn on the head) by a second user 404.

[0054] 4A , as shown in an overhead view 410, the first electronic device 101a and the second electronic device 101b are co-located in the physical environment 400. For example, the first electronic device 101a and the second electronic device 101b are both located in the same room, which includes a houseplant 408 and a window 409. In some examples, the determination that the first electronic device 101a and the second electronic device 101b are co-located in the physical environment 400 is based on the distance between the first electronic device 101a and the second electronic device 101b. For example, in FIG. 4A , the first electronic device 101a and the second electronic device 101b are co-located in the physical environment 400 because the first electronic device 101a is within a threshold distance (e.g., 0.1, 0.5, 1, 2, 3, 5, 10, 15, 20 meters, etc.) of the second electronic device 101b. In some examples, the determination that the first electronic device 101a and the second electronic device 101b are co-located in the physical environment 400 is based on communication between the first electronic device 101a and the second electronic device 101b. For example, in FIG. 4A , the first electronic device 101a and the second electronic device 101b are configured to communicate (e.g., wirelessly via Bluetooth, Wi-Fi, a server (e.g., a wireless communication terminal), etc.). In some examples, the first electronic device 101a and the second electronic device 101b are connected to the same wireless network in the physical environment 400. In some examples, the determination that the first electronic device 101a and the second electronic device 101b are co-located in the physical environment 400 is based on the strength of wireless signals transmitted between the electronic devices 101a and 101b. 4A , the first electronic device 101a and the second electronic device 101b are co-located in the physical environment 400 because the strength of the Bluetooth signal (or other wireless signal) transmitted between the electronic devices 101a and 101b is greater than a threshold strength. In some examples, the determination that the first electronic device 101a and the second electronic device 101b are co-located in the physical environment 400 is based on visual detection of the electronic devices 101a and 101b within the physical environment 400.For example, as shown in FIG. 4A, the second electronic device 101b is positioned within the field of view of the first electronic device 101a (e.g., because the second user 404 is standing within the field of view of the first electronic device 101a), which enables the first electronic device 101a to visually detect (e.g., identify or scan, such as via object detection or other image processing techniques) the second electronic device 101b (e.g., within one or more images captured by the first electronic device 101a, such as via external image sensors 114b-i and 114c-i). Similarly, as shown in FIG. 4A, the first electronic device 101a is optionally positioned within the field of view of the second electronic device 101b (e.g., because the first user 402 is standing within the field of view of the second electronic device 101b), thereby enabling the second electronic device 101b to visually detect the first electronic device 101a (e.g., in one or more images captured by the second electronic device 101b, such as via external image sensors 11b-ii and 114c-ii).

[0055] In some examples, the three-dimensional environment 450A / 450B includes a captured portion of the physical environment 400 in which the electronic device 460 / 470 is located. For example, because the first electronic device 101a and the second electronic device 101b are co-located in the physical environment 400, the three-dimensional environments 450A and 450B include a houseplant 408 (e.g., a representation of a houseplant) or a window 409 (e.g., a representation of a window) based on the perspectives of the first electronic device 101a and the second electronic device 101b, as shown in FIG. 4A . In some examples, the representation may include a portion of the physical environment 400 that is viewable through a transparent or translucent display of the electronic devices 101a and 101b. In some examples, the three-dimensional environment 450A / 450B has one or more characteristics of the three-dimensional environment 350A / 350B described above with reference to FIG. 3 .

[0056] As described above with reference to Figure 3, while electronic devices are communicatively linked in a multi-user communication session, users may be represented by avatars corresponding to the users of the electronic devices. In Figure 4A, first electronic device 101a and second electronic device 101b are co-located in physical environment 400, so that the users of electronic devices 101a and 101b are represented in the multi-user communication session via their physical personas (e.g., bodies) that are visible in a pass-through of physical environment 400 (e.g., rather than via virtual avatars). For example, as shown in Figure 4A, while first electronic device 101a and second electronic device 101b are in the multi-user communication session, second user 404 is visible in the field of view of first electronic device 101a, and first user 402 is visible in the field of view of second electronic device 101b. As described in more detail below, when a third user (e.g., a remote user) who is not co-located within physical environment 400 joins the multi-user communication session, the third user is represented via an avatar within three-dimensional environments 450A and 450B.

[0057] As also described above with reference to FIG. 3, while a first user 402 of a first electronic device 101a and a second user 404 of a second electronic device 101b are co-located in physical environment 400, and while the first electronic device 101a is in a multi-user communication session with the second electronic device 101b, the first user 402 and the second user 404 may be in a first spatial group within the multi-user communication session. In some examples, the first spatial group has one or more characteristics of spatial group 340 described above with reference to FIG. 3. As also described above, while the first user 402 and the second user 404 are in the first spatial group within the multi-user communication session, the users have a first spatial arrangement within the shared three-dimensional environment (e.g., represented by the locations of and / or the distance between users 402 and 404 within the overhead view 410 of FIG. 4A ) that is determined by the physical locations of electronic devices 101a and 101b within physical environment 440. Specifically, the first electronic device 101a and the second electronic device 101b experience spatial reality within the first spatial group as dictated by the physical locations and / or orientations of the first user 402 and the second user 404, respectively.

[0058] In some examples, as similarly described above with reference to FIG. 3, content viewable and / or interactive by a first user 402 (e.g., via the first electronic device 101a) and a second user 404 (e.g., via the second electronic device 101b) may be shared while the first electronic device 101a and the second electronic device 101b are in a multi-user communication session. For example, in FIG. 4A , the shared three-dimensional environment includes a virtual object 430 corresponding to a gaming user interface associated with a gaming application. In some examples, the virtual object 430 is a shared virtual object, such that, as shown in FIG. 4A , the virtual object 430 is displayed in and interactive within both the three-dimensional environment 450A and the three-dimensional environment 450B. In some examples, as shown in FIG. 4A , the virtual object 430 is displayed with a grabber bar 435 that is selectable to initiate movement of the virtual object 430 within the three-dimensional environment 450A / 450B. In some examples, the virtual object 430 has one or more characteristics of the shared virtual object 310 described above with reference to FIG.

[0059] 4B , while the first electronic device 101a is co-located with the second electronic device 101b in the physical environment 400 (e.g., and optionally while the first electronic device 101a is in a multi-user communication session with the second electronic device 110b), the first electronic device 101a detects an input corresponding to a request to move a virtual object 430 within the three-dimensional environment 450A. For example, as shown in FIG. 4B , the first electronic device 101a optionally detects that the hand 403 of the first user 402 performs an air pinch gesture (e.g., the index finger and thumb of the hand 403 come together to form a pinch hand shape) while the line of sight 425 of the first user 402 is directed toward the grabber bar 435 in the three-dimensional environment 450A, and then moves the hand 403 in space (e.g., toward the right relative to the body of the first user 402). It should be understood that additional or alternative inputs, such as an air tap gesture, gaze dwell, verbal commands, etc., may be provided to cause movement of the virtual object 430 within the three-dimensional environment 450A. Additionally, it should be understood that while such inputs (e.g., air gestures) performed by the first user 402 are not shown in FIG. 4B as being visible within the three-dimensional environment 450B presented at the second electronic device 101b, in some examples the inputs are visible within the three-dimensional environment 450B from the perspective of the second electronic device 101b (e.g., because the first user 402 is positioned within the field of view of the second electronic device 101b, as described above).

[0060] 4C , in response to detecting the input provided by the hand 403, the first electronic device 101a moves the virtual object 430 in the three-dimensional environment 450A according to the input. For example, as shown in the overhead view 410 of FIG. 4C , the first electronic device 101a moves the virtual object 430 to the right relative to the viewpoint of the first electronic device 101a in response to the first user 402's hand 403 moving to the right.

[0061] In some examples, as described above, movement of a virtual object 430, which is a shared virtual object within a shared three-dimensional environment, is based on whether the multi-user communication session includes co-located users, non-co-located users, or both. As described above, in the examples of FIGS. 4A-4C, a first user 402 at a first electronic device 101a and a second user 404 at a second electronic device 101b are co-located within a physical environment 400. When all participants in a multi-user communication session are co-located users, as in FIGS. 4A-4C, movement of a shared virtual object at an individual electronic device associated with one of the co-located users causes the shared virtual object to correspondingly move at the electronic devices associated with the other co-located users. Thus, as shown in FIG. 4C, when the first electronic device 101a moves the virtual object 430 according to input provided by the hand 403, the second electronic device 101b also moves the virtual object 430. For example, as shown in Figure 4C, the second electronic device 101b moves the virtual object 430 leftward within the three-dimensional environment 450B relative to the viewpoint of the second electronic device 101b, which mirrors the rightward movement of the virtual object 430 in the three-dimensional environment 450A at the first electronic device 101a. Additionally, as shown in Figure 4C, the first electronic device 101a and the second electronic device 101b move the virtual object 430 within the three-dimensional environment 450A / 450B according to the above-described inputs without updating the presentation of the pass-through representations of the first user 402 and the second user 404. For example, as shown in FIG. 4C, because neither first user 402 nor second user 404 physically moved within physical environment 400 when input provided by first user 402 was detected, the representation of second user 404 visible within three-dimensional environment 450A is optionally not updated, and the representation of first user 402 visible within three-dimensional environment 450B is optionally not updated.

[0062] 4D-4F illustrate exemplary interactions within a multi-user communication session including non-co-located users. In some examples, rather than being co-located within physical environment 400 as described above, a first user 402 of first electronic device 101a and a second user 404 of second electronic device 101b may not be co-located. For example, in FIG. 4D , first user 402 of first electronic device 101a is located within physical environment 400 (e.g., corresponding to physical environment 400 described above), and second user 404 of second electronic device 101b is located within physical environment 440 (e.g., including table 405) that is different from the physical environment 400 in which first electronic device 101a is located. In some examples, while second electronic device 101b is within physical environment 440, second electronic device 101b is located beyond a threshold distance (e.g., as described above) from first electronic device 101a. Additionally, in some instances, as shown in FIG. 4D, the second electronic device 101b is not within the field of view of the first electronic device 101a (or vice versa).

[0063] In some examples, while the first electronic device 101a and the second electronic device 101b are in a multi-user communication session, the first user 402 and the second user 404 may be visually represented using avatars in a shared three-dimensional environment, similar to that described above with reference to FIG. 3, because the first user 402 and the second user 404 are not co-located. For example, as shown in FIG. 4D, the first electronic device 101a displays an avatar 411 corresponding to the second user 404 of the second electronic device 101b in the three-dimensional environment 450A, and the second electronic device 101b displays an avatar 413 corresponding to the first user 402 of the first electronic device 101a in the three-dimensional environment 450B. In some examples, the avatars 411 and 413 have one or more characteristics of the avatars 315 and 317 described above with reference to FIG. 3.

[0064] Additionally, in some examples, while first electronic device 101a and second electronic device 101b are in a multi-user communication session, three-dimensional environments 450A and 450B include the above-described virtual object 430. In Figure 4D, virtual object 430 corresponds to a shared virtual object, as described above.

[0065] 4E, while the first electronic device 101a and the second electronic device 101b are in a multi-user communication session, the first electronic device 101a detects an input corresponding to a request to move a virtual object 430 within the three-dimensional environment 450A. For example, similar to above, the first electronic device 101a detects that the hand 403 of the first user 402 performs an air pinch gesture, and then moves the hand 403 to the right in space, optionally while the gaze 425 of the first user 402 is directed toward the grabber bar 435.

[0066] In some examples, when electronic devices are in a multi-user communication session and electronic devices such as first electronic device 101a and second electronic device 101b are not co-located, movement of a shared virtual object (e.g., virtual object 430) triggers spatial elaboration within the shared three-dimensional environment of the multi-user communication session. In some examples, the spatial elaboration corresponds to movement and / or rearrangement of avatars and / or shared objects (e.g., triggered by movement of the shared object) that enables spatial authenticity to be maintained within the first spatial group of first user 402 and second user 404. In FIG. 4E , because first electronic device 101a and second electronic device 101b are not co-located, input provided by first user 402 directed at virtual object 430 optionally triggers spatial elaboration at first electronic device 101a. 4F, in response to detecting the above-described input, the first electronic device 101a not only moves the virtual object 430 within the three-dimensional environment 450A according to the input, but also moves the avatar 411 corresponding to the second user 404 within the three-dimensional environment 450A according to the input. For example, as shown in FIG. 4F, the first electronic device 101a moves the virtual object 430 and the avatar 411 rightward (e.g., by equal amounts) within the three-dimensional environment 450A relative to the viewpoint of the first electronic device 101a in accordance with the rightward movement of the hand 403 of the first user 402.

[0067] In some examples, when spatial refinement is triggered at the first electronic device 101a, the movement of the virtual object 430 and the avatar 411 in the three-dimensional environment 450A is applied only to the avatar 413 corresponding to the first user 402 in the three-dimensional environment 450B at the second electronic device 101b. For example, as shown in FIG. 4F, the second electronic device 101b moves the avatar 413 to the right within the three-dimensional environment 450B, without moving the virtual object 430, relative to the viewpoint of the second electronic device 101b, in accordance with the input provided by the first user 402 at the first electronic device 101a, as reflected in the overhead view 412. Thus, as shown via the overhead views 410 and 412 of FIG. 4F, spatial authenticity from the perspectives of the first electronic device 101a and the second electronic device 101b is maintained following input provided by the first user 402 (e.g., the first user 402 looks to the right via the first electronic device 101a at the avatar 411 corresponding to the second user 404, and the second user 404 looks to the right via the second electronic device 101b at the avatar 413 corresponding to the first user 402).

[0068] 4G-4J illustrate exemplary interactions within a multi-user communication session including collocated and non-collocated users. In FIG. 4G, a first electronic device 101a is in a multi-user communication session with a second electronic device 101b and a third electronic device 101c. In some examples, as shown in overhead view 410 of FIG. 4G, the first electronic device 101a is collocated with the third electronic device 101c within the physical environment 400 described above. Additionally, in FIG. 4G, the second electronic device 101b is not collocated with the first electronic device 101a and the third electronic device 101c within the physical environment 400. For example, as shown in overhead view 412, the second electronic device is located within a physical environment 440 that is different from the physical environment 400 (e.g., as described above).

[0069] In some examples, while first electronic device 101a, second electronic device 101b, and third electronic device 101c are in a multi-user communication session, co-located (e.g., with respect to the individual electronic devices) users are represented in the shared three-dimensional environment via their physical bodies, as described above, and non-co-located (e.g., with respect to the individual electronic devices) users are represented in the shared three-dimensional environment via virtual representations (e.g., avatars), as described above. For example, in FIG. 4G, because second electronic device 101b is not co-located with first electronic device 101a and third electronic device 101c in physical environment 400, first electronic device 101a displays avatar 411 corresponding to second user 404 of second electronic device 101b, and third user 406 of third electronic device 101c is visible (e.g., in a pass-through or via a computer-generated representation) within three-dimensional environment 450A. Thus, the second electronic device 101b optionally displays within the three-dimensional environment 450B (e.g., as the second user 404 of the second electronic device 101b is located by itself within the physical environment 440) an avatar 413 corresponding to the first user 402 of the first electronic device 101a and an avatar 415 corresponding to the third user 406 of the third electronic device 101c, as shown in the overhead view 412 of FIG. 4G. Additionally, as also discussed above, the shared three-dimensional environment includes a virtual object 430 corresponding to the shared virtual object, as shown in FIG. 4G.

[0070] 4H, while first electronic device 101a, second electronic device 101b, and third electronic device 101c are in a multi-user communication session, first electronic device 101a detects an input corresponding to a request to move virtual object 430 within three-dimensional environment 450A. For example, similar to above, first electronic device 101a optionally detects that hand 403 of first user 402 performs an air pinch gesture while line of sight 425 of first user 402 is directed toward grabber bar 435 within three-dimensional environment 450A, and then hand 403 moves leftward in space, as shown in FIG.

[0071] In some examples, while electronic devices 101a, 101b, and 101c are in a multi-user communication session, it may be advantageous to provide a method for compensating for (e.g., reducing and / or preventing) lag between the detection of an input on one of the electronic devices and the execution of one or more corresponding actions on the other electronic devices. For example, when the first electronic device 101a detects the input performed by the hand 403 in FIG. 4H, the second electronic device 101b and the third electronic device 101c, depending on the data corresponding to the input transmitted by the first electronic device 101a (e.g., directly or indirectly via a server (e.g., a wireless communication terminal)), perform one or more actions based on the input detected by the first electronic device 101a, optionally creating a delay between when the first electronic device 101a responds to the input performed by the hand 403 of the first user 402 and when the second electronic device 101b and the third electronic device 101c perform one or more actions based on the input detected by the first electronic device 101a (e.g., this further creates a delay in each user's perception of the interaction).

[0072] To reduce the delays described above, one or more of the other electronic devices (e.g., the second electronic device 101b and / or the third electronic device 101c) may utilize computer vision techniques (e.g., object detection and / or tracking) in addition to the data transmitted by the first electronic device 101a to infer and / or predict the outcome of the input being detected by the first electronic device 101a. For example, as shown in FIG. 4H, when the first electronic device 101a detects an air pinch gesture performed by the hand 403 of the first user 402, the third electronic device 101c also detects (e.g., using the external image sensors 114b-iii and 114c-iii) that the first user 402 is performing an air pinch gesture using the hand 403 within the three-dimensional environment 450C presented via the display 120c of the third electronic device 101c. Additionally, the third electronic device 101c may detect that the hand 403 of the first user 402 moves leftward relative to the body of the first user 402 within the three-dimensional environment 450C (e.g., corresponding to the movement of the hand 403 away from the viewpoint of the third electronic device 101c). In some examples, the third electronic device 101c (e.g., and / or the second electronic device 101b) may therefore utilize the detected movement of the hand 403 to predict and / or infer the outcome of the input detected by the first electronic device 101a, as discussed in more detail below. In some examples, because the second electronic device 101b is not co-located with the first electronic device 101a and the third electronic device 101c in the physical environment 400, the third electronic device 101c may transmit data corresponding to the detected movement of the hand 403 to the second electronic device 101b, which allows the second electronic device 101b to predict and / or infer consequences of the input detected by the first electronic device 101a.

[0073] In some examples, as shown in FIG. 4I, in response to detecting an input performed by the hand 403 of the first user 402, the first electronic device 101a triggers spatial refinement in a similar manner as described above. For example, as shown in FIG. 4I, the first electronic device 101a moves the virtual object 430 and avatar 411 corresponding to the second user 404 leftward (e.g., by an equal amount) relative to the viewpoint of the first electronic device 101a within the three-dimensional environment 450A in accordance with the leftward movement of the hand 403 in FIG. 4H. In addition, as shown in FIG. 4I, the first electronic device 101a ceases updating the presentation of the third user 406 within the three-dimensional environment 450A (e.g., because spatial refinement applies only to virtual content displayed within the shared three-dimensional environment).

[0074] In some examples, when spatial refinement is triggered at the first electronic device 101a, the movement of the virtual object 430 and the avatar 411 in the three-dimensional environment 450A is applied only to the avatar 413 corresponding to the first user 402 and the avatar 415 corresponding to the third user 406 in the three-dimensional environment 450B at the second electronic device 101b. For example, as shown in FIG. 4I , the second electronic device 101b moves the avatars 413 and 415 leftward (e.g., by equal amounts) within the three-dimensional environment 450B, relative to the viewpoint of the second electronic device 101b, without moving the virtual object 430, in accordance with the input provided by the first user 402 at the first electronic device 101a, as reflected in the overhead view 412. Thus, as shown via the overhead views 410 and 412 of FIG. 4I, spatial authenticity from the perspectives of the first electronic device 101a, the second electronic device 101b, and the third electronic device 101c is maintained according to the input provided by the first user 402 (e.g., the first user 402 sees, via the first electronic device 101a, an avatar 411 corresponding to the second user 404 to his left, and the second user 404 sees, via the second electronic device 101b, an avatar 413 corresponding to the first user 402 and an avatar 415 corresponding to the third user 406 to his left).

[0075] Alternatively, in some examples, as shown in FIG. 4J, in response to detecting an input performed by the hand 403 of the first user 402 in FIG. 4H, the first electronic device 101a moves the virtual object 430 in the three-dimensional environment 450A according to the input (e.g., without triggering spatial refinement, in contrast to the above). For example, as shown in FIG. 4J, the first electronic device 101a moves the virtual object 430 leftward in the three-dimensional environment 450A relative to the viewpoint of the first electronic device 101a according to the leftward movement of the hand 403 in FIG. 4H, without moving the avatar 411 and without updating the presentation of the third user 406 in the three-dimensional environment 450A as shown in the overhead view 410. Furthermore, as shown in FIG. 4J, when the first electronic device 101a moves the virtual object 430 according to the input provided by the hand 403, the second electronic device 101b also moves the virtual object 430. 4J, the second electronic device 101b moves the virtual object 430 to the right in the three-dimensional environment 450B relative to the viewpoint of the second electronic device 101b, which mirrors the leftward movement of the virtual object 430 in the three-dimensional environment 450A on the first electronic device 101a. Furthermore, the second electronic device 101b stops moving the avatar 413 corresponding to the first user 402 and the avatar 415 corresponding to the third user 406 in the three-dimensional environment 450B. As shown in the overhead views 410 and 412, the alternative responses shown in FIG. 4J also allow spatial authenticity to be maintained from the perspectives of the first electronic device 101a, the second electronic device 101b, and the third electronic device 101c following input provided by the first user 402 (e.g., the first user 402 continues to see, through the first electronic device 101a, the avatar 411 corresponding to the second user 404 directly across from him and the third user 406 to his right, while the second user 404 continues to see, through the second electronic device 101b, the avatar 413 corresponding to the first user 402 directly across from him and the avatar 415 corresponding to the third user 406 to his left).

[0076] Thus, as outlined above, facilitating movement of shared virtual objects in a shared three-dimensional environment within a multi-user communication session based on whether the users in the multi-user communication session are co-located or non-co-located allows spatial authenticity to be maintained between the viewpoints of the users in the multi-user communication session, which improves user interaction and experience of the shared virtual object. Attention is now directed to examples of moving and / or repositioning shared virtual objects within a shared three-dimensional environment based on activation of one or more modes that control the movement of the shared virtual object within the shared three-dimensional environment.

[0077] In some examples, movement of the virtual object 430 within the shared three-dimensional environment is defined according to one or more (e.g., user-selected) modes. In some examples, the one or more modes include a first mode that, when activated, triggers spatial refinement when the virtual object 430 is moved within the shared three-dimensional environment (e.g., in response to detecting an input directed at the virtual object 430). For example, as shown in FIG. 4K, the virtual object 430 may be displayed with a selectable toggle 432 to activate (or deactivate) the first movement mode. In the example of FIG. 4K, the first mode is active, as indicated by the toggle 432 in the three-dimensional environment 450A, indicating that the first electronic device 101a triggers spatial refinement when moving the virtual object 430, such as moving the virtual object 430 discussed above with reference to FIG. 4I, in response to detecting an input corresponding to a request to move the virtual object 430, such as the input described above. Alternatively, when the first mode is not active, the first electronic device 101a optionally does not trigger spatial refinement within the three-dimensional environment 450A in response to detecting an input corresponding to a request to move the virtual object 430, such as moving the virtual object 430 described above with reference to FIG. 4J. In some examples, the toggle 432 is displayed along with the virtual object 430 in response to detecting an input corresponding to a request to display the toggle 432. For example, the first electronic device 101a displays the toggle 432 within the three-dimensional environment 450A in response to detecting a selection (e.g., an air pinch gesture directed thereat) of the grabber bar 435 of the virtual object 430 (e.g., without detecting a request to move the virtual object 430, such as moving the hand while in a pinch hand shape). It should be noted that the first mode may be activated or deactivated by any of the participants in the multi-user communication session, such as any of the first user 402, second user 404, and third user 406 (e.g., via input detected by their respective electronic devices).

[0078] In some examples, the one or more modes include a second mode that, when activated, triggers private movement when the virtual object 430 is moved within the shared three-dimensional environment (e.g., in response to detecting input directed at the virtual object 430). In some examples, the private movement of the virtual object 430 is similar to the movement of a private object, such as the private application window 330 of FIG. 3, even though the virtual object 430 is a shared virtual object, as described in more detail below.

[0079] 4L, the virtual object 430 may be displayed with a selectable toggle 434 to activate (or deactivate) the second movement mode. The example of FIG. 4L indicates that the second mode is active, as indicated by the toggle 434 in the three-dimensional environment 450A, and that in response to detecting an input corresponding to a request to move the virtual object 430, the first electronic device 101a moves the virtual object 430 privately for the first user 402, as discussed below. In some examples, the toggle 434 is displayed with the virtual object 430 in response to detecting an input corresponding to a request to display the toggle 434. For example, the first electronic device 101a displays the toggle 434 in the three-dimensional environment 450A in response to detecting a selection (e.g., an air pinch gesture directed thereat) of the grabber bar 435 of the virtual object 430 (without detecting a request to move the virtual object 430, such as, for example, moving the hand while in a pinch hand shape). It should be noted that, as also described above, the second mode may be activated or deactivated by any of the participants in the multi-user communication session, such as any of the first user 402, second user 404, and third user 406 (e.g., via input detected by their respective electronic devices).

[0080] 4L, while the second movement mode is active and the first electronic device 101a, the second electronic device 101b, and the third electronic device 101c are in a multi-user communication session, the first electronic device 101a detects an input corresponding to a request to move the virtual object 430 within the three-dimensional environment 450A. For example, as shown in FIG. 4L, the first electronic device 101a detects an air pinch gesture performed by the hand 403 of the first user 402, optionally while the line of sight 425 is directed toward the grabber bar 435, and then the hand moves forward in space (e.g., away from the body of the first user 402).

[0081] 4M, in response to detecting an input provided by the hand 403, the first electronic device 101a moves the virtual object 430 in the three-dimensional environment 450A according to the input. For example, as shown in FIG. 4M, the first electronic device 101a moves the virtual object 430 in the three-dimensional environment 450A away from the viewpoint of the first electronic device 101a according to the forward movement of the hand 403 in space, as shown in the overhead view 410. In some examples, because the second movement mode is active (e.g., private movement), the first electronic device 101a moves the virtual object 430 without performing spatial refinement. For example, as shown in FIG. 4M, when the first electronic device 101a moves the virtual object 430 within the three-dimensional environment 450A, the first electronic device 101a stops moving the avatar 411 corresponding to the second user 404 and stops updating the presentation of the third user 406 within the three-dimensional environment 450A. In addition, as shown in FIG. 4M, because the second movement mode is active when the above-mentioned input is detected by the first electronic device 101a, the movement of the virtual object 430 is private to the first user 402. In other words, the movement of the virtual object 430 is perceptible only by the first user 402 from the perspective of the first electronic device 101a within the three-dimensional environment 450A. Therefore, as shown in the overhead view 412, the second electronic device 101b stops updating the presentation of the three-dimensional environment 450B in response to the input detected by the first electronic device 101a. For example, the second electronic device 101b ceases moving the virtual object 430 within the three-dimensional environment 450B according to the input detected by the first electronic device 101a.

[0082] It should be understood that additional or alternative interactions directed at virtual object 430, other than movement, are similarly private to the individual user performing the interaction while the second mode (e.g., private mode) is active for virtual object 430. For example, inputs to rotate and / or resize virtual object 430 within three-dimensional environment 450A are similarly private to first user 402 of first electronic device 101a when the second mode is active.

[0083] 4N-4T illustrate exemplary interactions within a multi-user communication session involving co-located users. As shown in overhead view 410, a first user 402 of a first electronic device 101a, a second user 404 of a second electronic device 101b, and a third user 406 of a third electronic device 101c are, optionally, in a multi-user communication session. In some examples, as shown in overhead view 410 and as described herein, first user 402 (e.g., first electronic device 101a), second user 404 (e.g., second electronic device 101b), and third user 406 (e.g., third electronic device 101c) are co-located within physical environment 400 (e.g., corresponding to physical environment 400 described above). Thus, similar to above, views of the shared three-dimensional environment (e.g., including physical environment 400) are provided to (e.g., visible to) first user 402, second user 404, and third user 406 from the unique perspectives of first electronic device 101a, second electronic device 101b, and third electronic device 101c, respectively. Additionally, as shown in overhead view 410 of Figure 4N, the shared three-dimensional environment includes virtual object 430 (e.g., a game user interface associated with a game application), similar to above.

[0084] In some examples, during a multi-user communication session involving co-located users, the electronic devices facilitate movement of a virtual object according to a user-centric movement model, as described below. In FIG. 4N, while displaying a virtual object 430 within a shared three-dimensional environment, the first electronic device 101a detects an input corresponding to the initiation of movement of the virtual object 430. For example, as shown in FIG. 4N, the first electronic device 101a detects an air pinch gesture provided by the hand 403 of the first user 402 (e.g., and optionally, while the gaze of the first user 402 is directed toward the virtual object 430).

[0085] In some examples, facilitating the movement of a virtual object according to a user-centric model of movement includes grouping co-located users together in a multi-user communication session (e.g., based on the viewpoints of their respective electronic devices). For example, as shown in the overhead view 410 of FIG. 4O, a boundary 445 is defined around a first user 402, a second user 404, and a third user 406 in the physical environment 400 (e.g., based on the positions of the viewpoints of the first electronic device 101a, the second electronic device 101b, and the third electronic device 101c). In some examples, as shown in FIG. 4O, the boundary 445 corresponds to a “best fit” grouping of the co-located users in the multi-user communication session. For example, as shown in the overhead view of FIG. 4O, the size (e.g., including dimensions), shape, and / or location of the boundary 445 is based on the positions of the first user 402, the second user 404, and the third user 406 in the physical environment 400.

[0086] In some examples, the boundary 445 is determined by electronic devices associated with the first user 402, the second user 404, and / or the third user 406 based on position and attitude data provided by the electronic devices. For example, the first electronic device 101 a, the second electronic device 101 b, and / or the third electronic device 101 c may exchange data corresponding to the location of the electronic devices (e.g., and thereby the users) within the physical environment 400 and / or data corresponding to the orientation (e.g., including the forward looking direction) of the electronic devices (e.g., and thereby the users) within the physical environment 400. In some examples, the position and / or attitude data is determined by the electronic devices relative to a reference or center of reference for the spatial group of users and / or relative to each other, as also discussed herein.

[0087] In Figure 4O, the first electronic device 101a detects movement of the hand 403 in space while maintaining the air pinch gesture provided in Figure 4N. For example, as shown in Figure 4O, the first electronic device 101a detects movement of the hand 403 to the right relative to the perspective of the first electronic device 101a, corresponding to a request to move the virtual object 430 to the right within the shared three-dimensional environment from the perspective of the first electronic device 101a.

[0088] In some examples, as shown in FIG. 4P, in response to detecting the movement of the hand 403, the first electronic device 101a moves the virtual object 430 in a rightward direction from the perspective of the first electronic device 101a within the shared three-dimensional environment. Specifically, as shown in the overhead view 410 of FIG. 4P, the first electronic device 101a moves the virtual object 430 toward the group of co-located users within the shared three-dimensional environment. In some examples, facilitating the movement of the virtual object according to a user-centric model of movement includes moving the virtual object relative to the group of co-located users (e.g., defined by and / or based on boundary 445). For example, the movement of the virtual object according to the user-centric movement model is limited and / or constrained by boundary 445. In some examples, the degree to which the movement is constrained is based on the virtual object type. For example, in FIG. 4N, when an input to move the virtual object 430 is initially detected, the virtual object 430 is a first type object. In some examples, the first type of object is or includes a virtual object having a horizontal orientation within the shared three-dimensional environment, including two-dimensional and three-dimensional (e.g., volumetric) virtual objects having a horizontal orientation and / or surface (e.g., a horizontal top or bottom surface of a three-dimensional virtual object). In the example of FIG. 4P , because the virtual object 430 is a first type of object, in response to detecting the movement of the hand 403 of the first user 402 described above, the first electronic device 101a moves the virtual object 430 according to the movement of the hand 403 without specifically limiting the movement of the virtual object 430 to outside the boundary 445. For example, as shown in the overhead view 410 of FIG. 4P , the first electronic device 101a moves the virtual object 430 at least partially within the boundary 445 because the virtual object 430 is a horizontally oriented virtual object. In some examples, as described below, in the case of movement of a second type of virtual object different from the first type, the movement of the virtual object is limited to remaining outside the boundary 445 within the shared three-dimensional environment.

[0089] FIG. 4Q illustrates an example of a multi-user communication session including co-located users and a second type of virtual object (e.g., vertically oriented object) different from the first type (e.g., horizontally oriented object) described above. For example, as shown in overhead view 410 of FIG. 4Q, the multi-user communication session includes a first user 402 (e.g., and first electronic device 101a), a second user 404 (e.g., and second electronic device 101b), and a third user 406 (e.g., and third electronic device 101c) co-located within physical environment 400. Additionally, as shown in overhead view 410 of FIG. 4Q, the multi-user communication session includes shared virtual content within a shared three-dimensional environment of the multi-user communication session. For example, as shown in overhead view 410, the shared three-dimensional environment includes virtual object 436 that corresponds to a vertically oriented virtual object, as described in more detail below.

[0090] 4Q, while displaying virtual object 436 within the shared three-dimensional environment, first electronic device 101a detects an input corresponding to the initiation of movement of virtual object 436. For example, as shown in FIG. 4Q, first electronic device 101a detects an air pinch gesture provided by hand 403 of first user 402 (e.g., and optionally while first user 402's gaze is directed toward virtual object 436).

[0091] In some examples, as shown in FIG. 4R , similar to above, movement of the virtual object 436 is initiated within a multi-user communication session according to a user-centric model of movement (optionally because the multi-user communication session includes co-located users, as described above). Thus, in FIG. 4R , a boundary 445 is defined around the co-located users in the multi-user communication session, as previously described above. For example, as shown in the overhead view 410 of FIG. 4R , the boundary 445 is defined based on the positions of the viewpoints of the first electronic device 101 a, the second electronic device 101 b, and the third electronic device 101 c within the physical environment 400.

[0092] In Figure 4R, the first electronic device 101a detects movement of the hand 403 in space while maintaining the air pinch gesture provided in Figure 4Q. For example, as shown in Figure 4R, the first electronic device 101a detects movement of the hand 403 to the right relative to the viewpoint of the first electronic device 101a, corresponding to a request to move the virtual object 436 to the right within the shared three-dimensional environment from the viewpoint of the first electronic device 101a.

[0093] In some examples, as shown in Figure 4S, in response to detecting the movement of the hand 403, the first electronic device 101a moves the virtual object 436 within the shared three-dimensional environment to the right from the perspective of the first electronic device 101a (e.g., toward the perspective of the third electronic device 101c). Specifically, as shown in the overhead view 410 of Figure 4S, the first electronic device 101a moves the virtual object 436 toward the group of co-located users within the shared three-dimensional environment. In some examples, as described above, facilitating the movement of the virtual object according to a user-centric model of movement includes moving the virtual object relative to the group of co-located users (e.g., defined by and / or based on boundary 445). In some examples, because virtual object 436 is an object with a vertical orientation (e.g., a virtual window or user interface with a vertically oriented forward-facing surface), when an input to move virtual object 436 is first detected in FIG. 4Q, virtual object 436 is determined to be (e.g., classified as) a second type of object (e.g., a horizontally oriented virtual object) that is different from the first type of object described above. In the example of FIG. 4S, because virtual object 436 is a second type of object (e.g., not the first type described above), in response to detecting the movement of hand 403 of first user 402 described above, first electronic device 101a moves virtual object 436 in accordance with the movement of hand 403 (e.g., toward the group of co-located users) but restricts (e.g., stops) the movement of virtual object 436 to be outside boundary 445. For example, as shown in the overhead view 410 of FIG. 4S, even though and / or even if the movement of the hand 403 of the first user 402 corresponds to the movement of the virtual object 436 to a location within the boundary 445, the first electronic device 101a moves the virtual object 436 to a location within the shared three-dimensional environment that is at or outside the boundary 445 because the virtual object 436 is a vertically oriented virtual object.

[0094] 4S, when the virtual object 436 is moved into the boundary 445, the virtual object 436 "snaps" to a point on the boundary 445. For example, as shown in the overhead view 410, the first electronic device 101a aligns the center of the virtual object 436 with the point 437 on the boundary 445 (e.g., because the movement of the hand 403 corresponds to the movement of the virtual object 436 to a location within the boundary 445). In some examples, the movement of the virtual object 436 locks (e.g., stops) at the boundary 445 such that the virtual object 436 is prevented from being moved into the boundary 445. In some examples, the location at which the virtual object 436 is displayed in response to the movement input is a location offset (e.g., by a predetermined distance) from the boundary 445. 4S, when the center of virtual object 436 is aligned with point 437 on boundary 445, virtual object 436 is perpendicular to point 437, as indicated by the double-headed arrow in overhead view 410. In some examples, the orientation of virtual object 436 is perpendicular to the forward direction of the head or torso of the user providing the movement input, such as the forward direction of first user 402.

[0095] 4S, the first electronic device 101a detects further movement directed toward the virtual object 436 in the shared three-dimensional environment. For example, as shown in FIG. 4S, the first electronic device 101a detects an air pinch and drag gesture directed toward the virtual object 436 in the shared three-dimensional environment, such as an air pinch provided by the hand 403 followed by movement of the hand 403 in space relative to the viewpoint of the first electronic device 101a. Alternatively, in some examples, the movement input directed toward the virtual object 436 is or includes an air swing or flick gesture provided by the hand 403 of the first user 402. For example, the first electronic device 101a detects an air pinch gesture provided by the hand 403, followed by a "throwing" or "tossing" motion by the hand 403 in the direction of the arrow shown in FIG. 4S. In either example, the first electronic device 101a optionally also detects the first user's 402 gaze directed towards the virtual object 436 during input.

[0096] In some examples, as shown in FIG. 4T , in response to detecting a movement input directed at the virtual object 436, the first electronic device 101 a moves the virtual object 436 according to and / or based on the movement input. For example, as shown in the overhead view 410 of FIG. 4T , the first electronic device 101 a moves the virtual object 436 in a rightward direction relative to the group of co-located users in the shared three-dimensional environment so that the virtual object 436 is located farther from the viewpoint of the first electronic device 101 a and closer to the viewpoint of the second electronic device 101 b. Additionally, as also discussed above, as shown in the overhead view 410 of FIG. 4T , as the virtual object 436 is moved within the shared three-dimensional environment according to the user-centric movement model, the virtual object 436 “snaps” or locks to a second point on the boundary 445. For example, as also described above, the first electronic device 101a aligns the center of the virtual object 436 with a second point on the boundary 445 that is in the direction of movement of the virtual object 436.

[0097] In some examples, moving the virtual object 436 according to the user-centric movement model includes updating the orientation of the virtual object 436 within the shared three-dimensional environment. For example, as shown in the overhead view 410 of FIG. 4T, the first electronic device 101a rotates the virtual object 436 (e.g., about a vertical axis passing through the center of the virtual object 436) as the virtual object 436 is moved within the shared three-dimensional environment. In some examples, the amount (e.g., number of degrees) by which the virtual object 436 is rotated within the shared three-dimensional environment is based on the average forward direction of the first electronic device 101a (e.g., and first user 402), the second electronic device 101b (e.g., and second user 404), and the third electronic device 101c (e.g., and third user 406) within the physical environment 400. For example, as shown in the overhead view 410 of Figure 4T, an average forward direction 452 is determined based on averaging the forward directions (e.g., orientations) of the first electronic device 101a, the second electronic device 101b, and the third electronic device 101c. In some examples, the forward-facing surface of the virtual object 436 is angled to face toward (e.g., perpendicular or nearly perpendicular to) the average forward direction 452, as shown in the overhead view 410 of Figure 4T.

[0098] 4U-4AA illustrate exemplary interactions within a multi-user communication session including collocated and non-collocated users. As shown in overhead view 410, a first user 402 of first electronic device 101a discussed above and a third user 406 of third electronic device 101c discussed above, who are collocated within physical environment 400, are in a multi-user communication session with two non-collocated users (e.g., users not located within physical environment 400) who are visually represented within overhead view 410 as avatars 411 and 413 (it should be understood that alternative representations, such as those described herein above, are possible). In some examples, similar to the above, views of a shared three-dimensional environment (e.g., including physical environment 400) are provided to (e.g., visible to) first user 402 and third user 406 from the unique perspectives of first electronic device 101a and third electronic device 101c, respectively. Additionally, as shown in overhead view 410 of FIG. 4U, the shared three-dimensional environment includes virtual objects 430 (eg, game user interfaces associated with game applications) similar to those described above.

[0099] In some examples, during a multi-user communication session including co-located and non-co-located users, electronic devices facilitate movement of virtual objects according to the user-centric movement model described above. For example, in FIG. 4U, the third electronic device 101c detects an input corresponding to a request to initiate movement of a virtual object 430 within the shared three-dimensional environment from the perspective of the third electronic device 101c, such as via an air pinch gesture provided by the hand 407 of the third user 406, as also described above. In some examples, as shown in FIG. 4V, in response to detecting the input provided by the hand 407 of the third user 406, the virtual content displayed within the shared three-dimensional environment of the co-located users is grouped together for the co-located users. For example, as shown in the overhead view 410 of FIG. 4V, for the co-located first user 402 and third user 406, the virtual content of the shared three-dimensional environment includes the virtual object 430 and avatars 411 and 413. Thus, as shown in the overhead view 410 via the first boundary 445A, the virtual object 430 and the avatars 411 and 413 are grouped into a first group within the shared three-dimensional environment. In some examples, the first boundary 445A has one or more characteristics of the boundary 445 described above. For example, the first boundary 445A is based on (e.g., has a size and / or shape based on) the positions of the virtual object 430 and the avatars 411 and 413 within the shared three-dimensional environment. Thus, as described below, a movement input directed at the virtual object 430 optionally causes the virtual object 430 and the avatars 411 and 413 to move as a group (e.g., in unison) in accordance with the movement input and as defined by the first boundary 445A, similar to performing scene elaboration on the virtual content as described previously herein.

[0100] In Figure 4V, the third electronic device 101c detects movement of the hand 407 while maintaining the air pinch gesture detected in Figure 4U. For example, as shown in Figure 4V, the third electronic device 101c detects the hand 407 moving leftward in space relative to the perspective of the third electronic device 101c, corresponding to a request to move the virtual object 430 leftward within the shared three-dimensional environment from the perspective of the third electronic device 101c. In some examples, as described below, movement of the virtual object 430, and therefore the avatars 411 and 413 as discussed above, is implemented for the co-located users within the shared three-dimensional environment. 4V, the first user 402 and the third user 406 are grouped together as co-located users in the shared three-dimensional environment, as indicated by the second boundary 445B, and the virtual object 430 and the avatars 411 and 413 are moved accordingly within the shared three-dimensional environment, as discussed below. In some examples, the second boundary 445B has one or more characteristics of the boundary 445 described above. For example, the second boundary 445B is based on (e.g., has a size and / or shape based on) the positions of the viewpoints of the first electronic device 101a and the third electronic device 101c within the physical environment 400.

[0101] In some examples, as shown in FIG. 4W, in response to detecting the movement of the hand 407, the third electronic device 101c moves the virtual content bounded by the first boundary 445A within the shared three-dimensional environment based on the hand movement. In some examples, as shown in the overhead view 410 of FIG. 4W, lateral (e.g., leftward) movement of the hand 407 of the third user 406 (e.g., indicated by the hand 407 in FIG. 4V) causes the virtual content bounded by the first boundary 445A to move radially within the shared three-dimensional environment relative to the viewpoint of the third electronic device 101c according to a user-centric movement model. For example, in the overhead view 410, the virtual object 430 and the avatars 411 and 413 are moved radially to the left (e.g., counterclockwise) as a group along a circle or curve centered on the third electronic device 101c (e.g., as indicated by the center 448). 4W , the orientation of virtual object 430 and avatars 411 and 413 is updated within the shared three-dimensional environment according to the radial (e.g., counterclockwise) movement of virtual object 430 and avatars 411 and 413. It should be understood that in some examples, the center 448 according to which the radial movement is defined is the center of the group of co-located users (e.g., first user 402 and third user 406), which, optionally, corresponds to the center of second boundary 445B rather than the center of the electronic device that detects the movement input.

[0102] In Figure 4W, the third electronic device 101c detects further movement input directed toward the virtual object 430 in the shared three-dimensional environment. For example, as shown in Figure 4W, the third electronic device 101c detects that the hand 407 moves toward the viewpoint of the third electronic device 101c (e.g., toward the body of the third user 406) while maintaining the air pinch gesture described above.

[0103] 4X, in response to detecting the movement of the hand 407, the third electronic device 101c moves the virtual content bounded by the first boundary 445A based on the movement of the hand 407 within the shared three-dimensional environment. For example, as shown in the overhead view 410, the virtual object 430 and the avatars 411 and 413 are moved as a group (e.g., in unison) toward the group of co-located users (e.g., the first user 402 and the third user 406) in accordance with the movement of the hand 407. In some examples, similar to the above, the movement of the virtual content bounded by the first boundary 445A relative to the group of co-located users (e.g., the first user 402 and the third user 406) is selectively restricted by the second boundary 445B. For example, similar to the above, a first type of shared object (e.g., a horizontally oriented object) is permitted to cross the second boundary 445B, while a second type of shared object (e.g., a vertically oriented object) and the avatar are not permitted to cross the second boundary 445B. Thus, as shown by way of example in the overhead view 410 of FIG. 4X, when the virtual object 430 and the avatars 411 and 413 are moved in accordance with the movement of the hand 407, the amount of movement (e.g., the distance of movement) of the virtual object 430 and the avatars 411 and 413 relative to the viewpoints of the first electronic device 101a and the third electronic device 101c is constrained by the second boundary 445B. As a result, the virtual object 430 is permitted to at least partially cross the second boundary 445B as shown, but the avatar 413 (and therefore the avatar 411) is, optionally, not permitted to at least partially cross the second boundary 445B. In the example of FIG. 4X, it should be understood that movement of virtual object 430 (e.g., an object of the first type, as described above) toward second boundary 445B ceases when movement input causes avatar 413 to reach second boundary 445B, as shown in overhead view 410.

[0104] 4Y illustrates an example of a multi-user communication session including collocated and non-collocated users and a second type of virtual object different from the first type described above. For example, as shown in overhead view 410 of FIG. 4Y, the multi-user communication session includes a first user 402 (e.g., and first electronic device 101a) and a third user 406 (e.g., and third electronic device 101c collocated within physical environment 400), and includes a second user (e.g., represented by avatar 411) and a fourth user (e.g., represented by avatar 413) that are not collocated with first user 402 and third user 406 within physical environment 400. Additionally, as shown in overhead view 410 of FIG. 4Y, the multi-user communication session includes shared virtual content within the shared three-dimensional environment of the multi-user communication session. For example, as shown in overhead view 410, the shared three-dimensional environment includes virtual object 436, which corresponds to a vertically oriented virtual object, as described in more detail below.

[0105] 4Y, while displaying virtual object 436 within the shared three-dimensional environment, third electronic device 101c detects an input corresponding to the initiation of movement of virtual object 436. For example, as shown in FIG. 4Y, third electronic device 101c detects an air pinch gesture provided by hand 407 of third user 406 (e.g., and optionally while third user 406's gaze is directed toward virtual object 436).

[0106] In some examples, as shown in FIG. 4Z , similar to above, movement of virtual content (e.g., virtual object 436 and avatars 411 and 413) is initiated within the multi-user communication session according to a user-centric movement model (optionally because the multi-user communication session includes co-located users, as described above). Thus, in FIG. 4Z , as described above, first boundary 445A is defined around the virtual content with respect to the co-located users in the multi-user communication session. For example, as shown in overhead view 410 of FIG. 4Z , first boundary 445A is defined based on the positions of virtual object 436 and avatars 411 and 413.

[0107] In Figure 4Z, the third electronic device 101c detects movement of the hand 407 in space while maintaining the air pinch gesture provided in Figure 4Y. For example, as shown in Figure 4Z, the third electronic device 101c detects movement of the hand 407 toward the viewpoint of the third electronic device 101c (e.g., toward the body of the third user 406), corresponding to a request to move the virtual object 436 from the viewpoint of the third electronic device 101c toward the viewpoint of the third electronic device 101c in the shared three-dimensional environment. In some examples, similar to above, the movement of the hand 407, corresponding to the request to move the virtual object 436, moves the virtual object 436 and the avatars 411 and 413 within the shared three-dimensional environment. In particular, as described in more detail below, movement of virtual content bounded by a first boundary 445A according to a user-centric movement model is performed for a group of co-located users (e.g., a first user 402 and a third user 406) as defined by a second boundary 445B. In some examples, the second boundary 445B corresponds to the second boundary 445B described above. For example, as shown in the overhead view 410 of FIG. 4Z, the second boundary 445B is defined based on the positions of the viewpoints of the first electronic device 101a and the third electronic device 101c within the physical environment 400.

[0108] 4AA, in response to detecting the movement of hand 407, third electronic device 101c moves virtual content bounded by first boundary 445A based on the movement of hand 407 within the shared three-dimensional environment. For example, as shown in overhead view 410, virtual object 436 and avatars 411 and 413 are moved as a group (e.g., in unison) toward the group of co-located users (e.g., first user 402 and third user 406) in accordance with the movement of hand 407. In some examples, similar to the above, the movement of virtual content bounded by first boundary 445A relative to the group of co-located users (e.g., first user 402 and third user 406) is selectively constrained by second boundary 445B. For example, similar to the above, a first type of shared object (e.g., a horizontally oriented object) is allowed to cross the second boundary 445B, while a second type of shared object (e.g., a vertically oriented object) and avatars are not allowed to cross the second boundary 445B. Thus, as shown by way of example in the overhead view 410 of FIG. 4X, when a virtual object 436 (e.g., a second type of object) and avatars 411 and 413 are moved in accordance with the movement of the hand 407, the amount of movement (e.g., movement distance) of the virtual object 436 and the avatars 411 and 413 relative to the viewpoints of the first electronic device 101 a and the third electronic device 101 c is constrained by the second boundary 445B such that the virtual object 436 and the avatars 411 and 413 are optionally not allowed to at least partially cross the second boundary 445B. In the example of FIG. 4AA, it should be understood that when the movement input causes avatar 411 to reach second boundary 445B, as shown in overhead view 410, movement of virtual object 436 (e.g., and therefore avatar 413) toward the group of co-located users (e.g., first user 402 and third user 406) is discontinued.

[0109] Thus, as outlined above, providing one or more user-selectable modes that define the movement of a shared virtual object within a multi-user communication session provides users participating in the multi-user communication session with more control over interactions directed at the shared virtual object, which helps to increase user privacy and thus improves the user experience. We now turn our attention to additional interactions within a multi-user communication session involving co-located and non-co-located users.

[0110] 5A-5E illustrate exemplary interactions within a multi-user communication session including collocated and non-collocated users, according to some examples of the present disclosure. In FIG. 5A, a first electronic device 101a (e.g., associated with a first user 502), a second electronic device 101b (e.g., associated with a second user 504), and a third electronic device 101c (e.g., associated with a third user 506) are in a multi-user communication session. In some examples, the first user 502, the second user 504, and the third user 506 correspond to the first user 402, the second user 404, and the third user 406, respectively, of FIGS. 4A-4M.

[0111] As shown in overhead view 510 of Figure 5A, first electronic device 101a and second electronic device 101b are co-located in physical environment 500. Additionally, as shown in overhead view 512 of Figure 5A, third electronic device 101c is located in a physical environment 540 that is different from physical environment 500. Thus, as also described above, third electronic device 101c is not co-located with first electronic device 101a and second electronic device 101b (e.g., as previously described above, a spatial group including first user 502, second user 504, and third user 506 is a hybrid spatial group). In some examples, first electronic device 101a and second electronic device 101b thus display avatar 515 corresponding to third user 506, as shown in overhead view 512, and third electronic device 101c displays avatar 511 corresponding to first user 502 and avatar 513 corresponding to second user 504, as shown in overhead view 510. In some examples, avatars 511, 513, and 515 correspond to avatars 411, 413, and 415, respectively, of FIGS. 4A-4M. Additionally, as shown in FIG. 5A and similarly discussed above, the shared three-dimensional environment includes virtual object 530 corresponding to the shared virtual object. In some examples, virtual object 530 corresponds to virtual object 430 described above. In FIG. 5A, the viewpoints of electronic devices 101a, 101b, and 101c and the locations of virtual object 530 in overhead views 510 and 512 optionally correspond to the original locations of viewpoints and virtual object 530 in the spatial group.

[0112] 5A-5B, the spatial arrangement of users within a spatial group in a multi-user communication session is updated based on changes in the positions of the viewpoints of electronic devices 101a, 101b, and 101c relative to the shared three-dimensional environment. For example, as shown in the overhead view 510 of FIG. 5B, the first electronic device 101a has moved to a first updated position (e.g., relative to a previous position 536c (e.g., the original position of the viewpoint of the first electronic device 101a within the spatial group)) caused by the movement of the first user 502 within the physical environment 500, and the second electronic device 101b has moved to a second updated position (e.g., relative to a previous position 536c) caused by the movement of the second user 504 within the physical environment 500. Similarly, as shown in overhead view 512 of Figure 5B, the third electronic device 101c has moved to a third updated position (e.g., relative to previous position 536b) caused by movement of the third user 506 within the physical environment 540. As also described above with reference to Figure 3, movement of the first electronic device 101a and the second electronic device 101b causes avatars 511 and 513, respectively, to move relative to the viewpoint of the third electronic device 101c, as shown in overhead view 512, and movement of the third electronic device 101c causes avatar 515 to move relative to the viewpoints of the first electronic device 101a and the second electronic device 101b, as shown in overhead view 510 of Figure 5B.

[0113] In some examples, the spatial arrangement within a spatial group of a multi-user communication session is configured to be reset (e.g., recentered relative to the viewpoint of an individual user within the multi-user communication session). For example, resetting the spatial arrangement within a spatial group within a multi-user communication session causes the virtual content (e.g., avatars and shared objects) to be redisplayed relative to the individual user's current viewpoint (e.g., repositioned (e.g., moved by an equal amount) to be within the individual user's current field of view) from the individual user's (e.g., the user providing input to reset the spatial arrangement) perspective.

[0114] 5B, the first electronic device 101a detects an input corresponding to a request to reset the spatial arrangement of spatial groups in the multi-user communication session. For example, as shown in FIG. 5B, the first electronic device 101a detects an input directed at a physical button on the first electronic device 101a, such as via the hand 503 of the first user 502. In some examples, the input corresponds to a tap or series of taps on the physical button, a rotation of the physical button, a swipe of the physical button, etc. In some examples, the input corresponds to a selection of a virtual button associated with resetting the spatial arrangement displayed on the first electronic device 101a.

[0115] In some examples, as shown in Figure 5C, in response to detecting an input corresponding to a request to reset the spatial arrangement, the first electronic device 101a repositions the virtual object 530 and avatar 515 corresponding to the third user 506 relative to the viewpoint of the first electronic device 101a. For example, as shown in the overhead view 510, the first electronic device 101a moves the virtual object 530 and avatar 515 relative to the viewpoint of the first electronic device 101a such that the viewpoint of the first electronic device 101a is positioned at a previous position 536c relative to the virtual object 530 in Figure 5B.

[0116] In some examples, rather than moving the virtual object 530 and the avatar 515 by equal amounts when resetting the spatial arrangement (e.g., similar to spatial refinement as described above), the virtual object 530 and the avatar 515 are moved by different amounts relative to the perspective of the first electronic device 101a because the multi-user communication session includes co-located and non-co-located users. As shown in the overhead view 510 of FIG. 5C, the first electronic device 101a positions the avatar 515 so that it is positioned at a previous position 536b relative to the virtual object 530 of FIG. 5B. In some examples, the first electronic device 101a ceases updating the presentation of the second user 504 relative to the perspective of the first electronic device 101a (e.g., because the second user 504 has not physically moved within the physical environment 500 when the input was detected in FIG. 5B). Thus, as outlined above, in instances where a multi-user communication session includes co-located and non-co-located users, resetting the spatial arrangement causes the virtual content (e.g., avatars and shared virtual objects) to be individually repositioned relative to the virtual content's previous / original location within the shared three-dimensional environment, rather than relative to the perspective of the individual user providing the input to reset the spatial arrangement.

[0117] In some examples, the above techniques for resetting the spatial arrangement may shift the content relative to the viewpoint of another electronic device (e.g., the second electronic device 101b and / or the third electronic device 101b). For example, as shown in the overhead view 510 of FIG. 5C, the virtual object 530 and the avatar 515 are shifted farther from the viewpoint of the second electronic device 101b compared to FIG. 5B when the spatial arrangement of the spatial group is reset by the first electronic device 101a. In addition, as shown in the overhead view 512 of FIG. 5C, the avatar 511 corresponding to the first user 502 is shifted toward and positioned at a previous position 536c relative to the virtual object 530 from the viewpoint of the third electronic device 101c (e.g., based on the movement of the virtual object 530 relative to the viewpoint of the first electronic device 101a in the overhead view 510). Similarly, as described above, the third electronic device 101c updates the position of the avatar 513 without updating the position of the avatar 511 corresponding to the second user 504 (e.g., because the second electronic device 101b does not physically change its location within the physical environment 500 when the input is detected by the first electronic device 101a of FIG. 5B).

[0118] 5D-5E illustrate an example of a multi-user communication session in which one of the electronic devices in the multi-user communication session does not support a head-mounted display. In FIG. 5D, a first electronic device 101a is in a multi-user communication session with a second electronic device 101b and a mobile electronic device 570. For example, as shown in FIG. 5D, the first electronic device 101a displays a three-dimensional environment 550A including an avatar 511 corresponding to a second user of the second electronic device 101b and a third user 506 holding the mobile electronic device 570. As shown in FIG. 5D, the mobile electronic device 570 does not support a head-mounted display (e.g., the mobile electronic device 570 corresponds to a tablet computer or smartphone held by the third user 506). In some examples, mobile electronic device 570 has one or more components of electronic device 260 / 270 of FIG. 2, such as location sensor(s) 204A / 204B, image sensor(s) 206A / 206B, touch-sensitive surface(s) 209A / 209B, orientation sensor(s) 210A / 210B, microphone(s) 213A / 213B, display generating component(s) 214A / 214B, speaker(s) 216A / 216B, processor(s) 218A / 218B, memory 220A / 220B, communication circuitry 222A / 222B, and / or communication bus(es) 208A / 208B.

[0119] In the example of Figure 5D, the multi-user communication session includes co-located and non-co-located users, as similarly described above. For example, as shown in Figure 5D, first electronic device 101a (e.g., including a first user of first electronic device 101a) and mobile electronic device 570 (e.g., including a third user 506) are both located within physical environment 500, while a second user of second electronic device 101b is not located within physical environment 500 (e.g., is located within a different physical environment, such as physical environment 540 described above). Thus, as shown in Figure 5D, the second user is visually represented within three-dimensional environment 550A via avatar 511 corresponding to the second user. Additionally, as shown in Figure 5D, the shared three-dimensional environment of the multi-user communication session optionally includes shared virtual content, specifically, the aforementioned virtual object 530.

[0120] 5D , the multi-user communication session includes a non-head-mounted device (non-HMD) device, specifically a mobile electronic device 570. In some such examples, virtual content shared within the multi-user communication session may be viewable and / or interactive to a third user 506 via the mobile electronic device 570, but may be viewable and / or interactive as two-dimensional content rather than as a virtual object within a shared three-dimensional environment. For example, as shown in FIG. 5D , a virtual object 530 is shared among a first user, a second user, and a third user 506, and thus the mobile electronic device 570 is configured to display, via a display 571 (e.g., a touchscreen), a user interface 575 corresponding to the virtual object 530. As previously described herein, the virtual object 530 optionally corresponds to a gaming user interface (e.g., a virtual board game). Thus, the user interface 575 is the same or a similar gaming user interface that enables the third user 506 (e.g., via the mobile electronic device 570) to participate in a shared activity that is a virtual board game. For example, third user 506 may interact with virtual objects 530 in the shared three-dimensional environment via input detected by mobile electronic device 570 directed at user interface 575. However, in this case, third user 506 optionally has little understanding of the spatial arrangement of spatial groups in the multi-user communication session. For example, because only user interface 575 is displayed by mobile electronic device 570, third user 506 is optionally not provided with and / or is provided with limited (e.g., by mobile electronic device 570) an indication of the location of virtual objects 530 and avatars 511 in the shared three-dimensional environment (e.g., relative to the viewpoint of mobile electronic device 570). In such a case, third user 506 optionally does not experience spatial authenticity with the first and second users in the multi-user communication session.

[0121] Alternatively, in some examples, during a multi-user communication session, the mobile electronic device 570 may be configured to provide the third user 506 with a view of the shared three-dimensional environment, particularly from the perspective of the mobile electronic device 570 within the spatial group. For example, as shown in FIG. 5E , the mobile electronic device 570 provides the third user 506 with an augmented reality (AR) or mixed reality (MR) experience, such as using one or more external cameras of the mobile electronic device 570. As shown in FIG. 5E , the first user 502, including a portion of the physical environment 500 and the first electronic device 101a, is optionally presented on the display 571 (e.g., based on the camera view of the mobile electronic device 570). Furthermore, in some examples, the mobile electronic device 570 displays virtual objects 530 and an avatar 511 on the display 571. In some examples, as shown in FIG. 5E , the virtual objects 530 and the avatar 511 are displayed at locations on the display 571 based on the spatial arrangement of the spatial group from the perspective of the mobile electronic device 570. In some examples, the third user 506 may interact with the virtual object 530 via input detected by the mobile electronic device 570. For example, the mobile electronic device 570 may be configured to perform one or more actions, such as moving, rotating, and / or resizing the virtual object 530, and / or one or more interactions with the game user interface of the virtual object 530, in response to detecting touch input (e.g., a tap or swipe) on the display 571 and / or a hand-based air gesture (e.g., an air pinch gesture, an air tap gesture, etc.) detected by one or more cameras of the mobile electronic device 570. In such cases, the third user 506 optionally experiences spatial verisimilitude with the first user 502 and the second user in the multi-user communication session. Thus, as outlined above, even users not associated with an HMD-type device can participate in the multi-user communication session and actively interact with content shared among users in the multi-user communication session.

[0122] It is understood that the examples shown and described herein are merely illustrative, and that additional and / or alternative elements may be provided within the three-dimensional environment for interacting with the example content. It should be understood that the appearance, shape, form, and size of each of the various user interface elements and objects shown and described herein are exemplary, and that alternative appearances, shapes, forms, and / or sizes may be provided. For example, virtual objects (e.g., shared virtual object 310, private application window 330, and / or virtual objects 430 and 530) may be provided in alternative shapes other than rectangular, such as circular shapes, triangular shapes, etc. In some examples, the various selectable options (e.g., toggles 432 and 434), user interface elements (e.g., grabber bars 435 and 535), control elements, etc. described herein may be verbally selected via a user's verbal command (e.g., a "select an option" verbal command). Additionally or alternatively, in some examples, various options, user interface elements, controls, etc. described herein may be selected and / or manipulated via user input received via one or more separate input devices in communication with the electronic device(s). For example, selection input may be received via a physical input device, such as a mouse, trackpad, keyboard, etc., in communication with the electronic device(s).

[0123] It should also be noted that additional or alternative forms of content may be provided and / or interacted with within the shared three-dimensional environment in the examples provided above. For example, user interfaces of other types of applications may be provided within the shared three-dimensional environment, such as user interfaces for web browsing applications, media player applications, text editing applications, image viewing applications, video conferencing applications, etc. As another example, immersive content may be provided within the shared three-dimensional environment, such as a three-dimensional virtual environment that occupies a predetermined portion of the field of view of the individual electronic devices (e.g., 100%, 90%, 80%, 75%, 50%, etc. immersion). The virtual environment optionally corresponds to a virtual scene or setting (e.g., at a particular location and / or at a particular time), such as a virtual beach, a virtual park, a virtual theater, a virtual forest, etc. In some examples, the virtual environment includes virtual objects, such as virtual seats or benches, virtual rocks, virtual water, virtual clouds, virtual grass, virtual animals, etc. In instances where a virtual environment is presented within a shared three-dimensional environment, each user in a multi-user communication session can experience a portion of the virtual environment from a separate perspective (e.g., via each user's separate electronic device). For example, a virtual object in the virtual environment may be located in one location relative to one user's perspective, but in a different location relative to another user's perspective.

[0124] 6 shows a flow diagram illustrating an example process for moving an object within a three-dimensional environment within a multi-user communication session based on whether the multi-user communication session includes co-located or non-co-located users, according to some examples of the present disclosure. In some examples, process 600 begins with one or more displays, one or more input devices, and a first electronic device in communication with a second electronic device, the first electronic device being in a communication session with the second electronic device. In some examples, the first electronic device and the second electronic device are, optionally, head-mounted displays similar to or corresponding to electronic devices 260 / 270 of FIG. 2, respectively. As shown in FIG. 6, in some examples, at 602, the first electronic device presents, via one or more displays, a three-dimensional environment including a visual representation of a first object of a first type and a user of the second electronic device. For example, as shown in FIG. 4A, a first electronic device 101a presents a three-dimensional environment 450A including a virtual object 430 (e.g., a shared virtual object) and a visual representation (e.g., a pass-through representation or a computer-generated representation) of a second user 404 of a second electronic device 101b.

[0125] In some examples, at 604, while presenting a three-dimensional environment including a visual representation of a first object of a first type and a user of a second electronic device, the first electronic device receives, via one or more input devices, a first input corresponding to a request to move the first object within the three-dimensional environment. For example, as shown in FIG. 4B , the first electronic device 101a optionally detects an air pinch gesture provided by the hand 403 of the first user 402 while the line of sight 425 of the first user 402 is directed toward a virtual object 430 (e.g., a grabber bar 435 of the virtual object 430), followed by a movement of the hand 403 in space (e.g., toward the right relative to the body of the first user 402).

[0126] In some examples, in response to receiving the first input at 606, and in accordance with a determination at 608 that one or more criteria are met, including criteria that are met when the second electronic device is co-located with the first electronic device in the first physical environment, the first electronic device moves a first object of a first type within the three-dimensional environment relative to a viewpoint of the first electronic device in accordance with the first input without updating a presentation of a visual representation of a user of the second electronic device. For example, as shown in FIG. 4C , because the second electronic device 101b is co-located with the first electronic device 101a in the physical environment 400, the first electronic device 101a moves the virtual object 430 in accordance with the input, as shown in the overhead view 410, without updating a presentation of a visual representation of a second user 404 of the second electronic device 101b in the three-dimensional environment 450A. In some examples, in accordance with a determination at 610 that one or more criteria are not met because the second electronic device is not co-located with the first electronic device in the first physical environment, the method moves a visual representation of a first object of a first type and a user of the second electronic device in the three-dimensional environment relative to a viewpoint of the first electronic device in accordance with the first input. For example, as shown in FIG. 4F , because the second electronic device 101b is not co-located with the first electronic device 101a in the physical environment 400, the first electronic device 101a moves a virtual object 430 and an avatar 411 corresponding to a second user 404 of the second electronic device 101b in the three-dimensional environment 450A in accordance with the input, as shown in overhead views 410 and 412.

[0127] It will be understood that process 600 is an example, and that more, fewer, or different operations may be performed in the same or a different order. Additionally, the operations of process 600 described above are, optionally, implemented by executing one or more functional modules of an information processing device, such as a general-purpose processor (e.g., as described with respect to FIG. 2) or an application-specific chip, and / or by other components of FIG. 2.

[0128] 7A-7G illustrate exemplary interactions within a multi-user communication session involving co-located users, according to some examples of the present disclosure. In FIG. 7A, a first electronic device 101a (e.g., associated with a first user 702) and a second electronic device 101b (e.g., associated with a second user 704) are engaged in a multi-user communication session. In some examples, the first user 702 and the second user 704 correspond to the first user 402 and the second user 404, respectively, of FIGS. 4A-4AA and / or the first user 502 and the second user 504, respectively, of FIGS. 5A-5E.

[0129] 7A , the first electronic device 101a and the second electronic device 101b are co-located in a physical environment 700 that includes a houseplant 708 and a window 709. For example, a first user 702 wearing the first electronic device 101a is positioned directly across from a second user 704 wearing the second electronic device 101b in the physical environment 700. Thus, the second user 704 (e.g., and the second electronic device 101b) is visible in a three-dimensional environment 750A presented by the first electronic device 101a (e.g., via the display 120a), and the first user 702 (e.g., and the first electronic device 101a) is visible in a three-dimensional environment 750B presented by the second electronic device 101b (e.g., via the display 120b). 7A , the shared three-dimensional environment includes a virtual object 736 that corresponds to the shared virtual object. In some examples, the virtual object 736 corresponds to the virtual object 436 described above. For example, as shown in FIG. 7A , the virtual object 736 is or includes a user interface, such as a media player user interface, associated with a respective application running on the electronic devices 101a and 101b. In the example of FIG. 7A , the virtual object 736 is shared between the first electronic device 101a and the second electronic device 101b (e.g., and therefore viewable by and interactive with the first user 702 and the second user 704), but the virtual object 736 is not currently visible within the field of view of the second electronic device 101b from the current viewpoint of the second electronic device 101b. Additionally, in FIG. 7A, the viewpoints of electronic devices 101a and 101b and the location of virtual object 736 within overhead view 710 optionally represent the viewpoints and the location of virtual object 736 within the shared three-dimensional environment of the multi-user communication session.

[0130] 7A , virtual object 736 includes and / or is displayed with interactive controls for controlling the display of content in virtual object 736. For example, in FIG. 7A , virtual object 736 is displayed with playback controls 737 that are selectable to control playback of a content item currently displayed in virtual object 736, such as a movie, a television program episode, a music video, and / or other media-based content. In some examples, while first electronic device 101 a and second electronic device 101 b are in a multi-user communication session involving co-located users, interacting with virtual object 736 causes playback controls 737 to cease being displayed in the shared three-dimensional environment, as discussed below.

[0131] 7B , the first electronic device 101a detects an input corresponding to a request to move a virtual object 736 within the three-dimensional environment 750A. For example, as shown in FIG. 7B , the first electronic device 101a optionally detects an air pinch gesture performed by the hand 703 of the first user 702 while the line of sight 725 of the first user 702 is directed toward a grabber bar 739 within the three-dimensional environment 750A. In some examples, the grabber bar 739 is selectable to initiate movement of the virtual object 736 within the three-dimensional environment 750A. Additionally, following detection of the air pinch gesture, in some examples, the first electronic device 101a detects movement of the hand 703. For example, as shown in FIG. 7B , the first electronic device 101a detects that the hand 703 moves leftward in space relative to the viewpoint of the first electronic device 101a.

[0132] In some examples, as shown in Figure 7C, in response to detecting the movement of the hand 703 of the first user 702, the first electronic device 101a moves the virtual object 736 in accordance with the movement of the hand 703 in the three-dimensional environment 750A. For example, as shown in Figure 7C, the virtual object 736 is moved leftward in the three-dimensional environment 750A relative to the viewpoint of the first electronic device 101a in accordance with the leftward movement of the hand 703. In some examples, as shown in Figure 7C, the virtual object 736 is a shared virtual object in the multi-user communication session, so that the movement of the virtual object 736 in the three-dimensional environment 750A correspondingly moves the virtual object 736 in the three-dimensional environment 750B presented on the second electronic device 101b. For example, as shown in FIG. 7C , in response to receiving input data provided by the first electronic device 101a corresponding to movement of the virtual object 736, the second electronic device 101b moves the virtual object 736 in a rightward direction relative to the viewpoint of the second electronic device 101b, causing the virtual object 736 to be at least partially visible within the three-dimensional environment 750B from the viewpoint of the second electronic device 101b.

[0133] 7C , while the virtual object 736 is being moved within the three-dimensional environment 750A in accordance with the movement of the hand 703, the first electronic device 101a ceases displaying the playback controls 737 associated with the virtual object 736. Additionally, in some examples, as shown in FIG. 7C , the playback controls 737 cease displaying along with the virtual object 736 in the three-dimensional environment 750B presented on the second electronic device 101b. In particular, the playback controls 737 cease displaying while an interaction with the virtual object 736 (e.g., movement of the virtual object 736) is in progress, which, as one advantage, helps to avoid and / or prevent (e.g., unintended) interactions with the playback controls 737 that may interrupt a current interaction with the virtual object 736 and / or overload the electronic devices 101a and 101b when responding to potentially conflicting inputs.

[0134] 7C , while the virtual object 736 is being moved within the three-dimensional environment 750A according to the movement of the hand 703 detected by the first electronic device 101a, the second electronic device 101b updates the visual appearance of the virtual object 736 within the three-dimensional environment 750B during the movement of the virtual object 736 within the three-dimensional environment 750B. For example, as shown in FIG. 7C , the second electronic device 101b reduces the visual emphasis and / or visual fidelity of the virtual object 736 by, for example, increasing the transparency, decreasing the brightness, changing the color scheme, decreasing the saturation, and / or ceasing to display the content of the virtual object 736 (e.g., the user interface of the virtual object 736) during the movement of the virtual object 736 caused by the input provided by the first user 702 at the first electronic device 101a. Additionally or alternatively, in some examples, the second electronic device 101b displays a visual indication 726 (e.g., a notification, alert, or message) of the input being provided by the first user 702 at the first electronic device 101a causing the virtual object 736 at the second electronic device 101b to move within the three-dimensional environment 750B. For example, as shown in FIG. 7C , the second electronic device 101b provides a visual indication 726 informing the second user 704 that the first user 702 is currently providing movement input directed at the virtual object 736 within the three-dimensional environment 750B. Changing the visual appearance of the virtual object 736 and / or providing the visual indication 726 during movement of the virtual object 736 caused by input provided by the first user 702 visually informs the second user 704 that the virtual object 736 is currently being interacted with, which, as another benefit, helps to avoid and / or discourage further interaction with the virtual object 736 while the virtual object 736 is still being moved and / or helps to avoid user confusion as to the cause of the movement of the virtual object 736.

[0135] 7C, the first electronic device 101a detects further (e.g., continued) movement of the hand 703 of the first user 702 while the hand 703 maintains the air pinch gesture described above. For example, as shown in FIG. 7C, the first electronic device 101a detects that the hand 703 continues to move leftward relative to the viewpoint of the first electronic device 101a, corresponding to a request to move the virtual object 736 further leftward within the three-dimensional environment 750A relative to the viewpoint of the first electronic device 101a.

[0136] 7D , in response to detecting continued movement of the hand 703 of the first user 702, the first electronic device 101a moves the virtual object 736 further leftward within the three-dimensional environment 750A relative to the viewpoint of the first electronic device 101a in accordance with the movement of the hand 703. In some examples, as shown in FIG. 7D and similar to the above, as the first electronic device 101a moves the virtual object 736 in accordance with the movement of the hand 703, the second electronic device 101b correspondingly moves the virtual object 736 within the three-dimensional environment 750B. For example, in FIG. 7D , the second electronic device 101b moves the virtual object 736 further leftward within the three-dimensional environment 750B relative to the viewpoint of the second electronic device 101b based on the input data provided by the first electronic device 101a corresponding to the movement of the virtual object 736 within the three-dimensional environment 750A.

[0137] In some examples, when the first electronic device 101a detects the conclusion of the movement input provided by the hand 703 of the first user 702, such as the release of an air pinch gesture and / or the relaxation of the hand 703, as shown in Figure 7D, the first electronic device 101a redisplays the playback control 737 within the three-dimensional environment 750A (e.g., because the virtual object 736 is no longer being interacted with). Additionally, when the first electronic device 101a redisplays the playback control 737 with the virtual object 736 because the interaction with the virtual object 736 has ended, as shown in Figure 7D, the second electronic device 101b redisplays the playback control 737 with the virtual object 736 within the three-dimensional environment 750B (e.g., in response to receiving an indication from the first electronic device 101a that the input has ended). 7D , the second electronic device 101b also, optionally, restores the visual appearance of the virtual object 736 within the three-dimensional environment 750B in response to receiving an indication that the input directed at the virtual object 736 at the first electronic device 101a has ended. For example, in FIG. 7D , the second electronic device 101b increases and / or restores the visual emphasis and / or visual fidelity of the content of the virtual object 736, such as decreasing the transparency, increasing the brightness, restoring the saturation and / or color scheme of the user interface of the virtual object 736. Redisplaying the playback controls 737 and restoring the visual appearance of the virtual object 736 facilitates the user's discovery that their interaction with the virtual object 736 has ended, thereby providing the user with a visual indication that the playback controls 737 are now available for interaction, which, as one benefit, enhances and / or improves the overall user experience within the multi-user communication session.

[0138] In some examples, the orientation of the virtual object 736 can be manipulated with respect to the viewpoint of the individual electronic devices independently (e.g., separately) from the movement of the virtual object 736 with respect to the viewpoint of the individual electronic devices. Notably, in some examples, a rotation affordance may be provided that allows the user to directly rotate the virtual object 736 to update the orientation of the virtual object 736, without requiring and / or moving the virtual object 736. For example, in FIG. 7E , as described above, the virtual object 736 is currently displayed with a grabber bar 739 (e.g., a movement affordance) within the three-dimensional environment 750A. In FIGS. 7E-7F , the first electronic device 101a detects that the line of sight 725 of the first user 702 moves to be directed toward a predetermined portion of the virtual object 736. In some examples, the predetermined portion of the virtual object 736 corresponds to a side or edge of the virtual object 736, such as the right side of the virtual object 736, as shown in FIG. 7F . In some examples, as shown in FIG. 7F , in response to detecting a line of sight 725 directed toward a predetermined portion of the virtual object 736, the first electronic device 101a displays a rotate affordance 742 within the three-dimensional environment 750A. In some examples, as described below, the rotate affordance 742 is selectable to initiate a rotation of the virtual object 736 relative to the viewpoint of the first electronic device 101a. Additionally, in some examples, when the rotate affordance 742 is displayed with the virtual object 736 in the three-dimensional environment 750A, the first electronic device 101a ceases displaying the grabber bar 739 within the three-dimensional environment 750A, as shown in FIG.

[0139] 7F , while displaying a rotation affordance 742 within the three-dimensional environment 750A, the first electronic device 101a detects an input provided by the hand 703 of the first user 702 directed toward the rotation affordance 742 within the three-dimensional environment 750A. For example, as shown in FIG. 7F , the first electronic device 101a optionally detects an air pinch gesture provided by the hand 703 of the first user 702 following movement of the hand 703 in space relative to the viewpoint of the first electronic device 101a while the line of sight 725 is directed toward the rotation affordance 742.

[0140] 7G, in response to detecting an input provided by the hand 703, the first electronic device 101a rotates the virtual object 736, thereby changing the orientation of the virtual object 736 in the three-dimensional environment 750A relative to the viewpoint of the first electronic device 101a in accordance with the movement of the hand 703. For example, as shown in FIG. 7G, the first electronic device 101a rotates the virtual object 736 clockwise (e.g., about a vertical axis passing through the center of the virtual object 736) in the three-dimensional environment 750A in accordance with the leftward movement of the hand 703. As shown in FIG. 7G, when the first electronic device 101a rotates the virtual object 736, the first electronic device 101a ceases moving the virtual object 736 in the three-dimensional environment 750A in accordance with the movement of the hand 703. 7F and 7G, the virtual object 736 is rotated within the three-dimensional environment 750A, but the virtual object 736 remains positioned at the same location within the three-dimensional environment 750A from the perspective of the first electronic device 101a (e.g., because the input described above is directed toward the rotation affordance 742 rather than the grabber bar 739 within the three-dimensional environment 750A). As shown in FIG. 7G, when the input directed toward the rotation affordance 742 ends (e.g., when the first electronic device 101a detects the release of the air pinch gesture and / or the relaxation of the hand 703) and / or when the line of sight 725 is no longer directed toward the predetermined portion of the virtual object 736, the first electronic device 101a ceases displaying the rotation affordance 742 and re-displays the grabber bar 739 within the three-dimensional environment 750A.

[0141] Thus, in accordance with the above, some examples of the present disclosure include a first electronic device in communication with one or more displays, one or more input devices, and a second electronic device, the first electronic device being in a communication session with the second electronic device, presenting, at the first electronic device, via the one or more displays, a three-dimensional environment including a visual representation of a first object of a first type and a user of the second electronic device; receiving, while presenting the three-dimensional environment including the first object of the first type and the visual representation of the user of the second electronic device, via the one or more input devices, a first input corresponding to a request to move the first object within the three-dimensional environment; and, in response to receiving the first input, in accordance with a determination that one or more criteria are met, including criteria that are met when the second electronic device is co-located with the first electronic device in the first physical environment, moving a first object of a first type within the three-dimensional environment relative to a viewpoint of the first electronic device in accordance with the first input without updating a presentation of a visual representation of a user of the second electronic device; and in accordance with a determination that the one or more criteria are not met because the second electronic device is not co-located with the first electronic device in the first physical environment, moving a first object of a first type within the three-dimensional environment relative to a viewpoint of the first electronic device in accordance with the first input.

[0142] Additionally or alternatively, in some examples, the first type of object includes a virtual object shared between a user of the first electronic device and a user of the second electronic device within the communication session. Additionally or alternatively, in some examples, the three-dimensional environment further includes a second object of a second type different from the first type, and the method further includes, in response to receiving the first input, ceasing to move the second object of the second type within the three-dimensional environment relative to a viewpoint of the first electronic device according to the first input. Additionally or alternatively, in some examples, the second type of object includes a virtual object private to a user of the first electronic device within the communication session. Additionally or alternatively, in some examples, the first electronic device being co-located with the second electronic device within the first physical environment follows a determination that the second electronic device is within a threshold distance of the first electronic device within the first physical environment. Additionally or alternatively, in some examples, the second electronic device being co-located with the first electronic device in the first physical environment pursuant to a determination that the second electronic device is located within a field of view of the first electronic device. Additionally or alternatively, in some examples, pursuant to a determination that the second electronic device is co-located with the first electronic device in the first physical environment, a visual representation of a user of the second electronic device corresponds to a pass-through representation of the user of the second electronic device.

[0143] Additionally or alternatively, in some examples, in accordance with a determination that the second electronic device is not co-located with the first electronic device in the first physical environment, the visual representation of the user of the second electronic device corresponds to a virtual avatar of the user of the second electronic device. Additionally or alternatively, in some examples, the movement of the first object of the first type is associated with one or more modes in the three-dimensional environment, and the one or more criteria include a second criterion that is satisfied when a first mode of the one or more modes is not active. Additionally or alternatively, in some examples, the method includes detecting, via one or more input devices, a movement of a viewpoint of the first electronic device while presenting a three-dimensional environment including a visual representation of a first object of a first type and a user of a second electronic device; and in response to detecting the movement of the viewpoint of the first electronic device, updating the presentation of the three-dimensional environment based on an updated viewpoint of the first electronic device, where the visual representation of the first object of the first type and the user of the second electronic device are no longer visible within the field of view of the first electronic device from the updated viewpoint; and, while the visual representation of the first object of the first type and the user of the second electronic device are no longer visible within the field of view of the first electronic device, detecting, via the one or more input devices, a movement of a viewpoint of the first electronic device while the visual representation of the first object of the first type and the user of the second electronic device are no longer visible within the field of view of the first electronic device. receiving a second input corresponding to a request to update the spatial arrangement of the three-dimensional environment; and updating the spatial arrangement of the three-dimensional environment in response to receiving the second input, wherein the updating includes, in accordance with a determination that the one or more criteria are satisfied, moving a first object of a first type within the three-dimensional environment to be repositioned within a field of view of the first electronic device from an updated viewpoint of the first electronic device without updating a presentation of a visual representation of the user of the second electronic device; and, in accordance with a determination that the one or more criteria are not satisfied, moving a first object of a first type within the three-dimensional environment to be repositioned within a field of view of the first electronic device from an updated viewpoint of the first electronic device without updating a presentation of a visual representation of the user of the second electronic device.

[0144] Additionally or alternatively, in some examples, the movement of the first object of the first type is associated with one or more modes within the three-dimensional environment, including individual modes that define the movement of the first object of the first type relative to a viewpoint of the first electronic device. The method further includes, while displaying the first object of the first type and while the individual mode is active, receiving, via the one or more input devices, a second input corresponding to a request to move the first object within the three dimensional environment; and, in response to receiving the second input, moving the first object of the first type within the three dimensional environment relative to a point of view of the first electronic device in accordance with the first input without updating a presentation of a visual representation of a user of the second electronic device in accordance with a determination that the one or more criteria are met because the second electronic device is co-located with the first electronic device in the first physical environment; and, in accordance with a determination that the one or more criteria are not met because the second electronic device is not co-located with the first electronic device in the first physical environment, moving the first object of the first type within the three dimensional environment and the visual representation of the user of the second electronic device relative to a point of view of the first electronic device in accordance with the first input. Additionally or alternatively, in some examples, the method further includes, while displaying the first object of the first type and while the respective mode is not active, receiving, via the one or more input devices, a second input corresponding to a request to move the first object within the three-dimensional environment, and, in response to receiving the second input, moving the first object of the first type within the three-dimensional environment relative to a viewpoint of the first electronic device in accordance with the first input without updating a presentation of a visual representation of the user of the second electronic device.

[0145] Some examples of the present disclosure are directed to electronic devices comprising one or more processors, a memory, and one or more programs, the one or more programs being stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any of the methods described above.

[0146] Some examples of the present disclosure are directed to a non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of the first electronic device, cause the first electronic device to perform any of the methods described above.

[0147] Some examples of the present disclosure are directed to a first electronic device comprising one or more processors, a memory, and means for performing any of the above methods.

[0148] Some examples of the present disclosure are directed to an information processing apparatus for use in a first electronic device, the information processing apparatus comprising means for performing any of the above methods.

[0149] The foregoing description has been set forth with reference to specific embodiments for purposes of explanation. However, the illustrative description above is not intended to be exhaustive or to limit the disclosure to the precise form disclosed. Many modifications and variations are possible in light of the above teachings. The examples were chosen and described to best explain the principles of the disclosure and its practical application, thereby enabling those skilled in the art to best utilize the disclosure and various described examples, with various modifications suited to the particular use contemplated.

Claims

1. 1. A method comprising: A first electronic device in communication with one or more displays, one or more input devices, and a second electronic device, the first electronic device being in a communication session with the second electronic device, presenting, via the one or more displays, a three-dimensional environment including a first object of a first type and a visual representation of a user of the second electronic device; receiving, via the one or more input devices, a first input corresponding to a request to move the first object within the three-dimensional environment while presenting the three-dimensional environment including the first object of the first type and the visual representation of the user of the second electronic device; In response to receiving the first input, and moving the first object of the first type within the three-dimensional environment relative to a viewpoint of the first electronic device in accordance with the first input without updating a presentation of the visual representation of the user of the second electronic device in accordance with a determination that one or more criteria are met, including criteria that are met when the second electronic device is co-located with the first electronic device in a first physical environment. and in accordance with a determination that the one or more criteria are not met because the second electronic device is not co-located with the first electronic device in the first physical environment, moving the visual representation of the first object of the first type in the three-dimensional environment and the user of the second electronic device relative to the viewpoint of the first electronic device in accordance with the first input.

2. The method of claim 1 , wherein the first type of object comprises a virtual object shared between the user of the first electronic device and the user of the second electronic device within the communication session.

3. The three-dimensional environment further includes a second object of a second type different from the first type, and the method further comprises:

10. The method of claim 1, further comprising, in response to receiving the first input, ceasing to move the second object of the second type within the three-dimensional environment relative to the viewpoint of the first electronic device in accordance with the first input.

4. The method of claim 3 , wherein the second type of object comprises a virtual object that is private to the user of the first electronic device within the communication session.

5. 10. The method of claim 1, wherein the first electronic device being co-located with the second electronic device in the first physical environment is pursuant to a determination that the second electronic device is within a threshold distance of the first electronic device in the first physical environment.

6. 10. The method of claim 1, wherein the second electronic device being co-located with the first electronic device in the first physical environment is pursuant to a determination that the second electronic device is located within a field of view of the first electronic device.

7. 10. The method of claim 1, wherein, in accordance with the determination that the second electronic device is co-located with the first electronic device in the first physical environment, the visual representation of the user of the second electronic device corresponds to a pass-through representation of the user of the second electronic device.

8. 10. The method of claim 1, wherein, in accordance with the determination that the second electronic device is not co-located with the first electronic device in the first physical environment, the visual representation of the user of the second electronic device corresponds to a virtual avatar of the user of the second electronic device.

9. movement of the first object of the first type is associated with one or more modes within the three-dimensional environment; The method of claim 1 , wherein the one or more criteria includes a second criterion that is met when a first mode of the one or more modes is not active.

10. detecting, via the one or more input devices, movement of the viewpoint of the first electronic device while presenting the three-dimensional environment including the visual representation of the first object of the first type and the user of the second electronic device; updating the presentation of the three-dimensional environment based on an updated viewpoint of the first electronic device, in response to detecting the movement of the viewpoint of the first electronic device, wherein the visual representation of the first object of the first type and the user of the second electronic device are no longer visible within a field of view of the first electronic device from the updated viewpoint; receiving a second input via the one or more input devices corresponding to a request to update the spatial arrangement of the three-dimensional environment while the visual representation of the first object of the first type and the user of the second electronic device is not visible within the field of view of the first electronic device; and updating the spatial arrangement of the three-dimensional environment in response to receiving the second input, wherein the updating comprises: according to a determination that the one or more criteria are satisfied, moving the first object of the first type within the three-dimensional environment so that it is repositioned within the field of view of the first electronic device from the updated viewpoint of the first electronic device without updating a presentation of the visual representation of the user of the second electronic device; and moving the visual representation of the first object of the first type within the three-dimensional environment and the user of the second electronic device so as to be repositioned within the field of view of the first electronic device from the updated viewpoint of the first electronic device according to a determination that the one or more criteria are not satisfied.

11. 2. The method of claim 1, wherein movement of the first object of the first type is associated with one or more modes within the three-dimensional environment, including individual modes that define movement of the first object of the first type relative to the viewpoint of the first electronic device.

12. receiving, while displaying the first object of the first type and while the respective mode is active, a second input via the one or more input devices corresponding to a request to move the first object within the three-dimensional environment; In response to receiving the second input, according to a determination that the one or more criteria are satisfied because the second electronic device is co-located with the first electronic device in the first physical environment, moving the first object of the first type within the three-dimensional environment relative to the viewpoint of the first electronic device in accordance with the first input without updating a presentation of the visual representation of the user of the second electronic device; 12. The method of claim 11, further comprising: in accordance with a determination that the one or more criteria are not met because the second electronic device is not co-located with the first electronic device in the first physical environment, moving the visual representation of the first object of the first type in the three-dimensional environment and the user of the second electronic device relative to the viewpoint of the first electronic device in accordance with the first input.

13. receiving, while displaying the first object of the first type and while the individual mode is not active, a second input via the one or more input devices corresponding to a request to move the first object within the three-dimensional environment; In response to receiving the second input, 12. The method of claim 11 , further comprising: moving the first object of the first type within the three-dimensional environment relative to the viewpoint of the first electronic device in accordance with the first input without updating a presentation of the visual representation of the user of the second electronic device.

14. a first electronic device, one or more processors; Memory and and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any one of claims 1 to 13.

15. 14. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a first electronic device, cause the first electronic device to perform the method of any one of claims 1 to 13.

Citation Information

Patent Citations

  • Shared and private holographic objects

    CN105393158A

  • System and method for three-dimensional placement and refinement in multi-user communication sessions

    CN116668658A

  • Shared Holographic Objects and Private Holographic Objects

    JP2016525741A

  • 3D Object Annotation

    JP2023513747A

  • System and method of three-dimensional placement and refinement in multi-user communication sessions

    US20230273706A1