User interface for managing live communication sessions

By integrating display generation components and sensors in computer systems, a new user interface is provided, which solves the problems of low efficiency and high complexity in managing live communication sessions in the prior art, and achieves more efficient interaction and energy consumption management.

CN120075386APending Publication Date: 2025-05-30APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510471256.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-09-12
Filing Date
2023-09-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art methods and interfaces for managing live communication sessions have problems such as low efficiency, high complexity and error-prone, resulting in a large cognitive burden on users and wasted energy consumption.

Method used

By integrating display generation components and sensors in a computer system, a new user interface is provided that allows users to interact with the virtual/augmented reality environment through multiple input methods (such as gestures, gazes), simplifying the management of live communication sessions.

Benefits of technology

This method and interface reduces the number and complexity of user input, improves interaction efficiency, reduces energy consumption, and enhances the experience of virtual/augmented reality environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075386A_ABST
    Figure CN120075386A_ABST
Patent Text Reader

Abstract

The invention relates to a user interface for managing live communication sessions. The present disclosure generally relates to managing live communication sessions. A computer system optionally displays an option to invite a respective user to join an ongoing communication session. A computer system optionally displays one or more options to modify an appearance representing an avatar of the user of the computer system. A computer system optionally transitions a communication session from a spatial communication session to a non-spatial communication session. A computer system optionally displays information about participants in a communication session.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent application with the application number 202380065246.4, the application date of September 21, 2023, and the title of "User Interface for Managing Live Communication Sessions".

[0002] Cross - Reference to Related Applications

[0003] This application claims priority to U.S. Patent Application No. 18 / 367,418, titled "USER INTERFACES FOR MANAGING LIVE COMMUNICATION SESSIONS", filed on September 12, 2023; U.S. Provisional Patent Application No. 63 / 470,882, titled "USER INTERFACES FOR MANAGING LIVE COMMUNICATION SESSIONS", filed on June 3, 2023; and U.S. Provisional Patent Application No. 63 / 409,583, titled "USER INTERFACES FOR MANAGING LIVE COMMUNICATION SESSIONS", filed on September 23, 2022. The entire content of each of these patent applications is incorporated herein by reference. Technical Field

[0004] The present disclosure generally relates to computer systems that communicate with display generation components and optionally one or more sensors to provide computer-generated experiences, including but not limited to electronic devices that provide virtual reality experiences and mixed reality experiences via a display. Background Art

[0005] In recent years, the development of computer systems for augmented reality has increased significantly. Example augmented reality environments include virtual elements that at least partially replace or enhance the physical world. Input devices for computer systems and other electronic computing devices (such as cameras, controllers, joysticks, touch-sensitive surfaces, and touchscreen displays) are used to interact with virtual / augmented reality environments. Example virtual elements include virtual objects such as digital images, videos, text, icons, and control elements (such as buttons and other graphics). Summary of the Invention

[0006] Some methods and interfaces for managing live communication sessions, such as those including at least some virtual elements (e.g., applications, augmented reality environments, mixed reality environments, and virtual reality environments), are cumbersome, inefficient, and limited. For example, systems that provide insufficient control for performing actions associated with virtual objects, systems that require a series of inputs to achieve a desired result in an augmented reality environment, and systems in which virtual object manipulation is complex, tedious, and error-prone impose a significant cognitive burden on users and detract from the experience of the virtual / augmented reality environment. In addition, these methods take longer than necessary, wasting the energy of the computer system. This latter consideration is particularly important in battery-powered devices.

[0007] Accordingly, there is a need for computer systems that are more efficient and intuitive for users and have improved methods and interfaces for managing live communication sessions. Such methods and interfaces optionally supplement or replace conventional methods for managing live communication sessions. Such methods and interfaces form a more effective human-machine interface by helping users understand the connection between the inputs provided and the device's response to those inputs, thereby reducing the quantity, degree, and / or nature of the inputs from the user.

[0008] The above-mentioned deficiencies and other problems associated with the user interface of a computer system are reduced or eliminated by the disclosed system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a laptop computer, a tablet computer, or a handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device such as a watch or a head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also referred to as a "touch screen" or "touch screen display"). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, in addition to the display generation component, the computer system further has one or more output devices, which include one or more haptic output generators and / or one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, a memory, and one or more modules, programs, or instruction sets stored in the memory for performing multiple functions. In some embodiments, the user interacts with the GUI through contact and gestures of a stylus and / or finger on a touch-sensitive surface, movements of the user's eyes and hands in space relative to the GUI (and / or the computer system) or the user's body (as captured by cameras and other motion sensors), and / or voice input (as captured by one or more audio input devices). In some embodiments, the functions performed through the interaction optionally include image editing, drawing, presentation, word processing, spreadsheet creation, gaming, making and receiving phone calls, video conferencing, sending and receiving emails, instant messaging, test support, digital photography, digital video recording, web browsing, digital music playback, note-taking, and / or digital video playback. The executable instructions for performing these functions are optionally included in a transient and / or non-transient computer-readable storage medium or other computer program product configured to be executed by one or more processors.

[0009] There is a need for an electronic device having improved methods and interfaces for managing live communication sessions. Such methods and interfaces can supplement or replace conventional methods for managing live communication sessions. Such methods and interfaces reduce the amount, degree, and / or nature of input from the user and result in a more efficient human-machine interface. For battery-powered computing devices, such methods and interfaces save power and increase the time interval between battery charges.

[0010] In some embodiments, a computer system displays a set of controls (e.g., transport controls and / or other types of controls) associated with controlling playback of media content in response to detecting a user's gaze and / or gesture. In some embodiments, the computer system initially displays a first set of controls in a reduced salience state (e.g., having reduced visual salience) in response to detecting a first input, and then displays a second set of controls (which optionally includes additional controls) in an increased salience state in response to detecting a second input. In this way, the computer system optionally provides the user with feedback that the user has initiated display of the controls without unduly distracting the user from the content (e.g., by initially displaying the controls in a visually less salient manner), and then, based on detecting user input indicating that the user wishes to further interact with the controls, displays the controls in a visually more salient manner to allow for easier and more accurate interaction with the computer system.

[0011] Example methods are described herein. An example method includes: at a computer system in communication with a display generation component and one or more sensors: displaying, via the display generation component, representations of a plurality of users; receiving, via the one or more sensors, a selection of a representation of a corresponding user of the plurality of users; and in response to receiving the selection of the representation of the corresponding user: displaying, via the display generation component, an option to invite the corresponding user to join an ongoing communication session based on determining that an ongoing communication session exists; and foregoing displaying the option to invite the corresponding user to join the ongoing communication session based on determining that no ongoing communication session exists.

[0012] An example method includes: at a computer system in communication with a display generation component and one or more sensors: displaying, via the display generation component, a communication user interface for communicating with other users in a real-time communication session, wherein during the real-time communication session, a user of the computer system is represented by an avatar that moves during the real-time communication session in accordance with movement of the user of the computer system detected by the one or more sensors; displaying, via the display generation component, selectable user interface objects while displaying the communication user interface; detecting, via the one or more sensors, one or more inputs including a selection input directed to the selectable user interface object; and in response to detecting the one or more inputs including the selection input directed to the selectable user interface object, simultaneously displaying, via the display generation component, an avatar editing user interface that includes: an avatar representing the user of the computer system; and one or more options for modifying an appearance of the avatar representing the user of the computer system.

[0013] An example method includes: at a computer system in communication with a display generation component: while participating in a communication session that is a spatial communication session, where the spatial communication session includes displaying representations of multiple participants in the communication session in a spatially distributed manner in a 3D environment via the display generation component, and where displaying the multiple participants in the spatially distributed manner includes displaying: the representations of the multiple participants spaced apart from each other and from a user of the computer system by at least a threshold amount in a first non-vertical direction in the 3D environment; and the representations of the multiple participants spaced apart from each other and from the user by at least the threshold amount in a second non-vertical direction different from the first non-vertical direction; while displaying the representations of the multiple participants distributed in the 3D environment, detecting an event; and in response to detecting the event, transitioning the communication session from the spatial communication session to a non-spatial communication session, the non-spatial communication session including displaying representations of at least one subset of the multiple participants in the communication session in a grouped arrangement via the display generation component, where in the grouped arrangement: the representations of the multiple participants are spaced apart from each other by less than the threshold amount in the first non-vertical direction in the 3D environment; the representation of a first participant in the grouped arrangement has a different positioning from the representation of the first participant in the spatially distributed arrangement; and the representation of a second participant in the grouped arrangement has a different positioning from the representation of the second participant in the spatially distributed arrangement.

[0014] An example method includes: at a computer system in communication with a display generation component and one or more sensors: while in a communication session with one or more participants in the communication session, detecting a gaze input of a user of the computer system via the one or more sensors; and in response to detecting the gaze input: displaying information about a first participant in the communication session via the display generation component based on a determination that the gaze input meets a set of one or more gaze criteria; and withholding display of the information about the first participant in the communication session based on a determination that the gaze input does not meet the set of one or more gaze criteria.

[0015] This document describes an example non-transitory computer-readable storage medium. An example non-transitory computer-readable storage medium stores one or more programs that are configured to be executed by one or more processors of a computer system that communicates with a display generation component and one or more sensors and includes instructions for: displaying, via the display generation component, representations of multiple users; receiving, via the one or more sensors, a selection of a representation of a corresponding user among the multiple users; and in response to receiving the selection of the representation of the corresponding user: displaying, via the display generation component, an option to invite the corresponding user to join an ongoing communication session based on determining that an ongoing communication session exists; and refraining from displaying the option to invite the corresponding user to join the ongoing communication session based on determining that no ongoing communication session exists.

[0016] An example non-transitory computer-readable storage medium stores one or more programs that are configured to be executed by one or more processors of a computer system that communicates with a display generation component and one or more sensors and includes instructions for: displaying, via the display generation component, a communication user interface for communicating with other users in a real-time communication session, wherein during the real-time communication session, a user of the computer system is represented by an avatar that moves during the real-time communication session in accordance with movement of the user of the computer system detected by the one or more sensors; displaying, via the display generation component, selectable user interface objects while displaying the communication user interface; detecting, via the one or more sensors, one or more inputs including a selection input that points to the selectable user interface object; and in response to detecting the one or more inputs including the selection input that points to the selectable user interface object, simultaneously displaying, via the display generation component, an avatar editing user interface that includes: the avatar representing the user of the computer system; and one or more options for modifying an appearance of the avatar representing the user of the computer system.

[0017] An example non-transitory computer-readable storage medium stores one or more programs that are configured to be executed by one or more processors of a computer system that communicates with a display generation component and that include instructions for: when participating in a communication session that is a spatial communication session, where the spatial communication session includes displaying representations of a plurality of participants in the communication session in a spatially distributed manner in a 3D environment via the display generation component, where displaying the plurality of participants in the spatially distributed manner includes displaying: the representations of the plurality of participants that are spaced apart from each other and from a user of the computer system by at least a threshold amount in a first non-vertical direction in the 3D environment; and the representations of the plurality of participants that are spaced apart from each other and from the user by at least the threshold amount in a second non-vertical direction that is different from the first non-vertical direction; when displaying the representations of the plurality of participants distributed in the 3D environment, detecting an event; and in response to detecting the event, transitioning the communication session from the spatial communication session to a non-spatial communication session, the non-spatial communication session including displaying representations of at least a subset of the plurality of participants in the communication session in a grouped arrangement via the display generation component, where in the grouped arrangement: the representations of the plurality of participants are spaced apart from each other by less than the threshold amount in the first non-vertical direction in the 3D environment; the representation of a first participant in the grouped arrangement has a different positioning from the representation of the first participant in the spatially distributed arrangement; and the representation of a second participant in the grouped arrangement has a different positioning from the representation of the second participant in the spatially distributed arrangement.

[0018] An example non-transitory computer-readable storage medium stores one or more programs that are configured to be executed by one or more processors of a computer system that communicates with a display generation component and one or more sensors and that include instructions for: when in a communication session with one or more participants in the communication session, detecting a gaze input of a user of the computer system via the one or more sensors; and in response to detecting the gaze input: displaying information about a first participant in the communication session via the display generation component based on a determination that the gaze input meets a set of one or more gaze criteria; and withholding display of the information about the first participant in the communication session based on a determination that the gaze input does not meet the set of one or more gaze criteria.

[0019] This document describes an example transient computer-readable storage medium. An example transient computer-readable storage medium stores one or more programs that are configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more sensors, and includes instructions for: displaying, via the display generation component, representations of multiple users; receiving, via the one or more sensors, a selection of a representation of a corresponding user among the multiple users; and in response to receiving the selection of the representation of the corresponding user: displaying, via the display generation component, an option to invite the corresponding user to join an ongoing communication session based on a determination that an ongoing communication session exists; and refraining from displaying the option to invite the corresponding user to join the ongoing communication session based on a determination that no ongoing communication session exists.

[0020] An example transient computer-readable storage medium stores one or more programs that are configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more sensors and includes instructions for: displaying, via the display generation component, a communication user interface for communicating with other users in a real-time communication session, wherein during the real-time communication session, a user of the computer system is represented by an avatar that moves during the real-time communication session in accordance with movements of the user of the computer system detected by the one or more sensors; displaying, via the display generation component, selectable user interface objects when displaying the communication user interface; detecting, via the one or more sensors, one or more inputs including a selection input pointing to the selectable user interface object; and in response to detecting the one or more inputs including the selection input pointing to the selectable user interface object, simultaneously displaying, via the display generation component, an avatar editing user interface that includes: an avatar representing the user of the computer system; and one or more options for modifying an appearance of the avatar representing the user of the computer system.

[0021] An example transient computer-readable storage medium stores one or more programs that are configured to be executed by one or more processors of a computer system that communicates with a display generation component and includes instructions for: when participating in a communication session that is a spatial communication session, where the spatial communication session includes displaying representations of multiple participants in the communication session in a spatially distributed manner in a 3D environment via the display generation component, and where displaying the multiple participants in the spatially distributed manner includes displaying: the representations of the multiple participants that are spaced apart from each other and from a user of the computer system by at least a threshold amount in a first non-vertical direction in the 3D environment; and the representations of the multiple participants that are spaced apart from each other and from the user by at least the threshold amount in a second non-vertical direction different from the first non-vertical direction; when displaying the representations of the multiple participants distributed in the 3D environment, detecting an event; and in response to detecting the event, transitioning the communication session from the spatial communication session to a non-spatial communication session, the non-spatial communication session including displaying representations of at least one subset of the multiple participants in the communication session in a grouped arrangement via the display generation component, where in the grouped arrangement: the representations of the multiple participants are spaced apart from each other by less than the threshold amount in the first non-vertical direction in the 3D environment; the representation of a first participant in the grouped arrangement has a different positioning from the representation of the first participant in the spatially distributed arrangement; and the representation of a second participant in the grouped arrangement has a different positioning from the representation of the second participant in the spatially distributed arrangement.

[0022] An example transient computer-readable storage medium stores one or more programs that are configured to be executed by one or more processors of a computer system that communicates with a display generation component and one or more sensors and includes instructions for: when in a communication session with one or more participants in the communication session, detecting a gaze input of a user of the computer system via the one or more sensors; and in response to detecting the gaze input: displaying information about a first participant in the communication session via the display generation component based on a determination that the gaze input meets a set of one or more gaze criteria; and withholding display of the information about the first participant in the communication session based on a determination that the gaze input does not meet the set of one or more gaze criteria.

[0023] This document describes an example computer system. An example computer system is configured to communicate with a display generation component and one or more sensors and includes: one or more processors; and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying representations of multiple users via the display generation component; receiving a selection of a representation of a corresponding user among the multiple users via the one or more sensors; and in response to receiving the selection of the representation of the corresponding user: displaying, via the display generation component, an option to invite the corresponding user to join an ongoing communication session based on a determination that an ongoing communication session exists; and refraining from displaying the option to invite the corresponding user to join the ongoing communication session based on a determination that no ongoing communication session exists.

[0024] An example computer system is configured to communicate with a display generation component and one or more sensors and includes: one or more processors; and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying a communication user interface for communicating with other users in a real-time communication session via the display generation component, wherein during the real-time communication session, a user of the computer system is represented by an avatar that moves during the real-time communication session in accordance with movement of the user of the computer system detected by the one or more sensors; displaying selectable user interface objects via the display generation component when displaying the communication user interface; detecting, via the one or more sensors, one or more inputs including a selection input directed to the selectable user interface object; and in response to detecting the one or more inputs including the selection input directed to the selectable user interface object, simultaneously displaying, via the display generation component, an avatar editing user interface that includes: the avatar representing the user of the computer system; and one or more options for modifying an appearance of the avatar representing the user of the computer system.

[0025] An example computer system is configured to communicate with a display generation component and includes: one or more processors; and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: when participating in a communication session that is a spatial communication session, where the spatial communication session includes displaying representations of multiple participants in the communication session in a spatially distributed manner in a 3D environment via the display generation component, and where displaying the multiple participants in a spatially distributed manner includes displaying: representations of the multiple participants spaced apart from each other and from a user of the computer system by at least a threshold amount in a first non-vertical direction in the 3D environment; and representations of the multiple participants spaced apart from each other and from the user by at least the threshold amount in a second non-vertical direction different from the first non-vertical direction; when displaying the representations of the multiple participants distributed in the 3D environment, detecting an event; and in response to detecting the event, transitioning the communication session from the spatial communication session to a non-spatial communication session, the non-spatial communication session including displaying representations of at least one subset of the multiple participants in the communication session in a grouped arrangement via the display generation component, where in the grouped arrangement: the representations of the multiple participants are spaced apart from each other by less than the threshold amount in the first non-vertical direction in the 3D environment; the representation of a first participant in the grouped arrangement has a different positioning from the representation of the first participant in the spatially distributed arrangement; and the representation of a second participant in the grouped arrangement has a different positioning from the representation of the second participant in the spatially distributed arrangement.

[0026] An example computer system is configured to communicate with a display generation component and one or more sensors and includes: one or more processors; and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: when in a communication session with one or more participants in the communication session, detecting a gaze input of a user of the computer system via the one or more sensors; and in response to detecting the gaze input: displaying information about a first participant in the communication session via the display generation component based on a determination that the gaze input meets a set of one or more gaze criteria; and withholding display of the information about the first participant in the communication session based on a determination that the gaze input does not meet the set of one or more gaze criteria.

[0027] An example computer system is configured to communicate with a display generation component and one or more sensors and includes: means for displaying representations of a plurality of users via the display generation component; means for receiving a selection of a representation of a corresponding user among the plurality of users via the one or more sensors; and means for, in response to receiving the selection of the representation of the corresponding user, performing the following: displaying, via the display generation component, an option to invite the corresponding user to join the ongoing communication session based on a determination that an ongoing communication session exists; and refraining from displaying the option to invite the corresponding user to join the ongoing communication session based on a determination that no ongoing communication session exists.

[0028] An example computer system is configured to communicate with a display generation component and one or more sensors and includes: means for displaying a communication user interface for communicating with other users in a real-time communication session via the display generation component, wherein during the real-time communication session, a user of the computer system is represented by an avatar that moves during the real-time communication session in accordance with movement of the user of the computer system detected by the one or more sensors; means for displaying selectable user interface objects via the display generation component when displaying the communication user interface; means for detecting, via the one or more sensors, one or more inputs including a selection input directed to the selectable user interface object; and means for, in response to detecting the one or more inputs including the selection input directed to the selectable user interface object, performing the following: simultaneously displaying, via the display generation component, an avatar editing user interface that includes: a representation of the avatar of the user of the computer system; and one or more options for modifying an appearance of the avatar of the user of the computer system.

[0029] An example computer system is configured to communicate with a display generation component and includes: means for, when participating in a communication session that is a spatial communication session, where the spatial communication session includes displaying representations of a plurality of participants in the communication session in a spatially distributed manner in a 3D environment via the display generation component, where displaying the plurality of participants in the spatially distributed manner includes displaying: the representations of the plurality of participants spaced apart from each other and from a user of the computer system by at least a threshold amount in a first non-vertical direction in the 3D environment; and the representations of the plurality of participants spaced apart from each other and from the user by at least the threshold amount in a second non-vertical direction different from the first non-vertical direction; means for, when displaying the representations of the plurality of participants distributed in the 3D environment, detecting an event; and means for, in response to detecting the event, transitioning the communication session from the spatial communication session to a non-spatial communication session, the non-spatial communication session including displaying representations of at least a subset of the plurality of participants in the communication session in a grouped arrangement via the display generation component, where in the grouped arrangement: the representations of the plurality of participants are spaced apart from each other by less than the threshold amount in the first non-vertical direction in the 3D environment; the representation of a first participant in the grouped arrangement has a different positioning from the representation of the first participant in the spatially distributed arrangement; and the representation of a second participant in the grouped arrangement has a different positioning from the representation of the second participant in the spatially distributed arrangement.

[0030] An example computer system is configured to communicate with a display generation component and one or more sensors and includes: means for, when in a communication session with one or more participants in the communication session, detecting a gaze input of a user of the computer system via the one or more sensors; and means for, in response to detecting the gaze input, performing the following: displaying information about a first participant in the communication session via the display generation component based on a determination that the gaze input meets a set of one or more gaze criteria; and withholding display of the information about the first participant in the communication session based on a determination that the gaze input does not meet the set of one or more gaze criteria.

[0031] This document describes an example computer program product. An example computer program product includes one or more programs that are configured to be executed by one or more processors of a computer system that communicates with a display generation component and one or more sensors. The one or more programs include instructions for: displaying representations of multiple users via the display generation component; receiving a selection of a representation of a corresponding user among the multiple users via the one or more sensors; and in response to receiving the selection of the representation of the corresponding user: displaying, via the display generation component, an option to invite the corresponding user to join an ongoing communication session based on a determination that an ongoing communication session exists; and foregoing displaying the option to invite the corresponding user to join the ongoing communication session based on a determination that no ongoing communication session exists.

[0032] An example computer program product includes one or more programs that are configured to be executed by one or more processors of a computer system that communicates with a display generation component and one or more sensors. The one or more programs include instructions for: displaying, via the display generation component, a communication user interface for communicating with other users in a real-time communication session, where during the real-time communication session, a user of the computer system is represented by an avatar that moves during the real-time communication session based on movement of the user of the computer system detected by the one or more sensors; displaying, via the display generation component, selectable user interface objects when displaying the communication user interface; detecting, via the one or more sensors, one or more inputs including a selection input that points to the selectable user interface object; and in response to detecting the one or more inputs including the selection input that points to the selectable user interface object, simultaneously displaying, via the display generation component, an avatar editing user interface that includes: the avatar representing the user of the computer system; and one or more options for modifying an appearance of the avatar representing the user of the computer system.

[0033] An example computer program product includes one or more programs that are configured to be executed by one or more processors of a computer system that communicates with a display generation component. The one or more programs include instructions for: when participating in a communication session that is a spatial communication session, where the spatial communication session includes displaying representations of a plurality of participants in the communication session in a spatially distributed manner in a 3D environment via the display generation component, and where displaying the plurality of participants in the spatially distributed manner includes displaying: the representations of the plurality of participants that are spaced apart from each other and from a user of the computer system by at least a threshold amount in a first non-vertical direction in the 3D environment; and the representations of the plurality of participants that are spaced apart from each other and from the user by at least the threshold amount in a second non-vertical direction different from the first non-vertical direction; when displaying the representations of the plurality of participants distributed in the 3D environment, detecting an event; and in response to detecting the event, transitioning the communication session from the spatial communication session to a non-spatial communication session, the non-spatial communication session including displaying representations of at least one subset of the plurality of participants in the communication session in a grouped arrangement via the display generation component, where in the grouped arrangement: the representations of the plurality of participants are spaced apart from each other by less than the threshold amount in the first non-vertical direction in the 3D environment; the representation of a first participant in the grouped arrangement has a different positioning from the representation of the first participant in the spatially distributed arrangement; and the representation of a second participant in the grouped arrangement has a different positioning from the representation of the second participant in the spatially distributed arrangement.

[0034] An example computer program product includes one or more programs that are configured to be executed by one or more processors of a computer system that communicates with a display generation component and one or more sensors. The one or more programs include instructions for: when in a communication session with one or more participants in the communication session, detecting a gaze input of a user of the computer system via the one or more sensors; and in response to detecting the gaze input: displaying information about a first participant in the communication session via the display generation component based on a determination that the gaze input meets a set of one or more gaze criteria; and refraining from displaying the information about the first participant in the communication session based on a determination that the gaze input does not meet the set of one or more gaze criteria.

[0035] Note that the various embodiments described above can be combined with any other embodiments described herein. The features and advantages described in this specification are not exhaustive, and in particular, many additional features and advantages will be apparent to those of ordinary skill in the art from the accompanying drawings, the specification, and the claims. Additionally, it should be noted that the language used in this specification has been selected for readability and guidance purposes and may not have been selected to depict or define the subject matter of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] To better understand the various embodiments, reference should be made to the following detailed description in conjunction with the accompanying drawings, in which like reference numerals refer to corresponding parts in all the figures.

[0037] Figure 1A is a block diagram illustrating an operating environment of a computer system for providing an XR experience according to some embodiments.

[0038] Figures 1B to 1P is for providing an example of a computer system for an XR experience in the Figure 1A operating environment.

[0039] Figure 2 is a block diagram illustrating a controller of a computer system configured to manage and coordinate a user's XR experience according to some embodiments.

[0040] Figure 3 is a block diagram illustrating a display generation component of a computer system configured to provide a visual component of an XR experience to a user according to some embodiments.

[0041] Figure 4 is a block diagram illustrating a hand tracking unit of a computer system configured to capture a user's gesture input according to some embodiments.

[0042] Figure 5 is a block diagram illustrating an eye tracking unit of a computer system configured to capture a user's gaze input according to some embodiments.

[0043] Figure 6 is a flowchart illustrating a flash-assisted gaze tracking pipeline according to some embodiments.

[0044] Figures 7A to 7Q Illustrates an example technique for managing a live communication session according to some embodiments.

[0045] Figure 8 is a flowchart of a method for managing a live communication session according to various embodiments.

[0046] Figure 9 is a flowchart of a method for providing an avatar in a live communication session according to various embodiments.

[0047] Figures 10A to 10E Illustrates an example technique for providing a representation in a live communication session according to some embodiments.

[0048] Figure 11 is a flowchart of a method for providing a representation in a live communication session according to various embodiments.

[0049] Figures 12A to 12F Illustrative example techniques for providing information in a live communication session are presented in accordance with some embodiments.

[0050] Figure 13 is a flowchart of a method for providing information in a live communication session in accordance with various embodiments. DETAILED DESCRIPTION

[0051] In accordance with some embodiments, the present disclosure relates to a user interface for providing an extended reality (XR) experience to a user.

[0052] The systems, methods, and GUIs described herein improve user interface interactions with virtual / augmented reality environments in a variety of ways.

[0053] In some embodiments, a computer system permits live communication between users. The computer system displays representations of multiple users and receives a selection of a representation of a corresponding user among the multiple users. In response to receiving the selection of the representation of the corresponding user, the computer system displays an option to invite the corresponding user to join an ongoing communication session if it is determined that an ongoing communication session exists, and the computer system refrains from displaying the option to invite the corresponding user to join the ongoing communication session if it is determined that no ongoing communication session exists.

[0054] In some embodiments, the computer system provides an option for a user to change the appearance of the user's avatar. The computer system displays a communication user interface for communicating with other users during a live communication session. During the live communication session, the user is represented by an avatar that moves in accordance with the movement of the user of the computer system during the live communication session. When the communication user interface is displayed, the computer system simultaneously displays selectable user interface objects. When the communication user interface and the selectable user interface objects are simultaneously displayed, the computer system detects one or more inputs including a selection input that points to a selectable user interface object. In response to detecting the one or more inputs including the selection input that points to the selectable user interface object, the computer system simultaneously displays an avatar editing user interface that includes an avatar representing the user of the computer system and one or more options for modifying the appearance of the avatar representing the user of the computer system.

[0055] In some embodiments, a computer system switches between a spatial communication session and a non-spatial communication session. When participating in a communication session that is a spatial communication session, where the spatial communication session includes the computer system displaying representations of multiple participants in the communication session in a spatially distributed arrangement in a 3D environment. Displaying the multiple participants in a spatially distributed arrangement includes displaying representations of multiple participants that are spaced apart from each other and from the user by at least a threshold amount in a first non-vertical direction in the 3D environment, and representations of multiple participants that are spaced apart from each other and from the user by at least a threshold amount in a second non-vertical direction different from the first non-vertical direction. When displaying representations of multiple participants distributed in the 3D environment, the computer system detects an event, and in response to detecting the event, the computer system transitions the communication session from the spatial communication session to a non-spatial communication session. Transitioning to the non-spatial communication session includes displaying representations of at least one subset of the multiple participants in the communication session in a grouped arrangement. In the grouped arrangement, the representations of the multiple participants are spaced apart from each other by less than the threshold amount in the first non-vertical direction in the 3D environment, the representation of the first participant in the grouped arrangement has a different positioning from the representation of the first participant in the spatially distributed arrangement, and the representation of the second participant in the grouped arrangement has a different positioning from the representation of the second participant in the spatially distributed arrangement.

[0056] In some embodiments, a computer system provides information during a live communication session based on a user's gaze. While in a communication session with one or more participants in the communication session, the computer system detects a gaze input of a user of the computer system. In response to detecting the gaze input, based on a determination that the gaze input meets a set of one or more gaze criteria, the computer system displays information about a first participant in the communication session, and based on a determination that the gaze input does not meet the set of one or more gaze criteria, the computer system refrains from displaying information about the first participant in the communication session.

[0057] In some embodiments, a computer system displays content in a first region of a user interface. In some embodiments, while the computer system is displaying the content and while a first set of controls is not displayed in a first state, the computer system detects a first input from a first portion of the user. In some embodiments, in response to detecting the first input, and based on a determination that the user's gaze points to a second region of the user interface when the first input is detected, the computer system displays the first set of one or more controls in the first state in the user interface, and based on a determination that the user's gaze does not point to the second region of the user interface when the first input is detected, the computer system refrains from displaying the first set of one or more controls in the first state.

[0058] In some embodiments, a computer system displays content in a user interface. In some embodiments, while displaying the content, the computer system detects a first input based on movement of a first portion of a user of the computer system. In some embodiments, in response to detecting the first input, the computer system displays a first set of one or more controls in the user interface, wherein the first set of one or more controls is displayed in a first state and within a first region of the user interface. In some embodiments, while displaying the first set of one or more controls in the first state: based on a determination that one or more first criteria are satisfied, the one or more first criteria including criteria satisfied when directing a user's attention to the first region of the user interface based on movement of a second portion of the user that is different from the first portion of the user, the computer system transitions from displaying the first set of one or more controls in the first state to displaying a second set of one or more controls in a second state, wherein the second state is different from the first state.

[0059] Figures 1A to 6 A description of an example computer system for providing an XR experience to a user is provided. Figures 7A to 7Q Example techniques for managing a live communication session in accordance with some embodiments are illustrated. Figure 8 Is a flowchart of a method for managing a live communication session in accordance with various embodiments. Figure 9 Is a flowchart of a method for providing an avatar in a live communication session in accordance with various embodiments. Figures 7A to 7Q The user interface in is used to illustrate Figure 8 and Figure 9 the processes in. Figures 10A to 10E Example techniques for providing a representation in a live communication session in accordance with some embodiments are illustrated. Figure 11 Is a flowchart of a method for providing a representation in a live communication session in accordance with various embodiments.

[0060] Figures 10A to 10E The user interface in is used to illustrate Figure 11 the processes in. Figures 12A to 12F Example techniques for providing information in a live communication session in accordance with some embodiments are illustrated. Figure 13 Is a flowchart of a method for providing information in a live communication session in accordance with various embodiments. Figures 10A to 10E The user interface in is used to illustrate Figure 11 the processes in.

[0061] The processes described below enhance the operability of a device and make the user-device interface more efficient through various techniques (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), including by providing the user with improved visual feedback, reducing the number of inputs required to perform an operation, providing additional control options without cluttering the user interface with additional display controls, performing an operation when a set of conditions has been met without further user input, improving privacy and / or security, providing a richer, more detailed, and / or more realistic user experience while saving storage space, and / or additional techniques. These techniques also reduce power usage and extend the battery life of the device by enabling the user to use the device faster and more effectively. Saving battery power, and thus weight, improves the ergonomics of the device. These techniques also enable real-time communication, allow for the use of fewer and / or less precise sensors, resulting in a more compact, lighter, and cheaper device, and enable the device to be used in various lighting conditions. These techniques reduce energy usage, thereby reducing the heat emitted by the device, which is particularly important for wearable devices where it can become uncomfortable for the user to wear the device if the device generates too much heat while operating entirely within the operating parameters of the device components.

[0062] In addition, in methods described herein where one or more steps depend on one or more conditions having been met, it should be understood that the method may be repeated in multiple iterations such that, during the course of the repetition, all conditions that determine the steps in the method have been met in different iterations of the method. For example, if a method requires performing a first step (if a condition is met) and performing a second step (if the condition is not met), one of ordinary skill in the art will understand that the stated steps are repeated until both the condition being met and the condition not being met (in no particular order) have occurred. Thus, a method described as having one or more steps that depend on one or more conditions having been met can be rewritten as a method that repeats until each condition described in the method has been met. However, this does not require a system or computer-readable medium to state that the system or computer-readable medium includes instructions for performing conditional operations based on the satisfaction of corresponding one or more conditions and is thus capable of determining whether a possible scenario has been met without explicitly repeating the steps of the method until all conditions that determine the steps in the method have been met. One of ordinary skill in the art will also understand that, similar to a method with conditional steps, a system or computer-readable storage medium can repeat the steps of the method as many times as needed to ensure that all conditional steps have been performed.

[0063] In some embodiments, as Figure 1AAs shown, an XR experience is provided to a user via an operating environment 100 that includes a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted device (HMD), a display, a projector, a touch screen, etc.), one or more input devices 125 (e.g., an eye tracking device 130, a hand tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a haptic output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a haptic sensor, an orientation sensor, a proximity sensor, a temperature sensor, a position sensor, a motion sensor, a speed sensor, etc.), and optionally one or more peripheral devices 195 (e.g., home appliances, wearable devices, etc.). In some embodiments, one or more of the input device 125, the output device 155, the sensor 190, and the peripheral device 195 are integrated with the display generation component 120 (e.g., in a head-mounted device or a handheld device).

[0064] When describing an XR experience, various terms are used to distinctively refer to several related but different environments that a user can sense and / or interact with (e.g., interact using inputs detected by the computer system 101 that generates the XR experience, where these inputs cause the computer system generating the XR experience to generate audio, visual, and / or haptic feedback corresponding to the various inputs provided to the computer system 101). The following is a subset of these terms:

[0065] Physical environment: The physical environment refers to the physical world that people can sense and / or interact with without the help of an electronic system. Physical environments such as a physical park include physical objects such as physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment, such as through vision, touch, hearing, taste, and smell.

[0066] Extended Reality: In contrast, an extended reality (XR) environment refers to a fully or partially simulated environment in which people sense and / or interact via an electronic system. In XR, a subset of a person's physical movements or their representations are tracked, and in response, one or more characteristics of one or more virtual objects simulated in the XR environment are adjusted in a manner that complies with at least one physical law. For example, an XR system can detect a person's head rotation, and in response, adjust the graphical content and sound field presented to the person in a manner similar to how such views and sounds would change in a physical environment. In some cases (e.g., for accessibility reasons), the adjustment of the characteristics of virtual objects in the XR environment can be made in response to a representation of a physical movement (e.g., a voice command). A person can use any of their senses to sense and / or interact with XR objects, including vision, hearing, touch, taste, and smell. For example, a person can sense and / or interact with an audio object that creates a 3D or spatial audio environment that provides the perception of point audio sources in 3D space. Also, an audio object can implement audio transparency that selectively introduces ambient sounds from the physical environment with or without computer-generated audio. In some XR environments, a person can sense and / or interact only with audio objects.

[0067] Examples of XR include virtual reality and mixed reality.

[0068] Virtual Reality: A virtual reality (VR) environment refers to a simulated environment that is designed to be completely computer-generated sensory input for one or more senses. A VR environment includes multiple virtual objects with which a person can sense and / or interact. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can sense and / or interact with the virtual objects in the VR environment by a simulation of the person's presence within the computer-generated environment and / or by a simulation of a subset of the person's physical movements within the computer-generated environment.

[0069] Mixed Reality: Compared to a VR environment that is designed to be completely computer-generated sensory input, a mixed reality (MR) environment refers to a simulated environment that is designed to incorporate sensory input or its representation from the physical environment in addition to including computer-generated sensory input (e.g., virtual objects). On the virtual continuum, a mixed reality environment is any condition between a fully physical environment at one end and a virtual reality environment at the other end, but excluding these two ends. In some MR environments, the computer-generated sensory input can respond to changes in the sensory input from the physical environment. Additionally, some electronic systems for presenting an MR environment can track the position and / or orientation relative to the physical environment so that virtual objects can interact with real objects (i.e., physical items from the physical environment or their representations). For example, the system can cause movement such that a virtual tree appears stationary relative to the physical ground.

[0070] Examples of mixed reality include augmented reality and augmented virtuality.

[0071] Augmented Reality: An augmented reality (AR) environment is a simulated environment in which one or more virtual objects are superimposed over a physical environment or a representation of a physical environment. For example, an electronic system for presenting an AR environment may have a transparent or translucent display through which a person can directly view the physical environment. The system can be configured to present virtual objects on the transparent or translucent display such that the person, using the system, perceives the virtual objects superimposed over the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture images or video of the physical environment, which are representations of the physical environment. The system combines the images or video with the virtual objects and presents the combination on the opaque display. The person, using the system, indirectly views the physical environment via the images or video of the physical environment and perceives the virtual objects superimposed over the physical environment. As used herein, a video of a physical environment displayed on an opaque display is referred to as a "passthrough video," meaning that the system uses one or more image sensors to capture images of the physical environment and uses those images when presenting the AR environment on the opaque display. Further alternatively, the system may have a projection system that projects virtual objects into the physical environment, such as as a hologram or on a physical surface, such that the person, using the system, perceives the virtual objects superimposed over the physical environment. An augmented reality environment is also a simulated environment in which a representation of a physical environment is transformed by computer-generated sensory information. For example, in providing a passthrough video, the system may transform one or more sensor images to impose an alternative perspective (e.g., viewpoint) different from the perspective captured by the imaging sensors. As another example, a representation of a physical environment may be transformed by graphically modifying (e.g., magnifying) portions thereof such that the modified portions may be representative but not a photorealistic version of the original captured image. As yet another example, a representation of a physical environment may be transformed by graphically eliminating portions thereof or blurring portions thereof.

[0072] Augmented Virtuality: An augmented virtuality (AV) environment is a simulated environment in which a virtual environment or computer-generated environment incorporates one or more sensory inputs from a physical environment. The sensory input may be a representation of one or more characteristics of the physical environment. For example, an AV park may have virtual trees and virtual buildings, but a person's face is a photorealistic reproduction of an image of a physical person. As another example, a virtual object may adopt the shape or color of a physical item imaged by one or more imaging sensors. As yet another example, a virtual object may adopt a shadow that conforms to the positioning of the sun in the physical environment.

[0073] In an augmented reality, mixed reality, or virtual reality environment, a view of a three-dimensional environment is visible to a user. The view of the three-dimensional environment is typically visible to the user through a virtual viewport via one or more display generation components (e.g., a display or a pair of display modules that provide stereoscopic content to different eyes of the same user), the virtual viewport having a viewport boundary that defines the extent of the three-dimensional environment visible to the user via the one or more display generation components. In some embodiments, the region defined by the viewport boundary is smaller than the user's visual field in one or more dimensions (e.g., based on the user's visual field, the size, optical properties, or other physical characteristics of the one or more display generation components, and / or the position and / or orientation of the one or more display generation components relative to the user's eyes). In some embodiments, the region defined by the viewport boundary is larger than the user's visual field in one or more dimensions (e.g., based on the user's visual field, the size, optical properties, or other physical characteristics of the one or more display generation components, and / or the position and / or orientation of the one or more display generation components relative to the user's eyes). The viewport and the viewport boundary typically move as the one or more display generation components move (e.g., for a head-mounted device as the user's head moves, or for a handheld device such as a tablet or smartphone as the user's hand moves). The user's viewpoint determines the content visible in the viewport, the viewpoint typically specifying a position and orientation relative to the three-dimensional environment, and as the viewpoint shifts, the view of the three-dimensional environment will also shift in the viewport. For a head-mounted device, the viewpoint is typically based on the position and orientation of the user's head, face, and / or eyes to provide a perceptually accurate view of the three-dimensional environment and an immersive experience while the user is using the head-mounted device. For a handheld or stationary device, the viewpoint shifts as the handheld or stationary device moves and / or as the user's positioning relative to the handheld or stationary device changes (e.g., the user moves towards, away from, up, down, right, and / or left). For a device that includes a display generation component with virtual passthrough, the portions of the physical environment visible (e.g., displayed and / or projected) via the one or more display generation components are based on the field of view of one or more cameras in communication with the display generation component, the one or more cameras typically moving as the display generation component moves (e.g., for a head-mounted device as the user's head moves, or for a handheld device such as a tablet or smartphone as the user's hand moves), because the user's viewpoint moves as the field of view of the one or more cameras moves (and the appearance of one or more virtual objects displayed via the one or more display generation components is updated based on the user's viewpoint (e.g., the display position and pose of the virtual object are updated based on the movement of the user's viewpoint)).For a display generation component with optical passthrough, portions of the physical environment that are visible via one or more display generation components (e.g., optically visible through one or more partial or fully transparent portions of the display generation component) are based on the user's field of view through the partial or fully transparent portion of the display generation component (e.g., for a head-mounted device, moving as the user's head moves, or for a handheld device such as a tablet or smartphone, moving as the user's hand moves), because the user's viewpoint moves as the user's field of view through the partial or fully transparent portion of the display generation component moves (and the appearance of one or more virtual objects is updated based on the user's viewpoint).

[0074] In some embodiments, the representation of the physical environment (e.g., via virtual passthrough or optical passthrough display) may be partially or fully occluded by the virtual environment. In some embodiments, the amount of the virtual environment displayed (e.g., the amount of the physical environment not displayed) is based on the immersion level of the virtual environment (e.g., relative to the representation of the physical environment). For example, increasing the immersion level optionally causes more of the virtual environment to be displayed, replacing and / or occluding more of the physical environment, and decreasing the immersion level optionally causes less of the virtual environment to be displayed, thereby revealing portions of the physical environment that were previously not displayed and / or occluded. In some embodiments, at a particular immersion level, one or more first background objects (e.g., in the representation of the physical environment) are visually de-emphasized (e.g., dimmed, blurred, displayed with increased transparency) more than one or more second background objects, and one or more third background objects cease to be displayed. In some embodiments, the immersion level includes the associated degree to which virtual content (e.g., virtual environment and / or virtual content) displayed by a computer system occludes background content (e.g., content other than the virtual environment and / or virtual content) around / behind the virtual environment, optionally including the number of items of the displayed background content and / or the displayed visual characteristics (e.g., color, contrast, and / or opacity) of the background content, the angular range of the virtual content displayed by a display generation component (e.g., 60 degrees for content displayed at low immersion, 120 degrees for content displayed at medium immersion, or 180 degrees for content displayed at high immersion), and / or the proportion of the field of view displayed by the display generation component occupied by the virtual content (e.g., 33% of the field of view occupied by the virtual content at low immersion, 66% of the field of view occupied by the virtual content at medium immersion, or 100% of the field of view occupied by the virtual content at high immersion). In some embodiments, the background content is included in the background on which the virtual content is displayed (e.g., background content in the representation of the physical environment). In some embodiments, the background content includes a user interface (e.g., a user interface corresponding to an application generated by a computer system), virtual objects not associated with and / or not included in the virtual environment and / or virtual content (e.g., files or representations of other users generated by a computer system, etc.), and / or real-world objects (e.g., passthrough objects representing real-world objects in the physical environment around the user, these passthrough objects being visible such that they are displayed by the display generation component and / or visible via a transparent or translucent component of the display generation component because the computer system does not occlude / hinder their visibility through the display generation component). In some embodiments, at a low immersion level (e.g., a first immersion level), the background, virtual, and / or real-world objects are displayed in a non-occluded manner. For example, a virtual environment with a low immersion level is optionally displayed simultaneously with background content, which is optionally displayed at full brightness, color, and / or semi-transparency.In some embodiments, at a higher immersion level (e.g., a second immersion level higher than the first immersion level), background, virtual, and / or real objects are displayed in an occluded manner (e.g., dimmed, blurred, or removed from the display). For example, a corresponding virtual environment with a high immersion level is displayed without simultaneously displaying background content (e.g., in full screen or full immersion mode). As another example, a virtual environment displayed at a medium immersion level is displayed simultaneously with background content that is dimmed, blurred, or otherwise de-emphasized. In some embodiments, the visual characteristics of background objects vary among the background objects. For example, at a particular immersion level, one or more first background objects are more visually de-emphasized (e.g., dimmed, blurred, and / or displayed with increased transparency) than one or more second background objects, and one or more third background objects cease to be displayed. In some embodiments, zero immersion or a zero immersion level corresponds to a virtual environment that ceases to be displayed, and instead a representation of the physical environment is displayed (optionally with one or more virtual objects, such as applications, windows, or virtual three-dimensional objects), and the representation of the physical environment is not occluded by the virtual environment. Adjusting the immersion level using physical input elements provides a quick and efficient way to adjust the immersion, which enhances the operability of the computer system and makes the user-device interface more efficient.

[0075] Viewpoint-locked virtual objects: When a computer system displays a virtual object at the same location and / or orientation in the user's viewpoint, the virtual object is viewpoint-locked even if the user's viewpoint shifts (e.g., changes). In embodiments in which the computer system is a head-mounted device, the user's viewpoint is locked to the forward direction of the user's head (e.g., when the user looks straight ahead, the user's viewpoint is at least a portion of the user's field of view); thus, without moving the user's head, the user's viewpoint remains fixed even when the user's gaze shifts. In embodiments in which the computer system has a display generation component (e.g., a display screen) that is repositionable relative to the user's head, the user's viewpoint is the augmented reality view presented to the user on the computer system's display generation component. For example, a viewpoint-locked virtual object that is displayed in the upper left corner of the user's viewpoint when the user's viewpoint is in a first orientation (e.g., the user's head is facing north) continues to be displayed in the upper left corner of the user's viewpoint even when the user's viewpoint changes to a second orientation (e.g., the user's head is facing west). In other words, the location and / or orientation at which a viewpoint-locked virtual object is displayed in the user's viewpoint is independent of the user's location and / or orientation in the physical environment. In embodiments in which the computer system is a head-mounted device, the user's viewpoint is locked to the orientation of the user's head such that the virtual object is also referred to as a "head-locked virtual object".

[0076] Environment-Locked Visual Objects: When a computer system displays a virtual object at a location and / or orientation in the user's line of sight, the virtual object is environment-locked (alternatively, "world-locked"), where the location and / or orientation is based on a location and / or object in a three-dimensional environment (e.g., a physical environment or a virtual environment) (e.g., selected and / or anchored with reference to the location and / or object). As the user's line of sight shifts, the location and / or object in the environment relative to the user's line of sight changes, which causes the environment-locked virtual object to be displayed at a different location and / or orientation in the user's line of sight. For example, an environment-locked virtual object locked to a tree directly in front of the user is displayed at the center of the user's line of sight. When the user's line of sight shifts to the right (e.g., the user's head turns to the right) such that the tree is now to the left of center in the user's line of sight (e.g., the location of the tree in the user's line of sight has shifted), the environment-locked virtual object locked to the tree is displayed to the left of center in the user's line of sight. In other words, the location and / or orientation at which the environment-locked virtual object is displayed in the user's line of sight depends on the location and / or orientation of the location and / or object in the environment to which the virtual object is locked. In some embodiments, the computer system uses a stationary reference frame (e.g., a coordinate system anchored to a fixed location and / or object in a physical environment) to determine the orientation at which the environment-locked virtual object is displayed in the user's line of sight. The environment-locked virtual object can be locked to a stationary part of the environment (e.g., the floor, a wall, a table, or other stationary object), or can be locked to a movable part of the environment (e.g., a vehicle, an animal, a person, or even a representation of a part of the user's body such as the user's hand, wrist, arm, or foot that moves independently of the user's line of sight) such that the virtual object moves as the line of sight or that part of the environment moves to maintain a fixed relationship between the virtual object and that part of the environment.

[0077] In some embodiments, an environment-locked or viewpoint-locked virtual object exhibits a lazy follow behavior that reduces or delays the movement of the environment-locked or viewpoint-locked virtual object relative to the movement of a reference point that the virtual object follows. In some embodiments, when exhibiting the lazy follow behavior, when the computer system detects movement of a reference point (e.g., a portion of the environment, a viewpoint, or a point fixed relative to the viewpoint, such as a point between 5 cm and 300 cm from the viewpoint) that the virtual object is following, the computer system intentionally delays the movement of the virtual object. For example, when the reference point (e.g., a portion of the environment or a viewpoint) moves at a first speed, the virtual object is moved by the device to remain locked to the reference point but moves at a second speed that is slower than the first speed (e.g., until the reference point stops moving or slows down, at which point the virtual object begins to catch up to the reference point). In some embodiments, when the virtual object exhibits the lazy follow behavior, the device ignores small movements of the reference point (e.g., ignores movements of the reference point below a threshold amount of movement, such as moving 0 to 5 degrees or moving 0 to 50 cm). For example, when the reference point (e.g., a portion of the environment or a viewpoint to which the virtual object is locked) moves a first amount, the distance between the reference point and the virtual object increases (e.g., because the virtual object is being displayed to remain fixed or substantially fixed relative to a different viewpoint or portion of the environment than the reference point to which the virtual object is locked), and when the reference point (e.g., a portion of the environment or a viewpoint to which the virtual object is locked) moves a second amount that is greater than the first amount, the distance between the reference point and the virtual object initially increases (e.g., because the virtual object is being displayed to remain fixed or substantially fixed relative to a different viewpoint or portion of the environment than the reference point to which the virtual object is locked), and then decreases when the amount of movement of the reference point increases above a threshold (e.g., a "lazy follow" threshold) because the virtual object is moved by the computer system to remain fixed or substantially fixed relative to the reference point. In some embodiments, the virtual object remaining fixed or substantially fixed relative to the reference point includes the virtual object being displayed within a threshold distance (e.g., 1 cm, 2 cm, 3 cm, 5 cm, 15 cm, 20 cm, 50 cm) of the reference point in one or more dimensions (e.g., up / down, left / right, and / or forward / backward relative to the positioning of the reference point).

[0078] Hardware: There are many different types of electronic systems that enable a person to sense and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields integrated with display capabilities, windows integrated with display capabilities, displays formed as lenses designed to be placed on a person's eyes (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablet devices, and desktop / laptop computers. A head-mounted system may include speakers and / or other audio output devices integrated into the head-mounted system for providing audio output. A head-mounted system may have one or more speakers and an integrated opaque display. Alternatively, a head-mounted system may be configured to accept an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing images or video of the physical environment and / or one or more microphones for capturing audio of the physical environment. A head-mounted system may have a transparent or semi-transparent display instead of an opaque display. The transparent or semi-transparent display may have a medium through which light representing an image is directed to a person's eyes. The display may utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light sources, or any combination of these technologies. The medium may be an optical waveguide, holographic medium, optical combiner, optical reflector, or any combination thereof. In one embodiment, the transparent or semi-transparent display may be configured to selectively become opaque. Projection-based systems may employ retinal projection techniques that project graphic images onto a person's retina. The projection system may also be configured to project virtual objects into the physical environment, such as as a hologram or on a physical surface. In some embodiments, the controller 110 is configured to manage and coordinate a user's XR experience. In some embodiments, the controller 110 includes a suitable combination of software, firmware, and / or hardware. Below with respect to Figure 2Controller 110 is described in more detail. In some embodiments, controller 110 is a computing device that is local or remote relative to scene 105 (e.g., a physical environment). For example, controller 110 is a local server located within scene 105. As another example, controller 110 is a remote server (e.g., a cloud server, a central server, etc.) located outside of scene 105. In some embodiments, controller 110 is communicatively coupled to display generation component 120 (e.g., an HMD, a display, a projector, a touch screen, etc.) via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE802.11x, IEEE 802.16x, IEEE 802.3x, etc.). In another example, controller 110 is included within the housing (e.g., a physical enclosure) of one or more of display generation component 120 (e.g., an HMD or a portable electronic device including a display and one or more processors, etc.), input device 125, output device 155, one or more sensors of sensor 190, and / or one or more peripheral devices of peripheral device 195, or shares the same physical housing or support structure with one or more of the foregoing devices.

[0079] In some embodiments, display generation component 120 is configured to provide an XR experience (e.g., at least a visual component of the XR experience) to a user. In some embodiments, display generation component 120 includes a suitable combination of software, firmware, and / or hardware. Display generation component 120 is described in more detail below with respect to Figure 3 In some embodiments, the functionality of controller 110 is provided by and / or in combination with display generation component 120.

[0080] According to some embodiments, when a user is virtually and / or physically present within scene 105, display generation component 120 provides an XR experience to the user.

[0081] In some embodiments, the display generation component is worn on a part of the user's body (e.g., on his / her head, on his / her hand, etc.). Thus, the display generation component 120 includes one or more XR displays provided for displaying XR content. For example, in various embodiments, the display generation component 120 surrounds the user's field of view. In some embodiments, the display generation component 120 is a handheld device (such as a smart phone or a tablet device) configured to present XR content, and the user holds the device having a display pointing to the user's field of view and a camera pointing to the scene 105. In some embodiments, the handheld device is optionally placed in a housing worn on the user's head. In some embodiments, the handheld device is optionally placed on a support (e.g., a tripod) in front of the user. In some embodiments, the display generation component 120 is an XR chamber, housing, or room configured to present XR content, where the user does not wear or hold the display generation component 120. Many user interfaces described with reference to one type of hardware for displaying XR content (e.g., a handheld device or a device on a tripod) can be implemented on another type of hardware for displaying XR content (e.g., an HMD or other wearable computing device). For example, a user interface showing an interaction with XR content triggered based on an interaction occurring in the space in front of a handheld device or a tripod-mounted device can be similarly implemented with an HMD, where the interaction occurs in the space in front of the HMD, and the response to the XR content is displayed via the HMD. Similarly, a user interface showing an interaction with XR content triggered based on the movement of a handheld device or a tripod-mounted device relative to the physical environment (e.g., the scene 105 or a part of the user's body (e.g., the user's eyes, head, or hand)) can be similarly implemented with an HMD, where the movement is caused by the movement of the HMD relative to the physical environment (e.g., the scene 105 or a part of the user's body (e.g., the user's eyes, head, or hand)).

[0082] Although relevant features of the operating environment 100 are shown in Figure 1A for the sake of brevity and to not obscure more relevant aspects of the example embodiments disclosed herein, various other features are not illustrated.

[0083] Figures 1A to 1PIllustrates various examples of computer systems for performing methods and providing audio, visual, and / or tactile feedback as part of the user interfaces described herein. In some embodiments, the computer system includes one or more display generation components (e.g., a first display component 1-120a and a second display component 1-120b and / or a first optical module 11.1.1-104a and a second optical module 11.1.1-104b) for displaying a representation of a virtual element and / or a physical environment to a user of the computer system, the representation optionally being generated based on detected events and / or user input detected by the computer system. The user interface generated by the computer system is optionally corrected by one or more corrective lenses 11.3.2-216, the one or more corrective lenses optionally being removably attached to one or more of the optical modules such that the user interface is more easily viewable by users who would otherwise use glasses or contact lenses to correct their vision. Although many of the user interfaces illustrated herein show a single view of the user interface, the user interface in the HMD is optionally displayed using two optical modules (e.g., a first display assembly 1-120a and a second display assembly 1-120b and / or a first optical module 11.1.1-104a and a second optical module 11.1.1-104b), one optical module for the user's right eye and a different optical module for the user's left eye, and presenting slightly different images to the two different eyes to create an illusion of stereoscopic depth, the single view of the user interface is typically the right-eye view or the left-eye view, and the depth effect is explained in the text or using other schematic diagrams or views. In some embodiments, the computer system includes one or more external displays (e.g., a display assembly 1-108) for displaying status information of the computer system to a user of the computer system (when the computer system is not being worn) and / or to others in the vicinity of the computer system, the status information optionally being generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more audio output components (e.g., an electronic component 1-112) for generating audio feedback, the audio feedback optionally being generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting information about the physical environment of the device (e.g., a sensor assembly 1-356 and / or Figure 1I one or more of the sensors in), the one or more sensors being usable (optionally in combination with one or more illuminators, such as Figure 1IThe illuminator) generates a digital pass-through image, captures visual media corresponding to the physical environment (e.g., photos and / or videos), or determines the pose (e.g., location and / or orientation) of physical objects and / or surfaces in the physical environment such that virtual objects can be placed based on the detected pose of the physical objects and / or surfaces. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting hand location and / or movement (e.g., sensor assemblies 1-356 and / or Figure 1I one or more sensors therein), which can be used (optionally in combination with one or more illuminators, such as Figure 1I the illuminator 6-124 described therein) to determine when one or more air gestures have been performed. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting eye movement (e.g., Figure 1I the eye tracking and gaze tracking sensors therein), which can be used (optionally in combination with one or more lights, such as Figure 1OThe lights in (11.3.2-110) determine the attention or gaze localization and / or gaze movement, which can optionally be used to detect gaze-only input based on gaze movement and / or dwell. The combination of the various sensors described above can be used to determine the user's facial expression and / or hand movement for generating a user avatar or representation, such as an anthropomorphic avatar or representation for a real-time communication session, where the avatar has facial expressions, hand movements, and / or body movements based on or similar to the detected facial expressions, hand movements, and / or body movements of the user of the device. The gaze and / or attention information is optionally combined with hand tracking information to determine the interaction between the user and one or more user interfaces based on direct and / or indirect input, such as an air gesture or input using one or more hardware input devices, such as one or more buttons (e.g., the first button 1-128, button 11.1.1-114, second button 1-132, and / or dial or button 1-328), knobs (e.g., the first button 1-128, button 11.1.1-114, and / or dial or button 1-328), digital crowns (e.g., the first button 1-128, button 11.1.1-114, and / or dial or button 1-328 that can be pressed and twisted or rotated), touchpads, touchscreens, keyboards, mice, and / or other input devices. One or more buttons (e.g., the first button 1-128, button 11.1.1-114, second button 1-132, and / or dial or button 1-328) are optionally used to perform system operations, such as re-centering the content in the three-dimensional environment visible to the user of the device, displaying the main user interface for launching an application, starting a real-time communication session, or initiating the display of a virtual three-dimensional background. A knob or digital crown (e.g., the first button 1-128, button 11.1.1-114, and / or dial or button 1-328 that can be pressed and twisted or rotated) is optionally rotatable to adjust parameters of the visual content, such as the immersion level of the virtual three-dimensional environment (e.g., the extent to which the virtual content occupies the user's viewport in the three-dimensional environment) or other parameters associated with the three-dimensional environment and the virtual content displayed via an optical module (e.g., the first display component 1-120a and the second display component 1-120b and / or the first optical module 11.1.1-104a and the second optical module 11.1.1-104b).

[0084] Figure 1BIllustrates a front view, a top view, and a perspective view of an example of a head-mounted display (HMD) device 1-100 configured to be worn by a user and provide virtual and change / mixed reality (VR / AR) experiences. The HMD 1-100 may include a display unit 1-102 or component, an electronic strip component 1-104 connected to and extending from the display unit 1-102, and a strap component 1-106 fixed to the electronic strip component 1-104 at either end. The electronic strip component 1-104 and the strap 1-106 may be part of a retention component configured to wrap around the user's head to hold the display unit 1-102 against the user's face.

[0085] In at least one example, the strap component 1-106 may include a first strap 1-116 configured to wrap around the backside of the user's head and a second strap 1-117 configured to extend over the top of the user's head. As shown, the second strap may extend between a first electronic strip 1-105a and a second electronic strip 1-105b of the electronic strip component 1-104. The strip component 1-104 and the strap component 1-106 may be part of a fixation mechanism that extends rearward from the display unit 1-102 and is configured to hold the display unit 1-102 against the user's face.

[0086] In at least one example, the fixation mechanism includes a first electronic strip 1-105a that includes a first proximal end 1-134 coupled to the display unit 1-102 (e.g., the housing 1-150 of the display unit 1-102) and a first distal end 1-136 opposite the first proximal end 1-134. The fixation mechanism may also include a second electronic strip 1-105b that includes a second proximal end 1-138 coupled to the housing 1-150 of the display unit 1-102 and a second distal end 1-140 opposite the second proximal end 1-138. The fixation mechanism may also include a first strap 1-116 and a second strap 1-117, the first strap including a first end 1-142 coupled to the first distal end 1-136 and a second end 1-144 coupled to the second distal end 1-140, and the second strap extending between the first electronic strip 1-105a and the second electronic strip 1-105b. The strips 1-105a-b and the strap 1-116 may be coupled via a connection mechanism or component 1-114. In at least one example, the second strap 1-117 includes a first end 1-146 coupled to the first electronic strip 1-105a between the first proximal end 1-134 and the first distal end 1-136 and a second end 1-148 coupled to the second electronic strip 1-105b between the second proximal end 1-138 and the second distal end 1-140.

[0087] In at least one example, the first and second electronic strips 1-105a-b comprise plastic, metal, or other structural materials forming the shape of the substantially rigid strips 1-105a-b. In at least one example, the first strip 1-116 and the second strip 1-117 are formed of an elastic flexible material including woven textiles, rubber, and the like. The first strip 1-116 and the second strip 1-117 can be flexible to conform to the shape of the user's head when the HMD 1-100 is worn.

[0088] In at least one example, one or more of the first and second electronic strips 1-105a-b can define an internal strip volume and include one or more electronic components disposed within the internal strip volume. In one example, as Figure 1B shown, the first electronic strip 1-105a can include an electronic component 1-112. In one example, the electronic component 1-112 can include a speaker. In one example, the electronic component 1-112 can include a computing component, such as a processor.

[0089] In at least one example, the housing 1-150 defines a first front opening 1-152. The front opening is Figure 1B marked as 1-152 in dashed lines in, because the display component 1-108 is arranged to occlude the first opening 1-152 from the field of view when the HMD 1-100 is assembled. The housing 1-150 can also define a second rear opening 1-154. The housing 1-150 also defines an internal volume between the first opening 1-152 and the second opening 1-154. In at least one example, the HMD 1-100 includes a display component 1-108, which can include a front cover and a display screen (shown in other figures) disposed in or across the front opening 1-152 to occlude the front opening 1-152. In at least one example, the display screen of the display component 1-108, and generally the display component 1-108, has a curvature configured to follow the curvature of the user's face. The display screen of the display component 1-108 can be curved as shown to complement the user's facial features and the overall curvature from one side of the face to the other, e.g., from left to right and / or from top to bottom, where the display unit 1-102 is pressed.

[0090] In at least one example, the housing 1-150 may define a first aperture 1-126 between a first opening 1-152 and a second opening 1-154, and a second aperture 1-130 between the first opening 1-152 and the second opening 1-154. The HMD 1-100 may also include a first button 1-126 disposed in the first aperture 1-128, and a second button 1-132 disposed in the second aperture 1-130. The first button 1-128 and the second button 1-132 can be pressed through the respective apertures 1-126, 1-130. In at least one example, the first button 1-126 and / or the second button 1-132 can be a twist dial and a push button. In at least one example, the first button 1-128 is a pushable and twistable dial button, and the second button 1-132 is a push button.

[0091] Figure 1C Illustrated is a rear perspective view of the HMD 1-100. The HMD 1-100 may include a light seal 1-110 extending rearwardly from the housing 1-150 of the display assembly 1-108 around the perimeter of the housing 1-150, as shown. The light seal 1-110 may be configured to extend from the housing 1-150 to the user's face, surrounding the user's eyes, to block external light from being visible. In one example, the HMD 1-100 may include a first display assembly 1-120a and a second display assembly 1-120b, which are disposed at or within the second opening 1-154 defined by the housing 1-150 that faces rearward and / or disposed within the internal volume of the housing 1-150 and configured to project light through the second opening 1-154. In at least one example, each display assembly 1-120a-b may include a respective display screen 1-122a, 1-122b, which are configured to project light in a rearward direction through the second opening 1-154 toward the user's eyes.

[0092] In at least one example, referring Figure 1B and Figure 1C both, the display assembly 1-108 can be a front forward display assembly including a display screen configured to project light in a first forward direction, and the rear display screens 1-122a-b can be configured to project light in a second rearward direction opposite the first direction. As described above, the light seal 1-110 can be configured to block light external to the HMD 1-100 from reaching the user's eyes, including light projected by the forward display screen of the display assembly 1-108 shown in the front perspective view of Figure 1B In at least one example, the HMD 1-100 may also include a curtain 1-124 that occludes the second opening 1-154 between the housing 1-150 and the rear display assemblies 1-120a-b. In at least one example, the curtain 1-124 can be elastic or at least partially elastic.

[0093] Figure 1B and Figure 1C Any of the features, components, and / or parts shown, including their arrangements and configurations, may be included individually or in any combination in Figures 1D to 1F any other example of the devices, features, components, and parts shown and described herein. Similarly, with reference to Figures 1D to 1F any of the features, components, and / or parts shown and described, including their arrangements and configurations, may be included individually or in any combination in Figure 1B and Figure 1C the examples of the devices, features, components, and parts shown.

[0094] Figure 1D An exploded view of an example of the HMD 1-200 is illustrated, which includes various sections or parts separated according to the modular and selective coupling of those parts. For example, the HMD 1-200 may include a strap 1-216, which may be selectively coupled to a first electronic strip 1-205a and a second electronic strip 1-205b. The first fixed strip 1-205a may include a first electronic component 1-212a, and the second fixed strip 1-205b may include a second electronic component 1-212b. In at least one example, the first and second strips 1-205a-b are capable of being removably coupled to the display unit 1-202.

[0095] In addition, the HMD 1-200 may include a light seal 1-210 configured to be removably coupled to the display unit 1-202. The HMD 1-200 may also include a lens 1-218, which may be removably coupled to the display unit 1-202, for example, on a first component including a display screen and a second display component. The lens 1-218 may include a custom prescription lens configured for vision correction. As noted, each of the parts shown in Figure 1D the exploded view and described above can be removably coupled, attached, reattached, and replaced to update the parts or swap out parts for different users. For example, straps such as the strap 1-216, light seals such as the light seal 1-210, lenses such as the lens 1-218, and electronic strips such as the electronic strips 1-205a-b can be swapped out according to the user, such that these parts are customized to fit and correspond to a single user of the HMD 1-200.

[0096] Figure 1D Any of the features, components, and / or parts shown, including their arrangements and configurations, may be included individually or in any combination in Figure 1B , Figure 1C and Figures 1E to 1Fin any other example of the devices, features, components, and parts shown and described herein. Similarly, reference Figure 1B , Figure 1C and Figures 1E to 1F any one of the features, components, and / or parts shown or described, including their arrangements and configurations, may be included individually or in any combination in Figure 1D the example of the device, features, components, and parts shown.

[0097] Figure 1E FIG. illustrates an exploded view of an example of the display unit 1-306 of the HMD. The display unit 1-306 may include a front display assembly 1-308, a frame / housing assembly 1-350, and a curtain assembly 1-324. The display unit 1-306 may also include a sensor assembly 1-350, a logic board assembly 1-358, and a cooling assembly 1-360 disposed between the frame assembly 1-356 and the front display assembly 1-308. In at least one example, the display unit 1-306 may also include a rear display assembly 1-320, which includes a first rear display screen 1-322a and a second rear display screen 1-322b disposed between the frame 1-350 and the curtain assembly 1-324.

[0098] In at least one example, the display unit 1-306 may also include a motor assembly 1-362 configured as an adjustment mechanism for adjusting the positioning of the display screens 1-322a-b of the display assembly 1-320 relative to the frame 1-350. In at least one example, the display assembly 1-320 is mechanically coupled to the motor assembly 1-362, and each display screen 1-322a-b has at least one motor such that the motors can translate the display screens 1-322a-b to match the pupil spacing of the user's eyes.

[0099] In at least one example, the display unit 1-306 may include a dial or button 1-328 that can be pressed relative to the frame 1-350 and accessed by a user outside the frame 1-350. The button 1-328 may be electrically connected to the motor assembly 1-362 via a controller such that the button 1-328 can be manipulated by the user to cause the motors of the motor assembly 1-362 to adjust the positioning of the display screens 1-322a-b.

[0100] Figure 1E any one of the features, components, and / or parts shown, including their arrangements and configurations, may be included individually or in any combination in Figures 1B to 1D and Figure 1F any other example of the devices, features, components, and parts shown and described herein. Similarly, reference Figures 1B to 1D and Figure 1FAny of the features, components, and / or parts shown and described, including their arrangements and configurations, may be included individually or in any combination in Figure 1E the examples of the devices, features, components, and parts shown.

[0101] Figure 1F An exploded view of another example of the display unit 1-406 of an HMD device similar to other HMD devices described herein is illustrated. The display unit 1-406 may include a front display assembly 1-402, a sensor assembly 1-456, a logic board assembly 1-458, a cooling assembly 1-460, a frame assembly 1-450, a rear display assembly 1-421, and a curtain assembly 1-424. The display unit 1-406 may also include a motor assembly 1-462 for adjusting the positioning of a first display sub-assembly 1-420a and a second display sub-assembly 1-420b of the rear display assembly 1-421, including a first corresponding display screen and a second corresponding display screen for inter-pupillary adjustment, as described above.

[0102] Figure 1F The various parts, systems, and components shown in the exploded view are described in more detail herein with reference to Figures 1B to 1E and the subsequent figures referred to in this disclosure. Figure 1F The display unit 1-406 shown may be assembled and integrated with Figures 1B to 1E the shown fixing mechanisms, which include electronic strips, bands, and other components including light seals, connection components, etc.

[0103] Figure 1F Any of the features, components, and / or parts shown and described, including their arrangements and configurations, may be included individually or in any combination in Figures 1B to 1E any other example of the devices, features, components, and parts shown and described herein. Similarly, with reference to Figures 1B to 1E Any of the features, components, and / or parts shown and described, including their arrangements and configurations, may be included individually or in any combination in Figure 1F the examples of the devices, features, components, and parts shown.

[0104] Figure 1G A perspective exploded view of the front cover assembly 3-100 of the HMD device described herein is illustrated, such as Figure 1G the front cover assembly 3-1 of the HMD 3-100 or any other HMD device shown and described herein. Figure 1GThe front cover assembly 3-100 shown may include a transparent or translucent cover 3-102, a shield 3-104 (or "awning"), an adhesive layer 3-106, a display assembly 3-108 including a bi-convex lens panel or array 3-110, and a structural trim 3-112. The adhesive layer 3-106 may secure the shield 3-104 and / or the transparent cover 3-102 to the display assembly 3-108 and / or the trim 3-112. The trim 3-112 may secure the various components of the front cover assembly 3-100 to the frame or base of the HMD device.

[0105] In at least one example, as Figure 1G shown, the transparent cover 3-102, the shield 3-104, and the display assembly 3-108 including the bi-convex lens array 3-110 may be bent to conform to the curvature of the user's face. The transparent cover 3-102 and the shield 3-104 may be bent in two or three dimensions, e.g., vertically in the Z direction inside and outside the Z-X plane, and horizontally in the X direction inside and outside the Z-X plane. In at least one example, the display assembly 3-108 may include a bi-convex lens array 3-110 and a display panel having pixels configured to project light through the shield 3-104 and the transparent cover 3-102. The display assembly 3-108 may be bent in at least one direction (e.g., the horizontal direction) to conform to the curvature of the user's face from one side of the face (e.g., the left side) to the other side (e.g., the right side). In at least one example, each layer or component of the display assembly 3-108 (which will be shown and described in more detail in subsequent figures, but which may include the bi-convex lens array 3-110 and the display layer) may be similarly or concentrically bent in the horizontal direction to conform to the curvature of the user's face.

[0106] In at least one example, the shield 3-104 may include a transparent or translucent material through which the display assembly 3-108 projects light. In one example, the shield 3-104 may include one or more opaque portions, such as an opaque ink printed portion or other opaque film portion on the back surface of the shield 3-104. When the HMD device is worn, the back surface may be the surface of the shield 3-104 facing the user's eyes. In at least one example, the opaque portion may be on the front surface of the shield 3-104 opposite the back surface. In at least one example, one or more opaque portions of the shield 3-104 may include a peripheral portion that visually hides any components around the outer perimeter of the display screen of the display assembly 3-108. In this way, the opaque portions of the shield hide any other components of the HMD device that would otherwise be visible through the transparent or translucent cover 3-102 and / or the shield 3-104, including electronic components, structural components, etc.

[0107] In at least one example, the shield 3-104 may define one or more apertured transparent portions 3-120 through which the sensor may transmit and receive signals. In one example, portion 3-120 is an aperture through which the sensor may extend or through which the sensor may transmit and receive signals. In one example, portion 3-120 is a transparent portion, or a portion that is more transparent than the surrounding translucent or opaque portion of the shield, through which the sensor may transmit and receive signals through the shield and through the transparent cover 3-102. In one example, the sensor may include a camera, an IR sensor, a LUX sensor, or any other visual or non-visual environmental sensor of the HMD device.

[0108] Figure 1G Any one of the features, components, and / or parts shown, including their arrangement and configuration, may be included, either alone or in any combination, in any other example of the devices, features, components, and parts described herein. Similarly, any one of the features, components, and / or parts shown and described herein, including their arrangement and configuration, may be included, either alone or in any combination, in Figure 1G the examples of the devices, features, components, and parts shown.

[0109] Figure 1H An exploded view of an example of the HMD device 6-100 is illustrated. The HMD device 6-100 may include a sensor array or system 6-102 that includes one or more sensors, cameras, projectors, etc. mounted to one or more components of the HMD 6-100. In at least one example, the sensor system 6-102 may include a bracket 1-338 to which one or more sensors of the sensor system 6-102 may be fixed / fastened.

[0110] Figure 1I A portion of the HMD device 6-100 including the front transparent cover 6-104 and the sensor system 6-102 is illustrated. The sensor system 6-102 may include a plurality of different sensors, transmitters, receivers, including cameras, IR sensors, projectors, etc. The transparent cover 6-104 is illustrated in front of the sensor system 6-102 to illustrate the relative positioning of the various sensors and transmitters and the orientation of each sensor / transmitter of the system 6-102. As referred to herein, "lateral", "side", "transverse", "horizontal", and other similar terms refer to the orientation or direction indicated by the X-axis as Figure 1J shown. Terms such as "vertical", "upward", "downward", and similar terms refer to the orientation or direction indicated by the Z-axis as Figure 1J shown. Terms such as "forward", "backward", "frontward", "backward", and similar terms refer to the orientation or direction indicated by the Y-axis as Figure 1J shown.

[0111] In at least one example, the transparent cover 6-104 may define the front outer surface of the HMD device 6-100, and a sensor system 6-102 including various sensors and their components may be disposed behind the cover 6-104 in the Y-axis / direction. The cover 6-104 may be transparent or translucent to allow light to pass through the cover 6-104, including both light detected by the sensor system 6-102 and light emitted therefrom.

[0112] As described elsewhere herein, the HMD device 6-100 may include one or more controllers that include a processor for electrically coupling the various sensors and transmitters of the sensor system 6-102 to one or more motherboards, processing units, and other electronic devices such as display screens. Additionally, as will be shown in more detail with reference to other figures below, the various sensors, transmitters, and other components of the sensor system 6-102 may be coupled to Figure 1I various structural frame members, brackets, etc. of the HMD device 6-100 not shown. For purposes of illustration clarity, Figure 1I the components of the sensor system 6-102 are shown unattached and unelectrically coupled to other components.

[0113] In at least one example, the device may include one or more controllers having a processor configured to execute instructions stored on a memory component electrically coupled to the processor. The instructions may include one or more algorithms or cause the processor to execute the one or more algorithms for self-correcting the angles and positions of the various cameras described herein as the initial positioning, angle, or orientation of the camera changes over time due to accidental drop events or other events that cause collision or deformation.

[0114] In at least one example, the sensor system 6-102 may include one or more scene cameras 6-106. The system 6-102 may include two scene cameras 6-102 disposed on either side of the bridge or arch structure of the HMD device 6-100 such that each of the two cameras 6-106 generally corresponds to the positioning of the user's left and right eyes behind the cover 6-103. In at least one example, the scene cameras 6-106 are generally oriented forward in the Y-direction to capture images in front of the user during use of the HMD 6-100. In at least one example, the scene cameras are color cameras and provide images and content for MR video passthrough to the display screen facing the user's eyes when the HMD device 6-100 is in use. The scene cameras 6-106 may also be used for environment and object reconstruction.

[0115] In at least one example, the sensor system 6-102 can include a first depth sensor 6-108 that generally points forward in the Y direction. In at least one example, the first depth sensor 6-108 can be used for environment and object reconstruction and user hand and body tracking. In at least one example, the sensor system 6-102 can include a second depth sensor 6-110 that is centered along the width of the HMD device 6-100 (e.g., along the X axis). For example, the second depth sensor 6-110 can be disposed above the central nose bridge or on an adapter structure above the user's nose when wearing the HMD 6-100. In at least one example, the second depth sensor 6-110 can be used for environment and object reconstruction and hand and body tracking. In at least one example, the second depth sensor can include a LIDAR sensor.

[0116] In at least one example, the sensor system 6-102 can include a depth projector 6-112 that generally faces forward to project electromagnetic waves (e.g., in the form of a pre-determined pattern of light points) into the field of view or within the field of view of the user and / or the scene camera 6-106, or into a field of view that includes and extends beyond the field of view of the user and / or the scene camera 6-106. In at least one example, the depth projector is capable of projecting electromagnetic waves of light in the form of a pattern of light points that are reflected from an object and back into the aforementioned depth sensors, including depth sensors 6-108, 6-110. In at least one example, the depth projector 6-112 can be used for environment and object reconstruction and hand and body tracking.

[0117] In at least one example, the sensor system 6-102 can include a downward-facing camera 6-114 whose field of view generally points downward relative to the HDM device 6-100 along the Z axis. In at least one example, the downward camera 6-114 can be disposed on the left and right sides of the HMD device 6-100 as shown and is used for hand and body tracking, headset tracking, and facial avatar detection and creation for displaying a user avatar on the forward display screen of the HMD device 6-100 described elsewhere herein. For example, the downward camera 6-114 can be used to capture facial expressions and movements of the user's face below the HMD device 6-100, including the cheeks, mouth, and chin.

[0118] In at least one example, the sensor system 6-102 may include a jaw camera 6-116. In at least one example, the jaw camera 6-116 may be disposed on the left and right sides of the HMD device 6-100 as shown and used for hand and body tracking, head-mounted headset tracking, and face avatar detection and creation for display of a user avatar on the forward display screen of the HMD device 6-100 described elsewhere herein. For example, the jaw camera 6-116 may be used to capture facial expressions and movements of the user's face below the HMD device 6-100, including the user's jaw, cheeks, mouth, and chin. For hand and body tracking, head-mounted headset tracking, and face avatar

[0119] In at least one example, the sensor system 6-102 may include a side camera 6-118. The side camera 6-118 may be oriented to capture left and right views in the X-axis or with respect to the direction of the HMD device 6-100. In at least one example, the side camera 6-118 may be used for hand and body tracking, head-mounted headset tracking, and face avatar detection and recreation.

[0120] In at least one example, the sensor system 6-102 may include a plurality of eye tracking and gaze tracking sensors for determining the identity, condition, and gaze direction of the user's eyes during and / or before use. In at least one example, the eye / gaze tracking sensors may include a nose-eye camera 6-120 that is disposed on either side of the user's nose and adjacent to the user's nose when the HMD device 6-100 is worn. The eye / gaze sensors may also include a bottom eye camera 6-122 disposed below the respective user eye for capturing an image of the eye for face avatar detection and creation, gaze tracking, and iris identification functions.

[0121] In at least one example, the sensor system 6-102 may include an infrared illuminator 6-124 that points outward from the HMD device 6-100 to illuminate the external environment and any objects therein with IR light for IR detection by one or more IR sensors of the sensor system 6-102. In at least one example, the sensor system 6-102 may include a flicker sensor 6-126 and an ambient light sensor 6-128. In at least one example, the flicker sensor 6-126 may detect the top light refresh rate to avoid display flicker. In one example, the infrared illuminator 6-124 may include a light-emitting diode and may be particularly used in low-light environments to illuminate the user's hand and other objects in low light for detection by the infrared sensors of the sensor system 6-102.

[0122] In at least one example, multiple sensors (including scene cameras 6-106, downward camera 6-114, jaw camera 6-116, side cameras 6-118, depth projector 6-112, and depth sensors 6-108, 6-110) can be used in combination with an electrically coupled controller to combine depth data with camera data for hand tracking and for size determination to better perform hand tracking as well as object recognition and tracking functions of the HMD device 6-100. In at least one example, the downward camera 6-114, jaw camera 6-116, and side cameras 6-118 described above and shown in Figure 1I can be wide-angle cameras capable of operating in the visible and infrared spectra. In at least one example, these cameras 6-114, 6-116, 6-118 can operate only in black and white light detection to simplify image processing and obtain sensitivity.

[0123] Figure 1I Any of the features, components, and / or parts shown, including their arrangement and configuration, can be included individually or in any combination in Figures 1J to 1L any other example of the devices, features, components, and parts shown and described herein. Similarly, any of the features, components, and / or parts shown and described with reference to Figures 1J to 1L can be included individually or in any combination in Figure 1I the examples of the devices, features, components, and parts shown.

[0124] Figure 1J A lower perspective view of an example of an HMD 6-200 including a cover or shroud 6-204 fixed to a frame 6-230 is illustrated. In at least one example, the sensors 6-202 of the sensor system 6-203 can be disposed around the perimeter of the HDM 6-200 such that the sensors 6-203 are disposed outwardly around the perimeter of the display area or zone 6-232 so as not to obstruct the viewing of the displayed light. In at least one example, the sensors can be disposed behind the shroud 6-204 and aligned with the transparent portion of the shroud, thereby allowing the sensors and projectors to allow light to pass back and forth through the shroud 6-204. In at least one example, an opaque ink or other opaque material or film / layer can be disposed on the shroud 6-204 around the display zone 6-232 to hide the components of the HMD 6-200 outside the display zone 6-232 rather than the transparent portion defined by the opaque portion through which the sensors and projectors transmit and receive light and electromagnetic signals during operation. In at least one example, the shroud 6-204 allows light to pass through from the display (e.g., within the display area 6-232), but does not allow light to pass radially outward from the display area around the perimeter of the display and the shroud 6-204.

[0125] In some examples, the shroud 6-204 includes a transparent portion 6-205 and an opaque portion 6-207, as described above and elsewhere herein. In at least one example, the opaque portion 6-207 of the shroud 6-204 may define one or more transparent regions 6-209 through which the sensors 6-203 of the sensor system 6-202 may transmit and receive signals. In the illustrated example, the sensors 6-203 of the sensor system 6-202 transmit and receive signals through the shroud 6-204, or more specifically through the transparent regions 6-209 (or defined thereby) of the opaque portion 6-207 of the shroud 6-204, and the sensors may include the same or similar sensors as those shown in the example of Figure 1I such as depth sensors 6-108 and 6-110, depth projectors 6-112, a first scene camera and a second scene camera 6-106, a first downward camera and a second downward camera 6-114, a first side camera and a second side camera 6-118, and a first infrared illuminator and a second infrared illuminator 6-124. These sensors are also shown in Figure 1K and Figure 1L Examples. Other sensors, sensor types, sensor quantities, and their relative positioning may be included in one or more other examples of the HMD.

[0126] Figure 1J Any of the features, components, and / or parts shown, including their arrangement and configuration, may be included individually or in any combination in any other example of the devices, features, components, and parts shown in Figure 1I and Figures 1K to 1L shown and described herein. Similarly, any of the features, components, and / or parts shown or described in reference to Figure 1I and Figures 1K to 1L including their arrangement and configuration, may be included individually or in any combination in the examples of the devices, features, components, and parts shown in Figure 1J shown.

[0127] Figure 1K Illustrates a front view of a portion of an example of the HMD device 6-300, including a display 6-334, brackets 6-336, 6-338, and a frame or housing 6-330. Figure 1K The example shown does not include a front cover or shroud to illustrate the brackets 6-336, 6-338. For example, Figure 1J the shroud 6-204 shown includes an opaque portion 6-207 that will visually cover / block the viewing of anything outside the display / display area 6-334 (e.g., radially / peripherally outside), including the sensor 6-303 and the bracket 6-338.

[0128] In at least one example, the various sensors of the sensor system 6-302 are coupled to brackets 6-336, 6-338. In at least one example, the scene cameras 6-306 include tight tolerances on the angles relative to each other. For example, the tolerance on the mounting angle between two scene cameras 6-306 can be 0.5 degrees or less, such as 0.3 degrees or less. To achieve and maintain such tight tolerances, in one example, the scene cameras 6-306 can be mounted to bracket 6-338 rather than the shroud. The bracket can include a cantilever on which the scene cameras 6-306 and other sensors of the sensor system 6-302 can be mounted to maintain their position and orientation unchanged in the event of a drop event that causes any deformation of the other brackets 6-226, the housing 6-330, and / or the shroud by the user.

[0129] Figure 1K Any of the features, components, and / or parts shown, including their arrangement and configuration, can be included singly or in any combination in Figures 1I to 1J and Figure 1L any other example of the devices, features, components, and parts shown and described herein. Similarly, reference Figures 1I to 1J and Figure 1L to any of the features, components, and / or parts shown or described, including their arrangement and configuration, can be included singly or in any combination in Figure 1K the examples of the devices, features, components, and parts shown.

[0130] Figure 1L Illustrates a bottom view of an example of an HMD 6-400 including a front display / cover assembly 6-404 and a sensor system 6-402. The sensor system 6-402 can be similar to other sensor systems described above and elsewhere herein, including reference Figures 1I to 1K to. In at least one example, the chin camera 6-416 can face downward to capture images of the lower facial features of the user. In one example, the chin camera 6-416 can be directly coupled to the frame or housing 6-430 or one or more internal brackets that are directly coupled to the shown frame or housing 6-430. The frame or housing 6-430 can include one or more holes / openings 6-415 through which the chin camera 6-416 can transmit and receive signals.

[0131] Figure 1L Any of the features, components, and / or parts shown, including their arrangement and configuration, can be included singly or in any combination in Figures 1I to 1K any other example of the devices, features, components, and parts shown and described herein. Similarly, reference Figures 1I to 1K to any of the features, components, and / or parts shown and described, including their arrangement and configuration, can be included singly or in any combination inFigure 1L in the examples of the devices, features, components, and parts shown.

[0132] Figure 1M A rear perspective view of a pupil distance (IPD) adjustment system 11.1.1-102 is illustrated, which includes first and second optical modules 11.1.1-104a-b that are slidably engaged / coupled to respective guide rods 11.1.1-108a-b and motors 11.1.1-110a-b of left and right adjustment subsystems 11.1.1-106a-b. The IPD adjustment system 11.1.1-102 can be coupled to a bracket 11.1.1-112 and includes buttons 11.1.1-114 that are in electrical communication with the motors 11.1.1-110a-b. In at least one example, the buttons 11.1.1-114 can be in electrical communication with the first and second motors 11.1.1-110a-b via a processor or other circuit components to activate the first and second motors 11.1.1-110a-b and cause the first and second optical modules 11.1.1-104a-b to change their positions relative to each other, respectively.

[0133] In at least one example, the first and second optical modules 11.1.1-104a-b can include respective display screens that are configured to project light toward the user's eyes when the HMD 11.1.1-100 is worn. In at least one example, the user can manipulate (e.g., press and / or rotate) the buttons 11.1.1-114 to activate the positioning adjustment of the optical modules 11.1.1-104a-b to match the pupil distance of the user's eyes. The optical modules 11.1.1-104a-b can also include one or more cameras or other sensor / sensor systems for imaging and measuring the user's IPD such that the optical modules 11.1.1-104a-b can be adjusted to match the IPD.

[0134] In one example, the user may manipulate button 11.1.1-114 to cause an automatic positioning adjustment of the first and second optical modules 11.1.1-104a-b. In one example, the user may manipulate button 11.1.1-114 to cause a manual adjustment such that the optical modules 11.1.1-104a-b move further away or closer (e.g., when the user rotates button 11.1.1-114 in one way or another) until the user visually matches her / his own IPD. In one example, the manual adjustment is communicated electronically via one or more circuits, and the power for moving the optical modules 11.1.1-104a-b via motors 11.1.1-110a-b is provided by a power source. In one example, the adjustment and movement of the optical modules 11.1.1-104a-b via manipulation of button 11.1.1-114 is mechanically actuated via movement of button 11.1.1-114.

[0135] Figure 1M Any one of the features, components, and / or parts shown, including their arrangement and configuration, may be included, either alone or in any combination, in any other example of the devices, features, components, and parts shown in any other figure and described herein. Similarly, any one of the features, components, and / or parts shown or described with reference to any other figure, including their arrangement and configuration, may be included, either alone or in any combination, in Figure 1M the examples of the devices, features, components, and parts shown.

[0136] Figure 1N A front perspective view illustrating a portion of the HMD 11.1.2-100, including an external structural frame 11.1.2-102 and an internal or intermediate structural frame 11.1.2-104 that define a first aperture 11.1.2-106a and a second aperture 11.1.2-106b. The apertures 11.1.2-106a-b are shown Figure 1N in dashed lines because the view of the apertures 11.1.2-106a-b may be blocked by one or more other components of the HMD 11.1.2-100 that are coupled to the internal frame 11.1.2-104 and / or the external frame 11.1.2-102, as shown. In at least one example, the HMD 11.1.2-100 may include a first mounting bracket 11.1.2-108 that is coupled to the internal frame 11.1.2-104. In at least one example, the mounting bracket 11.1.2-108 is coupled to the internal frame 11.1.2-104 between the first and second apertures 11.1.2-106a-b.

[0137] The mounting bracket 11.1.2-108 may include an intermediate or central portion 11.1.2-109 coupled to the internal frame 11.1.2-104. In some examples, the intermediate or central portion 11.1.2-109 may not be the geometric middle or center of the bracket 11.1.2-108. Instead, the intermediate / central portion 11.1.2-109 may be disposed between a first cantilevered extension arm and a second cantilevered extension arm that extend away from the intermediate portion 11.1.2-109. In at least one example, the mounting bracket 108 includes a first cantilever 11.1.2-112 and a second cantilever 11.1.2-114 that extend away from the intermediate portion 11.1.2-109 of the mounting bracket 11.1.2-108 that is coupled to the internal frame 11.1.2-104.

[0138] As Figure 1N shown, the outer frame 11.1.2-102 may define a curved geometry on its lower side to accommodate the user's nose when the user wears the HMD 11.1.2-100. The curved geometry may be referred to as the nose bridge 11.1.2-111 and is centered on the lower side of the HMD 11.1.2-100 as shown. In at least one example, the mounting bracket 11.1.2-108 may be connected to the internal frame 11.1.2-104 between the holes 11.1.2-106a-b such that the cantilevers 11.1.2-112, 11.1.2-114 extend downward and laterally outward away from the intermediate portion 11.1.2-109 to be geometrically complementary to the nose bridge 11.1.2-111 geometry of the outer frame 11.1.2-102. In this way, the mounting bracket 11.1.2-108 is configured to accommodate the user's nose as described above. The geometry of the nose bridge 11.1.2-111 accommodates the nose because the nose bridge 11.1.2-111 provides a curvature that conforms to the shape of the user's nose, providing a comfortable fit from above, over, and around.

[0139] The first cantilever 11.1.2-112 can extend away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-108 in a first direction, and the second cantilever 11.1.2-114 can extend away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-108 in a second direction opposite to the first direction. The first cantilever 11.1.2-112 and the second cantilever 11.1.2-114 are referred to as "cantilever" or "cantilevered" arms because each arm 11.1.2-112, 11.1.2-114 respectively includes free distal ends 11.1.2-116, 11.1.2-118 that are not attached to the inner frame 11.1.2-102 and the outer frame 11.1.2-104. In this way, the arms 11.1.2-112, 11.1.2-114 overhang from the middle portion 11.1.2-109, which can be connected to the inner frame 11.1.2-104, while the distal ends 11.1.2-102, 11.1.2-104 are not attached.

[0140] In at least one example, the HMD 11.1.2-100 can include one or more components coupled to the mounting bracket 11.1.2-108. In one example, the components include a plurality of sensors 11.1.2-110a-f. Each sensor of the plurality of sensors 11.1.2-110a-f can include various types of sensors, including cameras, IR sensors, etc. In some examples, one or more of the sensors 11.1.2-110a-f can be used for object recognition in three-dimensional space, such that it is important to maintain the precise relative positions of two or more of the plurality of sensors 11.1.2-110a-f. The cantilevered nature of the mounting bracket 11.1.2-108 can protect the sensors 11.1.2-110a-f from damage and misalignment in the event of an accidental drop by the user. Because the sensors 11.1.2-110a-f are cantilevered on the arms 11.1.2-112, 11.1.2-114 of the mounting bracket 11.1.2-108, the stresses and deformations of the inner frame and / or the outer frame 11.1.2-104, 11.1.2-102 are not transmitted to the cantilevers 11.1.2-112, 11.1.2-114 and thus do not affect the relative positioning of the sensors 11.1.2-110a-f coupled / mounted to the mounting bracket 11.1.2-108.

[0141] Figure 1NAny of the features, components, and / or parts shown, including their arrangement and configuration, may be included, either individually or in any combination, in any other example of the devices, features, components described herein. Similarly, any of the features, components, and / or parts shown and described herein, including their arrangement and configuration, may be included, either individually or in any combination, in Figure 1N the examples of the devices, features, components, and parts shown.

[0142] Figure 1O Examples are illustrated for an electronic device such as an HMD, including the optical module 11.3.2-100 in the HMD device described herein. As shown in one or more other examples described herein, the optical module 11.3.2-100 may be one of two optical modules within the HMD, where each optical module is aligned to project light towards the user's eyes. In this manner, the first optical module may project light towards the user's first eye via a display screen, and the second optical module of the same device may project light towards the user's second eye via another display screen.

[0143] In at least one example, the optical module 11.3.2-100 may include an optical frame or housing 11.3.2-102, which may also be referred to as a barrel or an optical module barrel. The optical module 11.3.2-100 may also include a display 11.3.2-104 coupled to the housing 11.3.2-102, the display including one or more display screens. The display 11.3.2-104 may be coupled to the housing 11.3.2-102 such that the display 11.3.2-104 is configured to project light towards the user's eyes when wearing the HMD to which the display module 11.3.2-100 belongs during use. In at least one example, the housing 11.3.2-102 may surround the display 11.3.2-104 and provide connection features for coupling other components of the optical module described herein.

[0144] In one example, the optical module 11.3.2-100 may include one or more cameras 11.3.2-106 coupled to a housing 11.3.2-102. The cameras 11.3.2-106 may be positioned relative to a display 11.3.2-104 and the housing 11.3.2-102 such that the cameras 11.3.2-106 are configured to capture one or more images of a user's eyes during use. In at least one example, the optical module 11.3.2-100 may further include a light strip 11.3.2-108 surrounding the display 11.3.2-104. In one example, the light strip 11.3.2-108 is disposed between the display 11.3.2-104 and the cameras 11.3.2-106. The light strip 11.3.2-108 may include a plurality of lights 11.3.2-110. The plurality of lights may include one or more light-emitting diodes (LEDs) or other lights configured to project light toward a user's eyes when the HMD is worn. Each of the lights 11.3.2-110 in the light strip 11.3.2-108 may be spaced apart around the light strip 11.3.2-108 and thus may be spaced evenly or unevenly around the display 11.3.2-104 at various positions on the light strip 11.3.2-108 and around the display 11.3.2-104.

[0145] In at least one example, the housing 11.3.2-102 defines a viewing opening 11.3.2-101 through which a user may view the display 11.3.2-104 when the HMD device is worn. In at least one example, the LEDs are configured and arranged to emit light through the viewing opening 11.3.2-101 onto a user's eyes. In one example, the cameras 11.3.2-106 are configured to capture one or more images of a user's eyes through the viewing opening 11.3.2-101.

[0146] As described above, Figure 1O each of the components and features of the illustrated optical module 11.3.2-100 may be replicated in another (e.g., second) optical module provided with the HMD to interact (e.g., project light and capture images) with the user's other eye.

[0147] Figure 1O Any of the illustrated features, components, and / or parts, including their arrangement and configuration, may be included alone or in any combination in Figure 1P any other example of the devices, features, components, and parts illustrated or otherwise described herein. Similarly, reference Figure 1P to or any of the features, components, and / or parts illustrated or otherwise described herein, including their arrangement and configuration, may be included alone or in any combination in Figure 1Oin the examples of the devices, features, components, and parts shown.

[0148] Figure 1P An example cross-sectional view of an example of an optical module 11.3.2-200 is illustrated, including a housing 11.3.2-202, a display assembly 11.3.2-204 coupled to the housing 11.3.2-202, and a lens 11.3.2-216 coupled to the housing 11.3.2-202. In at least one example, the housing 11.3.2-202 defines a first hole or passage 11.3.2-212 and a second hole or passage 11.3.2-214. The passages 11.3.2-212, 11.3.2-214 may be configured to slidably engage corresponding tracks or guide rods of the HMD device to allow the optical module 11.3.2-200 to adjust its position relative to the user's eyes to match the user's interpupillary distance (IPD). The housing 11.3.2-202 is capable of slidably engaging the guide rods to fix the optical module 11.3.2-200 in place within the HMD.

[0149] In at least one example, the optical module 11.3.2-200 may further include a lens 11.3.2-216 coupled to the housing 11.3.2-202 and disposed between the display assembly 11.3.2-204 and the user's eyes when the HMD is worn. The lens 11.3.2-216 may be configured to direct light from the display assembly 11.3.2-204 to the user's eyes. In at least one example, the lens 11.3.2-216 may be part of a lens assembly that includes a corrective lens removably attached to the optical module 11.3.2-200. In at least one example, the lens 11.3.2-216 is disposed above a light bar 11.3.2-208 and one or more eye tracking cameras 11.3.2-206 such that the cameras 11.3.2-206 are configured to capture images of the user's eyes through the lens 11.3.2-216, and the light bar 11.3.2-208 includes lights configured to project light through the lens 11.3.2-216 onto the user's eyes during use.

[0150] Figure 1P Any of the features, components, and / or parts shown, including their arrangements and configurations, may be included, either individually or in any combination, in any other example of the devices, features, components, and parts described herein. Similarly, any of the features, components, and / or parts shown and described herein, including their arrangements and configurations, may be included, either individually or in any combination, in Figure 1P the examples of the devices, features, components, and parts shown.

[0151] Figure 2FIG. 0 is a block diagram of an example of controller 110 in accordance with some embodiments. Although certain specific features are illustrated, those skilled in the art will appreciate from this disclosure that various other features have not been illustrated for the sake of brevity and in order not to obscure more relevant aspects of the embodiments disclosed herein. To that end, by way of non-limiting example, in some embodiments, controller 110 includes one or more processing units 202 (e.g., microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), graphics processing units (GPUs), central processing units (CPUs), processing cores, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., universal serial bus (USB), FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, global system for mobile communications (GSM), code division multiple access (CDMA), time division multiple access (TDMA), global positioning system (GPS), infrared (IR), Bluetooth, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these components and various other components.

[0152] In some embodiments, one or more communication buses 204 include circuitry for interconnecting and controlling communication between system components. In some embodiments, one or more I / O devices 206 include at least one of a keyboard, mouse, touchpad, joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.

[0153] Memory 220 includes high-speed random access memory such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate random access memory (DDR RAM), or other random access solid state memory devices. In some embodiments, memory 220 includes non-volatile memory such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid state storage devices. Memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. Memory 220 includes non-transitory computer-readable storage medium. In some embodiments, memory 220 or the non-transitory computer-readable storage medium of memory 220 stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 230 and an XR experience module 240.

[0154] The operating system 230 includes instructions for handling various basic system services and for performing hardware-related tasks. In some embodiments, the XR experience module 240 is configured to manage and coordinate single or multiple XR experiences of one or more users (e.g., a single XR experience of one or more users, or multiple XR experiences of corresponding groups of one or more users). To this end, in various embodiments, the XR experience module 240 includes a data acquisition unit 241, a tracking unit 242, a coordination unit 246, and a data transmission unit 248.

[0155] In some embodiments, the data acquisition unit 241 is configured to obtain data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the display generation component 120 of Figure 1A and optionally from one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. To this end, in various embodiments, the data acquisition unit 241 includes instructions and / or logic for instructions and heuristics and metadata for heuristics.

[0156] In some embodiments, the tracking unit 242 is configured to map the scene 105 and track the positioning / location of at least the display generation component 120 relative to Figure 1A the scene 105, and optionally track the location of one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. To this end, in various embodiments, the tracking unit 242 includes instructions and / or logic for instructions and heuristics and metadata for heuristics. In some embodiments, the tracking unit 242 includes a hand tracking unit 244 and / or an eye tracking unit 243. In some embodiments, the hand tracking unit 244 is configured to track the positioning / location of one or more parts of the user's hand and / or the movement of one or more parts of the user's hand relative to Figure 1A the scene 105, relative to the display generation component 120, and / or relative to a coordinate system (which is defined relative to the user's hand). The hand tracking unit 244 is described in more detail below with respect to Figure 4 . In some embodiments, the eye tracking unit 243 is configured to track the positioning or movement of the user's gaze (or more generally, the user's eyes, face, or head) relative to the scene 105 (e.g., relative to the physical environment and / or relative to the user (e.g., the user's hand)) or relative to the XR content displayed via the display generation component 120. The eye tracking unit 243 is described in more detail below with respect to Figure 5 .

[0157] In some embodiments, the coordination unit 246 is configured to manage and coordinate the XR experience presented to the user by the display generation component 120 and optionally by one or more of the output device 155 and / or the peripheral device 195. To this end, in various embodiments, the coordination unit 246 includes instructions and / or logic for the instructions and heuristics and metadata for the heuristics.

[0158] In some embodiments, the data sending unit 248 is configured to send data (e.g., presentation data, location data, etc.) to at least the display generation component 120 and optionally to one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. To this end, in various embodiments, the data sending unit 248 includes instructions and / or logic for the instructions and heuristics and metadata for the heuristics.

[0159] Although the data acquisition unit 241, the tracking unit 242 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the coordination unit 246, and the data sending unit 248 are shown as residing on a single device (e.g., the controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 241, the tracking unit 242 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the coordination unit 246, and the data sending unit 248 may be located in separate computing devices.

[0160] In addition, Figure 2 Rather than being a structural schematic of the embodiments described herein, it is more of a functional description of the various features that may be present in a particular implementation. As will be recognized by those of ordinary skill in the art, items shown separately may be combined and some items may be separated. For example, Figure 2 some of the functional modules shown separately in may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various embodiments. The actual number of modules and the specific division of functions and how the features are allocated therein will vary depending on the particular implementation and, in some embodiments, will depend in part on the particular combination of hardware, software, and / or firmware selected for the particular implementation.

[0161] Figure 3FIG. 0 is a block diagram of an example of a display generation component 120 according to some embodiments. Although certain specific features are illustrated, those skilled in the art will appreciate from this disclosure that various other features are not illustrated for the sake of brevity and to not obscure more relevant aspects of the embodiments disclosed herein. To that end, as a non-limiting example, in some embodiments, the display generation component 120 (e.g., an HMD) includes one or more processing units 302 (e.g., a microprocessor, ASIC, FPGA, GPU, CPU, processing core, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, Bluetooth, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more XR displays 312, one or more optional inwardly and / or outwardly facing image sensors 314, a memory 320, and one or more communication buses 304 for interconnecting these components and various other components.

[0162] In some embodiments, one or more communication buses 304 include circuitry for interconnecting and controlling communication between system components. In some embodiments, one or more I / O devices and sensors 306 include an inertial measurement unit (IMU), accelerometers, gyroscopes, thermometers, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, and / or one or more depth sensors (e.g., structured light, time-of-flight, etc.).

[0163] In some embodiments, one or more XR displays 312 are configured to provide an XR experience to a user. In some embodiments, one or more XR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conduction electron-emitter display (SED), field-emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical systems (MEMS), and / or similar display types. In some embodiments, one or more XR displays 312 correspond to diffractive, reflective, polarization, holographic, and other waveguide displays. For example, the display generation component 120 (e.g., an HMD) includes a single XR display. In another example, the display generation component 120 includes an XR display for each eye of the user. In some embodiments, one or more XR displays 312 are capable of presenting MR and VR content. In some embodiments, one or more XR displays 312 are capable of presenting MR or VR content.

[0164] In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face including the user's eyes (and may be referred to as an eye-tracking camera). In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hand and optionally the user's arm (and may be referred to as a hand-tracking camera). In some embodiments, one or more image sensors 314 are configured to face forward in order to acquire image data corresponding to the scene that the user would see in the absence of the display generation component 120 (e.g., an HMD) (and may be referred to as a scene camera). One or more optional image sensors 314 may include one or more RGB cameras (e.g., having a complementary metal-oxide semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), one or more infrared (IR) cameras, and / or one or more event-based cameras, etc.

[0165] Memory 320 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. Memory 320 includes non-transitory computer-readable storage medium. In some embodiments, memory 320 or the non-transitory computer-readable storage medium of memory 320 stores the following programs, modules, and data structures, or subsets thereof, including optionally operating system 330 and XR rendering module 340.

[0166] Operating system 330 includes procedures for handling various basic system services and for performing hardware-related tasks. In some embodiments, XR rendering module 340 is configured to present XR content to a user via one or more XR displays 312. To this end, in various embodiments, XR rendering module 340 includes data acquisition unit 342, XR rendering unit 344, XR mapping generation unit 346, and data transmission unit 348.

[0167] In some embodiments, data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) at least from Figure 1A controller 110. To this end, in various embodiments, data acquisition unit 342 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.

[0168] In some embodiments, XR rendering unit 344 is configured to present XR content via one or more XR displays 312. To this end, in various embodiments, XR rendering unit 344 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.

[0169] In some embodiments, XR mapping generation unit 346 is configured to generate an XR map (e.g., a 3D map of a mixed reality scene or a map of a physical environment in which computer-generated objects can be placed to generate extended reality) based on media content data. To this end, in various embodiments, XR mapping generation unit 346 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.

[0170] In some embodiments, the data sending unit 348 is configured to send data (e.g., presentation data, location data, etc.) to at least the controller 110, and optionally to one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. To this end, in various embodiments, the data sending unit 348 includes instructions and / or logic for the instructions and heuristics and metadata for the heuristics.

[0171] Although the data acquisition unit 342, the XR presentation unit 344, the XR mapping generation unit 346, and the data sending unit 348 are shown as residing on a single device (e.g., Figure 1A the display generation component 120 of), it should be understood that in other embodiments, any combination of the data acquisition unit 342, the XR presentation unit 344, the XR mapping generation unit 346, and the data sending unit 348 may be located in separate computing devices.

[0172] In addition, Figure 3 Rather, it is more of a functional description of the various features that may be present in a particular embodiment, as opposed to a schematic diagram of the structure of the embodiments described herein. As will be recognized by those of ordinary skill in the art, the items shown separately may be combined, and some items may be separated. For example, Figure 3 some of the functional modules shown separately in may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various embodiments. The actual number of modules and the specific division of functions and how the features are allocated therein will vary depending on the particular implementation and, in some embodiments, will depend in part on the particular combination of hardware, software, and / or firmware selected for the particular implementation.

[0173] Figure 4 is a schematic illustration of an example embodiment of the hand tracking device 140. In some embodiments, the hand tracking device 140 ( Figure 1A ) is controlled by the hand tracking unit 244 ( Figure 2 ) to track the positioning / location of one or more parts of the user's hand, and / or the movement of one or more parts of the user's hand relative to Figure 1A the scene 105 (e.g., relative to a part of the physical environment around the user, relative to the display generation component 120, or relative to a part of the user (e.g., the user's face, eyes, or head), and / or relative to a coordinate system that is defined relative to the user's hand). In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).

[0174] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras, etc.) that captures three-dimensional scene information including at least the hand 406 of a human user. The image sensor 404 captures hand images at a sufficient resolution such that the fingers and their corresponding positions can be distinguished. The image sensor 404 typically captures images of other parts of the user's body, and may also or possibly capture images of all parts of the body, and may have zoom capabilities or dedicated sensors with increased magnification to capture images of the hand at a desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in combination with other image sensors to capture the physical environment of the scene 105, or serves as the image sensor for capturing the physical environment of the scene 105. In some embodiments, the image sensor 404 is positioned relative to the user or the user's environment in such a way that the field of view of the image sensor 404 or a portion thereof is used to define an interaction space, in which hand movements captured by the image sensor are considered inputs to the controller 110.

[0175] In some embodiments, the image sensor 404 outputs a sequence of frames containing 3D map data (and in addition, possibly color image data) to the controller 110, which extracts high-level information from the map data. This high-level information is typically provided to an application running on the controller via an application programming interface (API), and the application accordingly drives the display generation component 120. For example, the user can interact with the software running on the controller 110 by moving his hand 406 and changing his hand pose.

[0176] In some embodiments, the image sensor 404 projects a speckle pattern onto the scene containing the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral offset of the speckles in the pattern. This method is advantageous because it does not require the user to hold or wear any kind of beacon, sensor, or other marker. This method gives the depth coordinates of points in the scene relative to a pre-determined reference plane at a specific distance from the image sensor 404. In the present disclosure, it is assumed that the image sensor 404 defines an orthogonal set of x-axis, y-axis, and z-axis, such that the depth coordinates of points in the scene correspond to the z-component measured by the image sensor. Alternatively, the image sensor 404 (e.g., the hand tracking device) can use other 3D mapping methods, such as stereoscopy or time-of-flight measurement, based on a single or multiple cameras or other types of sensors.

[0177] In some embodiments, the hand tracking device 140 captures and processes a time series of depth maps of the user's hand as the user moves his hand (e.g., the entire hand or one or more fingers). Software running on a processor in the image sensor 404 and / or the controller 110 processes the 3D map data to extract hand patch descriptors in these depth maps. The software may match these descriptors with the patch descriptors stored in the database 408 based on a previous learning process to estimate the pose of the hand in each frame. The pose generally includes the 3D positions of the user's hand joints and finger tips.

[0178] The software may also analyze the trajectories of the hand and / or fingers over multiple frames in the sequence to identify gestures. The pose estimation function described herein may alternate with the motion tracking function such that the patch-based pose estimation is only performed once every two (or more) frames, while tracking is used to find changes in the pose that occur on the remaining frames. Pose, motion, and gesture information is provided to an application running on the controller 110 via the API described above. The program may, for example, move and modify the image presented on the display generation component 120 in response to the pose and / or gesture information, or perform other functions.

[0179] In some embodiments, the gestures include air gestures. An air gesture is detected when the user does not touch an input element (or independent of an input element that is part of a device (e.g., the computer system 101, one or more input devices 125, and / or the hand tracking device 140)) and is based on the detected movement of a part of the user's body (e.g., the head, one or more arms, one or more hands, one or more fingers, and / or one or more legs) through the air (including the movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), the movement relative to another part of the user's body (e.g., the movement of the user's hand relative to the user's shoulder, the movement of one of the user's hands relative to the other of the user's hands, and / or the movement of the user's finger relative to another finger or part of the hand), and / or the absolute movement of a part of the user's body (e.g., a tap gesture that includes the hand moving a predetermined amount and / or speed in a predetermined pose, or a shake gesture that includes a predetermined speed or amount of rotation of a part of the user's body)).

[0180] In some embodiments, according to some embodiments, the input gestures used in the various examples and embodiments described herein include air gestures performed by the movement of a user's finger relative to other fingers (or a part of the user's hand) for interacting with an XR environment (e.g., a virtual or mixed reality environment). In some embodiments, the air gesture is detected without the user touching an input element that is part of the device (or independent of an input element that is part of the device) and is based on the detected movement of a part of the user's body through the air, including the movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), the movement relative to another part of the user's body (e.g., the movement of the user's hand relative to the user's shoulder, the movement of one hand of the user relative to the other hand of the user, and / or the movement of the user's finger relative to another finger or part of the user's hand), and / or the absolute movement of a part of the user's body (e.g., a tap gesture that includes the hand moving a predetermined amount and / or speed in a predetermined posture, or a shake gesture that includes a predetermined speed or amount of rotation of a part of the user's body).

[0181] In some embodiments where the input gesture is an air gesture (e.g., in the absence of physical contact with an input device that provides information to the computer system about which user interface element is the target of the user input, such as contact with a user interface element displayed on a touch screen or contact with a mouse or touchpad to move a cursor to a user interface element), the gesture takes into account the user's attention (e.g., gaze) to determine the target of the user input (e.g., for direct input, as described below). Thus, in embodiments involving air gestures, for example, the input gesture is detected in combination with (e.g., simultaneously) the movement of the user's finger and / or hand and the attention (e.g., gaze) towards a user interface element to perform a pinch and / or tap input, as described in more detail below.

[0182] In some embodiments, input gestures directed to user interface objects are performed with direct or indirect reference to the user interface objects. For example, for direct input, the user input is performed directly on the user interface object by performing an input at a location corresponding to the location of the user interface object in a three-dimensional environment (e.g., as determined based on the user's current viewpoint). In some embodiments, for indirect input, the input gesture is performed on the user interface object when the user's attention (e.g., gaze) to the user interface object is detected and the location of the user's hand while performing the input gesture is not at the location corresponding to the location of the user interface object in the three-dimensional environment. For example, for a direct input gesture, the user is enabled to direct their input to the user interface object by initiating a gesture at or near a location corresponding to the display location of the user interface object (e.g., within a distance of 0.5 cm, 1 cm, 5 cm, or between 0 and 5 cm measured from the outer edge of the option or the central portion of the option). For an indirect input gesture, the user is enabled to direct their input to the user interface object by focusing on the user interface object (e.g., by gazing at the user interface object), and while focusing on the option, the user initiates an input gesture (e.g., at any location detectable by the computer system) (e.g., at a location not corresponding to the display location of the user interface object).

[0183] In some embodiments, according to some embodiments, input gestures (e.g., air gestures) used in the various examples and embodiments described herein include pinch inputs and tap inputs for interacting with virtual or mixed reality environments. For example, the pinch inputs and tap inputs described below are performed as air gestures.

[0184] In some embodiments, the pinch input is part of an air gesture that includes one or more of the following: a pinch gesture, a long pinch gesture, a pinch and drag gesture, or a double pinch gesture. For example, a pinch gesture as an air gesture includes the movement of two or more fingers of a hand to contact each other, that is, optionally, followed by an immediate interruption of contact with each other (e.g., within 0 seconds to 1 second). A long pinch gesture as an air gesture includes the movement of two or more fingers of a hand contacting each other for at least a threshold amount of time (e.g., at least 1 second) before detecting an interruption of contact with each other. For example, a long pinch gesture includes a user maintaining a pinch gesture (e.g., where two or more fingers are in contact), and the long pinch gesture continues until an interruption of contact between two or more fingers is detected. In some embodiments, a double pinch gesture as an air gesture includes two (e.g., or more) pinch inputs (e.g., performed by the same hand) detected consecutively and immediately with each other (e.g., within a predefined time period). For example, a user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., interrupts contact between two or more fingers), and performs a second pinch input within a predefined time period after releasing the first pinch input (e.g., within 1 second or within 2 seconds).

[0185] In some embodiments, a pinching and dragging gesture as an air gesture includes a pinching gesture (e.g., a pinching gesture or a long - pinching gesture) performed in combination with (e.g., following) a dragging input that changes the positioning of the user's hand from a first positioning (e.g., the starting positioning of the drag) to a second positioning (e.g., the ending positioning of the drag). In some embodiments, the user maintains the pinching gesture while performing the dragging input and releases the pinching gesture (e.g., opens two or more of their fingers) to end the dragging gesture (e.g., at the second positioning). In some embodiments, the pinching input and the dragging input are performed by the same hand (e.g., the user pinches two or more fingers together and moves the same hand into a second positioning in the air using the dragging gesture). In some embodiments, the pinching input is performed by the user's first hand and the dragging input is performed by the user's second hand (e.g., while the user continues the pinching input with the user's first hand, the user's second hand moves from a first positioning to a second positioning in the air). In some embodiments, an input gesture as an air gesture includes an input performed using both of the user's hands (e.g., a pinching and / or tapping input). For example, the input gesture includes two (e.g., or more) pinching inputs performed in combination with each other (e.g., simultaneously or within a predefined time period). For example, a first pinching gesture (e.g., a pinching input, a long - pinching input, or a pinching and dragging input) is performed using the user's first hand, and in combination with performing the pinching input with the first hand, a second pinching input is performed using the other hand (e.g., the second of the user's two hands). In some embodiments, there is movement between the user's two hands (e.g., increasing and / or decreasing the distance or relative orientation between the user's two hands).

[0186] In some embodiments, a tapping input performed as an air gesture (e.g., pointing to a user interface element) includes movement of the user's finger towards the user interface element, movement of the user's hand towards the user interface element (optionally, the user's finger extends towards the user interface element), downward movement of the user's finger (e.g., mimicking a mouse click movement or a tap on a touchscreen), or other predefined movement of the user's hand. In some embodiments, a tapping input performed as an air gesture is detected based on the movement characteristics of the finger or hand performing the tapping gesture movement, which is a movement of the finger or hand away from the user's viewing point and / or towards an object that is the target of the tapping input, followed by an end of the movement. In some embodiments, the end of the movement is detected based on a change in the movement characteristics of the finger or hand performing the tapping gesture (e.g., the end of the movement away from the user's viewing point and / or towards the object that is the target of the tapping input, a reversal of the direction of movement of the finger or hand, and / or a reversal of the acceleration direction of the movement of the finger or hand).

[0187] In some embodiments, the user's attention being directed to a portion of a three-dimensional environment is determined based on detection of a gaze directed to the portion of the three-dimensional environment (optionally, without requiring other conditions). In some embodiments, the user's attention being directed to a portion of a three-dimensional environment is determined based on detection of a gaze directed to the portion of the three-dimensional environment using one or more additional conditions, such as requiring the gaze to be directed to the portion of the three-dimensional environment for at least a threshold duration (e.g., dwell duration) and / or requiring the gaze to be directed to the portion of the three-dimensional environment when the user's viewpoint is within a distance threshold from the portion of the three-dimensional environment, such that the device determines that the user's attention is directed to the portion of the three-dimensional environment, where if one of these additional conditions is not met, the device determines that the attention is not directed to the portion of the three-dimensional environment to which the gaze is directed (e.g., until the one or more additional conditions are met).

[0188] In some embodiments, detection of a readiness state configuration of the user or a portion of the user is detected by a computer system. Detection of a readiness state configuration of a hand is used by the computer system as an indication that the user may be preparing to interact with the computer system using one or more air gesture inputs (e.g., pinch, tap, pinch and drag, double pinch, long pinch, or other air gestures described herein) performed by the hand. For example, based on whether the hand has a pre-determined hand shape (e.g., a pre-pinch shape where the thumb and one or more fingers are extended and spaced apart to prepare for a pinch or grab gesture, or a pre-tap where one or more fingers are extended and the palm is facing away from the user), based on whether the hand is in a pre-determined positioning relative to the user's viewpoint (e.g., below the user's head and above the user's waist and extending at least 15 cm, 20 cm, 25 cm, 30 cm, or 50 cm from the body), and / or based on whether the hand has moved in a particular manner (e.g., moving towards an area in front of the user that is above the user's waist and below the user's head or moving away from the user's body or legs) to determine the readiness state of the hand. In some embodiments, the readiness state is used to determine whether interactive elements of the user interface respond to attention (e.g., gaze) inputs.

[0189] In scenarios where input is described with reference to air gestures, it should be understood that a hardware input device attached to one or more of the user's hands or held by one or more of the user's hands can be used to detect similar gestures, where optical tracking, one or more accelerometers, one or more gyroscopes, one or more magnetometers, and / or one or more inertial measurement units can be used to track the positioning of the hardware input device in space, and the positioning and / or movement of the hardware input device is used to replace the positioning and / or movement of one or more hands in the corresponding air gesture. In scenarios where input is described with reference to air poses, it should be understood that a hardware input device attached to one or more of the user's hands or held by one or more of the user's hands can be used to detect similar poses. User input can be detected using controls contained within the hardware input device, such controls as one or more touch-sensitive input elements, one or more pressure-sensitive input elements, one or more buttons, one or more knobs, one or more dials, one or more joysticks, one or more hand or finger overlays that can detect the positioning or change in positioning of parts of the hand and / or fingers relative to each other, relative to the user's body, and / or relative to the user's physical environment, and / or other hardware input device controls, where user input using the controls contained within the hardware input device is used to replace hand and / or finger gestures such as an air tap or an air pinch in the corresponding air gesture. For example, a selection input described as being performed using an air tap or an air pinch input can alternatively be detected using a button press, a tap on a touch-sensitive surface, a press on a pressure-sensitive surface, or other hardware input. As another example, a movement input described as being performed using an air pinch and drag can alternatively be detected based on an interaction with a hardware input control (such as a button press and hold, a touch on a touch-sensitive surface, a press on a pressure-sensitive surface, or other hardware input after the movement of the hardware input device (e.g., along with the hand associated with the hardware input device) through space). Similarly, two-handed input including movement of the hands relative to each other can be performed using one air gesture and one hardware input device not performing an air gesture in a hand, two hardware input devices held in different hands, or two air gestures performed using various combinations of an air gesture and / or input detected by one or more of the above hardware input devices.

[0190] In some embodiments, the software can be downloaded electronically to the controller 110, for example, over a network, or alternatively can be provided on a tangible non-transitory medium such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is similarly stored in the memory associated with the controller 110. Alternatively or in addition, some or all of the described functions of the computer can be implemented in dedicated hardware (such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP)). Although in Figure 4Controller 110 is shown, but by way of example, as a unit separate from image sensor 404. Some or all of the processing functions of the controller may be performed by a suitable microprocessor and software or by dedicated circuitry within the housing of image sensor 404 (e.g., a hand tracking device) or by other devices associated with image sensor 404. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with display generation component 120 (e.g., in a television receiver, a handheld device, or a head-mounted device) or integrated with any other suitable computerized device (such as a gaming console or a media player). The sensing function of image sensor 404 may likewise be integrated into a computer or other computerized device that will be controlled by the sensor output.

[0191] Figure 4 Also shown is a schematic illustration of depth map 410 captured by image sensor 404 according to some embodiments. As described above, the depth map includes a matrix of pixels having corresponding depth values. Pixels 412 corresponding to hand 406 have been segmented from the background and wrist in the figure. The brightness of each pixel within depth map 410 is inversely proportional to its depth value (i.e., the measured z-distance from image sensor 404), where the gray shading becomes darker as the depth increases. Controller 110 processes these depth values to identify and segment the components of the image having human hand characteristics (i.e., adjacent pixel groupings). These characteristics may include, for example, overall size, shape, and motion from frame to frame in a sequence of depth maps.

[0192] Figure 4 Also schematically illustrated is hand skeleton 414 that controller 110 ultimately extracts from depth map 410 of hand 406 according to some embodiments. In Figure 4 , hand skeleton 414 is superimposed on hand background 416 that has been segmented from the original depth map. In some embodiments, key feature points on the hand and optionally on the wrist or arm connected to the hand (e.g., points corresponding to knuckles, finger tips, palm center, the end of the hand connected to the wrist, etc.) are identified and located on hand skeleton 414. In some embodiments, controller 110 uses the positions and movements of these key feature points across multiple image frames to determine, according to some embodiments, the gesture being performed by the hand or the current state of the hand.

[0193] Figure 5 An example embodiment of eye tracking device 130 ( Figure 1A ) is illustrated. In some embodiments, eye tracking device 130 is comprised of eye tracking unit 243 ( Figure 2)Control is used to track the positioning and movement of the user's gaze relative to the scene 105 or relative to the XR content displayed via the display generation component 120. In some embodiments, the eye tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, when the display generation component 120 is a head-mounted device (such as, a head-mounted headset, helmet, goggles, or glasses) or a hand-held device placed in a wearable frame, the head-mounted device includes both components for generating XR content for the user to view and components for tracking the user's gaze relative to the XR content. In some embodiments, the eye tracking device 130 is separate from the display generation component 120. For example, when the display generation component is a hand-held device or an XR room, the eye tracking device 130 is optionally a device separate from the hand-held device or the XR room. In some embodiments, the eye tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, the head-mounted eye tracking device 130 is optionally used in combination with a display generation component that is also head-mounted or a display generation component that is not head-mounted. In some embodiments, the eye tracking device 130 is not a head-mounted device and is optionally used in combination with a head-mounted display generation component. In some embodiments, the eye tracking device 130 is not a head-mounted device and is optionally part of a non-head-mounted display generation component.

[0194] In some embodiments, the display generation component 120 uses a display mechanism (e.g., a left near-eye display panel and a right near-eye display panel) to display a frame including a left image and a right image in front of the user's eyes, thereby providing the user with a 3D virtual view. For example, a head-mounted display generation component may include a left optical lens and a right optical lens (referred to herein as eye lenses) located between the display and the user's eyes. In some embodiments, the display generation component may include or be coupled to one or more external cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or semi-transparent display, and virtual objects are displayed on the transparent or semi-transparent display, and the user can directly view the physical environment through the transparent or semi-transparent display. In some embodiments, the display generation component projects virtual objects into the physical environment. The virtual objects may be projected, for example, onto a physical surface or projected as a hologram such that an individual using the system observes the virtual objects superimposed over the physical environment. In this case, separate display panels and image frames for the left and right eyes may not be required.

[0195] As Figure 5As shown, in some embodiments, the eye tracking device 130 (e.g., a gaze tracking device) includes at least one eye tracking camera (e.g., an infrared (IR) or near-infrared (NIR) camera), and an illumination source (e.g., an IR or NIR light source, such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eyes. The eye tracking camera can be directed at the user's eyes to receive IR or NIR light directly reflected from the eyes by the light source, or alternatively can be directed at a "hot" mirror located between the user's eyes and the display panel, which reflects the IR or NIR light from the eyes to the eye tracking camera while allowing visible light to pass through. The eye tracking device 130 optionally captures images of the user's eyes (e.g., as a video stream captured at 60 frames per second - 120 frames per second (fps)), analyzes the images to generate gaze tracking information, and transmits the gaze tracking information to the controller 110. In some embodiments, both of the user's eyes are tracked separately by corresponding eye tracking cameras and illumination sources. In some embodiments, only one of the user's eyes is tracked by corresponding eye tracking cameras and illumination sources.

[0196] In some embodiments, the eye tracking device 130 is calibrated using a device-specific calibration process to determine the parameters of the eye tracking device for a particular operating environment 100, such as the 3D geometric relationships and parameters of the LEDs, cameras, hot mirrors (if any), eye lenses, and display screens. The device-specific calibration process can be performed at the factory or another facility before the AR / VR equipment is delivered to the end user. The device-specific calibration process can be an automatic calibration process or a manual calibration process. According to some embodiments, the user-specific calibration process can include an estimation of the eye parameters of a particular user, such as pupil position, fovea position, optical axis, visual axis, eye separation, etc. According to some embodiments, once the device-specific parameters and user-specific parameters are determined for the eye tracking device 130, a flash-assisted method can be used to process the images captured by the eye tracking camera to determine the current visual axis and the user's fixation point relative to the display.

[0197] As Figure 5As shown, the eye tracking device 130 (e.g., 130A or 130B) includes an eye lens 520 and a gaze tracking system that includes at least one eye tracking camera 540 (e.g., an infrared (IR) or near-infrared (NIR) camera) positioned on the side of the user's face where eye tracking is to be performed, and an illumination source 530 (e.g., an IR or NIR light source, such as an array or ring of NIR light-emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eye 592. The eye tracking camera 540 may be pointed at a mirror 550 (which reflects IR or NIR light from the eye 592 while allowing visible light to pass through) (e.g., as shown in the top portion of Figure 5 ), or alternatively may be pointed at the user's eye 592 to receive reflected IR or NIR light from the eye 592 (e.g., as shown in the bottom portion of Figure 5 ).

[0198] In some embodiments, the controller 110 renders AR or VR frames 562 (e.g., left and right frames for a left display panel and a right display panel) and provides the frames 562 to the display 510. The controller 110 uses the gaze tracking input 542 from the eye tracking camera 540 for various purposes, such as for processing the frames 562 for display. The controller 110 optionally estimates the user's gaze point on the display 510 based on the gaze tracking input 542 obtained from the eye tracking camera 540 using a flash assist method or other suitable method. The gaze point estimated from the gaze tracking input 542 is optionally used to determine the direction the user is currently looking.

[0199] The following describes several possible use cases of the current gaze direction of a user and is not intended to be limiting. As an example use case, the controller 110 may render virtual content differently based on the determined direction of the user's gaze. For example, the controller 110 may generate virtual content at a higher resolution in the foveal region determined according to the user's current gaze direction than in the peripheral region. As another example, the controller may position or move virtual content in the view at least in part based on the user's current gaze direction. As another example, the controller may display specific virtual content in the view at least in part based on the user's current gaze direction. As another example use case in an AR application, the controller 110 may direct an external camera for capturing the physical environment of the XR experience to focus in the determined direction. Then, the autofocus mechanism of the external camera may focus on an object or surface in the environment that the user is currently looking at on the display 510. As another example use case, the eye lens 520 may be a focusable lens, and the controller uses the gaze tracking information to adjust the focus of the eye lens 520 such that the virtual object that the user is currently looking at has an appropriate vergence to match the convergence of the user's eyes 592. The controller 110 may utilize the gaze tracking information to direct the eye lens 520 to adjust the focus such that a nearby object that the user is looking at appears at the correct distance.

[0200] In some embodiments, the eye tracking device is part of a head-mounted device that includes a display (e.g., display 510) mounted in a wearable housing, two eye lenses (e.g., eye lens 520), an eye tracking camera (e.g., eye tracking camera 540), and a light source (e.g., illumination source 530 (e.g., IR or NIR LED)). The light source emits light (e.g., IR or NIR light) towards the user's eyes 592. In some embodiments, the light source may be arranged in a ring or circle around each of the lenses, as Figure 5 shown. In some embodiments, as an example, eight illumination sources 530 (e.g., LEDs) are arranged around each lens 520. However, more or fewer illumination sources 530 may be used, and other arrangements and positions of the illumination sources 530 may be used.

[0201] In some embodiments, the display 510 emits light in the visible light range and does not emit light in the IR or NIR range, and thus does not introduce noise in the gaze tracking system. Note that the position and angle of the eye tracking camera 540 are given by way of example and are not intended to be limiting. In some embodiments, a single eye tracking camera 540 is located on each side of the user's face. In some embodiments, two or more NIR cameras 540 may be used on each side of the user's face. In some embodiments, cameras 540 with a wider field of view (FOV) and cameras 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, cameras 540 operating at one wavelength (e.g., 850 nm) and cameras 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.

[0202] As Figure 5 The illustrated embodiments of the gaze tracking system can be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide a computer-generated reality, virtual reality, augmented reality, and / or augmented virtual experience to the user.

[0203] Figure 6 An example of a flash-assisted gaze tracking pipeline according to some embodiments is illustrated. In some embodiments, the gaze tracking pipeline is implemented by a flash-assisted gaze tracking system (e.g., the eye tracking device 130 as Figure 1A and Figure 5 illustrated). The flash-assisted gaze tracking system can maintain a tracking state. Initially, the tracking state is off or "no". When in the tracking state, when analyzing the current frame to track the pupil contour and flash in the current frame, the flash-assisted gaze tracking system uses the previous information from the previous frame. When not in the tracking state, the flash-assisted gaze tracking system attempts to detect the pupil and flash in the current frame, and if successful, initializes the tracking state to "yes" and continues with the next frame in the tracking state.

[0204] As Figure 6 shown, the gaze tracking camera can capture left and right images of the user's left and right eyes. The captured images are then input into the gaze tracking pipeline for processing starting at 610. As indicated by the arrow returning to element 600, the gaze tracking system can continue to capture images of the user's eyes, for example, at a rate of 60 to 120 frames per second. In some embodiments, each set of the captured images can be input into the pipeline for processing. However, in some embodiments or under some conditions, not all of the captured frames are processed by the pipeline.

[0205] At 610, for the currently captured image, if the tracking state is yes, the method proceeds to element 640. At 610, if the tracking state is no, then as indicated at 620, the image is analyzed to detect the user's pupil and flash in the image. At 630, if the pupil and flash are successfully detected, the method proceeds to element 640. Otherwise, the method returns to element 610 to process the next image of the user's eye.

[0206] At 640, if proceeding from element 610, the current frame is analyzed to track the pupil and flash based in part on previous information from a previous frame. At 640, if proceeding from element 630, the tracking state is initialized based on the pupil and flash detected in the current frame. The processing result at element 640 is checked to verify that the result of the tracking or detection can be trusted. For example, the result can be checked to determine whether the pupil and a sufficient number of flashes for performing gaze estimation are successfully tracked or detected in the current frame. At 650, if the result cannot be trusted, then at element 660, the tracking state is set to no, and the method returns to element 610 to process the next image of the user's eye. At 650, if the result is trusted, the method proceeds to element 670. At 670, the tracking state is set to yes (if it is not already yes), and the pupil and flash information is passed to element 680 to estimate the user's gaze point.

[0207] Figure 6 It is intended to be used as an example of an eye tracking technique that can be used for a particular specific implementation. As will be appreciated by those of ordinary skill in the art, according to various embodiments, in the computer system 101 for providing an XR experience to a user, other eye tracking techniques that currently exist or are developed in the future can be used to replace the flash-assisted eye tracking technique described herein or used in combination with the flash-assisted eye tracking technique.

[0208] In some embodiments, a captured portion of the real-world environment 602 is used to provide an XR experience to the user, such as a mixed reality environment in which one or more virtual objects are superimposed over a representation of the real-world environment 602.

[0209] Accordingly, the description herein describes some implementations of a three-dimensional environment (e.g., an XR environment) that includes representations of real-world objects and representations of virtual objects. For example, the three-dimensional environment optionally includes a representation of a table that exists in a physical environment, which is captured and displayed in the three-dimensional environment (e.g., actively displayed via a camera and a display of a computer system or passively displayed via a transparent or semi-transparent display of the computer system). As previously described, the three-dimensional environment is optionally a mixed reality system, where the three-dimensional environment is based on the physical environment captured by one or more sensors of the computer system and displayed via a display generation component. As a mixed reality system, the computer system is optionally capable of selectively displaying portions and / or objects of the physical environment such that the corresponding portions and / or objects of the physical environment appear as if they exist in the three-dimensional environment displayed by the computer system. Similarly, the computer system is optionally capable of displaying virtual objects in the three-dimensional environment at corresponding positions that have corresponding positions in the real world such that the virtual objects appear as if they exist in the real world (e.g., the physical environment). For example, the computer system optionally displays a vase such that the vase appears as if a real vase is placed on top of a table in the physical environment. In some implementations, the corresponding positions in the three-dimensional environment have corresponding positions in the physical environment. Thus, when the computer system is described as displaying a virtual object at a corresponding position relative to a physical object (e.g., a position at or near the user's hand or a position at or near a physical table), the computer system displays the virtual object at a specific position in the three-dimensional environment such that it appears as if the virtual object is at or near the physical object in the physical environment (e.g., the virtual object is displayed at a position in the three-dimensional environment that corresponds to the position in the physical environment where the virtual object would be displayed if the virtual object were a real object at that specific position).

[0210] In some implementations, real-world objects that exist in the physical environment and are displayed in the three-dimensional environment (e.g., and / or visible via a display generation component) can interact with virtual objects that only exist in the three-dimensional environment. For example, the three-dimensional environment can include a table and a vase placed on top of the table, where the table is a view (or representation) of a physical table in the physical environment and the vase is a virtual object.

[0211] In a three-dimensional environment (e.g., a physical environment, a virtual environment, or an environment that includes a mixture of physical and virtual objects), an object is sometimes referred to as having depth or simulated depth, or an object is referred to as being visible, displayed, or placed at different depths. In this context, depth refers to a dimension that is different from height or width. In some embodiments, depth is defined relative to a fixed set of coordinates (e.g., where a room or object has a height, depth, and width defined relative to a fixed set of coordinates). In some embodiments, depth is defined relative to the position or viewpoint of a user, in which case the depth dimension varies based on the position of the user and / or the position and orientation of the user's viewpoint. In some embodiments in which depth is defined relative to the position of the user relative to a surface of the environment (e.g., the surface of the floor or ground of the environment), an object that is further from the user along a line extending parallel to the surface is considered to have a greater depth in the environment, and / or the depth of the object is measured along an axis that extends outward from the position of the user and is parallel to the surface of the environment (e.g., depth is defined in a cylindrical or substantially cylindrical coordinate system, where the user's position is at the center of a cylinder that extends from the user's head toward the user's feet). In some embodiments in which depth is defined relative to the user's viewpoint (e.g., the direction relative to a point in space that determines which portion of the environment is visible via a head-mounted device or other display), an object that is further from the user's viewpoint along a line extending parallel to the user's viewpoint is considered to have a greater depth in the environment, and / or the depth of the object is measured along an axis that extends outward from the user's viewpoint and is parallel to the direction of the user's viewpoint (e.g., depth is defined in a spherical or substantially spherical coordinate system, where the origin of the viewpoint is at the center of a sphere that extends outward from the user's head). In some embodiments, depth is defined relative to a user interface container (e.g., a window or application in which an application and / or system content is displayed), where the user interface container has a height and / or width, and depth is a dimension that is orthogonal to the height and / or width of the user interface container. In some embodiments in which depth is defined relative to a user interface container, when the container is placed in a three-dimensional environment or is initially displayed (e.g., such that the depth dimension of the container extends outward away from the user or the user's viewpoint), the height and / or width of the container is typically orthogonal or substantially orthogonal to a line that extends from the position of the user (e.g., the user's viewpoint or the position of the user) to the user interface container (e.g., the center of the user interface container or another feature point of the user interface container). In some embodiments in which depth is defined relative to a user interface container, the depth of an object relative to the user interface container refers to the positioning of the object along the depth dimension of the user interface container. In some embodiments, multiple different containers may have different depth dimensions (e.g., different depth dimensions that extend in different directions and / or from different starting points away from the user or the user's viewpoint).In some embodiments, when defining depth relative to a user interface container, the direction of the depth dimension remains constant for the user interface container as the position of the user interface container, the user, and / or the user's viewing point changes (e.g., or when multiple different viewers are viewing the same container in a three-dimensional environment, such as during an in-person collaboration session and / or when multiple participants are in a real-time communication session with shared virtual content including the container). In some embodiments, for a curved container (e.g., including a container having a curved surface or a curved content area), the depth dimension optionally extends into the surface of the curved container. In some cases, z-spacing (e.g., the spacing between two objects in the depth dimension), z-height (e.g., the distance of one object from another in the depth dimension), z-positioning (e.g., the positioning of one object in the depth dimension), z-depth (e.g., the positioning of one object in the depth dimension), or the simulated z-dimension (e.g., depth used as a dimension of an object, a dimension of the environment, a direction in space, and / or a simulated direction in space) are used to refer to the concept of depth as described above.

[0212] In some embodiments, the user can optionally interact with virtual objects in a three-dimensional environment using one or both hands as if the virtual objects were real objects in a physical environment. For example, as described above, one or more sensors of the computer system optionally capture one or more hands of the user and display a representation of the user's hands in the three-dimensional environment (e.g., in a manner similar to displaying real-world objects in the three-dimensional environment described above), or in some embodiments, the user's hands can be seen via the display generation component, via the ability to see the physical environment through the user interface, due to the transparency / translucency of a portion of the user interface being displayed by the display generation component, or due to the projection of the user interface onto a transparent / translucent surface or onto the user's eyes or into the user's field of view. Thus, in some embodiments, the user's hands are displayed at corresponding positions in the three-dimensional environment and are treated as if they were objects in the three-dimensional environment that can interact with virtual objects in the three-dimensional environment as if those virtual objects were physical objects in a physical environment. In some embodiments, the computer system is able to update the display of the representation of the user's hands in the three-dimensional environment in conjunction with the movement of the user's hands in the physical environment.

[0213] In some of the embodiments described below, the computer system optionally can determine an "effective" distance between a physical object in the physical world and a virtual object in a three-dimensional environment, e.g., for determining whether the physical object is directly interacting with the virtual object (e.g., whether a hand is touching, grasping, holding, etc. the virtual object or is within a threshold distance of the virtual object). For example, a hand directly interacting with a virtual object optionally includes one or more of the following: a finger of the hand pressing a virtual button, a hand of the user grasping a virtual vase, the hands of the user brought together and pinching / holding a user interface of an application, and two fingers performing any of the other types of interactions described herein. For example, when determining whether a user is interacting with a virtual object and / or how the user is interacting with the virtual object, the computer system optionally determines the distance between the user's hand and the virtual object. In some embodiments, the computer system determines the distance between the user's hand and the virtual object by determining the distance between the position of the hand in the three-dimensional environment and the position of the virtual object of interest in the three-dimensional environment. For example, the one or more hands of the user are located at a particular location in the physical world, and the computer system optionally captures the one or more hands and displays the one or more hands at a particular corresponding location in the three-dimensional environment (e.g., the location where the hand would be displayed in the three-dimensional environment if the hand were a virtual hand rather than a physical hand). Optionally, the location of the hand in the three-dimensional environment is compared with the location of the virtual object of interest in the three-dimensional environment to determine the distance between the one or more hands of the user and the virtual object. In some embodiments, the computer system optionally determines the distance between the physical object and the virtual object by comparing locations in the physical world (e.g., rather than comparing locations in the three-dimensional environment). For example, when determining the distance between one or more hands of the user and a virtual object, the computer system optionally determines the corresponding position of the virtual object in the physical world (e.g., the location where the virtual object would be located in the physical world if the virtual object were a physical object rather than a virtual object), and then determines the distance between the corresponding physical location and the one or more hands of the user. In some embodiments, the same techniques are optionally used to determine the distance between any physical object and any virtual object. Thus, as described herein, when determining whether a physical object is in contact with a virtual object or whether a physical object is within a threshold distance of a virtual object, the computer system optionally performs any of the techniques described above to map the position of the physical object to the three-dimensional environment and / or to map the position of the virtual object to the physical environment.

[0214] In some embodiments, the same or similar techniques are used to determine where and what the user's gaze is directed at, and / or where and what a physical stylus held by the user is directed at. For example, if the user's gaze is directed at a particular location in the physical environment, the computer system optionally determines the corresponding location in the three-dimensional environment (e.g., the virtual location of the gaze), and if a virtual object is located at the corresponding virtual location, the computer system optionally determines that the user's gaze is directed at the virtual object. Similarly, the computer system is optionally able to determine the direction in the physical environment that the stylus is directed based on the orientation of the physical stylus. In some embodiments, based on the determination, the computer system determines the corresponding virtual location in the three-dimensional environment that corresponds to the location in the physical environment where the stylus is directed, and optionally determines that the stylus is directed at the corresponding virtual location in the three-dimensional environment.

[0215] Similarly, the embodiments described herein may refer to the position of the user (e.g., the user of the computer system) in the three-dimensional environment and / or the position of the computer system in the three-dimensional environment. In some embodiments, the user of the computer system is holding, wearing, or otherwise located at or near the computer system. Thus, in some embodiments, the position of the computer system serves as a proxy for the position of the user. In some embodiments, the position of the computer system and / or the user in the physical environment corresponds to the corresponding position in the three-dimensional environment. For example, the position of the computer system would be the position in the physical environment (and its corresponding position in the three-dimensional environment) from which the user would see the objects in the physical environment at the same location, orientation, and / or size (e.g., in an absolute sense and / or relative to each other) as the objects that are displayed or are visible in the three-dimensional environment by the display generation component of the computer system if the user were standing at that position and facing the corresponding portion of the physical environment visible via the display generation component. Similarly, if the virtual objects displayed in the three-dimensional environment are physical objects in the physical environment (e.g., physical objects placed at the same positions in the physical environment as the positions of these virtual objects in the three-dimensional environment, and physical objects having the same size and orientation in the physical environment as when in the three-dimensional environment), then the position of the computer system and / or the user is the location from which the user would see the virtual objects in the physical environment at the same location, orientation, and / or size (e.g., in an absolute sense and / or relative to each other and real-world objects) as the virtual objects displayed in the three-dimensional environment by the display generation component of the computer system.

[0216] In this disclosure, various input methods are described relative to interaction with a computer system. When one input device or input method is used to provide an example and another input device or input method is used to provide another example, it should be understood that each example can be compatible with and optionally utilize the input device or input method described relative to the other example. Similarly, various output methods are described relative to interaction with a computer system. When one output device or output method is used to provide an example and another output device or output method is used to provide another example, it should be understood that each example can be compatible with and optionally utilize the output device or output method described relative to the other example. Similarly, various methods are described relative to interaction with a virtual environment or a mixed reality environment via a computer system. When interaction with a virtual environment is used to provide an example and a mixed reality environment is used to provide another example, it should be understood that each example can be compatible with and optionally utilize the methods described relative to the other example. Accordingly, this disclosure discloses embodiments that are combinations of features of multiple examples without exhaustively listing all features of the embodiments in the description of each example embodiment.

[0217] User Interface and Associated Processes

[0218] Attention is now turned to embodiments of a user interface (“UI”) and associated processes that can be implemented on a computer system such as a portable multifunctional device or a head-mounted device that communicates with a display generation component and (optionally) one or more sensors.

[0219] The examples described herein illustrate ways in which a user of a computer system (e.g., device 700) can initiate and / or modify a live communication session in which the user communicates with one or more users of other corresponding computer systems. In some embodiments, the live communication session is an audio communication session (e.g., a voice call or a telephone call). In some embodiments, the live communication session is a video communication session (e.g., a video call and / or a video conference). In some embodiments, the live communication session is an XR communication session, such as a spatial communication session or a non-spatial communication session. During a spatial communication session, one or more users are represented in the XR environment by three-dimensional (3D) representations (e.g., avatars) corresponding to the users. In some embodiments, the 3D representations have spatial agency such that the 3D representations can move within the XR environment relative to other elements and / or users in the XR environment. During a non-spatial communication session, one or more users are represented in the XR environment by two-dimensional (2D) representations corresponding to the users. In some embodiments, the 2D representations include a video feed of the user and optionally have a fixed position (e.g., location) within the XR environment.

[0220] Figures 7A to 7Q Illustrates an example of managing a live communication session. Figure 8 Is a flowchart of an exemplary method 800 for managing a live communication session. Figure 9 Is a flowchart of an exemplary method 900 for providing an avatar in a live communication session. Figures 7A to 7Q The user interface in is used to illustrate the processes described below, which include Figure 8 and / or Figure 9 the processes in.

[0221] Although Figures 7A to 7Q Device 700 is illustrated as a handheld device (e.g., a tablet, smartphone, or laptop) having a display 702, in some embodiments, device 700 is a head-mounted device (HMD). The HMD is configured to be worn on the head of a user of device 700 and includes a display 702 on and / or within an interior portion of the HMD. When device 700 is worn on the user's head, display 702 is visible to the user. For example, in some embodiments, the HMD at least partially covers the user's eyes when worn on the user's head such that display 702 is positioned above and / or in front of the user's eyes. In such embodiments, display 702 is configured to display an XR environment during a live communication session in which the user of the HMD is participating.

[0222] In Figure 7A Device 700 displays an XR environment 704 on display 702 that includes elements (e.g., virtual elements and / or physical elements) such as table 704a and couch 704b. When displaying XR environment 704, device 700 receives a request to display a communication interface. In some embodiments, the request to display the communication interface is pressing a button 703 of device 700. As Figure 7B shown, in response to receiving the request, device 700 displays communication interface 710. In some embodiments, communication interface 710 is displayed within XR environment 704.

[0223] Typically, the communication interface 710 can be used to initiate and / or modify a live communication session (e.g., an audio communication session, a video communication session, or an XR communication session). The communication interface 710 includes pinned contacts 712 (e.g., pinned contacts 712a - 712g) and recent contacts 714 (e.g., recent contacts 714a - 714i). In some embodiments, the pinned contacts 712 are a set of contacts selected by the user of device 700 to be included in the communication interface 710 (e.g., contacts favorited or pinned by the user). In some embodiments, the recent contacts 714 are contacts with whom the user of device 700 has recently communicated (e.g., via text, phone, and / or live communication sessions) using device 700 and optionally one or more other devices associated with the user of device 700. In some embodiments, the recent contacts 714 are arranged (e.g., sorted or ranked) based on the recency of the communication between the recent contacts 714 and the user of device 700.

[0224] In some embodiments, one or more of the pinned contacts 712 and / or recent contacts 714 correspond to defined contact groups. As an example, pinned contact 712d corresponds to the contact group "Surfers". As another example, recent contact 714d corresponds to the contact group "Lake Crew".

[0225] In some embodiments, the pinned contacts 712 and / or recent contacts 714 indicate the most recent communication between the user of device 700 and various contacts. As an example, pinned contact 712b ("John") indicates that the contact last sent a text message 1 minute ago. Optionally, the communication interface 710 includes a preview 716b that indicates the content of the text message sent by pinned user 712b. As another example, pinned user 712c indicates that the contact last sent a text reaction (e.g., a "heart" reaction) at 2:10. As yet another example, recent contact 714a ("Mom") indicates that the user of device 700 most recently communicated with contact 714a in an XR communication session (e.g., a spatial live communication session or a non - spatial live communication session) at 3:32. As yet another example, recent contact 714e ("Uncle Bob") indicates that the user of device 700 most recently communicated with contact 714e in an audio communication session (e.g., a phone call) at 9:41.

[0226] In some embodiments, the pinned contacts 712 and / or the recent contacts 714 indicate a pending invitation to a live communication session. As an example, the recent contact 714b ("Dad") indicates that the user of device 700 can join a live communication session with the recent contact 714b. As another example, the recent contact 714d ("Lake Crew") indicates that three members of the group are currently in an ongoing live communication session that the user of device 700 has been invited to join.

[0227] In some embodiments, the contacts 712 and 714 of the communication interface 710 can be used to manage contacts. By way of example, when the communication interface 710 is displayed, device 700 detects a selection of the contact 712e ("Jo"). In some embodiments, the selection of the contact 712e is a tap gesture 705b on the contact 712e. In some embodiments, the selection of the contact 712e is an air gesture, for example, indicating the selection of the contact 712e. As Figure 7C1 and / or Figure 7C2 shown, in response to detecting the selection of the contact 712e, device 700 displays a contact menu 720 associated with the contact 712e.

[0228] The contact menu 720 includes an invite option 720a and an expand option 720b. When the invite option 720a is selected, it causes device 700 to invite the contact 712e to an XR communication session. When the expand option 720b is selected, it causes device 700 to display one or more additional options for managing the contact 712e. For example, when the contact menu 720 is displayed, device 700 detects a selection of the expand option 720b. In some embodiments, the selection of the expand option 720b is a tap gesture 705c on the expand option 720b. In some embodiments, the selection of the expand option 720b is an air gesture, for example, indicating the selection of the expand option 720b. As Figure 7D shown, in response to detecting the selection of the expand option 720b, device 700 expands the contact menu 720 to display (e.g., replaces the display of the expand option 720b with) one or more additional options (e.g., options 720c - 720f).

[0229] In some embodiments, when the contacts menu 720 is expanded, it includes an audio option 720c, a messaging option 720d, an information option 720e, and an edit option 720f. Selecting the audio option 720c causes the device 700 to initiate an audio communication session with the contact 712e (e.g., without a live video component). In some embodiments, the device 700 is not capable of communicating over a cellular network and / or is configured to use an external device for audio calls. Thus, in some examples, the device 700 uses a nearby device (e.g., a mobile phone and / or a tablet) that is capable of communicating over a cellular network to initiate the audio communication session. Selecting the edit option 720f allows a user of the device 700 to remove the contact 712e from the top contacts 712 (or add the contact 712e to the top contacts in embodiments where the contact 712e is not already a top contact). Selecting the messaging option 720d allows the user to send a message to the contact 712e. For example, when the contacts menu 720 is displayed, the device 700 detects a selection of the messaging option 720d. In some embodiments, the selection of the messaging option 720d is a tap gesture 705d on the messaging option 720d. In some embodiments, the selection of the messaging option 720d is an air gesture that indicates, for example, the selection of the messaging option 720d. As Figure 7E shown, in response to detecting the selection of the messaging option 720d, the device 700 displays (e.g., replaces the display of the communication interface 710 with) a messaging interface 730. Thereafter, the messaging interface 730 can be used to send a message to the contact 712e.

[0230] Referring again to Figure 7D , selecting the information option 720e causes the device 700 to display information corresponding to the contact 712e (e.g., without displaying additional information corresponding to other contacts). For example, when the contacts menu 720 is displayed, the device 700 detects a selection of the information option 720e. In some embodiments, the selection of the information option 720e is a tap gesture 707d on the information option 720e. In some embodiments, the selection of the information option 720e is an air gesture that indicates, for example, the selection of the information option 720e. As Figure 7F shown, in response to detecting the selection of the information option 720e, the device 700 displays an information interface 740. The information interface 740 includes various details corresponding to the contact 712e, including but not limited to a name and contact information.

[0231] In some embodiments, the contact's device is unable to participate in an XR communication session with device 700. Thus, in some embodiments, one or more options of the contact menu may be omitted, de-emphasized (e.g., grayed out or dimmed), and / or replaced to accurately reflect the capabilities of the contact's device. For example, referring again to Figure 7B , when displaying communication interface 710, device 700 detects a selection of contact 712g ("Sam"). In some embodiments, the selection of contact 712g is a tap gesture 709b on contact 712g. In some embodiments, the selection of contact 712g is an air gesture, for example, indicating the selection of contact 712g. As Figure 7C1 shown, in response to detecting the selection of contact 712g, device 700 displays contact menu 722 associated with contact 712g.

[0232] Since in some embodiments the device of contact 712g is unable to communicate with device 700 in an XR communication session, menu 722 does not include an invitation option (e.g., invitation option 720a), and instead includes an audio option 722a. The audio option 722a, when selected, causes device 700 to initiate an audio communication session with contact 712g. Menu 722 also includes an expand option 722b, which, when selected, causes device 700 to display one or more additional options for contact 712g.

[0233] In some embodiments, the contact menu associated with a contact includes one or more additional options based on the state of device 700. As an example, in some embodiments, in an instance where device 700 is participating in a live communication session (e.g., an XR communication session or an audio communication session), the contact menu includes an option to invite the contact to the live communication session. For example, referring to Figure 7B , when participating in an XR communication session and when displaying communication interface 710, device 700 detects a selection of contact 714g ("Dylan"). In some embodiments, the selection of contact 714g is a tap gesture 711b on contact 714g. In some embodiments, the selection of contact 714g is an air gesture, for example, indicating the selection of contact 714f. As Figure 7C1 shown, in response to detecting the selection of contact 714g, device 700 displays contact menu 724 associated with contact 714g.

[0234] The contact menu 724 includes an invite option 724a, an invite option 724b, and an expand option 724c. The invite option 724a, when selected (e.g., a tap gesture 709c), causes the device 700 to invite the contact 714g to a new live communication session. The invite option 724b, when selected, causes the device 700 to invite the contact 714g to the live communication session in which the device 700 is currently participating. The expand option 720c, when selected, causes the device 700 to display one or more additional options for the contact 714g.

[0235] In some embodiments, inviting a contact to a new live communication session (e.g., in response to a selection of option 724a) causes the device 700 to disconnect from and / or terminate the live communication session in which the device 700 is currently participating. In some embodiments, before terminating an existing live communication session in this manner, the device 700 confirms that the user wishes to disconnect from the current live communication session before initiating a new live communication session. For example, as Figure 7G shown, in response to a selection of the invite option 724a, the device 700 displays a confirmation interface 740 that includes a confirmation affordance 742. In response to a selection of the confirmation affordance 742, the device 700 terminates the current live communication session and invites the contact 714f to a new live communication session.

[0236] In some embodiments, the user optionally uses the communication interface 710 to send a message to a contact. For example, referring to Figure 7B , when the communication interface 710 is displayed, the device 700 detects a selection of the preview 716b associated with the pinned contact 712b. In some embodiments, the selection of the preview 716b is a tap gesture 707b on the preview 716b. In some embodiments, the selection of the preview 716b is an air gesture, for example, indicating the selection of the preview 716b. As Figure 7C1 shown, in response to detecting the selection of the preview 716b, the device 700 expands the preview 716b to display reply options 718.

[0237] The reply options 718, when selected, cause the device 700 to display a reply interface for sending a message to the contact 712b. For example, when the reply options 718 are displayed in the preview 716b, the device 700 detects a selection of the reply options 718. In some embodiments, the selection of the reply options 718 is a tap gesture 707c on the reply options 718. In some embodiments, the selection of the reply options 718 is an air gesture, for example, indicating the selection of the reply options 718. As Figure 7H shown, in response to detecting the selection of the reply options 718, the device 700 displays a reply interface 750 that can be used to send a message to the contact 712b.

[0238] In some embodiments, Figure 7C1 the techniques and user interfaces described in Figures 1A to 1P are provided by one or more of the devices described in Figure 7C2 Illustrated are embodiments in which, for example, as described in Figure 7B and Figure 7C1 a communication interface X710 is displayed on a display module X702 of a head-mounted device (HMD) X700. In some embodiments, the device X700 includes a pair of display modules that provide stereoscopic content to different eyes of the same user. For example, the HMD X700 includes a display module X702 that provides content to the user's left eye and a second display module that provides content to the user's right eye. In some embodiments, the second display module displays an image that is slightly different from the display module X702 to create an illusion of stereoscopic depth.

[0239] As Figure 7C2 shown, in response to detecting a selection of a contact X712e, the HMD X700 displays a contact menu X720 associated with the contact X712e. In some embodiments, the HMD X700 detects the selection of the contact X712e based on an air gesture performed by a user of the HMD X700. In some embodiments, the HMD X700 detects a hand X750a and / or X750b of a user of the HMD X700 and determines whether the movement of the hand X750a and / or X750b performs a pre-determined air gesture corresponding to the selection of the contact X712e. In some embodiments, the pre-determined air gesture for selecting the contact X712e includes a pinch gesture. In some embodiments, the pinch gesture includes detecting the movement of a finger X750c and a thumb X750d towards each other. In some embodiments, the HMD X700 detects the selection of the contact X712e based on gaze and air gesture input performed by a user of the HMD X700. In some embodiments, the gaze and air gesture input includes detecting that the user of the HMD X700 is looking at the contact X712e (e.g., for a duration greater than a pre-determined amount of time) and the hand X750a and / or X750b of the user of the HMD X700 performs a pinch gesture.

[0240] The contact menu X720 includes an invitation option X720a and an expansion option X720b. The invitation option X720a, when selected (e.g., via an air gesture such as a pinch gesture and / or via a gaze and pinch gesture), causes the HMD X700 to invite the contact X712e to an XR communication session. The expansion option X720b, when selected, causes the HMD X700 to display one or more additional options for managing the contact X712e. For example, when the contact menu X720 is displayed, the HMD X700 detects a selection of the expansion option X720b. In some embodiments, the selection of the expansion option X720b is, for example, an air gesture (e.g., a pinch gesture, and / or a gaze and pinch gesture) that indicates the selection of the expansion option X720b. In response to detecting the selection of the expansion option X720b, the HMD X700 expands the contact menu X720 to display (e.g., replaces the display of the expansion option X720b with) one or more additional options (e.g., options 720c - 720f, as Figure 7D illustrated).

[0241] In some embodiments, when expanded, the contact menu X720 includes an audio option (e.g., 720c), a messaging option (e.g., 720d), an information option (e.g., 720e), and an edit option (e.g., 720f), for example, as described with respect to Figure 7D The audio option, when selected, causes the HMD X700 to initiate an audio communication session with the contact X712e (e.g., without a live video component). In some embodiments, the HMD X700 is not capable of communicating via a cellular network and / or is configured to use an external device to make an audio call. Thus, in some examples, the HMD X700 uses a nearby device (e.g., a mobile phone and / or a tablet) that is capable of communicating via a cellular network to initiate the audio communication session. The edit option, when selected, allows a user of the HMD X700 to remove the contact X712e from the top contacts X712 (or add the contact X712e to the top contacts in embodiments where the contact X712e is not already a top contact), for example, as described with respect to Figure 7D The messaging option, when selected, allows the user to send a message to the contact X712e, for example, as described with respect to Figure 7D For example, when the expanded contact menu X720 is displayed, the HMD X700 detects a selection of the messaging option (e.g., 720d). In some embodiments, the selection of the messaging option is, for example, an air gesture (e.g., a pinch gesture, and / or a gaze and pinch gesture) that indicates the selection of the messaging option. In some embodiments, as Figure 7EAs shown, in response to detecting a selection of a message option (e.g., 720d), the HMD X700 displays (e.g., replaces the display of the communication interface X710 with) a message interface (e.g., 730). Thereafter, the message interface can be used to transmit a message to the contact X712e.

[0242] In some embodiments, the contact menu associated with a contact includes one or more additional options based on the state of the HMD X700. As an example, in some embodiments, in an instance where the HMD X700 is participating in a live communication session (e.g., an XR communication session or an audio communication session), the contact menu includes an option to invite the contact to the live communication session. For example, while participating in an XR communication session and while the communication interface X710 is being displayed, the HMD X700 detects a selection of the contact X714g ("Dylan"). In some embodiments, the selection of the contact X714g is, for example, an air gesture (e.g., a pinch gesture, and / or a gaze and pinch gesture) indicating a selection of the contact X714f. As Figure 7C2 shown, in response to detecting the selection of the contact X714g, the HMD X700 displays the contact menu X724 associated with the contact X714g.

[0243] The contact menu X724 includes an invite option X724a, an invite option X724b, and an expand option X724c. The invite option X724a, when selected (e.g., via an air gesture such as a pinch gesture, and / or a gaze and pinch gesture), causes the HMD X700 to invite the contact X714g to a new live communication session. The invite option X724b, when selected (e.g., via an air gesture, such as a pinch gesture, and / or a gaze and pinch gesture), causes the HMD X700 to invite the contact X714g to the live communication session in which the HMD X700 is currently participating. The expand option X724c, when selected, causes the HMD X700 to display one or more additional options for the contact X714g.

[0244] In some embodiments, inviting a contact to a new live communication session (e.g., in response to a selection of the option X724a) will cause the HMD X700 to disconnect from and / or terminate the live communication session in which the HMD X700 is currently participating. In some embodiments, before terminating an existing live communication session in this manner, the HMD X700 confirms that the user wishes to disconnect from the current live communication session before initiating a new live communication session. For example, as Figure 7GAs shown, in response to a selection of invitation option X724a, HMD X700 can display a confirmation interface 740 that includes a confirmation affordance representation 742. In response to a selection of the confirmation affordance representation 742 (e.g., via an air gesture such as a pinch gesture, and / or a gaze and pinch gesture), HMD X700 terminates the current live communication session and invites contact X714f to a new live communication session.

[0245] In some embodiments, the user optionally uses communication interface X710 to send a message to a contact. For example, when communication interface X710 is displayed, HMD X700 detects a selection of a preview 716b associated with a pinned contact X712b (e.g., as Figure 7B shown). In some embodiments, the selection of preview 716b is, for example, an air gesture (e.g., a pinch gesture, and / or a gaze and pinch gesture) that indicates the selection of preview 716b. As Figure 7C2 shown, in response to detecting the selection of preview 716b, HMD X700 expands preview 716b to display reply options X718.

[0246] Reply options X718, when selected, cause HMD X700 to display a reply interface for sending a message to contact X712b. For example, when reply options X718 are displayed in preview 716b, HMD X700 detects a selection of reply options X718. In some embodiments, the selection of reply options X718 is, for example, an air gesture (e.g., a pinch gesture, and / or a gaze and pinch gesture) that indicates the selection of reply options X718. As Figure 7H shown, in response to detecting the selection of reply options X718, HMD X700 can display a reply interface 750 that can be used to send a message to contact X712b.

[0247] Figures 1B to 1PAny of the features, components, and / or parts shown, including their arrangement and configuration, may be included in the HMD X700 individually or in any combination. For example, in some embodiments, the HMD X700 includes, individually or in any combination, any of the features, components, and / or parts of HMDs 1-100, 1-200, 3-100, 6-100, 6-200, 6-300, 6-400, 11.1.1-100, and / or 11.1.2-100. In some embodiments, the display module X702 includes, individually or in any combination, any of the features, components, and / or parts of display unit 1-102, display unit 1-202, display unit 1-306, display unit 1-406, display generation component 120, display screen 1-122a-b, first rear display screen 1-322a, and second rear display screen 1-322b, display 11.3.2-104, first display assembly 1-120a and second display assembly 1-120b, display assembly 1-320, display assembly 1-421, first display subassembly 1-420a and second display subassembly 1-420b, display assembly 3-108, display assembly 11.3.2-204, first optical module 11.1.1-104a and second optical module 11.1.1-104b, optical module 11.3.2-100, optical module 11.3.2-200, bi-convex lens array 3-110, display area or zone 6-232, and / or display / display area 6-334. In some embodiments, the HMD X700 includes a sensor that includes, individually or in any combination, any of the features, components, and / or parts of sensor 190, sensor 306, image sensor 314, image sensor 404, sensor assembly 1-356, sensor assembly 1-456, sensor system 6-102, sensor system 6-202, sensor 6-203, sensor system 6-302, sensor 6-303, sensor system 6-402, and / or sensor 11.1.2-110a-f. In some embodiments, the input device X703 includes, individually or in any combination, any of the features, components, and / or parts of first button 1-128, button 11.1.1-114, second button 1-132, and / or dial or button 1-328. In some embodiments, the HMD X700 includes one or more audio output components (e.g., electronic component 1-112) for generating audio feedback (e.g., audio output), which is optionally generated based on detected events and / or user input detected by the HMD X700.

[0248] In some embodiments, the communication interface 710 is used to generate an avatar. In some embodiments, the avatar serves as a representation (e.g., a 3D representation) of the user of device 700 in an XR communication session. For example, referring to Figure 7B , when displaying the communication interface 710, device 700 detects a selection of the avatar option 715. In some embodiments, the selection of the avatar option 715 is a tap gesture 713b on the avatar option 715. In some embodiments, the selection of the avatar option 715 is an air gesture, for example, indicating the selection of the avatar option 715. As Figure 7I shown, in response to detecting the selection of the avatar option 715, device 700 displays the avatar interface 760.

[0249] At Figure 7I , the avatar interface 760 includes a first option 762 (e.g., more realistic compared to the second option) and a second option 764 (e.g., less realistic compared to the first option). When the first option 762 is enabled, the avatar of the user of device 700 reflects the appearance of the user. For example, in some embodiments, when the first option 762 is enabled, the avatar includes one or more visual characteristics corresponding to one or more physical characteristics of the user. When the second option 764 is enabled, the avatar of the user of device 700 indicates the movement of the user (e.g., during a live communication session) without reflecting the appearance of the user. For example, in some embodiments, when the second option 764 is enabled, an avatar with a default appearance is used. In some embodiments, when the first option 762 (compared to the second option 764) is enabled, the avatar of the user of device 700 is represented by a first representation style, and the avatar is displayed at a first level of detail (e.g., a first level of detail with respect to the appearance of the user and / or one or more parts of the user), and indicates the positioning and movement of a first user part of the user relative to a second user part of the user in a first manner. In some embodiments, when the second option 764 (compared to the first option 762) is enabled, the avatar of the user of device 700 is represented by a second representation style different from the first representation style, and the avatar is displayed at a second level of detail lower than the first level of detail (e.g., less than the appearance of the user and / or mimicking the appearance of the user with less detail and / or a lower amount of detail), and indicates the positioning and movement of a first user part of the user relative to a second user part of the user in a second manner different from the first manner.

[0250] The avatar interface 760 further includes a menu option 766, which when selected causes device 700 to display an avatar menu, as Figure 7I shown. For example, at Figure 7IAt, when the avatar interface 760 is displayed, the device 700 detects a selection of the menu option 766. In some embodiments, the selection of the menu option 766 is a tap gesture 705i on the menu option 766. In some embodiments, the selection of the menu option 766 is an air gesture, for example, indicating the selection of the menu option 766. As Figure 7J shown, in response to detecting the selection of the menu option 766, the device 700 displays an avatar menu 768.

[0251] At Figure 7J this point, the avatar menu 768 includes an edit option 768a, a create option 768b, and / or a delete option 768c. In some embodiments, if the user of the device 700 has not created an avatar, the avatar menu 768 includes the create option 768b and does not include the edit option 768a and the delete option 768c. In some embodiments, if the user of the device 700 has created an avatar, the avatar menu 768 includes the edit option 768a and the delete option 768c and does not include the create option 768b.

[0252] At Figure 7J this point, when the avatar menu 768 is displayed, the device 700 detects a selection of the create option 768b. In some embodiments, the selection of the create option 768b is a tap gesture 705j on the create option 768b. In some embodiments, the selection of the create option 768b is an air gesture, for example, indicating the selection of the create option 768b. As Figure 7K shown, in response to detecting the selection of the create option 768b, the device 700 displays a settings interface 770.

[0253] At Figure 7K this point, the settings interface 770 includes a settings option 772 that, when selected, causes the device 700 to display an avatar editing interface. For example, when the settings interface 770 is displayed, the device 700 detects a selection of the settings affordance 772. In some embodiments, the selection of the settings affordance 772 is a tap gesture 705k on the settings affordance 772. In some embodiments, the selection of the settings affordance 772 is an air gesture, for example, indicating the selection of the settings affordance 772. As Figure 7L1 shown, in response to detecting the selection of the settings affordance 772, the device 700 displays an avatar editing interface 780.

[0254] At Figure 7L1At this location, the avatar editing interface 780 includes a live view 781 of the user's avatar of the device 700. In some embodiments, the device 700 displays the live view 781 of the avatar editing interface 780 based on the movement and / or demeanor of the user of the device 700 as detected by the device 700, and the live view is updated in real time. The avatar editing interface 780 also includes various settings and / or parameters for adjusting the visual characteristics of the avatar. As an example, the avatar editing interface 780 includes settings 782, and these settings include a brightness setting 782a and a warmth setting 782b. The brightness setting 782a and the warmth setting 782b are respectively used to adjust the simulated lighting and the temperature of the avatar's skin. As another example, the avatar editing interface 780 includes a color palette 783, and the color palette includes a set of one or more colors and / or shades from which to select the color of the avatar's skin.

[0255] In some embodiments, as Figure 7L1 illustrated, the avatar editing interface 780 includes a set of parameters 784, such as a shirt parameter 784a and a headgear parameter 784b. In some embodiments, selecting a parameter allows for the selection of the visual characteristics of one or more aspects of the avatar. For example, referring Figure 7M to, the selection of the shirt parameter 784a (e.g., a tap input 705l or an air gesture corresponding to the location of the parameter 784a) causes the device 700 to display a parameter menu 790, and the user can select from any number of options (e.g., options 790a - 790c) for the shirt parameter 784a according to this parameter menu. Referring Figure 7N to, once an option has been selected (e.g., a tap input 705m or an air gesture corresponding to the location of the option 790b), the user can select from the styles 792 (e.g., 792a - 792f) of the selected option, and accordingly update the visual characteristics of the avatar. Similarly, in some embodiments, the selection of the headgear parameter 784b causes the device 700 to display headgear options for the headgear parameter 784b, and the selection of an option causes the device 700 to display the type of the selected option.

[0256] Although the present disclosure is described with respect to parameters 784a and 784b corresponding to a shirt and a headgear, respectively, it will be understood that in some embodiments, the parameters of the avatar editing interface 780 optionally correspond to other / additional visual aspects of the avatar. By way of example, in some embodiments, parameter 784 is used to select one or more aspects of an eye-worn device of the avatar (e.g., parameter 784a corresponds to glasses and parameter 784b corresponds to an eye patch). In an example where the device 700 receives a user selection of parameter 784a corresponding to glasses, the device 700 displays options for various designs of the glasses (e.g., frameless, thin frame, thick frame, etc.). Once the device 600 receives the user selection of the design, the device 700 displays various styles of the selected design as type 792 for the user to select. In an example where the device 700 receives a user selection of an eye patch, the device 700 displays options for various designs of the eye patch (e.g., left eye patch or right eye patch). Once the device 600 receives the user selection of the design, the device 700 displays various styles of the selected design as type 792 for the user to select.

[0257] In some embodiments, the avatar editing interface 780 includes a set of parameters 786. As shown, in some embodiments, parameter 786 is used to select one or more aspects of hair. By way of example, parameter 786a corresponds to a hairstyle, parameter 786b corresponds to hair color, and parameter 786c corresponds to hair highlights. In other embodiments, parameter 786 is used to select one or more aspects of an accessibility feature. By way of example, in some embodiments, parameter 786a corresponds to a hand prosthesis, parameter 786b corresponds to a hearing aid, and parameter 786c corresponds to a wheelchair.

[0258] In some embodiments, Figure 7L1 the techniques and user interfaces described in Figures 1A to 1P are provided by one or more of the devices described in Figure 7L2 Illustrates an embodiment in which (e.g., as described in Figures 7L1 to 7N the avatar editing interface X780 is displayed on the display module X702 of a head-mounted device (HMD) X700. In some embodiments, the device X700 includes a pair of display modules that provide stereoscopic content to different eyes of the same user. For example, the HMD X700 includes a display module X702 that provides content to the left eye of the user and a second display module that provides content to the right eye of the user. In some embodiments, the second display module displays an image that is slightly different from the display module X702 to create an illusion of stereoscopic depth.

[0259] In Figure 7L2At location, the avatar editing interface X780 includes a live view X781 of the user's avatar of the HMD X700. In some embodiments, the HMD X700 displays the live view X781 of the avatar editing interface X780 according to the movement and / or demeanor of the user of the HMD X700 as detected by the HMD X700, and the live view is updated in real time. The avatar editing interface X780 also includes various settings and / or parameters for adjusting the visual characteristics of the avatar. As an example, the avatar editing interface X780 includes settings X782, which include a brightness setting X782a and a warmth setting X782b. The brightness setting X782a and the warmth setting X782b are used to adjust the simulated lighting and the temperature of the avatar's skin, respectively. As another example, the avatar editing interface X780 includes a color palette X783, which includes a set of one or more colors and / or shades from which to select the color of the avatar's skin.

[0260] In some embodiments, as Figure 7L2 illustrated, the avatar editing interface X780 includes a set of parameters X784, such as a shirt parameter X784a and a headgear parameter X784b. In some embodiments, selecting a parameter allows the selection of the visual characteristics of one or more aspects of the avatar. For example, the selection of the shirt parameter X784a (e.g., a gaze and pinch gesture, where the gaze is indicated by a gaze indicator X705L corresponding to the location of the parameter X784a) causes the HMD X700 to display a parameter menu 790, from which the user can select from any number of options (e.g., options 790a to 790c) for the shirt parameter X784a, as Figure 7M shown.

[0261] In some embodiments, the HMD X700 detects a selection of a shirt parameter X784a based on an air gesture performed by a user of the HMD X700. In some embodiments, the HMD X700 detects hands X750a and / or X750b of a user of the HMD X700 and determines whether the movement of hands X750a and / or X750b performs a predetermined air gesture corresponding to a selection of a shirt parameter X784a. In some embodiments, the predetermined air gesture for selecting a shirt parameter X784a includes a pinching gesture. In some embodiments, the pinching gesture includes detecting a movement of finger X750c and thumb X750d towards each other. In some embodiments, the HMD X700 detects a selection of a shirt parameter X784a based on gaze and air gesture inputs performed by a user of the HMD X700. In some embodiments, the gaze and air gesture inputs include detecting that the user of the HMD X700 is looking at a shirt parameter X784a (e.g., for a duration greater than a predetermined amount of time) and that hands X750a and / or X750b of the user of the HMD X700 perform a pinching gesture.

[0262] Reference Figure 7N , once an option has been selected (e.g., a tap input 705m or an air gesture corresponding to the location of option 790b), the user can select from the styles 792 (e.g., 792a - 792f) of the selected option and update the visual characteristics of the avatar accordingly. Similarly, in some embodiments, a selection of a headgear parameter 784b causes the device 700 to display headgear options for the headgear parameter 784b, and a selection of an option causes the device 700 to display the type of the selected option.

[0263] Although the present disclosure is described with respect to parameters X784a and X784b corresponding to a shirt and a headgear, respectively, it will be understood that in some embodiments, the parameters of the avatar editing interface X780 optionally correspond to other / additional visual aspects of the avatar. By way of example, in some embodiments, the parameter X784 is used to select one or more aspects of an eye-wear device of the avatar (e.g., the parameter X784a corresponds to glasses and the parameter X784b corresponds to an eye patch). In an example where the HMD X700 receives a user selection of the parameter X784a corresponding to glasses, the HMD X700 displays options for various designs of glasses (e.g., frameless, thin frame, thick frame, etc.). Once the HMD X700 receives a user selection of a design, the HMD X700 displays various styles of the selected design as type 792 for the user to select. In an example where the HMD X700 receives a user selection of an eye patch, the HMD X700 displays options for various designs of the eye patch (e.g., left eye patch or right eye patch). Once the HMD X700 receives a user selection of a design, the HMD X700 displays various styles of the selected design as type 792 for the user to select.

[0264] In some embodiments, the avatar editing interface X780 includes a set of parameters X786. As shown, in some embodiments, the parameter X786 is used to select one or more aspects of hair. By way of example, the parameter X786a corresponds to a hairstyle, the parameter X786b corresponds to a hair color, and the parameter X786c corresponds to hair highlights. In other embodiments, the parameter X786 is used to select one or more aspects of an accessibility feature. By way of example, in some embodiments, the parameter X786a corresponds to a hand prosthesis, the parameter X786b corresponds to a hearing aid, and the parameter X786c corresponds to a wheelchair.

[0265] Figures 1B to 1PAny of the features, components, and / or parts shown, including their arrangement and configuration, may be included in the HMD X700 either individually or in any combination. For example, in some embodiments, the HMD X700 includes, either individually or in any combination, any of the features, components, and / or parts of the HMDs 1-100, 1-200, 3-100, 6-100, 6-200, 6-300, 6-400, 11.1.1-100, and / or 11.1.2-100. In some embodiments, the display module X702 includes, either individually or in any combination, any of the features, components, and / or parts of the display unit 1-102, display unit 1-202, display unit 1-306, display unit 1-406, display generation component 120, display screen 1-122a-b, first rear display screen 1-322a, and second rear display screen 1-322b, display 11.3.2-104, first display assembly 1-120a and second display assembly 1-120b, display assembly 1-320, display assembly 1-421, first display subassembly 1-420a and second display subassembly 1-420b, display assembly 3-108, display assembly 11.3.2-204, first optical module 11.1.1-104a and second optical module 11.1.1-104b, optical module 11.3.2-100, optical module 11.3.2-200, bi-convex lens array 3-110, display area or zone 6-232, and / or display / display area 6-334. In some embodiments, the HMD X700 includes a sensor that includes, either individually or in any combination, any of the features, components, and / or parts of the sensor 190, sensor 306, image sensor 314, image sensor 404, sensor assembly 1-356, sensor assembly 1-456, sensor system 6-102, sensor system 6-202, sensor 6-203, sensor system 6-302, sensor 6-303, sensor system 6-402, and / or sensor 11.1.2-110a-f. In some embodiments, the input device X703 includes, either individually or in any combination, any of the features, components, and / or parts of the first button 1-128, button 11.1.1-114, second button 1-132, and / or dial or button 1-328. In some embodiments, the HMD X700 includes one or more audio output components (e.g., electronic component 1-112) for generating audio feedback (e.g., audio output), which is optionally generated based on detected events and / or user input detected by the HMD X700.

[0266] Reference Figure 7N, when the avatar editing interface 780 is displayed, the device 700 detects a selection of the save option 788. In some embodiments, the selection of the save option 788 is a tap gesture 705n on the save option 788. In some embodiments, the selection of the save option 788 is an air gesture, for example, indicating the selection of the save option 788. In response to detecting the selection of the save option 788, the device 700 stores (e.g., locally and / or remotely) the configuration of the avatar selected for the user of the device 700 for subsequent use in an XR communication session. As Figure 7O shown, further in response to detecting the selection of the settings enabling representation 772, the device 700 displays a completion interface 795, which indicates that the user has successfully created and / or updated the user's avatar.

[0267] In Figure 7P , the user of the device 700 is participating in an XR communication session with a contact 712f (“Ann” in Figure 7B ) within the XR environment 704. In some embodiments, the XR communication session is a spatial communication session. Thus, in some embodiments, the contact 712f and / or the user of the device 700 are represented in the XR environment 704 by 3D representations (e.g., avatars). For example, as Figure 7P shown, the user of the device 700 is represented by a representation 700A (as shown in the self-preview 706A), and the contact 712f is represented by a representation 701A.

[0268] In some embodiments, the view of the XR environment 704 for the user of the device 700 is provided from the perspective of the representation 700A within the XR environment 704. Since this may prevent the user from viewing the representation 700A in other ways, the device 700 displays a self-preview 706A that includes a live view of the representation 700A in the XR environment 704. Although the self-preview 706A is shown in the lower right corner of the display 702, it will be understood that the self-preview 706A may optionally be displayed at any location on the display 702. By way of example, in some embodiments, the self-preview 706A is positioned adjacent to the representation of the contact in the XR communication session. In some embodiments, the self-preview 706A is located, for example, at a position 708A close to the representation 701A.

[0269] In some embodiments, a participant represented by a 3D representation in an XR environment has spatial agency. Thus, during an XR communication session, the 3D representation optionally moves within the XR environment 704 such that the 3D representation moves relative to elements in the XR environment 704 (e.g., table 704a and couch 704b) and / or other participants. In some embodiments, the 3D representation moves in accordance with the movement of the corresponding device. By way of example, the 3D representation 700A can move within the XR environment 704 in response to the movement of device 700. In some embodiments, the 3D representation moves in a manner corresponding to the movement of the device. For example, if device 700 first moves in a first direction (e.g., left) and then in a second direction (e.g., right), the 3D representation 700A will move within the XR environment 704 in a similar manner in the first and second directions.

[0270] In some embodiments, when participating in an XR communication session, device 700 displays a set of controls 704A for managing one or more aspects of the XR communication session, as Figure 7P shown. The set of controls 704A includes a message option 704Aa, an information option 704Ab, a microphone option 704Ac, an avatar option 704Ad, a camera option 704Ae, and a termination option 704Af. The message option 704Aa, when selected, causes device 700 to display a message interface for transmitting a message to contact 712f. The information option 704Ab, when selected, causes device 700 to display an information interface corresponding to contact 712f. The microphone option 704Ac, when selected, toggles the state of the microphone of device 700 (e.g., enables or disables the microphone). In some embodiments, disabling the microphone of device 700 prevents device 700 from providing audio during the XR communication session. The camera option 704Ae, when selected, toggles the state of the camera of device 700 (e.g., enables or disables the camera). In some embodiments, disabling the camera of device 700 prevents device 700 from providing video during the XR communication session (e.g., a video feed of the user of device 700 and / or movement of the representation of the user of device 700). The termination option 704Af, when selected, causes device 700 to disconnect from the XR communication session.

[0271] The avatar option 704Ad, when selected, toggles (e.g., enables or disables) the use of the 3D representation in the XR environment 704. For example, when displaying the XR environment 704, device 700 detects a selection of the avatar option 704Ad. In some embodiments, the selection of the avatar option 704Ad is a tap gesture 705p on the avatar option 704Ad. In some embodiments, the selection of the avatar option 704Ad is an air gesture, for example, indicating the selection of the avatar option 704Ad. As Figure 7QAs shown, in response to the selection of avatar option 704Ad, device 700 disables the use of 3D representations in XR environment 704.

[0272] In some embodiments, when toggling the use of 3D representations in XR environment 704, device 700 causes the XR communication session to transition between a spatial communication session and a non-spatial communication session. Disabling the use of 3D representations causes device 700 to transition the XR communication session from a spatial communication session to a non-spatial communication session. Enabling the use of 3D representations causes device 700 to transition the XR communication session from a non-spatial communication session to a spatial communication session.

[0273] In some embodiments, in a non-spatial communication session, participants in the XR communication session are represented by 2D representations. By way of example, as Figure 7Q shown, the user of device 700 is represented by 2D representation 710A (as shown in self-preview 706A), and contact 712f is represented by 2D representation 712A.

[0274] In some embodiments, the 2D representation includes a video feed of the user (e.g., a live video feed). In some embodiments, if the video feed of the user is unavailable (e.g., the device's camera is disabled), then the 2D representation of the user instead includes an image associated with the user (e.g., a thumbnail), a monogram corresponding to the user, and / or another 2D representation. In some embodiments, the user represented by the 2D representation in the XR environment does not have a spatial proxy and is optionally located at one or more predetermined locations in XR environment 704. In some embodiments, device 600 is configured to move the 2D representation of a remote participant in the XR environment based on user input received at device 700 (e.g., an input to drag the representation from a first location to a second location). In some embodiments, device 600 is not configured to move the 3D representation of a remote participant in the XR environment based on user input received at device 700.

[0275] Additional description regarding Figures 7A to 7Q is provided below with reference to methods 800 and 900, each of which is described with respect to Figures 7A to 7Q

[0276] Figure 8 is a flowchart of an exemplary method 800 for managing a live communication session according to some embodiments. In some embodiments, method 800 is executed at a computer system (e.g., Figure 1A computer system 101, computer system 700, and / or HMD X700 in Figure 1A ,Figure 3 and Figure 4 the display generation component 120, the display 702, and / or the display X702) in (e.g., a visual output device, a 3D display, a display having at least a portion that is transparent or translucent onto which an image can be projected (e.g., a see-through display), a projector, a heads-up display, and / or a display controller) and one or more sensors (e.g., a touch-sensitive surface, a gyroscope, an accelerometer, a motion sensor, a movement sensor, a microphone, an infrared sensor, a camera sensor, a depth camera, a visible light camera, an eye tracking sensor, a gaze tracking sensor, a physiological sensor, and / or an image sensor). In some embodiments, method 800 is governed by instructions stored in a non-transitory (or transitory) computer-readable storage medium and executed by one or more processors of a computer system such as one or more processors 202 of computer system 101 (e.g., Figure 1A the control 110) in. Some operations in method 800 are optionally combined, and / or the order of some operations is optionally changed.

[0277] The computer system (e.g., 700 and / or X700) displays (802) representations (e.g., 712a - 712g and 714a - 714i) (e.g., static avatars, animated avatars, images, and / or letter combinations) of multiple users (e.g., users who do not operate the computer system (remote users) and / or users other than the user of the computer system) via the display generation component.

[0278] The computer system (e.g., 700 and / or X700) receives (804) a selection (e.g., 705b and / or 711b) (e.g., via a touch input on a touch-sensitive surface and / or via an air gesture) of a representation (e.g., 712e, X712e, 714g, and / or X714g) (e.g., static avatars, animated avatars, images, and / or letter combinations) of a corresponding user among the multiple users via one or more sensors.

[0279] In response to (806) receiving the selection of the representation (e.g., 712e, X712e, 714g, and / or X714g) and based on a determination that there is an ongoing (e.g., active and / or currently established) communication session (e.g., a video communication session, an audio communication session, an extended reality communication session, a spatial communication session, and / or a non-spatial communication session), the computer system (e.g., 700 and / or X700) displays (808) options (e.g., Figure 7C1 724b at and / or Figure 7C2 X724b at) for inviting the corresponding user to join the ongoing communication session via the display generation component (e.g., 702 and / or X702).

[0280] In response to receiving a selection of a representation of a corresponding user (e.g., 712e, X712e, 714g, and / or X714g) and based on determining that there is no ongoing communication session, a computer system (e.g., 700 and / or X700) forgoes displaying (810) an option (e.g., as in menu 720 at Figure 7C1 and / or Figure 7C2 menu X720 at to invite the corresponding user to join an ongoing communication session. Conditionally displaying an option to invite a corresponding user to join an ongoing communication session enables a user of the computer system to invite the corresponding user without navigating to a different user interface, thereby reducing the amount of input required to perform the invitation operation.

[0281] In some embodiments, in response to receiving a selection of a representation of a corresponding user (e.g., 712e, X712e, 714g, and / or X714g) (e.g., independent of a determination of whether there is an ongoing communication session), a computer system (e.g., 700 and / or X700) displays, via a display generation component (e.g., 702 and / or X702), an option (e.g., Figure 7C1 720a and / or 724a at Figure 7C2 and / or Figure 7C1 X720a and / or X724a at Figure 7C2 ) to initiate a new spatial communication session with the corresponding user, as well as options for additional features (e.g., Figure 7C1 720b and / or 724c at Figure 7C2 and / or Figure 7C1 X720b and / or X724c at Figure 7C2 ) (e.g., without displaying options to send a text message to the corresponding user and / or display additional information about the corresponding user). In some embodiments, when displaying options for additional features (e.g., Figure 7C1 720a and / or 724a at Figure 7C2 and / or Figure 7C1 X720a and / or X724a at Figure 7C2 ), the computer system (e.g., 700 and / or X700) receives a selection of an option for an additional feature (e.g., Figure 7C1 720a and / or 724a at Figure 7C2 and / or [[ID=269at 720c - 720f) (e.g., one or more options for communicating with a corresponding user, such as by sending a text message to the corresponding user and / or displaying additional information about the corresponding user). In some embodiments, a spatial communication session is a communication session having representations of at least some (e.g., less than all, multiple, and / or all) of the users participating in the communication session distributed in a 3D environment. Displaying options to initiate a new spatial communication session with a corresponding user and options to access additional features enables a user of a computer system to quickly access options to initiate a new spatial communication session without cluttering the user interface, while still providing access to additional (and potentially less frequently used) features, thereby improving the human - machine interface.

[0282] In some embodiments, in response to receiving a selection of a representation of a corresponding user (e.g., independent of a determination of whether there is an ongoing communication session), a computer system (e.g., 700 and / or X700) displays, via a display generation component (e.g., 702 and / or X702), options for additional features (e.g., ​ at 720b and / or ​ at X720b) (e.g., without displaying options for sending a text message to the corresponding user and / or displaying additional information about the corresponding user). In some embodiments, when displaying options for additional features (e.g., ​ at 720b and / or ​ at X720b), the computer system (e.g., 700 and / or X700) receives, via one or more sensors, a selection of the options for additional features (e.g., Figure 7C1 at 720b and / or Figure 7C2 at X720b) (e.g., 705c) (e.g., via a touch input on a touch - sensitive surface and / or via an air gesture). In some embodiments, in response to receiving a selection of the options for additional features (e.g., Figure 7C1 at 720b and / or Figure 7C2 at X720b), the computer system (e.g., 700 and / or X700) displays, via a display generation component (e.g., 702 and / or X702) (e.g., by replacing the display of options to invite the corresponding user to join an ongoing communication session), options to initiate an audio communication session with the corresponding user (e.g., without including a live visual representation of the participants and / or without including a video component) (e.g., Figure 7D at 720c) (e.g., as part of one or more options associated with the corresponding user). In some embodiments, the computer system receives, via one or more sensors, a selection of the options to initiate an audio communication session with the corresponding user (e.g., Figure 7DSelection of 720c) at (e.g., via a touch input on a touch-sensitive surface and / or via an air gesture). In some embodiments, in response to receiving an option to initiate an audio communication session with a corresponding user (e.g., Figure 7D selection of 720c) at, the computer system (e.g., 700 and / or X700) initiates an audio communication session with the corresponding user (e.g., without other users) (e.g., not including a live visual representation of the participant and / or not including a video component). Displaying the option to initiate an audio communication session enables a user of the computer system to start a communication session that does not include a live visual representation of the user without having to initiate a video communication session and separately disable the video portion, thereby reducing the amount of input required to initiate an audio communication session.

[0283] In some embodiments, the computer system initiating an audio communication session includes using an external electronic device (e.g., a smart phone and / or a cellular phone) within a predetermined range (e.g., distance and / or wireless range) of the computer system (e.g., 700 and / or X700) to initiate an audio call (e.g., a voice call and / or a telephone call). In some embodiments, an option to initiate an audio communication session is displayed for a corresponding user who does not have an account that includes a particular online service (or does not have an active account) (e.g., the user of the computer system has an account that includes a particular online service for video and / or extended reality communication, but the corresponding user does not have that account). Using an external electronic device to initiate an audio communication session via an audio call enables the computer system to use the resources of an external computer system (e.g., the cellular connection and / or CPU processing of that external computer system), thereby improving the functionality of the computer system while reducing the workload of the computer system.

[0284] In some embodiments, in response to receiving a selection of a representation of a corresponding user (e.g., 712e, X712e, 714g, and / or X714g) (e.g., 705c and / or 711b) (e.g., independent of a determination of whether there is an ongoing communication session), a computer system (e.g., 700 and / or X700) displays, via a display generation component (e.g., 702 and / or X702), options for additional features (e.g., 720b, X720b, 724c, and / or X724C) (e.g., without displaying options for initiating a process for transmitting a message to the corresponding user and / or displaying additional information about the corresponding user). In some embodiments, when displaying options for additional features (e.g., 720b, X720b, 724c, and / or X724c), the computer system (e.g., 700 and / or X700) receives, via one or more sensors, a selection of an option for an additional feature (e.g., 720b and / or X720b) (e.g., 705c) (e.g., via a touch input on a touch-sensitive surface and / or via an air gesture). In some embodiments, in response to receiving a selection of an option for an additional feature (e.g., 720b and / or X720b) (e.g., 705c), the computer system (e.g., 700 and / or X700) displays, via the display generation component (e.g., 702 and / or X702), options for a process for initiating a message transmission to the corresponding user (e.g., a live transmission excluding audio and / or video) (e.g., by replacing the display of an option for inviting the corresponding user to join an ongoing communication session) (e.g., as part of one or more options associated with the corresponding user). In some embodiments, the computer system receives, via one or more sensors, a selection of an option for a process for initiating a message transmission to the corresponding user (e.g., 720d) (e.g., 705d) (e.g., via a touch input on a touch-sensitive surface and / or via an air gesture). In some embodiments, in response to receiving a selection of an option for a process for initiating a message transmission to the corresponding user (e.g., 720d) (e.g., 705d), the computer system (e.g., 700 and / or X700) initiates a message transmission to the corresponding user (e.g., Figure 7E at 730 and / or Figure 7HProcess at 750) (e.g., instead of transmitting a message to another user) that (e.g., does not include a live visual representation of the participant and / or does not include a video component). In some embodiments, the process of initiating the transmission of a message to a corresponding user includes displaying a user interface including a conversation between the user of the computer system and the corresponding user, displaying a keyboard and / or displaying a text entry field for entering a message. Providing an option to initiate the process of transmitting a message to a corresponding user via additional feature selection enables the user of the computer system to quickly initiate the process without specifying a recipient, thereby reducing the amount of input required to transmit the message.

[0285] In some embodiments, in response to receiving a selection of a representation of a corresponding user (e.g., 712e, X712e, 714g, and / or X714g) (e.g., 705c and / or 711b) (e.g., independent of a determination of whether there is an ongoing communication session), a computer system (e.g., 700 and / or X700) displays, via a display generation component (e.g., 702 and / or X702), options for additional features (e.g., 720b, X720b, 724c, and / or X724c) (e.g., without displaying a process for initiating a message transmission to the corresponding user and / or options for displaying additional information about the corresponding user). In some embodiments, when displaying options for additional features (e.g., 720b, X720b, 724c, and / or X724C), the computer system receives, via one or more sensors, a selection of the options for additional features (e.g., 720b, X720b, 724c, and / or X724c) (e.g., via a touch input on a touch-sensitive surface and / or via an air gesture). In some embodiments, in response to receiving a selection of the options for additional features (e.g., 720b, X720b, 724c, and / or X724c), the computer system displays, via the display generation component, options (e.g., 720e) for displaying additional information about the corresponding user (e.g., by replacing the display of options for inviting the corresponding user to join an ongoing communication session) (e.g., as part of one or more options associated with the corresponding user). In some embodiments, the computer system receives, via one or more sensors, a selection of the options (e.g., 720e) for displaying additional information about the corresponding user (e.g., 707d) (e.g., via a touch input on a touch-sensitive surface and / or via an air gesture). In some embodiments, in response to receiving a selection of the options (e.g., 720e) for displaying additional information about the corresponding user (e.g., 707d), the computer system (e.g., 700 and / or X700) displays, via the display generation component (e.g., 702 and / or X702), additional information about the corresponding user (e.g., not displayed when the options for additional features are selected and / or not displayed when the options for additional information are selected) (e.g., Figure 7F 740 at) (e.g., the previous communication history of the corresponding user, the phone number of the corresponding user, and / or the email address of the corresponding user) (e.g., without displaying additional information about other users). Providing additional information about the corresponding user to the user of the computer system provides feedback about the corresponding user and / or the corresponding user's device, thereby providing improved visual feedback.

[0286] In some embodiments, in response to receiving a selection (e.g., 705c and / or 711b) of a representation of a corresponding user (e.g., 712e, X712e, 714g, and / or X714g) (e.g., independent of a determination of whether an ongoing communication session exists), a computer system (e.g., 700 and / or X700) displays, via a display generation component (e.g., 702 and / or X702), options for additional features (e.g., 720b, X720b, 724c, and / or X724c) (e.g., without displaying a process for initiating a message transfer to the corresponding user and / or options for displaying additional information about the corresponding user). In some embodiments, when displaying options for additional features (e.g., 720b, X720b, 724c, and / or X724c), the computer system (e.g., 700 and / or X700) receives, via one or more sensors, a selection of an option for an additional feature (e.g., 705c) (e.g., via a touch input on a touch-sensitive surface and / or via an air gesture). In some embodiments, in response to receiving a selection (e.g., 705c) of an option for an additional feature (e.g., 720b and / or X720b), the computer system displays, via the display generation component, an option (e.g., 720f) for removing the corresponding user from a favorites page (e.g., a list or grouping of favorite users) (e.g., by replacing a display of an option for inviting the corresponding user to join an ongoing communication session) (e.g., as part of one or more options associated with the corresponding user). In some embodiments, removing the corresponding user from the favorites page includes stopping displaying the representation of the corresponding user as part of a representation of multiple users (e.g., a static avatar, an animated avatar, an image, and / or a logo) (e.g., the representation is optionally displayed as part of a primary user interface). In some embodiments, the computer system receives, via one or more sensors, a selection (e.g., via a touch input on a touch-sensitive surface and / or via an air gesture) of an option (e.g., 720f) for removing the corresponding user from the favorites page (e.g., 712). In some embodiments, in response to receiving a selection of an option (e.g., 720f) for removing the corresponding user from the favorites page (e.g., 712), the computer system (e.g., 700 and / or X700) initiates a process for removing the corresponding user from the favorites page (e.g., requests confirmation for removing the corresponding user from the favorites page and / or removes the corresponding user from the favorites page). Initiating the process for removing the corresponding user from the favorites page enables the user to limit the users accessible via the favorites page and / or the primary user interface, thereby reducing visual clutter and enabling other users to be added to the favorites page, thus providing improved visual feedback.

[0287] In some embodiments, in response to receiving a selection (e.g., 711b) of a representation of a corresponding user (e.g., 714g and / or X714g) and based on a determination that there is an ongoing (e.g., active and / or currently established) communication session (e.g., a video communication session, an audio communication session, an extended reality communication session, a spatial communication session, and / or a non-spatial communication session), a computer system (e.g., 700 and / or X700) simultaneously displays, via a display generation component (e.g., 702 and / or X702), an option (e.g., 724a and / or X724a) to initiate a process to start a new communication session with the corresponding user and an option (e.g., 724b and / or X724b) to invite the corresponding user to join the ongoing communication session. In some embodiments, the computer system (e.g., 700 and / or X700) receives a selection (e.g., 709c) of the option (e.g., 724a and / or X724a) to initiate a process to start a new communication session with the corresponding user via one or more sensors (e.g., via a touch input on a touch-sensitive surface and / or via an air gesture). In some embodiments, in response to receiving a selection (e.g., 709c) of the option (e.g., 724a and / or X724a) to initiate a process to start a new communication session with the corresponding user, the computer system (e.g., 700 and / or X700) initiates a process to end the ongoing communication session (e.g., Figure 7G 740 at) and start a new communication session with the corresponding user. In some embodiments, in response to receiving a selection of the option to initiate a process to start a new communication session with the corresponding user, the computer system automatically (e.g., without requiring summation and / or receiving additional input from the user and / or without requesting user confirmation) ends the ongoing communication session and starts a new communication session with the corresponding user. Providing the option to start a new communication session with the corresponding user enables the computer system to end the ongoing communication session and start a new communication session without requiring a separate user input to end the ongoing communication session and start a new communication session, thereby reducing the number of inputs required to perform the operation.

[0288] In some embodiments, during the process of starting a new communication session with the corresponding user (e.g., Figure 7GDuring the 740) period at [location], the computer system (e.g., 700 and / or X700) prompts (e.g., 742) (e.g., via audio using a speaker and / or via a display using a display generation component) to confirm ending an ongoing communication session. In some embodiments, the computer system (e.g., 700 and / or X700) receives confirmation to end the ongoing communication session via one or more sensors (e.g., when the display prompts to confirm ending the ongoing communication session). In some embodiments, in response to receiving confirmation to end the ongoing communication session, the computer system ends the ongoing communication session (and optionally, starts a new communication session with the corresponding user). Requesting confirmation from the user of the computer system to end the ongoing communication session enables the computer system to avoid the user inadvertently ending the ongoing communication session, thus improving the human-machine interface.

[0289] In some embodiments, representations of multiple users (e.g., static avatars, animated avatars, images, and / or monograms) are displayed as part of the primary user interface (e.g., optionally including representations of recently communicated contacts) (e.g., as Figures 7B to 7D shown), and displaying options (e.g., 724b and / or X724b) to invite the corresponding user to join an ongoing communication session and / or displaying one or more options associated with the corresponding user (e.g., 720a - 720b, X720a - X720b, 724a - 724c, and / or X724a - X724c) includes obscuring the primary user interface (e.g., partially blocking the display of the primary user interface, blurring the primary user interface, and / or otherwise partially obscuring the primary user interface). In some embodiments, the primary user interface includes multiple user interface objects for displaying corresponding applications (e.g., a first user interface object that, when activated, causes the display of the user interface of a first application, and a second user interface object that, when activated, causes the display of the user interface of a second application different from the first application). In some embodiments, in response to detecting corresponding user input (e.g., detecting the pressing of a physical button and / or detecting a corresponding gesture, such as an air gesture), the computer system displays the primary user interface (e.g., regardless of what the computer system is displaying when the corresponding user input is received). In some embodiments, after the computer system exits a low-power mode (e.g., wakes up) and / or receives user input to unlock the computer system, the computer system automatically displays the primary user interface. Continuing to display the primary user interface (when obscured) provides the user with context about the content the user is accessing, including information about the corresponding user (e.g., name and / or contact information).

[0290] In some embodiments, one or more options associated with a corresponding user include an option (e.g., 720d) (e.g., as part of one or more options associated with the corresponding user) to initiate a process (e.g., excluding live transmission of audio and / or video) to transmit a message to the corresponding user. In some embodiments, a computer system (e.g., 700 and / or X700) receives, via one or more sensors, a selection (e.g., 705d) of an option (e.g., 720d) to initiate a process to transmit a message to the corresponding user. In some embodiments, in response to receiving (e.g., 705d) an option (e.g., 720d) to initiate a process to transmit a message to the corresponding user (and optionally, based on a determination of not displaying a message from the corresponding user when a selection of a representation of the corresponding user is received), the computer system (e.g., 700 and / or X700) displays, via a display generation component (e.g., 702 and / or X702), a messaging user interface (e.g., Figure 7E 730 at) (e.g., a messaging user interface having a first size and / or a messaging user interface including a displayed keyboard) having a first appearance for messaging with the corresponding user without displaying the main user interface. Displaying the messaging user interface having the first appearance without displaying the main user interface provides the user with feedback that the messaging user interface is in a first state, thereby providing the user with improved visual feedback.

[0291] In some embodiments, one or more options associated with a corresponding user include an option (e.g., 718 and / or X718) (e.g., as part of one or more options associated with the corresponding user) to initiate a process (e.g., excluding live transmission of audio and / or video) to transmit a message to the corresponding user. In some embodiments, a computer system (e.g., 700 and / or X700) receives, via one or more sensors, a selection (e.g., 707c) of an option (e.g., 718 and / or X718) to initiate a process to transmit a message to the corresponding user. In some embodiments, in response to receiving a selection (e.g., 707c) of an option (e.g., 718 and / or X718) to initiate a process to transmit a message to the corresponding user (and optionally based on a determination of being in the process of displaying a message from the corresponding user when a selection of a representation of the corresponding user is received), the computer system (e.g., 700 and / or X700) simultaneously displays, via a display generation component, a messaging user interface (e.g., 750) having a second appearance (e.g., different from the first appearance) for messaging with the corresponding user (e.g., a messaging user interface having a second size smaller than the first size and / or a messaging user interface not including a displayed keyboard) and at least a portion of the main user interface (e.g., as Figure 7H(as shown) (e.g., showing an obscured main user interface). Displaying a messaging user interface having a second appearance and a portion of the main user interface provides feedback to the user that the messaging user interface is in the second state, thereby providing improved visual feedback to the user.

[0292] In some embodiments, the main user interface is not movable by the user, and wherein the messaging user interface (e.g., 750) (e.g., having a first appearance and / or having a second appearance) is movable by the user. Enabling the user to move the messaging user interface without enabling the user to move the main user interface provides feedback to the user as to which elements are part of the main user interface and which elements are not part of the main user interface, thereby providing improved visual feedback to the user.

[0293] In some embodiments, displaying representations of multiple users (e.g., 712a - 712g and 714a - 714i) via a display generation component (e.g., 702 and / or X702) includes displaying a first grouping of first representations of a first plurality of users (e.g., 712) via the display generation component (e.g., 702 and / or X702), where the first plurality of users (e.g., 712a - 712g) are selected to be included as part of the representations of the multiple users independent of the recency of communication between the user of the computer system and the first plurality of users (e.g., included for display based on being manually selected as part of favorites contacts and / or based on the frequency of communication). In some embodiments, displaying representations of multiple users (e.g., 712a - 712g and 714a - 714i) via the display generation component includes displaying a second grouping of second representations of a second plurality of users (e.g., 714) via the display generation component (e.g., 702 and / or X702), where the second plurality of users (e.g., 714a - 714i) are selected to be included as part of the multiple users based on the recency of communication between the user of the computer system and the second plurality of users (e.g., included for display based on the most recent users communicated with). In some embodiments, the order of the second representations of the second plurality of users is based on the recency of communication between the user of the computer system and the second plurality of users. In some embodiments, displaying representations of multiple users via the display generation component includes: displaying a first grouping of first representations of a first plurality of users, where the first plurality of users are selected to be included as part of the representations of the multiple users independent of the recency of communication between the user of the computer system and the first plurality of users; and a second grouping of second representations of a second plurality of users, where: based on the determination that recent communication by the user of the computer system includes communication with a first set of one or more users and does not include communication with a second set of one or more users, the second plurality of users includes the first set of one or more users and does not include the second set of one or more users, and based on the determination that recent communication by the user of the computer system includes communication with a second set of one or more users and does not include communication with a first set of one or more users, the second plurality of users includes the second set of one or more users and does not include the first set of one or more users. Grouping the first plurality of users together and grouping the second plurality of users together enables the computer system to provide the user with feedback on which users are selected independent of communication recency and which users are included based on communication recency, thus providing improved visual feedback.

[0294] In some embodiments, the recency of communication between a user of a computer system (e.g., 700 and / or X700) and a second plurality of users is based on multiple communication modalities (e.g., text messaging, phone calls, and / or communication sessions (e.g., video communication sessions, audio communication sessions, extended reality communication sessions, spatial communication sessions, and / or non-spatial communication sessions)). Grouping the second plurality of users together based on the recency of communication using multiple communication modalities enables the computer system to group recent contacts regardless of how the communication occurs, thereby providing improved visual feedback.

[0295] In some embodiments, the computer system (e.g., 700 and / or X700) simultaneously displays an indication of recent communication activity (e.g., 714f and / or X714f) (e.g., information about a recent call or communication, information about an active call or communication, and / or information about a recent message) between a user of the computer system and a corresponding user via a display generation component (e.g., 702 and / or X702) and in association with an option (e.g., 724b and / or X724b) to invite the corresponding user to join an ongoing communication session and / or one or more options associated with the corresponding user (e.g., 724a, X724a, 724c, and / or X724c). Displaying the indication of recent communication activity and the one or more options enables the user to see what recent communication method was used and quickly select a communication method for another communication session, thereby reducing the amount of input required to initiate an appropriate typ...

Claims

1. A method, comprising: at a computer system in communication with a display generation component and one or more sensors: displaying representations of multiple users via the display generation component; receiving a selection of a representation of a corresponding user among the multiple users via the one or more sensors; and in response to receiving the selection of the representation of the corresponding user: displaying, via the display generation component, an option to invite the corresponding user to join the ongoing communication session based on determining that an ongoing communication session exists; and abandoning displaying the option to invite the corresponding user to join the ongoing communication session based on determining that no ongoing communication session exists.

2. The method according to claim 1, further comprising: in response to receiving the selection of the representation of the corresponding user, displaying via the display generation component: an option to initiate a new spatial communication session with the corresponding user; and options for additional features; receiving a selection of the options for additional features when the options for additional features are displayed; and in response to receiving the selection of the options for additional features, displaying, via the display generation component, one or more options associated with the corresponding user.

3. The method according to any one of claims 1 to 2, further comprising: in response to receiving the selection of the representation of the corresponding user, displaying, via the display generation component, options for additional features; receiving a selection of the options for additional features via the one or more sensors when the options for additional features are displayed; in response to receiving the selection of the options for additional features, displaying, via the display generation component, an option to initiate an audio communication session with the corresponding user; receiving a selection of the option to initiate an audio communication session with the corresponding user via the one or more sensors; and in response to receiving the selection of the option to initiate an audio communication session with the corresponding user, initiating an audio communication session with the corresponding user.

4. The method according to claim 3, wherein initiating an audio communication session comprises: using an external electronic device within a pre-determined range of the computer system to initiate an audio call.

5. The method according to any one of claims 1 to 2, further comprising: in response to receiving the selection of the representation of the corresponding user, displaying, via the display generation component, options for additional features; receiving a selection of the options for additional features via the one or more sensors when the options for additional features are displayed; in response to receiving the selection of the options for additional features, displaying, via the display generation component, an option to initiate a process of sending a message to the corresponding user; receiving a selection of the option to initiate a process of sending a message to the corresponding user via the one or more sensors; and in response to receiving the selection of the option to initiate a process of sending a message to the corresponding user, initiating a process of sending a message to the corresponding user.

6. The method according to any one of claims 1 to 2, further comprises: in response to receiving a selection of the representation of the corresponding user, displaying, via the display generating component, options for additional features; when displaying the options for additional features, receiving, via the one or more sensors, a selection of the options for additional features; in response to receiving a selection of the options for additional features, displaying, via the display generating component, options for displaying additional information about the corresponding user; receiving, via the one or more sensors, a selection of the options for displaying additional information about the corresponding user; and in response to receiving a selection of the options for displaying additional information about the corresponding user, displaying, via the display generating component, additional information about the corresponding user.

7. The method according to any one of claims 1 to 2, further comprises: in response to receiving a selection of the representation of the corresponding user, displaying, via the display generating component, options for additional features; when displaying the options for additional features, receiving, via the one or more sensors, a selection of the options for additional features; in response to receiving a selection of the options for additional features, displaying, via the display generating component, options for removing the corresponding user from the favorites page; receiving, via the one or more sensors, a selection of the options for removing the corresponding user from the favorites page; and in response to receiving a selection of the options for removing the corresponding user from the favorites page, initiating a process of removing the corresponding user from the favorites page.

8. The method according to any one of claims 1 to 2, further comprises: in response to receiving a selection of the representation of the corresponding user and based on determining that there is an ongoing communication session, displaying, via the display generating component and simultaneously with the options for inviting the corresponding user to join the ongoing communication session, options for initiating a process of starting a new communication session with the corresponding user; receiving, via the one or more sensors, a selection of the options for initiating a process of starting a new communication session with the corresponding user; and in response to receiving a selection of the options for initiating a process of starting a new communication session with the corresponding user, initiating the following process: ending the ongoing communication session; and starting the new communication session with the corresponding user.

9. The method according to claim 8, further comprises: during the process of starting the new communication session with the corresponding user, prompting for confirmation of ending the ongoing communication session; receiving, via the one or more sensors, confirmation of ending the ongoing communication session; and in response to receiving confirmation of ending the ongoing communication session, ending the ongoing communication session.

10. The method according to claim 2, wherein the representations of the plurality of users are displayed as part of a main user interface, and displaying the option to invite the corresponding user to join the ongoing communication session and / or displaying the one or more options associated with the corresponding user includes obscuring the main user interface.

11. The method according to claim 10, wherein the one or more options associated with the corresponding user includes an option to initiate a process of transmitting a message to the corresponding user, and the method further comprises: receiving, via the one or more sensors, a selection of the option to initiate a process of transmitting a message to the corresponding user; and in response to receiving the selection of the option to initiate a process of transmitting a message to the corresponding user, displaying, via the display generation component, a messaging user interface having a first appearance without displaying the main user interface.

12. The method according to claim 10, wherein the one or more options associated with the corresponding user includes an option to initiate a process of transmitting a message to the corresponding user, and the method further comprises: receiving, via the one or more sensors, a selection of the option to initiate a process of transmitting a message to the corresponding user; and in response to receiving the selection of the option to initiate a process of transmitting a message to the corresponding user, simultaneously displaying, via the display generation component, a messaging user interface having a second appearance and at least a portion of the main user interface.

13. The method according to claim 11, wherein the main user interface is not user - movable, and wherein the messaging user interface is user - movable.

14. The method according to any one of claims 1 to 2, wherein displaying the representations of the plurality of users via the display generation component comprises: displaying, via the display generation component, a first grouping of first pluralities of representations of a first plurality of users, wherein the first plurality of users are selected to be included as part of the representations of the plurality of users independently of the recency of communication between the user of the computer system and the first plurality of users; and displaying, via the display generation component, a second grouping of second pluralities of representations of a second plurality of users, wherein the second plurality of users are selected to be included as part of the plurality of users based on the recency of communication between the user of the computer system and the second plurality of users.

15. The method according to claim 14, wherein the recency of communication between the user of the computer system and the second plurality of users is based on multiple communication modes.

16. The method according to claim 2, further comprises: simultaneously displaying, via the display generation component and with the option to invite the corresponding user to join the ongoing communication session and / or with the one or more options associated with the corresponding user, an indication of recent communication activity between the user of the computer system and the corresponding user.

17. The method according to claim 2, wherein the one or more options associated with the corresponding user include an option to initiate a spatial communication session with the corresponding user, and the method further comprises: receiving, via the one or more sensors, a selection of the option to initiate a spatial communication session with the corresponding user; and in response to receiving the selection of the option to initiate a spatial communication session with the corresponding user, initiating a spatial communication session with the corresponding user.

18. The method according to any one of claims 1 to 2, further comprises: displaying, via the display generating component and simultaneously with the representations of the plurality of users, options for previewing and / or editing an avatar of a user of the computer system; receiving, via the one or more sensors, a selection of the options for previewing and / or editing the avatar of the user of the computer system; and in response to receiving the selection of the options for previewing and / or editing the avatar of the user of the computer system, displaying, via the display generating component, a user interface for previewing and / or editing the avatar of the user of the computer system.

19. The method according to any one of claims 1 to 2, further comprises: displaying, via the display generating component and simultaneously with the representations of the plurality of users: a first indication of the most recent communication between a user of the computer system and a first user among the plurality of users; and a second indication of the most recent communication between the user of the computer system and a second user among the plurality of users, different from the first user.

20. The method according to any one of claims 1 to 2, wherein a first representation corresponding to a first user among the representations of the plurality of users includes a recent message from the first user.

21. The method according to claim 20, further comprises: receiving, via the one or more sensors, a selection of the recent message from the first user; and in response to receiving the selection of the recent message from the first user, displaying, via the display generating component, an option to reply to the recent message.

22. The method according to claim 21, further comprises: receiving, via the one or more sensors, a selection of the option to reply to the recent message from the first user; and in response to receiving the selection of the option to reply to the recent message from the first user, displaying, via the display generating component, a user interface for replying to the recent message from the first user.

23. The method according to any one of claims 1 to 2, wherein the representations of the plurality of users are displayed as part of a main user interface, and wherein the main user interface is displayed in response to detecting an activation of a hardware button.

24. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more sensors, the one or more programs including instructions for performing the method according to any one of claims 1 to 23.

25. A computer system configured to communicate with a display generation component and one or more sensors, the computer system comprising: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 1 to 23.

26. A computer system configured to communicate with a display generation component and one or more sensors, the computer system comprising: means for performing the method according to any one of claims 1 to 23.

27. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more sensors, the one or more programs including instructions for performing the method according to any one of claims 1 to 23.