User interface for managing live communication sessions
By combining display generation components and sensors to dynamically adjust the user interface and avatar, the inefficiency of existing live communication sessions is solved, achieving more efficient and intuitive user interaction and energy savings.
Patent Information
- Application Number
- CN202511368896.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-09-12
- Filing Date
- 2023-09-21
- Publication Date
- 2025-11-07
AI Technical Summary
Existing live communication session management methods and interfaces are inefficient in augmented reality environments, with cumbersome and error-prone user input, leading to increased cognitive burden and wasted computer system energy.
By combining computer systems with display generation components and sensors, an intelligent user interface is provided that reduces the amount and nature of user input, dynamically displays controls and avatars using gaze and gesture detection, supports switching between spatial and non-spatial communication sessions, and displays information based on gaze input.
It improves the efficiency and accuracy of user interaction, reduces energy consumption, enhances the user experience, and especially improves the battery life of battery-powered devices.
Smart Images

Figure CN120915906A_ABST
Abstract
Description
[0001] This application is a continuation of the application entitled “USER INTERFACES FOR MANAGING LIVE COMMUNICATION SESSIONS” having application number 202380065246.4, filed September 21, 2023, having a filing date of September 21, 2023.
[0002] Cross-reference to related applications
[0003] This application claims priority to U.S. Patent Application No. 18 / 367,418, entitled “USER INTERFACES FOR MANAGING LIVE COMMUNICATION SESSIONS,” filed September 12, 2023, U.S. Provisional Patent Application No. 63 / 470,882, entitled “USER INTERFACES FOR MANAGING LIVE COMMUNICATION SESSIONS,” filed June 3, 2023, and U.S. Provisional Patent Application No. 63 / 409,583, entitled “USER INTERFACES FOR MANAGING LIVE COMMUNICATION SESSIONS,” filed September 23, 2022. The contents of each of these patent applications are incorporated herein by reference in their entirety. TECHNICAL FIELD
[0004] The present disclosure generally relates to computer systems that provide computer-generated experiences in communication with display generation components and optionally one or more sensors, including but not limited to electronic devices that provide virtual reality experiences and mixed reality experiences via a display. BACKGROUND
[0005] In recent years, there has been a significant increase in the development of computer systems for augmented reality. Example augmented reality environments include virtual elements that replace or augment at least some of the physical world. Input devices for computer systems and other electronic computing devices, such as cameras, controllers, joysticks, touch-sensitive surfaces, and touch-screen displays, are used to interact with virtual / augmented reality environments. Example virtual elements include virtual objects such as digital images, video, text, icons, and control elements such as buttons and other graphics. SUMMARY
[0006] Some methods and interfaces for managing live communication sessions, such as those that include at least some virtual elements (e.g., applications, augmented reality environments, mixed reality environments, and virtual reality environments), are cumbersome, inefficient, and limited. For example, systems that provide insufficient control for performing actions associated with virtual objects, systems that require a series of inputs to achieve a desired result in an augmented reality environment, and systems in which virtual object manipulation is complex, tedious, and error-prone, place a significant cognitive burden on users and detract from the experience of the virtual / augmented reality environment. Moreover, these methods take longer than necessary, wasting computer system resources. This latter consideration is of particular importance in battery-operated devices.
[0007] Accordingly, there is a need for computer systems that are more efficient and intuitive for users to have improved methods and interfaces for managing live communication sessions. Such methods and interfaces optionally complement or replace conventional methods for managing live communication sessions. Such methods and interfaces reduce the number, extent, and / or nature of the inputs from the user, by helping the user to understand the connection between the inputs provided and the response of the device to those inputs, thereby creating a more efficient human-machine interface.
[0008] The above-mentioned deficiencies and other problems associated with user interfaces of computer systems are reduced or eliminated by the disclosed systems. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a notebook computer, a tablet computer, or a handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device such as a watch or a head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also known as a “touch screen” or “touch-screen display”). In some embodiments, the computer system has one or more eye tracking components. In some embodiments, the computer system has one or more hand tracking components. In some embodiments, in addition to the display generation component, the computer system has one or more output devices including one or more tactile output generators and / or one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, memory, and one or more modules, programs or sets of instructions stored in memory for performing various functions. In some embodiments, the user interacts with the GUI by way of contact and gestures with a stylus and / or finger that are detected by a touch-sensitive surface (e.g., a touchpad or touch screen), user’s eyes and hands in space relative to the GUI (and / or computer system) or user’s body (as captured by cameras and other movement sensors), and / or voice input (as captured by one or more audio input devices). In some embodiments, functions performed by the interactions optionally include image editing, drawing, presenting, word processing, spreadsheet
[0009] There is a need for electronic devices with improved methods and interfaces for managing live communication sessions. Such methods and interfaces can supplement or replace conventional methods for managing live communication sessions. Such methods and interfaces reduce the number, extent, and / or nature of inputs from users and result in more efficient human-machine interfaces. For battery-powered computing devices, such methods and interfaces conserve power and increase the time between battery charges.
[0010] In some embodiments, the computer system displays a set of controls (e.g., transport controls and / or other types of controls) associated with controlling playback of media content in response to detecting a gaze and / or gesture of a user. In some embodiments, the computer system initially displays a first set of controls in a reduced salience state (e.g., with reduced visual salience) in response to detecting a first input, and then displays a second set of controls (which optionally includes additional controls) in an increased salience state in response to detecting a second input. In this way, the computer system optionally provides feedback to the user that the user has initiated a call for display of controls without unduly distracting the user from the content (e.g., by initially displaying the controls in a visually less salient manner), and then displays the controls in a visually more salient manner based on detecting user input that indicates that the user wishes to further interact with the controls to allow for easier and more accurate interaction with the computer system.
[0011] An example method is described herein. An example method includes, at a computer system in communication with a display generation component and one or more sensors: displaying, via the display generation component, representations of a plurality of users; receiving, via the one or more sensors, a selection of a representation of a respective user of the plurality of users; and in response to receiving the selection of the representation of the respective user: in accordance with a determination that there is an ongoing communication session, displaying, via the display generation component, an option to invite the respective user to join the ongoing communication session; and in accordance with a determination that there is not an ongoing communication session, forgoing displaying the option to invite the respective user to join the ongoing communication session.
[0012] An example method includes, at a computer system in communication with a display generation component and one or more sensors: displaying, via the display generation component, a communication user interface for communicating with other users in a live communication session, wherein during the live communication session, a user of the computer system is represented by an avatar that moves during the live communication session in accordance with movements of the user of the computer system that are detected by the one or more sensors; while displaying the communication user interface, displaying, via the display generation component, a selectable user interface object; detecting, via the one or more sensors, one or more inputs that include a selection input that is directed to the selectable user interface object; and in response to detecting the one or more inputs that include the selection input that is directed to the selectable user interface object, concurrently displaying, via the display generation component, an avatar editing user interface that includes: the avatar that represents the user of the computer system; and one or more options to modify an appearance of the avatar that represents the user of the computer system.
[0013] An example method includes, at a computer system in communication with a display generation component: while participating in a communication session as a spatial communication session, where the spatial communication session includes displaying representations of a plurality of participants in the communication session in a spatial distribution arrangement in a 3D environment via the display generation component, where displaying the plurality of participants in the spatial distribution arrangement includes displaying: the representations of the plurality of participants spaced apart from each other and a user of the computer system by at least a threshold amount in a first non-vertical direction in the 3D environment; and the representations of the plurality of participants spaced apart from each other and the user by at least the threshold amount in a second non-vertical direction different from the first non-vertical direction; while displaying the representations of the plurality of participants distributed in the 3D environment, detecting an event; and in response to detecting the event, transitioning the communication session from the spatial communication session to a non-spatial communication session, the non-spatial communication session including displaying representations of at least a subset of the plurality of participants of the communication session in a grouped arrangement via the display generation component, where in the grouped arrangement: the representations of the plurality of participants are spaced apart from each other by less than the threshold amount in the first non-vertical direction in the 3D environment; a representation of a first participant in the grouped arrangement has a different positioning than a representation of the first participant in the spatial distribution arrangement; and a representation of a second participant in the grouped arrangement has a different positioning than a representation of the second participant in the spatial distribution arrangement.
[0014] An example method includes, at a computer system in communication with a display generation component and one or more sensors: while in a communication session with one or more participants in the communication session, detecting, via the one or more sensors, a gaze input of a user of the computer system; and in response to detecting the gaze input: in accordance with a determination that the gaze input satisfies a set of one or more gaze criteria, displaying, via the display generation component, information about a first participant in the communication session; and in accordance with a determination that the gaze input does not satisfy the set of one or more gaze criteria, forgoing displaying the information about the first participant in the communication session.
[0015] An example non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more sensors, and including instructions for: displaying, via the display generation component, representations of a plurality of users; receiving, via the one or more sensors, a selection of a representation of a respective user of the plurality of users; and in response to receiving the selection of the representation of the respective user: in accordance with a determination that there is an ongoing communication session, displaying, via the display generation component, an option to invite the respective user to join the ongoing communication session; and in accordance with a determination that there is not an ongoing communication session, forgoing displaying the option to invite the respective user to join the ongoing communication session.
[0016] An example non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more sensors, and including instructions for: displaying, via the display generation component, a communication user interface for communicating with other users in a live communication session, wherein during the live communication session, a user of the computer system is represented by an avatar that moves during the live communication session in accordance with movements of the user of the computer system that are detected by the one or more sensors; while displaying the communication user interface, displaying, via the display generation component, a selectable user interface object; detecting, via the one or more sensors, one or more inputs that include a selection input that is directed to the selectable user interface object; and in response to detecting the one or more inputs that include the selection input that is directed to the selectable user interface object, concurrently displaying, via the display generation component, an avatar editing user interface that includes: the avatar that represents the user of the computer system; and one or more options to modify an appearance of the avatar that represents the user of the computer system.
[0017] One example non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and includes instructions for: while participating in a communication session as a spatial communication session, wherein the spatial communication session includes displaying representations of a plurality of participants in the communication session in a spatial distribution via the display generation component in a 3D environment, wherein displaying the plurality of participants in the spatial distribution includes displaying: the representations of the plurality of participants spaced apart from one another and a user of the computer system by at least a threshold amount in a first non-vertical direction in the 3D environment; and the representations of the plurality of participants spaced apart from one another and the user by at least the threshold amount in a second non-vertical direction different from the first non-vertical direction; while displaying the representations of the plurality of participants distributed in the 3D environment, detecting an event; and in response to detecting the event, transitioning the communication session from the spatial communication session to a non-spatial communication session, the non-spatial communication session including displaying representations of at least a subset of the plurality of participants of the communication session in a grouped arrangement via the display generation component, wherein in the grouped arrangement: the representations of the plurality of participants are spaced apart from one another by less than the threshold amount in the first non-vertical direction in the 3D environment; a representation of a first participant in the grouped arrangement has a different positioning than a representation of the first participant in the spatial distribution; and a representation of a second participant in the grouped arrangement has a different positioning than a representation of the second participant in the spatial distribution.
[0018] One example non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more sensors and includes instructions for: while in a communication session with one or more participants in the communication session, detecting, via the one or more sensors, a gaze input of a user of the computer system; and in response to detecting the gaze input: in accordance with a determination that the gaze input satisfies a set of one or more gaze criteria, displaying, via the display generation component, information about a first participant in the communication session; and in accordance with a determination that the gaze input does not satisfy the set of one or more gaze criteria, forgoing displaying the information about the first participant in the communication session.
[0019] An example transitory computer-readable storage medium is described herein. An example transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more sensors, and includes instructions for: displaying, via the display generation component, representations of a plurality of users; receiving, via the one or more sensors, a selection of a representation of a respective user of the plurality of users; and in response to receiving the selection of the representation of the respective user: in accordance with a determination that there is an ongoing communication session, displaying, via the display generation component, an option to invite the respective user to join the ongoing communication session; and in accordance with a determination that there is not an ongoing communication session, forgoing displaying the option to invite the respective user to join the ongoing communication session.
[0020] An example transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more sensors, and includes instructions for: displaying, via the display generation component, a communication user interface for communicating with other users in a live communication session, wherein during the live communication session, a user of the computer system is represented by an avatar that moves during the live communication session in accordance with movements of the user of the computer system that are detected by the one or more sensors; while displaying the communication user interface, displaying, via the display generation component, a selectable user interface object; detecting, via the one or more sensors, one or more inputs that include a selection input that is directed to the selectable user interface object; and in response to detecting the one or more inputs that include the selection input that is directed to the selectable user interface object, concurrently displaying, via the display generation component, an avatar editing user interface that includes: the avatar that represents the user of the computer system; and one or more options to modify an appearance of the avatar that represents the user of the computer system.
[0021] An example transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and includes instructions for: while participating in a communication session as a spatial communication session, wherein the spatial communication session includes displaying representations of a plurality of participants in the communication session in a spatial distribution arrangement in a 3D environment via the display generation component, wherein displaying the plurality of participants in the spatial distribution arrangement includes displaying: the representations of the plurality of participants spaced apart from each other and a user of the computer system by at least a threshold amount in a first non-vertical direction in the 3D environment; and the representations of the plurality of participants spaced apart from each other and the user by at least the threshold amount in a second non-vertical direction different from the first non-vertical direction; while displaying the representations of the plurality of participants distributed in the 3D environment, detecting an event; and in response to detecting the event, transitioning the communication session from the spatial communication session to a non-spatial communication session, the non-spatial communication session including displaying representations of at least a subset of the plurality of participants of the communication session in a grouped arrangement via the display generation component, wherein in the grouped arrangement: the representations of the plurality of participants are spaced apart from each other by less than the threshold amount in the first non-vertical direction in the 3D environment; a representation of a first participant in the grouped arrangement has a different positioning than a representation of the first participant in the spatial distribution arrangement; and a representation of a second participant in the grouped arrangement has a different positioning than a representation of the second participant in the spatial distribution arrangement.
[0022] An example transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more sensors and includes instructions for: while in a communication session with one or more participants in the communication session, detecting, via the one or more sensors, a gaze input of a user of the computer system; and in response to detecting the gaze input: in accordance with a determination that the gaze input satisfies a set of one or more gaze criteria, displaying, via the display generation component, information about a first participant in the communication session; and in accordance with a determination that the gaze input does not satisfy the set of one or more gaze criteria, forgoing displaying the information about the first participant in the communication session.
[0023] An example computer system is configured to communicate with a display generation component and one or more sensors and includes: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying, via the display generation component, representations of a plurality of users; receiving, via the one or more sensors, a selection of a representation of a respective user of the plurality of users; and in response to receiving the selection of the representation of the respective user: in accordance with a determination that there is an ongoing communication session, displaying, via the display generation component, an option to invite the respective user to join the ongoing communication session; and in accordance with a determination that there is not an ongoing communication session, forgoing displaying the option to invite the respective user to join the ongoing communication session.
[0024] An example computer system is configured to communicate with a display generation component and one or more sensors and includes: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying, via the display generation component, a communication user interface for communicating with other users in a live communication session, wherein during the live communication session, a user of the computer system is represented by an avatar that moves during the live communication session in accordance with movements of the user of the computer system that are detected by the one or more sensors; while displaying the communication user interface, displaying, via the display generation component, a selectable user interface object; detecting, via the one or more sensors, one or more inputs that include a selection input that is directed to the selectable user interface object; and in response to detecting the one or more inputs that include the selection input that is directed to the selectable user interface object, concurrently displaying, via the display generation component, an avatar editing user interface that includes: the avatar that represents the user of the computer system; and one or more options to modify an appearance of the avatar that represents the user of the computer system.
[0025] An example computer system is configured to communicate with a display generation component and includes: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: while participating in a communication session as a spatial communication session, wherein the spatial communication session includes displaying representations of a plurality of participants in the communication session in a spatial distribution in a 3D environment via the display generation component, wherein displaying the plurality of participants in the spatial distribution includes displaying: representations of a plurality of participants spaced apart from each other and a user of the computer system in a first non-vertical direction in the 3D environment by at least a threshold amount; and the representations of the plurality of participants spaced apart from each other and the user in a second non-vertical direction different from the first non-vertical direction by at least the threshold amount; while displaying the representations of the plurality of participants distributed in the 3D environment, detecting an event; and in response to detecting the event, transitioning the communication session from the spatial communication session to a non-spatial communication session, the non-spatial communication session including displaying representations of at least a subset of the plurality of participants in the communication session in a grouped arrangement via the display generation component, wherein in the grouped arrangement: the representations of the plurality of participants are spaced apart from each other in the first non-vertical direction in the 3D environment by less than the threshold amount; a representation of a first participant in the grouped arrangement has a different positioning than a representation of the first participant in the spatial distribution; and a representation of a second participant in the grouped arrangement has a different positioning than a representation of the second participant in the spatial distribution.
[0026] An example computer system is configured to communicate with a display generation component and one or more sensors and includes: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: while in a communication session with one or more participants in the communication session, detecting, via the one or more sensors, a gaze input of a user of the computer system; and in response to detecting the gaze input: in accordance with a determination that the gaze input satisfies a set of one or more gaze criteria, displaying, via the display generation component, information about a first participant in the communication session; and in accordance with a determination that the gaze input does not satisfy the set of one or more gaze criteria, forgoing displaying the information about the first participant in the communication session.
[0027] An example computer system is configured to communicate with a display generation component and one or more sensors and includes: means for displaying, via the display generation component, representations of a plurality of users; means for receiving, via the one or more sensors, a selection of a representation of a respective user of the plurality of users; and means for, in response to receiving the selection of the representation of the respective user: in accordance with a determination that there is an ongoing communication session, displaying, via the display generation component, an option to invite the respective user to join the ongoing communication session; and in accordance with a determination that there is not an ongoing communication session, forgoing displaying the option to invite the respective user to join the ongoing communication session.
[0028] An example computer system is configured to communicate with a display generation component and one or more sensors and includes: means for displaying, via the display generation component, a communication user interface for communicating with other users in a live communication session, wherein during the live communication session, a user of the computer system is represented by an avatar that moves during the live communication session in accordance with movements of the user of the computer system that are detected by the one or more sensors; means for, while displaying the communication user interface, displaying, via the display generation component, a selectable user interface object; means for detecting, via the one or more sensors, one or more inputs that include a selection input that is directed to the selectable user interface object; and means for, in response to detecting the one or more inputs that include the selection input that is directed to the selectable user interface object: concurrently displaying, via the display generation component, an avatar editing user interface that includes: the avatar that represents the user of the computer system; and one or more options to modify an appearance of the avatar that represents the user of the computer system.
[0029] An example computer system is configured to communicate with a display generation component and includes: means for, while participating in a communication session as a spatial communication session, wherein the spatial communication session includes displaying representations of a plurality of participants in the communication session in a spatial distribution arrangement in a 3D environment via the display generation component, wherein displaying the plurality of participants in the spatial distribution arrangement includes displaying: the representations of the plurality of participants spaced apart from each other and a user of the computer system by at least a threshold amount in a first non-vertical direction in the 3D environment; and the representations of the plurality of participants spaced apart from each other and the user by at least the threshold amount in a second non-vertical direction different from the first non-vertical direction; means for, while displaying the representations of the plurality of participants distributed in the 3D environment, detecting an event; and means for, in response to detecting the event, transitioning the communication session from the spatial communication session to a non-spatial communication session, the non-spatial communication session including displaying representations of at least a subset of the plurality of participants of the communication session in a grouped arrangement via the display generation component, wherein in the grouped arrangement: the representations of the plurality of participants are spaced apart from each other by less than the threshold amount in the first non-vertical direction in the 3D environment; a representation of a first participant in the grouped arrangement has a different positioning than a representation of the first participant in the spatial distribution arrangement; and a representation of a second participant in the grouped arrangement has a different positioning than a representation of the second participant in the spatial distribution arrangement.
[0030] An example computer system is configured to communicate with a display generation component and one or more sensors and includes: means for, while in a communication session with one or more participants in the communication session, detecting gaze input of a user of the computer system via the one or more sensors; and means for, in response to detecting the gaze input: in accordance with a determination that the gaze input satisfies a set of one or more gaze criteria, displaying information about a first participant in the communication session via the display generation component; and in accordance with a determination that the gaze input does not satisfy the set of one or more gaze criteria, forgoing displaying the information about the first participant in the communication session.
[0031] An example computer program product is described herein. An example computer program product includes one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more sensors, the one or more programs including instructions for: displaying, via the display generation component, representations of a plurality of users; receiving, via the one or more sensors, a selection of a representation of a respective user of the plurality of users; and in response to receiving the selection of the representation of the respective user: in accordance with a determination that there is an ongoing communication session, displaying, via the display generation component, an option to invite the respective user to join the ongoing communication session; and in accordance with a determination that there is not an ongoing communication session, forgoing displaying the option to invite the respective user to join the ongoing communication session.
[0032] An example computer program product includes one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more sensors, the one or more programs including instructions for: displaying, via the display generation component, a communication user interface for communicating with other users in a live communication session, wherein during the live communication session, a user of the computer system is represented by an avatar that moves during the live communication session in accordance with movements of the user of the computer system that are detected by the one or more sensors; while displaying the communication user interface, displaying, via the display generation component, a selectable user interface object; detecting, via the one or more sensors, one or more inputs that include a selection input that is directed to the selectable user interface object; and in response to detecting the one or more inputs that include the selection input that is directed to the selectable user interface object, concurrently displaying, via the display generation component, an avatar editing user interface that includes: the avatar that represents the user of the computer system; and one or more options to modify an appearance of the avatar that represents the user of the computer system.
[0033] An example computer program product includes one or more programs configured for execution by one or more processors of a computer system in communication with a display generation component, the one or more programs including instructions for: while participating in a communication session as a spatial communication session, wherein the spatial communication session includes displaying representations of a plurality of participants in the communication session in a spatial distribution in a 3D environment via the display generation component, wherein displaying the plurality of participants in the spatial distribution includes displaying: the representations of the plurality of participants spaced apart from each other and a user of the computer system by at least a threshold amount in a first non-vertical direction in the 3D environment; and the representations of the plurality of participants spaced apart from each other and the user by at least the threshold amount in a second non-vertical direction different from the first non-vertical direction; while displaying the representations of the plurality of participants distributed in the 3D environment, detecting an event; and in response to detecting the event, transitioning the communication session from the spatial communication session to a non-spatial communication session, the non-spatial communication session including displaying representations of at least a subset of the plurality of participants of the communication session in a grouped arrangement via the display generation component, wherein in the grouped arrangement: the representations of the plurality of participants are spaced apart from each other by less than the threshold amount in the first non-vertical direction in the 3D environment; a representation of a first participant in the grouped arrangement has a different positioning than a representation of the first participant in the spatial distribution; and a representation of a second participant in the grouped arrangement has a different positioning than a representation of the second participant in the spatial distribution.
[0034] An example computer program product includes one or more programs configured for execution by one or more processors of a computer system in communication with a display generation component and one or more sensors, the one or more programs including instructions for: while in a communication session with one or more participants in the communication session, detecting, via the one or more sensors, a gaze input of a user of the computer system; and in response to detecting the gaze input: in accordance with a determination that the gaze input satisfies a set of one or more gaze criteria, displaying, via the display generation component, information about a first participant in the communication session; and in accordance with a determination that the gaze input does not satisfy the set of one or more gaze criteria, forgoing displaying the information about the first participant in the communication session.
[0035] It is noted that the various embodiments described above can be combined with any of the other embodiments described herein. The features and advantages described in the specification are not all-inclusive and many additional features and advantages will be apparent to one of ordinary skill in the art in view of the drawings, specification, and claims. Moreover, it should be noted that the language used in the specification has been principally selected for readability and instructional purposes and can not have been selected to delineate or circumscribe the subject of the present application. BRIEF DESCRIPTION OF DRAWINGS
[0036] For a better understanding of the various described implementations, reference should be made to the following detailed description in conjunction with the accompanying drawings in which like reference numerals refer to corresponding parts throughout the figures.
[0037] Figure 1A is a block diagram illustrating an operating environment of a computer system for providing an XR experience in accordance with some embodiments.
[0038] Figures 1B to 1P is an example of a computer system for providing an XR experience in Figure 1A the operating environment.
[0039] Figure 2 is a block diagram illustrating a controller of a computer system configured to manage and coordinate an XR experience of a user in accordance with some embodiments.
[0040] Figure 3 is a block diagram illustrating a display generation component of a computer system configured to provide visual components of an XR experience to a user in accordance with some embodiments.
[0041] Figure 4 is a block diagram illustrating a hand tracking unit of a computer system configured to capture gesture inputs of a user in accordance with some embodiments.
[0042] Figure 5 is a block diagram illustrating an eye tracking unit of a computer system configured to capture gaze inputs of a user in accordance with some embodiments.
[0043] Figure 6 is a flow diagram illustrating a flash-assisted gaze tracking pipeline in accordance with some embodiments.
[0044] Figures 7A to 7Q example techniques for managing a live communication session in accordance with some embodiments are illustrated.
[0045] Figure 8 is a flow diagram of a method of managing a live communication session in accordance with various embodiments.
[0046] Figure 9 is a flow diagram of a method of providing an avatar in a live communication session in accordance with various embodiments.
[0047] Figures 10A to 10E example techniques for providing a representation in a live communication session in accordance with some embodiments are illustrated.
[0048] Figure 11 is a flow diagram of a method of providing a representation in a live communication session in accordance with various embodiments.
[0049] Figures 12A to 12F Example techniques for providing information in a live communication session are illustrated in accordance with some embodiments.
[0050] Figure 13 is a flow diagram of a method of providing information in a live communication session in accordance with various embodiments. DETAILED DESCRIPTION
[0051] According to some embodiments, the present disclosure relates to user interfaces for providing an extended reality (XR) experience to a user.
[0052] The systems, methods, and GUIs described herein improve user interface interactions with virtual / augmented reality environments in a variety of ways.
[0053] In some embodiments, a computer system allows for live communication between users. The computer system displays representations of a plurality of users and receives a selection of a representation of a respective user of the plurality of users. In response to receiving the selection of the representation of the respective user, in accordance with a determination that there is an ongoing communication session, the computer system displays an option to invite the respective user to join the ongoing communication session, and in accordance with a determination that there is not an ongoing communication session, the computer system forgoes displaying the option to invite the respective user to join the ongoing communication session.
[0054] In some embodiments, a computer system provides an option for a user to change an appearance of an avatar of the user. The computer system displays a communication user interface for communicating with other users in a real-time communication session. During the real-time communication session, the user is represented by an avatar that moves in accordance with movements of the user of the computer system during the real-time communication session. While displaying the communication user interface, the computer system concurrently displays a selectable user interface object. While concurrently displaying the communication user interface and the selectable user interface object, the computer system detects one or more inputs that include a selection input directed to the selectable user interface object. In response to detecting the one or more inputs that include the selection input directed to the selectable user interface object, the computer system concurrently displays an avatar editing user interface that includes the avatar representing the user of the computer system and one or more options to modify an appearance of the avatar representing the user of the computer system.
[0055] In some embodiments, a computer system switches between a spatial communication session and a non-spatial communication session. While participating in a communication session as a spatial communication session, in which the spatial communication session includes the computer system displaying representations of multiple participants in the communication session arranged in a spatial distribution in a 3D environment. Displaying the multiple participants in the spatial distribution includes displaying representations of the multiple participants spaced apart from each other and the user by at least a threshold amount in a first non-vertical direction in the 3D environment, and representations of the multiple participants spaced apart from each other and the user by at least the threshold amount in a second non-vertical direction different from the first non-vertical direction. While displaying the representations of the multiple participants distributed in the 3D environment, the computer system detects an event, and in response to detecting the event, the computer system transitions the communication session from the spatial communication session to a non-spatial communication session. Transitioning to the non-spatial communication session includes displaying representations of at least a subset of the multiple participants in the communication session in a grouped arrangement. In the grouped arrangement, the representations of the multiple participants are spaced apart from each other by less than the threshold amount in the first non-vertical direction in the 3D environment, a representation of a first participant in the grouped arrangement has a different positioning than a representation of the first participant in the spatial distribution arrangement, and a representation of a second participant in the grouped arrangement has a different positioning than a representation of the second participant in the spatial distribution arrangement.
[0056] In some embodiments, a computer system provides information during a live communication session based on a gaze of a user. While in a communication session with one or more participants in the communication session, the computer system detects a gaze input of a user of the computer system. In response to detecting the gaze input, in accordance with a determination that the gaze input satisfies a set of one or more gaze criteria, the computer system displays information about a first participant in the communication session, and in accordance with a determination that the gaze input does not satisfy the set of one or more gaze criteria, the computer system forgoes displaying the information about the first participant in the communication session.
[0057] In some embodiments, a computer system displays content in a first region of a user interface. In some embodiments, while the computer system is displaying the content and while a first set of controls is not displayed in a first state, the computer system detects a first input from a first portion of a user. In some embodiments, in response to detecting the first input, and in accordance with a determination that a gaze of the user was directed at a second region of the user interface at a time of detecting the first input, the computer system displays the first set of one or more controls in the first state in the user interface, and in accordance with a determination that the gaze of the user was not directed at the second region of the user interface at a time of detecting the first input, the computer system forgoes displaying the first set of one or more controls in the first state.
[0058] In some embodiments, the computer system displays content in a user interface. In some embodiments, while displaying the content, the computer system detects a first input based on movement of a first portion of a user of the computer system. In some embodiments, in response to detecting the first input, the computer system displays a first set of one or more controls in the user interface, where the first set of one or more controls is displayed in a first state and within a first region of the user interface. In some embodiments, while displaying the first set of one or more controls in the first state: in accordance with a determination that one or more first criteria are met, the one or more first criteria including a criterion that is met based on movement of a second portion of the user that is different from the first portion of the user directing attention of the user to the first region of the user interface, the computer system transitions from displaying the first set of one or more controls in the first state to displaying a second set of one or more controls in a second state, where the second state is different from the first state.
[0059] Figures 1A to 6 A description of an example computer system for providing XR experiences to users is provided. Figures 7A to 7Q Example techniques for managing a live communication session are illustrated in accordance with some embodiments. Figure 8 is a flow diagram of a method of managing a live communication session in accordance with various embodiments. Figure 9 is a flow diagram of a method of providing an avatar in a live communication session in accordance with various embodiments. Figures 7A to 7Q the user interface in Figure 8 and Figure 9 the process in Figures 10A to 10E Example techniques for providing a representation in a live communication session are illustrated in accordance with some embodiments. Figure 11 is a flow diagram of a method of providing a representation in a live communication session in accordance with various embodiments. Figures 10A to 10E the user interface in Figure 11 the process in Figures 12A to 12F Example techniques for providing information in a live communication session are illustrated in accordance with some embodiments. Figure 13 is a flow diagram of a method of providing information in a live communication session in accordance with various embodiments. Figures 10A to 10E the user interface in Figure 11 the process in
[0060] The processes described below, by various techniques, enhance the operability of devices and make user-device interfaces more efficient (e.g., by helping to provide appropriate inputs and reducing user mistakes when operating / interacting with the device) including by providing improved visual feedback to users, reducing the number of inputs needed to perform operations, providing additional control options without cluttering user interfaces with additional displayed controls, performing an operation when a set of conditions has been met without requiring further user input, improving privacy and / or security, providing a richer, more detailed, and / or more realistic user experience while conserving storage space, and / or additional techniques. These techniques also reduce power usage and improve battery life of the device by enabling users to use the device faster and more efficiently. Saving battery power, and thus weight, improves the ergonomics of the device. These techniques also enable real-time communications, allow the use of less accurate and / or fewer sensors, resulting in a more compact, lighter, and cheaper device, and enable the device to be used in various lighting conditions. These techniques reduce energy usage, which reduces the amount of heat emitted by the device, which is particularly important for wearable devices, where if the device generates too much heat within the operating parameters of the device components, it becomes uncomfortable for the user to wear the device.
[0061] Further, in methods described herein in which one or more steps depend on one or more conditions having been met, it should be understood that the method can be repeated in multiple iterations such that, over the course of the iterations, all of the conditions that determine steps in the method have been met in different iterations of the method. For example, if a method requires performing a first step if a condition is met, and performing a second step if the condition is not met, one of ordinary skill will appreciate that the declared steps can be repeated until both the condition is met and the condition is not met (in no particular order). Thus, a method that is described as having one or more steps that depend on one or more conditions having been met can be rewritten as a method that is repeated until every condition described in the method has been met. However, this need not require a system or computer-readable medium to declare that the system or computer-readable medium contains instructions for performing the conditional operations based on the satisfaction of the corresponding one or more conditions, and thus is able to determine whether the possible conditions have been met without explicitly repeating the steps of the method until all of the conditions that determine steps in the method have been met. One of ordinary skill in the art will also appreciate that, similar to a method having conditional steps, a system or computer-readable storage medium can repeat the steps of a method multiple times as necessary to ensure that all of the conditional steps have been performed.
[0062] In some embodiments, as Figure 1AAs shown, an XR experience is provided to a user via an operating environment 100 that includes a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted device (HMD), a display, a projector, a touchscreen, etc.), one or more input devices 125 (e.g., an eye tracking device 130, a hand tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a haptic output generator 170, and other output devices 180), one or more sensors 190 (e.g., image sensors, light sensors, depth sensors, tactile sensors, orientation sensors, proximity sensors, temperature sensors, position sensors, motion sensors, velocity sensors, etc.), and optionally one or more peripheral devices 195 (e.g., a home appliance, a wearable device, etc.). In some embodiments, one or more of the input devices 125, the output devices 155, the sensors 190, and the peripheral devices 195 are integrated with the display generation component 120 (e.g., in a head-mounted device or a handheld device).
[0063] In describing XR experiences, various terms are used to refer to several related but distinct environments that a user can sense and / or with which a user can interact (e.g., with input detected by the computer system 101 generating the XR experience that causes the computer system generating the XR experience to generate audio, visual, and / or haptic feedback corresponding to the various input provided to the computer system 101). The following is a subset of these terms:
[0064] Physical Environment: A physical environment refers to the physical world that people are able to sense and / or interact with without aid of electronic systems. A physical environment such as a physical park includes physical articles such as physical trees, physical buildings, and physical people. People are able to directly sense and / or interact with a physical environment such as through sight, touch, hearing, taste, and smell.
[0065] Extended Reality: In contrast, an extended reality (XR) environment refers to a fully or partially simulated environment that people sense and / or interact with via electronic systems. In XR, a subset of a person's physical motions, or representations thereof, are tracked, and, in response, one or more characteristics of one or more virtual objects simulated in the XR environment are adjusted in a manner that comports with at least one physical law. For example, an XR system can detect a person's head turn, and, in response, adjust graphical content and an acoustic field presented to the person in a manner that resembles how such views and sounds would change in a physical environment. In some situations (e.g., for accessibility reasons), adjustments to characteristics of virtual objects in an XR environment can be made in response to representations of physical motions (e.g., voice commands). People can sense and / or interact with XR objects with any of their senses, including sight, hearing, touch, taste, and smell. For example, a person can sense and / or interact with audio objects that create a 3D or spatial audio environment that provides a perception of point audio sources in 3D space. As another example, audio objects can enable audio transparency that selectively incorporates ambient sounds from a physical environment with or without computer-generated audio. In certain XR environments, people can sense and / or only interact with audio objects.
[0066] Examples of XR include virtual reality and mixed reality.
[0067] Virtual Reality: A virtual reality (VR) environment refers to a simulated environment that is designed to be entirely based on computer-generated sensory inputs for one or more senses. A VR environment includes multiple virtual objects that a person can sense and / or interact with. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can sense and / or interact with virtual objects in a VR environment through a simulation of a person's presence within the computer-generated environment and / or through a simulation of a subset of a person's physical movements within the computer-generated environment.
[0068] Mixed Reality: In contrast to a VR environment, which is designed to be entirely based on computer-generated sensory inputs, a mixed reality (MR) environment refers to a simulated environment that is designed to incorporate sensory inputs, or representations thereof, from a physical environment in addition to including computer-generated sensory inputs (e.g., virtual objects). On a virtual continuum, a mixed reality environment is anywhere between a fully physical environment on one end and a virtual reality environment on the other end, excluding the two extremes. In some MR environments, computer-generated sensory inputs can be responsive to changes in sensory inputs from the physical environment. Additionally, some electronic systems for presenting MR environments can track position and / or orientation with respect to a physical environment to enable virtual objects to interact with real objects (that is, physical articles from the physical environment or representations thereof). For example, a system can cause movement so that a virtual tree appears to be stationary with respect to a physical ground.
[0069] Examples of mixed reality include augmented reality and augmented virtuality.
[0070] Augmented reality: An augmented reality (AR) environment refers to a simulated environment in which one or more virtual objects are superimposed over a physical environment or a representation of a physical environment. For example, an electronic system for presenting an AR environment can have a transparent or translucent display through which a person can directly view a physical environment. The system can be configured to present virtual objects on the transparent or translucent display so that a person, using the system, perceives the virtual objects as superimposed over the physical environment. Alternatively, a system can have an opaque display and one or more imaging sensors that capture images or video of the physical environment, which are representations of the physical environment. The system combines the images or video with virtual objects and presents the combination on the opaque display. A person, using the system, views the physical environment indirectly via the images or video of the physical environment and perceives the virtual objects as superimposed over the physical environment. As used herein, video of a physical environment displayed on an opaque display is referred to as “pass-through video,” meaning that the system captures images of the physical environment using one or more image sensors and uses those images in presenting the AR environment on the opaque display. Further alternatively, a system can have a projection system that projects virtual objects into the physical environment, for example, as a hologram or on a physical surface, so that a person, using the system, perceives the virtual objects as superimposed over the physical environment. An augmented reality environment also refers to a simulated environment in which a representation of a physical environment is transformed by computer-generated sensory information. For example, in providing pass-through video, the system can transform one or more sensor images to impose a selected perspective (e.g., viewpoint) that is different from the perspective imaged by the imaging sensors. As another example, a representation of a physical environment can be transformed by graphically modifying (e.g., enlarging) portions thereof so that the modified portions can be representative but not photorealistic versions of the original captured images. As yet another example, a representation of a physical environment can be transformed by graphically eliminating or obscuring portions thereof.
[0071] Augmented virtuality: An augmented virtuality (AV) environment refers to a simulated environment in which a virtual or computer-generated environment is combined with one or more sensory inputs from a physical environment. The sensory input can be a representation of one or more characteristics of the physical environment. For example, an AV park can have virtual trees and virtual buildings, but people’s faces are photorealistically rendered from images taken of physical people. As another example, virtual objects can take on the shape or color of physical articles imaged by one or more imaging sensors. As yet another example, virtual objects can take on shadows consistent with the positioning of the sun in the physical environment.
[0072] In an augmented reality, mixed reality, or virtual reality environment, a view of a three-dimensional environment is visible to a user. The view of the three-dimensional environment is typically visible to the user via a virtual viewport, which has a viewport boundary that defines a range of the three-dimensional environment that is visible to the user via one or more display generation components (e.g., a display or a pair of display modules that provide stereoscopic content to different eyes of the same user), that view is typically visible to the user via one or more display generation components (e.g., a display or a pair of display modules that provide stereoscopic content to different eyes of the same user). In some embodiments, the area defined by the viewport boundary is smaller than the user’s field of view in one or more dimensions (e.g., based on the user’s field of view, the size, optical properties, or other physical characteristics of the one or more display generation components, and / or the position and / or orientation of the one or more display generation components relative to the user’s eyes). In some embodiments, the area defined by the viewport boundary is larger than the user’s field of view in one or more dimensions (e.g., based on the user’s field of view, the size, optical properties, or other physical characteristics of the one or more display generation components, and / or the position and / or orientation of the one or more display generation components relative to the user’s eyes). The viewport and the viewport boundary typically move with movement of the one or more display generation components (e.g., with the user’s head for a head-mounted device, or with the user’s hand for a handheld device such as a tablet or smartphone). The user’s point of view determines the content that is visible in the viewport, which typically specifies a position and direction relative to the three-dimensional environment, and as the point of view shifts, the view of the three-dimensional environment will also shift in the viewport. For a head-mounted device, the point of view is typically based on the position, direction of the user’s head, face, and / or eyes to provide a view of the three-dimensional environment that is perceptually accurate and provides an immersive experience while the user is using the head-mounted device. For a handheld or stationary device, the point of view shifts with movement of the handheld or stationary device and / or with changes in the user’s positioning relative to the handheld or stationary device (e.g., the user moves towards, away from, up, down, right, and / or left). For devices that include display generation components with virtual pass-through, portions of the physical environment that are visible (e.g., displayed and / or projected) via the one or more display generation components are based on the field of view of one or more cameras that are in communication with the display generation components, which typically move with movement of the display generation components (e.g., with the user’s head for a head-mounted device, or with the user’s hand for a handheld device such as a tablet or smartphone), as the user’s point of view moves with movement of the field of view of the one or more cameras (and the appearance of one or more virtual objects that are displayed via the one or more display generation components is updated based on the user’s point of view (e.g., the display positioning and pose of the virtual objects are updated based on movement of the user’s point of view)).For display generation components with optical see-through, portions of the physical environment that are visible via the one or more display generation components (e.g., optically visible through one or more partially or fully transparent portions of the display generation component) are based on the user’s field of view through the partially or fully transparent portions of the display generation component (e.g., moving with the user’s head for a head-mounted device, or moving with the user’s hand for a handheld device such as a tablet or smartphone) as the user’s point of view moves with the user’s movement through the user’s field of view of the partially or fully transparent portions of the display generation component (and the appearance of the one or more virtual objects is updated based on the user’s point of view).
[0073] In some embodiments, the representation of the physical environment (e.g., via virtual pass-through or optical pass-through display) can be partially or fully occluded by the virtual environment. In some embodiments, the amount of virtual environment displayed (e.g., the amount of physical environment not displayed) is based on a level of immersion of the virtual environment (e.g., relative to the representation of the physical environment). For example, increasing the level of immersion optionally causes more of the virtual environment to be displayed, replacing and / or occluding more of the physical environment, and decreasing the level of immersion optionally causes less of the virtual environment to be displayed, revealing portions of the physical environment that were previously not displayed and / or occluded. In some embodiments, at a particular level of immersion, one or more first background objects (e.g., in the representation of the physical environment) are more visually de-emphasized (e.g., darkened, blurred, displayed with increased transparency) than one or more second background objects, and one or more third background objects cease to be displayed. In some embodiments, the level of immersion includes an associated degree to which virtual content (e.g., virtual environment and / or virtual content) displayed by the computer system occludes background content (e.g., content other than the virtual environment and / or virtual content) behind / surrounding the virtual environment, optionally including a number of items of the displayed background content and / or displayed visual properties (e.g., color, contrast, and / or opacity) of the background content, an angular range of virtual content displayed via the display generation component (e.g., 60 degrees of content displayed at a low level of immersion, 120 degrees of content displayed at a medium level of immersion, or 180 degrees of content displayed at a high level of immersion), and / or a proportion of a field of view displayed via the display generation component that is occupied by the virtual content (e.g., 33% of the field of view occupied by the virtual content at a low level of immersion, 66% of the field of view occupied by the virtual content at a medium level of immersion, or 100% of the field of view occupied by the virtual content at a high level of immersion). In some embodiments, the background content is included in a background on which the virtual content is displayed (e.g., background content in the representation of the physical environment). In some embodiments, the background content includes a user interface (e.g., a user interface generated by the computer system corresponding to an application), virtual objects that are not associated with or included in the virtual environment and / or virtual content (e.g., representations of files or other users generated by the computer system, etc.), and / or real-world objects (e.g., pass-through objects representing real-world objects in the physical environment surrounding the user that are visible such that they are displayed via the display generation component and / or are visible via transparent or semi-transparent components of the display generation component because the computer system does not occlude / block their visibility through the display generation component). In some embodiments, at a low level of immersion (e.g., a first level of immersion), the background, virtual, and / or real-world objects are displayed in a manner that is not occluded. For example, a virtual environment with a low level of immersion is optionally displayed concurrently with the background content, which is optionally displayed at full brightness, color, and / or translucency.In some embodiments, at higher levels of immersion (e.g., a second level of immersion that is higher than a first level of immersion), background, virtual, and / or real objects are displayed in an occluded manner (e.g., dimmed, blurred, or removed from display). For example, a respective virtual environment is displayed with a high level of immersion without also displaying background content (e.g., in a full screen or fully immersive mode). As another example, a virtual environment is displayed with a medium level of immersion concurrently with background content that is dimmed, blurred, or otherwise de-emphasized. In some embodiments, the visual properties of background objects differ between background objects. For example, at a particular level of immersion, one or more first background objects are more visually de-emphasized (e.g., dimmed, blurred, and / or displayed with increased transparency) than one or more second background objects, and one or more third background objects cease to be displayed. In some embodiments, zero immersion or a zero level of immersion corresponds to a virtual environment ceasing to be displayed, and as an alternative, a representation of a physical environment is displayed (optionally with one or more virtual objects, such as an application, window, or virtual three-dimensional object), without the representation of the physical environment being occluded by the virtual environment. Adjusting the level of immersion using physical input elements provides a quick and efficient way to adjust the degree of immersion, which enhances the operability of the computer system and makes the user-device interface more efficient.
[0074] Viewpoint-locked virtual objects: A virtual object is viewpoint-locked when the computer system displays the virtual object at the same position and / or location in the user’s viewpoint, even if the user’s viewpoint shifts (e.g., changes). In embodiments in which the computer system is a head-mounted device, the user’s viewpoint is locked to the forward direction of the user’s head (e.g., the user’s viewpoint is at least a portion of the user’s field of view when the user looks straight ahead); thus, the user’s viewpoint remains fixed without moving the user’s head, even when the user’s gaze shifts. In embodiments in which the computer system has a display generation component (e.g., a display screen) that can be repositioned relative to the user’s head, the user’s viewpoint is the augmented reality view presented to the user on the display generation component of the computer system. For example, a viewpoint-locked virtual object that is displayed in the upper left corner of the user’s viewpoint when the user’s viewpoint is at a first orientation (e.g., the user’s head is facing north) continues to be displayed in the upper left corner of the user’s viewpoint even when the user’s viewpoint changes to a second orientation (e.g., the user’s head is facing west). In other words, the position and / or location of the viewpoint-locked virtual object in the user’s viewpoint is independent of the user’s location and / or orientation in the physical environment. In embodiments in which the computer system is a head-mounted device, the user’s viewpoint is locked to the orientation of the user’s head, such that the virtual object is also referred to as a “head-locked virtual object.”
[0075] Environment-locked visual objects: A virtual object is environment-locked (alternatively, "world-locked") when the computer system displays the virtual object at a position and / or location in the user's point of view that is based on a location and / or object in the three-dimensional environment (e.g., a physical environment or a virtual environment) that the virtual object is selected and / or anchored to (e.g., with reference to). As the user's point of view shifts, the location and / or object in the environment relative to the user's point of view changes, which causes the environment-locked virtual object to be displayed at a different position and / or location in the user's point of view. For example, an environment-locked virtual object that is locked to a tree immediately in front of the user is displayed at the center of the user's point of view. When the user's point of view shifts to the right (e.g., the user's head turns to the right) such that the tree is now to the left center of the user's point of view (e.g., the tree location in the user's point of view shifts), the environment-locked virtual object that is locked to the tree is displayed to the left center of the user's point of view. In other words, the position and / or location in the user's point of view at which the environment-locked virtual object is displayed depends on the location and / or orientation of the location and / or object in the environment that the virtual object is locked to. In some embodiments, the computer system uses a stationary frame of reference (e.g., a coordinate system that is anchored to a fixed location and / or object in the physical environment) in order to determine the location at which to display the environment-locked virtual object in the user's point of view. The environment-locked virtual object can be locked to a stationary portion of the environment (e.g., a floor, a wall, a table, or other stationary object), or can be locked to a movable portion of the environment (e.g., a vehicle, an animal, a person, or even a representation of a portion of the user's body such as the user's hand, wrist, arm, or foot that moves independently of the user's point of view) such that the virtual object moves with the portion of the environment to maintain a fixed relationship between the virtual object and the portion of the environment.
[0076] In some embodiments, an environment-locked or viewpoint-locked virtual object exhibits a lazy following behavior that reduces or delays movement of the environment-locked or viewpoint-locked virtual object relative to movement of a reference point that the virtual object is following. In some embodiments, when exhibiting a lazy following behavior, the computer system intentionally delays movement of the virtual object when movement of a reference point (e.g., a portion of the environment, a viewpoint, or a point fixed relative to a viewpoint, such as a point between 5 cm and 300 cm from the viewpoint) that the virtual object is following is detected. For example, when a reference point (e.g., a portion of the environment or a viewpoint) moves at a first speed, the virtual object is moved by the device to remain locked to the reference point but moves at a second speed that is slower than the first speed (e.g., until the reference point stops moving or slows down, at which point the virtual object begins to catch up to the reference point). In some embodiments, when a virtual object exhibits a lazy following behavior, the device ignores small movements of the reference point (e.g., ignores movements of the reference point that are below a threshold amount of movement, such as movements of 0 to 5 degrees or movements of 0 to 50 cm). For example, when a reference point (e.g., a portion of the environment or a viewpoint to which the virtual object is locked) moves a first amount, the distance between the reference point and the virtual object increases (e.g., because the virtual object is being displayed so as to remain fixed or substantially fixed in position relative to a viewpoint or portion of the environment that is different from the reference point to which the virtual object is locked), and when the reference point (e.g., a portion of the environment or a viewpoint to which the virtual object is locked) moves a second amount that is greater than the first amount, the distance between the reference point and the virtual object initially increases (e.g., because the virtual object is being displayed so as to remain fixed or substantially fixed in position relative to a viewpoint or portion of the environment that is different from the reference point to which the virtual object is locked) and then decreases when the amount of movement of the reference point increases above a threshold (e.g., a “lazy following” threshold) because the virtual object is moved by the computer system to remain fixed or substantially fixed in position relative to the reference point. In some embodiments, the virtual object remaining substantially fixed in position relative to the reference point includes the virtual object being displayed within a threshold distance (e.g., 1 cm, 2 cm, 3 cm, 5 cm, 15 cm, 20 cm, 50 cm) of the reference point in one or more dimensions (e.g., up / down, left / right, and / or forward / backward of the position relative to the reference point).
[0077] Hardware: There are many different types of electronic systems that enable a person to sense various XR environments and / or to interact with various XR environments. Examples include head-mounted systems, projection-based systems, heads-up displays (HUDs), vehicle windshields having integrated display capability, windows having integrated display capability, displays formed as lenses that are designed to be placed on a person’s eyes (e.g., similar to contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system can include speakers that are integrated into the head-mounted system for providing audio output and / or other audio output devices. A head-mounted system can have one or more speakers and an integrated opaque display. Alternatively, a head-mounted system can be configured to accept an external opaque display (e.g., a smartphone). A head-mounted system can incorporate one or more imaging sensors for capturing images or video of a physical environment and / or one or more microphones for capturing audio of the physical environment. A head-mounted system can have a transparent or semi-transparent display instead of an opaque display. A transparent or semi-transparent display can have a medium through which light representative of images is directed to a person’s eyes. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium can be an optical waveguide, hologram medium, optical combiner, optical reflector, or any combination thereof. In one embodiment, a transparent or semi-transparent display can be configured to selectively become opaque. A projection-based system can employ retinal projection technology that projects graphical images onto a person’s retina. Projection systems can also be configured to project virtual objects into the physical environment, e.g., as a hologram or on a physical surface. In some embodiments, controller 110 is configured to manage and coordinate a person’s XR experience. In some embodiments, controller 110 includes a suitable combination of software, firmware, and / or hardware. Reference is made to FIG. 1 for further details regarding the components of controller 110. Figure 2Controller 110 is described in more detail. In some embodiments, controller 110 is a computing device that is in a local or remote location relative to scene 105 (e.g., a physical environment). For example, controller 110 is a local server that is located within scene 105. As another example, controller 110 is a remote server (e.g., a cloud server, a central server, etc.) that is located outside of scene 105. In some embodiments, controller 110 is communicatively coupled with display generation component 120 (e.g., an HMD, a display, a projector, a touchscreen, etc.) via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802. l lx, IEEE 802.16x, IEEE 802.3x, etc.). In another example, controller 110 is included within a housing (e.g., a physical enclosure) of display generation component 120 (e.g., an HMD or a portable electronic device that includes a display and one or more processors, etc.), one or more input devices of input devices 125, one or more output devices of output devices 155, one or more sensors of sensors 190, and / or one or more peripheral devices of peripheral devices 195, or shares the same physical housing or support structure as one or more of the aforementioned devices.
[0078] In some embodiments, display generation component 120 is configured to provide an XR experience (e.g., at least a visual component of an XR experience) to a user. In some embodiments, display generation component 120 includes a suitable combination of software, firmware, and / or hardware. Reference is made to FIG. 2 for a description of the functionality of display generation component 120. Figure 3 Display generation component 120 is described in more detail. In some embodiments, the functionality of controller 110 is provided by and / or combined with display generation component 120.
[0079] According to some embodiments, display generation component 120 provides an XR experience to a user when the user is virtually and / or physically present within scene 105.
[0080] In some embodiments, the display generation component is worn on a portion of the user’s body (e.g., on his / her head, on his / her hand, etc.). As such, the display generation component 120 includes one or more XR displays provided for displaying XR content. For example, in various embodiments, the display generation component 120 encloses the user’s field of view. In some embodiments, the display generation component 120 is a handheld device (such as a smartphone or tablet device) configured to present XR content, and the user holds the device with a display directed at the user’s field of view and a camera directed at the scene 105. In some embodiments, the handheld device is optionally placed within a housing worn on the user’s head. In some embodiments, the handheld device is optionally placed on a support (e.g., a tripod) in front of the user. In some embodiments, the display generation component 120 is an XR chamber, housing, or room configured to present XR content, where the user does not wear or hold the display generation component 120. Many of the user interfaces described with reference to one type of hardware for displaying XR content (e.g., a handheld device or a device on a tripod) can be implemented on another type of hardware for displaying XR content (e.g., an HMD or other wearable computing device). For example, a user interface showing interactions with XR content triggered based on interactions occurring in a space in front of a handheld device or a tripod-mounted device can be similarly implemented with an HMD, where the interactions occur in a space in front of the HMD and responses to the XR content are displayed via the HMD. Similarly, a user interface showing interactions with XR content triggered based on movement of a handheld device or a tripod-mounted device relative to a physical environment (e.g., the scene 105 or a portion of the user’s body (e.g., the user’s eyes, head, or hand)) can be similarly implemented with an HMD, where the movement is caused by movement of the HMD relative to the physical environment (e.g., the scene 105 or a portion of the user’s body (e.g., the user’s eyes, head, or hand)).
[0081] While relevant features of the operating environment 100 are shown in Figure 1A the disclosure, those of ordinary skill in the art will appreciate from the disclosure that various other features have not been illustrated in the interest of brevity and so as not to obscure more pertinent aspects of the example embodiments disclosed herein.
[0082] Figures 1A to 1PVarious examples of computer systems for performing the methods and providing audio, visual, and / or haptic feedback as part of the user interfaces described herein are illustrated. In some embodiments, the computer system includes one or more display generation components (e.g., first and second display components 1-120a and 1-120b and / or first and second optical modules 11.1.1-104a and 11.1.1-104b) for displaying a representation of a virtual element and / or a physical environment to a user of the computer system, the representation optionally generated based on a detected event and / or user input detected by the computer system. The user interfaces generated by the computer system are optionally corrected by one or more corrective lenses 11.3.2-216, which are optionally removably attached to one or more of the optical modules, to make the user interfaces easier to view by users who would otherwise use eyeglasses or contact lenses to correct their vision. While many of the user interfaces illustrated herein show a single view of the user interface, the user interfaces in the HMD are optionally displayed using two optical modules (e.g., first and second display components 1-120a and 1-120b and / or first and second optical modules 11.1.1-104a and 11.1.1-104b), one optical module for the user’s right eye and a different optical module for the user’s left eye, and slightly different images are presented to the two different eyes to create the illusion of stereoscopic depth, the single view of the user interface typically being either the right eye view or the left eye view, the depth effects explained in the text or using other diagrams or views. In some embodiments, the computer system includes one or more external displays (e.g., display assembly 1-108) for displaying status information of the computer system to a user of the computer system (when the computer system is not being worn) and / or to other people in the vicinity of the computer system, the status information optionally generated based on a detected event and / or user input detected by the computer system. In some embodiments, the computer system includes one or more audio output components (e.g., electronic components 1-112) for generating audio feedback, the audio feedback optionally generated based on a detected event and / or user input detected by the computer system. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors (e.g., sensor assembly 1-356 and / or one or more sensors in Figure 1I Figure 1I The illuminator described herein generates a digital pass-through image, captures visual media corresponding to the physical environment (e.g., photographs and / or videos), or determines the pose (e.g., positioning and / or orientation) of physical objects and / or surfaces in the physical environment, enabling virtual objects to be placed based on the detected pose of the physical objects and / or surfaces. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors (e.g., sensor assemblies 1-356 and / or...) for detecting hand positioning and / or movement. Figure 1I One or more sensors), which can be used (optionally in conjunction with one or more illuminators, such as Figure 1I The illuminator 6-124 described herein determines when one or more air gestures are performed. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting eye movement (e.g., ...). Figure 1I (Eye-tracking and gaze-tracking sensors in the system), one or more of these sensors can be used (optionally in conjunction with one or more lights, such as...) Figure 10The gaze and / or attention information is optionally combined with hand tracking information to determine user interaction with one or more user interfaces based on direct and / or indirect input, such as air gestures or input using one or more hardware input devices, such as one or more buttons (e.g., first button 1-128, button 11.1.1-114, second button 1-132, and / or dial or button 1-328), knobs (e.g., first button 1-128, button 11.1.1-114, and / or dial or button 1-328), digital crowns (e.g., first button 1-128, button 11.1.1-114, and / or dial or button 1-328 that can be pressed and that can be twisted or rotated), touchpads, touchscreens, keyboards, mice, and / or other input devices. One or more buttons (e.g., first button 1-128, button 11.1.1-114, second button 1-132, and / or dial or button 1-328) are optionally used to perform system operations, such as re-centering content in a three-dimensional environment that is visible to the user of the device, displaying a home user interface for launching applications, starting a live communication session, or initiating display of a virtual three-dimensional background. A knob or digital crown (e.g., first button 1-128, button 11.1.1-114, and / or dial or button 1-328 that can be pressed and that can be twisted or rotated) is optionally rotatable to adjust a parameter of visual content, such as an immersion level of a virtual three-dimensional environment (e.g., a degree to which virtual content occupies a user’s viewport in the three-dimensional environment) or other parameters associated with the three-dimensional environment and virtual content displayed via the optical modules (e.g., first display assembly 1-120a and second display assembly 1-120b and / or first optical module 11.1.1-104a and second optical module 11.1.1-104b).
[0083] Figure 1BA front view, top view, perspective view of an example of a head-mounted display (HMD) device 1-100 configured to be worn by a user and provide virtual and altered / mixed reality (VR / AR) experiences is illustrated. The HMD 1-100 can include a display unit 1-102 or assembly, an electronic strap assembly 1-104 connected to and extending from the display unit 1-102, and a band assembly 1-106 secured to the electronic strap assembly 1-104 at either end. The electronic strap assembly 1-104 and the band 1-106 can be part of a retention assembly configured to wrap around a user’s head to hold the display unit 1-102 against the user’s face.
[0084] In at least one example, the band assembly 1-106 can include a first band 1-116 configured to wrap around a back side of the user’s head and a second band 1-117 configured to extend over a top of the user’s head. As shown, the second band can extend between a first electronic strap 1-105a and a second electronic strap 1-105b of the electronic strap assembly 1-104. The strap assembly 1-104 and the band assembly 1-106 can be part of a securing mechanism that extends rearward from the display unit 1-102 and is configured to hold the display unit 1-102 against the user’s face.
[0085] In at least one example, the securing mechanism includes a first electronic strap 1-105a that includes a first proximal end 1-134 coupled to the display unit 1-102 (e.g., a housing 1-150 of the display unit 1-102) and a first distal end 1-136 opposite the first proximal end 1-134. The securing mechanism can also include a second electronic strap 1-105b that includes a second proximal end 1-138 coupled to the housing 1-150 of the display unit 1-102 and a second distal end 1-140 opposite the second proximal end 1-138. The securing mechanism can also include a first band 1-116 that includes a first end 1-142 coupled to the first distal end 1-136 and a second end 1-144 coupled to the second distal end 1-140, and a second band 1-117 that extends between the first electronic strap 1-105a and the second electronic strap 1-105b. The straps 1-105a-b and the bands 1-116 can be coupled via a connection mechanism or assembly 1-114. In at least one example, the second band 1-117 includes a first end 1-146 coupled to the first electronic strap 1-105a between the first proximal end 1-134 and the first distal end 1-136 and a second end 1-148 coupled to the second electronic strap 1-105b between the second proximal end 1-138 and the second distal end 1-140.
[0086] In at least one example, the first and second electronic strips 1-105a-b comprise plastic, metal, or other structural materials forming a substantially rigid strip shape. In at least one example, the first strip 1-116 and the second strip 1-117 are formed of an elastic, flexible material including woven textiles, rubber, etc. The first strip 1-116 and the second strip 1-117 may be flexible enough to conform to the shape of the user's head when the HMD 1-100 is worn.
[0087] In at least one example, one or more of the first and second electronic strips 1-105a-b may define an inner strip volume and include one or more electronic components disposed within the inner strip volume. In one example, such as Figure 1B As shown, the first electronic strip 1-105a may include electronic components 1-112. In one example, electronic components 1-112 may include a speaker. In another example, electronic components 1-112 may include computing components, such as a processor.
[0088] In at least one example, the housing 1-150 defines a first front opening 1-152. The front opening is located in... Figure 1B The section marked 1-152 with dashed lines is because the display assembly 1-108 is configured to obscure the first opening 1-152 from the field of view when the HMD 1-100 is assembled. The housing 1-150 may also define a rearward second opening 1-154. The housing 1-150 also defines an internal volume between the first opening 1-152 and the second opening 1-154. In at least one example, the HMD 1-100 includes a display assembly 1-108, which may include a front cover disposed in or across the front opening 1-152 to obscure the front opening 1-152 and a display screen (shown in other figures). In at least one example, the display screen of the display assembly 1-108, and the display assembly 1-108 in general, has a curvature configured to follow the curvature of the user's face. The display screen of display unit 1-108 can be bent as shown to complement the user's facial features and the overall curvature from one side of the face to the other, such as from left to right and / or from top to bottom, wherein display unit 1-102 is pressed.
[0089] In at least one example, the housing 1-150 can define a first aperture 1-126 between the first opening 1-152 and the second opening 1-154 and a second aperture 1-130 between the first opening 1-152 and the second opening 1-154. The HMD 1-100 can also include a first button 1-126 disposed in the first aperture 1-128 and a second button 1-132 disposed in the second aperture 1-130. The first button 1-128 and the second button 1-132 can be pressed through the respective apertures 1-126, 1-130. In at least one example, the first button 1-126 and / or the second button 1-132 can be a twistable dial and a pressable button. In at least one example, the first button 1-128 is a pressable and twistable dial button and the second button 1-132 is a pressable button.
[0090] Figure 1C A rear perspective view of the HMD 1-100 is illustrated. The HMD 1-100 can include a light seal 1-110 extending rearward from the housing 1-150 of the display assembly 1-108 around a perimeter of the housing 1-150, as shown. The light seal 1-110 can be configured to extend from the housing 1-150 to a face of a user, around eyes of the user, to block external light from being visible. In one example, the HMD 1-100 can include a first display assembly 1-120a and a second display assembly 1-120b disposed at or in a rear-facing second opening 1-154 defined by the housing 1-150 and configured to project light through the second opening 1-154. In at least one example, each display assembly 1-120a-b can include a respective display screen 1-122a, 1-122b configured to project light in a rearward direction through the second opening 1-154 toward eyes of a user.
[0091] In at least one example, with reference to Figure 1B and Figure 1C both, the display assembly 1-108 can be a front-facing forward display assembly including a display screen configured to project light in a first forward direction and the rear-facing display screens 1-122a-b can be configured to project light in a second rearward direction opposite the first direction. As described above, the light seal 1-110 can be configured to block light external to the HMD 1-100 from reaching the eyes of a user, including light projected by the forward display screen of the display assembly 1-108 shown in the front perspective view of Figure 1B In at least one example, the HMD 1-100 can also include a curtain 1-124 occluding the second opening 1-154 between the housing 1-150 and the rear-facing display assemblies 1-120a-b. In at least one example, the curtain 1-124 can be elastic or at least partially elastic.
[0092] Figure 1B and Figure 1C any of the features, components, and / or parts, including their arrangements and configurations, shown in Figures 1D to 1F any of the other examples of apparatuses, features, components, and parts shown and described herein. Likewise, any of the features, components, and / or parts, including their arrangements and configurations, shown and described in Figures 1D to 1F any of the examples of apparatuses, features, components, and parts shown and described herein. Figure 1B and Figure 1C may be included individually or in any combination.
[0093] Figure 1D An exploded view illustrating an example of an HMD 1-200 that includes various portions or parts that are separated according to the modularization and selective coupling of those parts. For example, the HMD 1-200 can include a band 1-216 that is selectively coupleable to a first electronic band 1-205a and a second electronic band 1-205b. The first stationary band 1-205a can include first electronic components 1-212a, and the second stationary band 1-205b can include second electronic components 1-212b. In at least one example, the first and second bands 1-205a-b are removably coupleable to a display unit 1-202.
[0094] Further, the HMD 1-200 can include a light seal 1-210 that is configured to be removably coupled to the display unit 1-202. The HMD 1-200 can also include a lens 1-218 that is removably coupleable to the display unit 1-202, for example, on a first assembly that includes a display screen and a second display assembly. The lens 1-218 can include a custom prescription lens that is configured for correcting vision. As noted, in the exploded view of Figure 1D each of the parts shown in the exploded view and described above can be removably coupled, attached, reattached, and replaced to update the part or swap out the part for a different user. For example, bands such as the band 1-216, light seals such as the light seal 1-210, lenses such as the lens 1-218, and electronic bands such as the electronic bands 1-205a-b can be swapped out according to a user, such that these portions are customized to fit and correspond to a single user of the HMD 1-200.
[0095] Figure 1D any of the features, components, and / or parts, including their arrangements and configurations, shown in Figure 1B , Figure 1C and Figures 1E to 1Fany other example of the devices, features, components, and parts shown and described herein. Likewise, any of the features, components, and / or parts shown or described, including their arrangement and configuration, can be included in the examples of devices, features, components, and parts shown in Figure 1B , Figure 1C and Figures 1E to 1F may be included individually or in any combination in the examples of devices, features, components, and parts shown in Figure 1D .
[0096] Figure 1E An exploded view illustrating an example of a display unit 1-306 of an HMD is shown. The display unit 1-306 can include a front display assembly 1-308, a frame / housing assembly 1-350, and a curtain assembly 1-324. The display unit 1-306 can also include a sensor assembly 1-350, a logic board assembly 1-358, and a cooling assembly 1-360 disposed between the frame assembly 1-356 and the front display assembly 1-308. In at least one example, the display unit 1-306 can also include a rear display assembly 1-320 including a first rear display screen 1-322a and a second rear display screen 1-322b disposed between the frame 1-350 and the curtain assembly 1-324.
[0097] In at least one example, the display unit 1-306 can also include a motor assembly 1-362 configured as an adjustment mechanism for adjusting the positioning of the display screens 1-322a-b of the display assembly 1-320 relative to the frame 1-350. In at least one example, the display assembly 1-320 is mechanically coupled to the motor assembly 1-362, each display screen 1-322a-b having at least one motor such that the motors are able to translate the display screens 1-322a-b to match the interpupillary distance of the user’s eyes.
[0098] In at least one example, the display unit 1-306 can include a dial or button 1-328 that can be pressed relative to the frame 1-350 and that can be accessed by a user outside of the frame 1-350. The button 1-328 can be electrically connected to the motor assembly 1-362 via a controller such that the button 1-328 can be manipulated by a user to cause the motors of the motor assembly 1-362 to adjust the positioning of the display screens 1-322a-b.
[0099] Figure 1E any of the features, components, and / or parts shown or described, including their arrangement and configuration, can be included individually or in any combination in the examples of devices, features, components, and parts shown in Figures 1B to 1D and Figure 1F may be included individually or in any combination in the examples of devices, features, components, and parts shown in Figures 1B to 1D and Figure 1FAny of the features, components, and / or parts shown and described, including their arrangements and configurations, can be included alone or in any combination Figure 1E Examples of the devices, features, components, and parts shown.
[0100] Figure 1F An exploded view of another example of a display unit 1-406 of an HMD device similar to other HMD devices described herein is illustrated. The display unit 1-406 can include a front display assembly 1-402, a sensor assembly 1-456, a logic board assembly 1-458, a cooling assembly 1-460, a frame assembly 1-450, a rear display assembly 1-421, and a curtain assembly 1-424. The display unit 1-406 can also include a motor assembly 1-462 for adjusting the positioning of first and second display subassemblies 1-420a, 1-420b of the rear display assembly 1-421, including first and second respective display screens for inter-pupillary adjustment, as described above.
[0101] Figure 1F The various parts, systems, and assemblies shown in the exploded view of Figures 1B to 1E are described in greater detail herein with reference to Figure 1F The display unit 1-406 shown can be assembled and integrated with Figures 1B to 1E fixation mechanisms including electronic straps, bands, and other components including light seals, connection assemblies, and the like.
[0102] Figure 1F Any of the features, components, and / or parts shown and described, including their arrangements and configurations, can be included alone or in any combination Figures 1B to 1E Any of the other examples of devices, features, components, and parts shown and described herein. Likewise, reference is made to Figures 1B to 1E Any of the features, components, and / or parts shown and described, including their arrangements and configurations, can be included alone or in any combination Figure 1F Examples of the devices, features, components, and parts shown.
[0103] Figure 1G A perspective exploded view of a front cover assembly 3-100 of an HMD device described herein is illustrated, for example Figure 1G The front cover assembly 3-1 of an HMD 3-100 shown or any other HMD device shown and described herein. Figure 1GThe front cover assembly 3-100 shown may include a transparent or translucent cover 3-102, a shield 3-104 (or “canopy”), an adhesive layer 3-106, a display assembly 3-108 including a biconvex lens panel or array 3-110, and a structural decorative element 3-112. The adhesive layer 3-106 secures the shield 3-104 and / or the transparent cover 3-102 to the display assembly 3-108 and / or the decorative element 3-112. The decorative element 3-112 secures various components of the front cover assembly 3-100 to the frame or base of the HMD device.
[0104] In at least one example, such as Figure 1G As shown, a transparent cover 3-102, a shield 3-104, and a display assembly 3-108, including a biconvex lens array 3-110, can be bent to accommodate the curvature of a user's face. The transparent cover 3-102 and the shield 3-104 can be bent in two or three dimensions, for example, vertically in and out of the Z-plane along the Z direction, and horizontally in and out of the Z-plane along the X direction. In at least one example, the display assembly 3-108 may include the biconvex lens array 3-110 and a display panel with pixels configured to project light through the shield 3-104 and the transparent cover 3-102. The display assembly 3-108 can be bent in at least one direction (e.g., the horizontal direction) to accommodate the curvature of a user's face from one side (e.g., the left) to the other (e.g., the right). In at least one example, each layer or component of the display assembly 3-108 (which will be shown and described in more detail in the following figures, but may include the biconvex lens array 3-110 and the display layer) may be similarly or concentrically curved in the horizontal direction to accommodate the curvature of the user's face.
[0105] In at least one example, the cover 3-104 may include a transparent or translucent material through which the display component 3-108 projects light. In one example, the cover 3-104 may include one or more opaque portions, such as opaque ink-printed portions or other opaque film portions on the back of the cover 3-104. When the HMD device is worn, the rear surface may be the surface of the cover 3-104 facing the user's eyes. In at least one example, the opaque portion may be on the front surface of the cover 3-104 opposite the rear surface. In at least one example, one or more opaque portions of the cover 3-104 may include peripheral portions that visually conceal any components surrounding the outer periphery of the display screen of the display component 3-108. In this way, the opaque portions of the cover conceal any other components of the HMD device that would otherwise be visible through the transparent or translucent cover 3-102 and / or the cover 3-104, including electronic components, structural components, etc.
[0106] In at least one example, the shroud 3-104 can define one or more apertured transparent portions 3-120 through which the sensor can transmit and receive signals. In one example, the portions 3-120 are apertures through which the sensor can extend or through which the sensor can transmit and receive signals. In one example, the portions 3-120 are transparent portions, or portions that are more transparent than the surrounding translucent or opaque portions of the shroud, through which the sensor can transmit and receive signals through the shroud and through the transparent cover 3-102. In one example, the sensor can include a camera, an IR sensor, a LUX sensor, or any other visual or non-visual environmental sensor of the HMD device.
[0107] Figure 1G Any of the illustrated features, components, and / or parts, including their arrangement and configuration, can be included in any of the other examples of devices, features, components, and parts described herein, either alone or in any combination. Likewise, any of the features, components, and / or parts illustrated and described herein, including their arrangement and configuration, can be included in the examples of devices, features, components, and parts illustrated Figure 1G in the examples.
[0108] Figure 1H An exploded view of an example of an HMD device 6-100 is illustrated. The HMD device 6-100 can include a sensor array or system 6-102 that includes one or more sensors, cameras, projectors, etc. mounted to one or more components of the HMD 6-100. In at least one example, the sensor system 6-102 can include a cradle 1-338 to which one or more sensors of the sensor system 6-102 can be secured / fastened.
[0109] Figure 1I A portion of the HMD device 6-100 including a front transparent cover 6-104 and a sensor system 6-102 is illustrated. The sensor system 6-102 can include a number of different sensors, emitters, receivers, including cameras, IR sensors, projectors, etc. The transparent cover 6-104 is illustrated in front of the sensor system 6-102 to illustrate the relative positioning of the various sensors and emitters and the orientation of each sensor / emitter of the system 6-102. As referenced herein, "lateral," "side," "laterally," "horizontally," and other similar terms refer to the orientation or direction as Figure 1J indicated by the X-axis illustrated. Terms such as "vertical," "up," "down," and similar terms refer to the orientation or direction as Figure 1J indicated by the Z-axis illustrated. Terms such as "forward," "backward," "forwardly," "backwardly," and similar terms refer to the orientation or direction as Figure 1J indicated by the Y-axis illustrated.
[0110] In at least one example, the transparent cover 6-104 can define a front outer surface of the HMD device 6-100, and the sensor system 6-102 including various sensors and components thereof can be disposed behind the cover 6-104 in the Y axis / direction. The cover 6-104 can be transparent or translucent to allow light to pass through the cover 6-104, including both light detected by the sensor system 6-102 and light emitted thereby.
[0111] As described elsewhere herein, the HMD device 6-100 can include one or more controllers including processors for electrically coupling the various sensors and emitters of the sensor system 6-102 with one or more motherboards, processing units, and other electronic devices such as display screens, etc. Further, as will be shown in greater detail below with reference to other figures, the various sensors, emitters, and other components of the sensor system 6-102 can be coupled to various structural frame members, brackets, etc. of the HMD device 6-100 not shown in FIG. 6-1. Figure 1I For clarity of illustration, the components of the sensor system 6-102 are shown detached and not electrically coupled with other components. Figure 1I For clarity of illustration, the components of the sensor system 6-102 are shown detached and not electrically coupled with other components.
[0112] In at least one example, the device can include one or more controllers having processors configured to execute instructions stored on memory components electrically coupled to the processors. The instructions can include or cause the processors to execute one or more algorithms for self-correcting the angle and positioning of the various cameras described herein over time as the initial positioning, angle, or orientation of the cameras is impacted or distorted due to an accidental drop event or other event.
[0113] In at least one example, the sensor system 6-102 can include one or more scene cameras 6-106. The system 6-102 can include two scene cameras 6-102 disposed on either side of the bridge or arch structure of the HMD device 6-100 such that each of the two cameras 6-106 generally corresponds to the positioning of the left and right eyes of the user behind the cover 6-103. In at least one example, the scene cameras 6-106 are generally oriented forward in the Y direction to capture images of the user’s forward field of view during use of the HMD 6-100. In at least one example, the scene cameras are color cameras and provide images and content for MR video pass-through to a display screen facing the user’s eyes when using the HMD device 6-100. The scene cameras 6-106 can also be used for environment and object reconstruction.
[0114] In at least one example, sensor system 6-102 can include a first depth sensor 6-108 that is generally directed forward in the Y direction. In at least one example, first depth sensor 6-108 can be used for environment and object reconstruction as well as hand and body tracking of a user. In at least one example, sensor system 6-102 can include a second depth sensor 6-110 that is centrally disposed along the width of HMD device 6-100 (e.g., along the X axis). For example, second depth sensor 6-110 can be disposed on a central nose bridge or on a fitting structure above the nose of a user when donning HMD 6-100. In at least one example, second depth sensor 6-110 can be used for environment and object reconstruction as well as hand and body tracking. In at least one example, the second depth sensor can include a LIDAR sensor.
[0115] In at least one example, sensor system 6-102 can include a depth projector 6-112 that is generally forward facing to project electromagnetic waves (e.g., in the form of a predetermined pattern of light dots) into or within the field of view of the user and / or scene camera 6-106, or into or within a field of view that includes and extends beyond the field of view of the user and / or scene camera 6-106. In at least one example, the depth projector is capable of projecting electromagnetic waves of light in the form of a pattern of dots that reflect off of objects and back into the aforementioned depth sensors, including depth sensors 6-108, 6-110. In at least one example, depth projector 6-112 can be used for environment and object reconstruction as well as hand and body tracking.
[0116] In at least one example, sensor system 6-102 can include downward facing cameras 6-114 that are generally directed downward in the Z axis relative to HMD device 6-100. In at least one example, downward facing cameras 6-114 can be disposed on the left and right sides of HMD device 6-100 as shown and used for hand and body tracking, headset tracking, and facial avatar detection and creation for display of a user avatar on a front-facing display screen of HMD device 6-100 as described elsewhere herein. For example, downward facing cameras 6-114 can be used to capture facial expressions and movements of a user’s face below HMD device 6-100, including cheeks, mouth, and chin.
[0117] In at least one example, the sensor system 6-102 can include a chin camera 6-116. In at least one example, the chin camera 6-116 can be disposed on the left and right sides of the HMD device 6-100 as shown and used for hand and body tracking, headset tracking, and facial avatar detection and creation for display of a user avatar on the front-facing display screen of the HMD device 6-100 as described elsewhere herein. For example, the chin camera 6-116 can be used to capture facial expressions and movements of a user's face below the HMD device 6-100, including the user's chin, cheeks, mouth, and jaw. The hand and body tracking, headset tracking, and facial avatar detection and recreation
[0118] In at least one example, the sensor system 6-102 can include side cameras 6-118. The side cameras 6-118 can be oriented to capture left and right side views in the X-axis or direction relative to the HMD device 6-100. In at least one example, the side cameras 6-118 can be used for hand and body tracking, headset tracking, and facial avatar detection and recreation.
[0119] In at least one example, the sensor system 6-102 can include a plurality of eye tracking and gaze tracking sensors for determining identity, condition, and gaze direction of the user's eyes during and / or prior to use. In at least one example, the eye / gaze tracking sensors can include a nose-eye camera 6-120 disposed on either side of the user's nose and adjacent the user's nose when the HMD device 6-100 is donned. The eye / gaze sensors can also include bottom eye cameras 6-122 disposed below the respective user's eyes for capturing images of the eyes for facial avatar detection and creation, gaze tracking, and iris identification functions.
[0120] In at least one example, the sensor system 6-102 can include an infrared illuminator 6-124 that points outward from the HMD device 6-100 to illuminate the external environment and any objects therein with IR light for IR detection with one or more IR sensors of the sensor system 6-102. In at least one example, the sensor system 6-102 can include a flicker sensor 6-126 and an ambient light sensor 6-128. In at least one example, the flicker sensor 6-126 can detect a ceiling light refresh rate to avoid display flicker. In one example, the infrared illuminator 6-124 can include a light emitting diode and can be particularly used for low light environments for illuminating the user's hands and other objects in low light for detection by the infrared sensors of the sensor system 6-102.
[0121] In at least one example, the multiple sensors, including scene camera 6-106, downward camera 6-114, mandible camera 6-116, side camera 6-118, depth projector 6-112, and depth sensors 6-108, 6-110, can be used in combination with an electrically coupled controller to combine depth data with camera data for hand tracking and for size determination in order to better perform hand tracking and object recognition and tracking functions for HMD device 6-100. In at least one example, downward camera 6-114, mandible camera 6-116, and side camera 6-118 described above and shown in FIG. 6B can be wide angle cameras capable of working in the visible and infrared spectrum. In at least one example, these cameras 6-114, 6-116, 6-118 can work in black and white light detection only to simplify image processing and obtain sensitivity. Figure 1I
[0122] Figure 1I Any of the features, components, and / or parts shown, including their arrangement and configuration, can be included individually or in any combination in Figures 1J to 1L any other example of the devices, features, components, and parts shown and described herein. Likewise, any of the features, components, and / or parts shown and described can be included individually or in any combination in Figures 1J to 1L examples of the devices, features, components, and parts shown. Figure 1I
[0123] Figure 1J A lower perspective view of an example of HMD 6-200 including a cover or shroud 6-204 secured to frame 6-230 is illustrated. In at least one example, sensors 6-202 of sensor system 6-203 can be disposed around the perimeter of HMD 6-200 such that sensors 6-203 are disposed outward around the perimeter of display area or zone 6-232 so as not to obstruct viewing of displayed light. In at least one example, sensors can be disposed behind shroud 6-204 and aligned with transparent portions of the shroud, allowing sensors and projectors to allow light to pass back and forth through shroud 6-204. In at least one example, an opaque ink or other opaque material or film / layer can be disposed on shroud 6-204 around display zone 6-232 to hide components of HMD 6-200 outside of display zone 6-232 other than transparent portions defined by the opaque portions through which sensors and projectors transmit and receive light and electromagnetic signals during operation. In at least one example, shroud 6-204 allows light to pass from a display (e.g., within display area 6-232) but not radially outward from the display area around the perimeter of the display and shroud 6-204.
[0124] In some examples, the shield 6-204 includes a transparent portion 6-205 and an opaque portion 6-207, as described above and elsewhere herein. In at least one example, the opaque portion 6-204 of the shield 6-207 may define one or more transparent areas 6-209 through which the sensor 6-203 of the sensor system 6-202 transmits and receives signals. In the illustrated example, the sensor 6-203 of the sensor system 6-202 transmits and receives signals through the shield 6-204, or more specifically through (or defined by) the transparent area 6-209 of the opaque portion 6-207 of the shield 6-204, the sensor may include... Figure 1I The examples show the same or similar sensors, such as depth sensors 6-108 and 6-110, a depth projector 6-112, a first scene camera and a second scene camera 6-106, a first downward camera and a second downward camera 6-114, a first side camera and a second side camera 6-118, and a first infrared illuminator and a second infrared illuminator 6-124. These sensors also... Figure 1K and Figure 1L The example is shown. Other sensors, sensor types, number of sensors, and their relative positioning can be included in one or more other examples of the HMD.
[0125] Figure 1J Any of the features, components, and / or parts shown, including their arrangement and configuration, may be included individually or in any combination. Figure 1I and Figures 1K to 1L In any of the other examples of devices, features, components, and parts shown and described herein. Similarly, refer to... Figure 1I and Figures 1K to 1L Any of the features, components, and / or parts shown or described, including their arrangement and configuration, may be included individually or in any combination. Figure 1J Examples of devices, features, components, and parts are shown.
[0126] Figure 1K A front view of a portion of an example of an HMD device 6-300 is shown, including a display 6-334, brackets 6-336, 6-338, and a frame or housing 6-330. Figure 1K The examples shown do not include a front cover or shield to illustrate brackets 6-336 and 6-338. For example, Figure 1J The shield 6-204 shown includes an opaque portion 6-207 that visually covers / blocks the view of anything outside the display / display area 6-334 (e.g., radially / peripherally), including the sensor 6-303 and the bracket 6-338.
[0127] In at least one example, various sensors of the sensor system 6-302 are coupled to the cradle 6-336, 6-338. In at least one example, the scene cameras 6-306 include tight tolerances for the angles relative to one another. For example, the tolerance for the mounting angle between two scene cameras 6-306 can be 0.5 degrees or less, such as 0.3 degrees or less. To achieve and maintain such tight tolerances, in one example, the scene cameras 6-306 can be mounted to the cradle 6-338 rather than the shroud. The cradle can include a cantilever on which the scene cameras 6-306, as well as other sensors of the sensor system 6-302, can be mounted to remain positioned and oriented invariant to drop events that cause other cradles 6-226, the housing 6-330, and / or the shroud to deform.
[0128] Figure 1K Any of the illustrated features, components, and / or parts, including their arrangement and configuration, can be included alone or in any combination in Figures 1I to 1J and Figure 1L any other example of the devices, features, components, and parts illustrated and described herein. Likewise, reference to Figures 1I to 1J and Figure 1L any of the features, components, and / or parts illustrated or described, including their arrangement and configuration, can be included alone or in any combination in Figure 1K the examples of the devices, features, components, and parts illustrated.
[0129] Figure 1L A bottom view of an example of an HMD 6-400 including a front display / cover assembly 6-404 and a sensor system 6-402 is illustrated. The sensor system 6-402 can be similar to other sensor systems described above and elsewhere herein, including with reference to Figures 1I to 1K In at least one example, the mandible camera 6-416 can face downward to capture images of the lower facial features of the user. In one example, the mandible camera 6-416 can be directly coupled to the frame or housing 6-430 or one or more internal cradles that are directly coupled to the illustrated frame or housing 6-430. The frame or housing 6-430 can include one or more apertures / openings 6-415 through which the mandible camera 6-416 can transmit and receive signals.
[0130] Figure 1L Any of the illustrated features, components, and / or parts, including their arrangement and configuration, can be included alone or in any combination in Figures 1I to 1K any other example of the devices, features, components, and parts illustrated and described herein. Likewise, reference to Figures 1I to 1K any of the features, components, and / or parts illustrated and described, including their arrangement and configuration, can be included alone or in any combination in Figure 1L Examples of the devices, features, components, and parts shown.
[0131] Figure 1M A rear perspective view of an interpupillary distance (IPD) adjustment system 11.1.1-102 is illustrated, including first and second optical modules 11.1.1-104a-b slidably engaged / coupled to respective guide rods 11.1.1-108a-b and motors 11.1.1-110a-b of left and right adjustment subsystems 11.1.1-106a-b. The IPD adjustment system 11.1.1-102 can be coupled to a cradle 11.1.1-112 and include a button 11.1.1-114 in electrical communication with the motors 11.1.1-110a-b. In at least one example, the button 11.1.1-114 can be in electrical communication with the first and second motors 11.1.1-110a-b via a processor or other circuitry component to activate the first and second motors 11.1.1-110a-b and cause the first and second optical modules 11.1.1-104a-b to change positioning relative to one another, respectively.
[0132] In at least one example, the first and second optical modules 11.1.1-104a-b can include respective display screens configured to project light toward a user’s eyes when donning the HMD 11.1.1-100. In at least one example, a user can manipulate (e.g., press and / or rotate) the button 11.1.1-114 to activate positioning adjustment of the optical modules 11.1.1-104a-b to match the interpupillary distance of the user’s eyes. The optical modules 11.1.1-104a-b can also include one or more cameras or other sensors / sensor systems for imaging and measuring the IPD of a user, such that the optical modules 11.1.1-104a-b can be adjusted to match the IPD.
[0133] In one example, the user manipulates the button 11.1.1-114 to cause an automatic positioning adjustment of the first and second optical modules 11.1.1-104a-b. In one example, the user manipulates the button 11.1.1-114 to cause a manual adjustment such that the optical modules 11.1.1-104a-b move further apart or closer together (e.g., when the user rotates the button 11.1.1-114 one way or another) until the user visually matches her / his own IPD. In one example, the manual adjustment is communicated electronically via one or more circuits, and power to move the optical modules 11.1.1-104a-b via the motors 11.1.1-110a-b is provided by a power source. In one example, the adjustment and movement of the optical modules 11.1.1-104a-b via the manipulation of the button 11.1.1-114 is mechanically actuated via the movement of the button 11.1.1-114.
[0134] Figure 1M Any of the features, components, and / or parts shown, including their arrangement and configuration, can be included individually or in any combination in any of the other examples of devices, features, components, and parts shown in any of the other figures and described herein. Likewise, any of the features, components, and / or parts shown or described with reference to any of the other figures, including their arrangement and configuration, can be included individually or in any combination in the examples of devices, features, components, and parts shown. Figure 1M in the examples of devices, features, components, and parts shown.
[0135] Figure 1N A front perspective view illustrates a portion of an HMD 11.1.2-100, including an outer structural frame 11.1.2-102 and an inner or intermediate structural frame 11.1.2-104 that define a first aperture 11.1.2-106a and a second aperture 11.1.2-106b. The apertures 11.1.2-106a-b are shown in dashed lines in Figure 1N as the view of the apertures 11.1.2-106a-b can be obstructed by one or more other components of the HMD 11.1.2-100 that are coupled to the inner frame 11.1.2-104 and / or the outer frame 11.1.2-102, as shown. In at least one example, the HMD 11.1.2-100 can include a first mounting bracket 11.1.2-108 that is coupled to the inner frame 11.1.2-104. In at least one example, the mounting bracket 11.1.2-108 is coupled to the inner frame 11.1.2-104 between the first and second apertures 11.1.2-106a-b.
[0136] Mounting brackets 11.1.2-108 may include intermediate or central portions 11.1.2-109 coupled to the inner frame 11.1.2-104. In some examples, the intermediate or central portions 11.1.2-109 may not be the geometric center or middle of the bracket 11.1.2-108. Instead, the intermediate / central portions 11.1.2-109 may be positioned between a first cantilever extension arm and a second cantilever extension arm extending away from the intermediate portions 11.1.2-109. In at least one example, mounting bracket 108 includes first cantilever arms 11.1.2-112 and second cantilever arms 11.1.2-114 extending away from the intermediate portions 11.1.2-109 of the mounting bracket 11.1.2-108 coupled to the inner frame 11.1.2-104.
[0137] like Figure 1N As shown, the outer frame 11.1.2-102 may define a curved geometry on its underside to adapt to the user's nose when the user wears the HMD 11.1.2-100. This curved geometry may be referred to as the bridge of the nose 11.1.2-111 and is centrally located on the underside of the HMD 11.1.2-100 as shown. In at least one example, the mounting bracket 11.1.2-108 may be connected to the inner frame 11.1.2-104 between holes 11.1.2-106a-b, such that the cantilever 11.1.2-112, 11.1.2-114 extend downward and laterally outward away from the central portion 11.1.2-109 to complement the nose bridge geometry of the outer frame 11.1.2-102. In this way, the mounting bracket 11.1.2-108 is configured to adapt to the user's nose, as described above. The geometry of the bridge of the nose 11.1.2-111 adapts to the nose, as it provides a curvature that conforms to the shape of the user's nose, offering a comfortable fit from above, above, and around.
[0138] The first cantilever 11.1.2-112 can extend away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-108 in a first direction, and the second cantilever 11.1.2-114 can extend away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-108 in a second direction opposite the first direction. The first cantilever 11.1.2-112 and the second cantilever 11.1.2-114 are referred to as“cantilevered” or“cantilever” arms because each arm 11.1.2-112, 11.1.2-114 includes a free distal end 11.1.2-116, 11.1.2-118, respectively, that is not attached to the inner frame 11.1.2-102 and the outer frame 11.1.2-104. In this way, the arms 11.1.2-112, 11.1.2-114 cantilever from the middle portion 11.1.2-109, which can be connected to the inner frame 11.1.2-104, while the distal ends 11.1.2-102, 11.1.2-104 are unattached.
[0139] In at least one example, the HMD 11.1.2-100 can include one or more components coupled to the mounting bracket 11.1.2-108. In one example, the components include a plurality of sensors 11.1.2-110a-f. Each of the plurality of sensors 11.1.2-110a-f can include various types of sensors, including cameras, IR sensors, and the like. In some examples, one or more of the sensors 11.1.2-110a-f can be used for object recognition in three-dimensional space, such that it is important to maintain the precise relative positions of two or more of the plurality of sensors 11.1.2-110a-f. The cantilevered nature of the mounting bracket 11.1.2-108 can protect the sensors 11.1.2-110a-f from damage and repositioning in the event of an accidental drop by the user. Because the sensors 11.1.2-110a-f are overhanging on the arms 11.1.2-112, 11.1.2-114 of the mounting bracket 11.1.2-108, stresses and deformations of the inner and / or outer frames 11.1.2-104, 11.1.2-102 are not transmitted to the cantilevers 11.1.2-112, 11.1.2-114, and thus do not affect the relative positioning of the sensors 11.1.2-110a-f coupled / mounted to the mounting bracket 11.1.2-108.
[0140] Figure 1NAny of the illustrated features, components, and / or parts, including their arrangement and configuration, can be included in any of the other examples of devices, features, components described herein, alone or in any combination. Likewise, any of the features, components, and / or parts illustrated and described herein, including their arrangement and configuration, can be included in any of the other examples of devices, features, components described herein, alone or in any combination. Figure 1N Examples of the devices, features, components, and parts illustrated.
[0141] Figure 10 An example of an optical module 11.3.2-100 for an electronic device such as an HMD, including the HDM devices described herein, is illustrated. As shown in one or more other examples described herein, the optical module 11.3.2-100 can be one of two optical modules within an HMD, with each optical module aligned to project light toward an eye of a user. In this way, a first optical module can project light toward a first eye of a user via a display screen, and a second optical module of the same device can project light toward a second eye of the user via another display screen.
[0142] In at least one example, the optical module 11.3.2-100 can include an optical frame or housing 11.3.2-102, which can also be referred to as a barrel or optical module barrel. The optical module 11.3.2-100 can also include a display 11.3.2-104 coupled to the housing 11.3.2-102, which includes one or more display screens. The display 11.3.2-104 can be coupled to the housing 11.3.2-102 such that the display 11.3.2-104 is configured to project light toward an eye of a user when donned during use by an HMD to which the display module 11.3.2-100 belongs. In at least one example, the housing 11.3.2-102 can surround the display 11.3.2-104 and provide connection features for coupling other components of the optical module described herein.
[0143] In one example, optical module 11.3.2-100 can include one or more cameras 11.3.2-106 coupled to housing 11.3.2-102. Cameras 11.3.2-106 can be positioned relative to display 11.3.2-104 and housing 11.3.2-102 such that cameras 11.3.2-106 are configured to capture one or more images of a user’s eyes during use. In at least one example, optical module 11.3.2-100 can also include a light bar 11.3.2-108 that surrounds display 11.3.2-104. In one example, light bar 11.3.2-108 is disposed between display 11.3.2-104 and cameras 11.3.2-106. Light bar 11.3.2-108 can include a plurality of lights 11.3.2-110. The plurality of lights can include one or more light-emitting diodes (LEDs) or other lights configured to project light toward a user’s eyes when the HMD is donned. Individual lights 11.3.2-110 in light bar 11.3.2-108 can be spaced apart around light bar 11.3.2-108, and thus, evenly or unevenly, around display 11.3.2-104 at various locations on light bar 11.3.2-108 and around display 11.3.2-104.
[0144] In at least one example, housing 11.3.2-102 defines a viewing opening 11.3.2-101 through which a user can view display 11.3.2-104 when the HMD device is donned. In at least one example, the LEDs are configured and arranged to emit light through viewing opening 11.3.2-101 onto a user’s eyes. In one example, cameras 11.3.2-106 are configured to capture one or more images of a user’s eyes through viewing opening 11.3.2-101.
[0145] As described above, Figure 10 Each of the components and features of optical module 11.3.2-100 shown can be replicated in another (e.g., second) optical module disposed with the HMD to interact with (e.g., project light and capture images of) the other eye of the user.
[0146] Figure 10 Any of the features, components, and / or parts shown, including their arrangement and configuration, can be included individually or in any combination in Figure 1P Any of the other examples of devices, features, components, and parts shown or otherwise described herein. Likewise, reference to Figure 1P Any of the features, components, and / or parts shown or otherwise described herein, including their arrangement and configuration, can be included individually or in any combination in Figure 10Examples of the devices, features, components, and parts shown.
[0147] Figure 1P A cross-sectional view illustrating an example of an optical module 11.3.2-200 is shown, including a housing 11.3.2-202, a display assembly 11.3.2-204 coupled to the housing 11.3.2-202, and a lens 11.3.2-216 coupled to the housing 11.3.2-202. In at least one example, the housing 11.3.2-202 defines a first aperture or channel 11.3.2-212 and a second aperture or channel 11.3.2-214. The channels 11.3.2-212, 11.3.2-214 can be configured to slidably engage corresponding rails or guide rods of an HMD device to allow the optical module 11.3.2-200 to adjust positioning relative to a user’s eyes to match the user’s interpupillary distance (IPD). The housing 11.3.2-202 can be slidably engaged with the guide rods to secure the optical module 11.3.2-200 in place within the HMD.
[0148] In at least one example, the optical module 11.3.2-200 can also include a lens 11.3.2-216 coupled to the housing 11.3.2-202 and disposed between the display assembly 11.3.2-204 and the user’s eyes when the HMD is donned. The lens 11.3.2-216 can be configured to direct light from the display assembly 11.3.2-204 to the user’s eyes. In at least one example, the lens 11.3.2-216 can be part of a lens assembly, including a corrective lens removably attached to the optical module 11.3.2-200. In at least one example, the lens 11.3.2-216 is disposed above the light bar 11.3.2-208 and the one or more eye tracking cameras 11.3.2-206, such that the cameras 11.3.2-206 are configured to capture images of the user’s eyes through the lens 11.3.2-216, and the light bar 11.3.2-208 includes lights configured to project light to the user’s eyes through the lens 11.3.2-216 during use.
[0149] Figure 1P Any of the features, components, and / or parts shown, including their arrangement and configuration, can be included in any of the other examples of devices, features, components, and parts described herein, alone or in any combination. Likewise, any of the features, components, and / or parts shown and described herein, including their arrangement and configuration, can be included in the devices, features, components, and parts described in the examples shown Figure 1P Examples of the devices, features, components, and parts shown.
[0150] Figure 2is a block diagram of an example of a controller 110 according to some embodiments. While certain specific features are illustrated, those skilled in the art will appreciate from the present disclosure that various other features have not been illustrated for the sake of brevity and so as not to obscure more pertinent aspects of the embodiments disclosed herein. To that end, as a non-limiting example, in some embodiments, the controller 110 includes one or more processing units 202 (e.g., microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), graphics processing units (GPUs), central processing units (CPUs), processing cores, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., universal serial bus (USB), FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.1 lx, IEEE 802.16x, global system for mobile communications (GSM), code division multiple access (CDMA), time division multiple access (TDMA), global positioning system (GPS), infrared (IR), BLUETOOTH, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these and various other components.
[0151] In some embodiments, the one or more communication buses 204 include circuitry that interconnects and controls communications between system components. In some embodiments, the one or more I / O devices 206 include at least one of a keyboard, mouse, touchpad, joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.
[0152] The memory 220 includes high-speed random access memory, such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate random access memory (DDR RAM), or other random access solid state memory devices. In some embodiments, the memory 220 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid state storage devices. The memory 220 optionally includes one or more storage devices remotely located from the one or more processing units 202. The memory 220 comprises a non-transitory computer readable storage medium. In some embodiments, the memory 220, or the non-transitory computer readable storage medium of the memory 220, stores the following programs, modules, and data structures, or a subset thereof, including an optional operating system 230 and an XR experience module 240.
[0153] The operating system 230 includes instructions for handling various basic system services and for executing hardware-related tasks. In some embodiments, the XR experience module 240 is configured to manage and coordinate single or multiple XR experiences for one or more users (e.g., single XR experiences for one or more users, or multiple XR experiences for respective groupings of one or more users). To this end, in various embodiments, the XR experience module 240 includes a data acquisition unit 241, a tracking unit 242, a coordination unit 246, and a data transmission unit 248.
[0154] In some embodiments, the data acquisition unit 241 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the display generation component 120, and optionally from one or more of the input devices 125, the output devices 155, the sensors 190, and / or the peripheral devices 195. To this end, in various embodiments, the data acquisition unit 241 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics. Figure 1A
[0155] In some embodiments, the tracking unit 242 is configured to map the scene 105, and to track the positioning / location of at least the display generation component 120 relative to the scene 105, and optionally to track the location of one or more of the input devices 125, the output devices 155, the sensors 190, and / or the peripheral devices 195. To this end, in various embodiments, the tracking unit 242 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics. In some embodiments, the tracking unit 242 includes a hand tracking unit 244 and / or an eye tracking unit 243. In some embodiments, the hand tracking unit 244 is configured to track the positioning / location of one or more portions of a user’s hand, and / or the motion of one or more portions of a user’s hand relative to the scene 105, relative to the display generation component 120, and / or relative to a coordinate system that is defined relative to the user’s hand, with respect to the scene 105. Figure 1A Figure 1A The hand tracking unit 244 is described in greater detail below with respect to the scene 105. In some embodiments, the eye tracking unit 243 is configured to track the positioning or movement of a user’s gaze (or more broadly, the user’s eyes, face, or head) relative to the scene 105 (e.g., relative to the physical environment and / or relative to the user (e.g., the user’s hand)), or relative to XR content displayed via the display generation component 120. The eye tracking unit 243 is described in greater detail below with respect to the scene 105. Figure 4 Figure 5
[0156] In some embodiments, the coordination unit 246 is configured to manage and coordinate the XR experience presented to the user by the display generation component 120, and optionally by one or more of the output devices 155 and / or peripheral devices 195. To this end, in various embodiments, the coordination unit 246 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.
[0157] In some embodiments, the data sending unit 248 is configured to send data (e.g., presentation data, position data, etc.) to at least the display generation component 120, and optionally to one or more of the input devices 125, output devices 155, sensors 190, and / or peripheral devices 195. To this end, in various embodiments, the data sending unit 248 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.
[0158] While the data acquisition unit 241, tracking unit 242 (e.g., including eye tracking unit 243 and hand tracking unit 244), coordination unit 246, and data sending unit 248 are shown as residing on a single device (e.g., the controller 110), it will be appreciated that, in other embodiments, any combination of the data acquisition unit 241, tracking unit 242 (e.g., including eye tracking unit 243 and hand tracking unit 244), coordination unit 246, and data sending unit 248 can reside in separate computing devices.
[0159] Further, Figure 2 More functionally described as the various features that can be present in a particular implementation, as distinct from structural illustrations of embodiments described herein. As will be appreciated by those of ordinary skill in the art, items shown separately could be combined and some items could be separated. For example, Figure 2 Some of the functional modules shown separately in the can be implemented in a single module, and various functions of a single functional block can be implemented by one or more functional blocks in various embodiments. The actual number of modules and the division of particular functions between them will vary from implementation to implementation, and in some embodiments, depend in part on the particular combination of hardware, software, and / or firmware chosen to implement the features of the specific implementation.
[0160] Figure 3is a block diagram of an example of a display generation component 120 according to some embodiments. While certain specific features are illustrated, those skilled in the art will appreciate from the present disclosure that various other features have not been illustrated for the sake of brevity and so as not to obscure more pertinent aspects of the embodiments disclosed herein. To that end, as a non-limiting example, in some embodiments, the display generation component 120 (e.g., an HMD) includes one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, and / or the like), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.1 lx, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, BLUETOOTH, ZIGBEE, and / or the like types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more XR displays 312, one or more optional inward- and / or outward-facing image sensors 314, memory 320, and one or more communication buses for interconnecting these and various other components, and
[0161] In some embodiments, the one or more communication buses 304 include circuitry that interconnects and controls communications between system components. In some embodiments, the one or more I / O devices and sensors 306 include an inertial measurement unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., blood pressure monitor, heart rate monitor, blood oxygen sensor, blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptics engine, and / or one or more depth sensors (e.g., structured light, time-of-flight, etc.), etc.
[0162] In some embodiments, the one or more XR displays 312 are configured to provide an XR experience to a user. In some embodiments, the one or more XR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transitory (OLET), organic light-emitting diode (OLED), surface-conduction electron-emitter display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical system (MEMS), and / or similar display types. In some embodiments, the one or more XR displays 312 correspond to diffractive, reflective, polarized, holographic, etc. waveguide displays. For example, the display generation component 120 (e.g., HMD) includes a single XR display. In another example, the display generation component 120 includes an XR display for each eye of the user. In some embodiments, the one or more XR displays 312 are capable of presenting MR and VR content. In some embodiments, the one or more XR displays 312 are capable of presenting MR or VR content.
[0163] In some embodiments, the one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user’s face, including the user’s eyes (and can be referred to as eye tracking cameras). In some embodiments, the one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user’s hands, and optionally the user’s arms (and can be referred to as hand tracking cameras). In some embodiments, the one or more image sensors 314 are configured to face forward so as to acquire image data corresponding to a scene that the user would see in the absence of the display generation component 120 (e.g., HMD) (and can be referred to as a scene camera). The one or more optional image sensors 314 can include one or more RGB cameras (e.g., with a complementary metal-oxide-semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), one or more infrared (IR) cameras, and / or one or more event-based cameras, etc.
[0164] Memory 320 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM or other random access solid state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid state storage devices. Memory 320 optionally includes one or more storage devices remotely located from one or more processing units 302. Memory 320 comprises a non-transitory computer readable storage medium. In some embodiments, memory 320 or the non-transitory computer readable storage medium of memory 320 stores the following programs, modules, and data structures, among others:
[0165] Operating system 330 includes procedures for handling various basic system services and for performing hardware dependent tasks. In some embodiments, XR presentation module 340 is configured to present XR content to a user via one or more XR displays 312. To this end, in various embodiments, XR presentation module 340 includes data acquisition unit 342, XR presentation unit 344, XR mapping generation unit 346, and data transmission unit 348.
[0166] In some embodiments, data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least Figure 1A controller 110. To this end, in various embodiments, data acquisition unit 342 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.
[0167] In some embodiments, XR presentation unit 344 is configured to present XR content via one or more XR displays 312. To this end, in various embodiments, XR presentation unit 344 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.
[0168] In some embodiments, XR mapping generation unit 346 is configured to generate an XR map (e.g., a 3D map of a mixed reality scene or a map of a physical environment in which computer-generated objects can be placed to generate an extended reality) based on media content data. To this end, in various embodiments, XR mapping generation unit 346 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.
[0169] In some embodiments, the data transmission unit 348 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the controller 110, and optionally one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the data transmission unit 348 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.
[0170] Although the data acquisition unit 342, the XR rendering unit 344, the XR mapping generation unit 346, and the data sending unit 348 are shown residing in a single device (e.g., Figure 1A The data acquisition unit 342, the XR rendering unit 344, the XR mapping generation unit 346, and the data sending unit 348 are located on the display generation component 120, but it should be understood that in other embodiments, any combination of the data acquisition unit 342, the XR rendering unit 344, the XR mapping generation unit 346, and the data sending unit 348 may be located in a separate computing device.
[0171] also, Figure 3 This is more of a functional description of various features that may exist in a particular implementation, and differs from the structural diagrams of the implementations described herein. As those skilled in the art will recognize, items shown individually can be combined, and some items can be separated. For example, Figure 3 Some functional modules shown individually may be implemented in a single module, and the various functions of a single functional block may be implemented in various implementations through one or more functional blocks. The actual number of modules and the division of specific functions and how features are allocated therein will vary depending on the specific implementation, and in some implementations, it depends in part on the specific combination of hardware, software and / or firmware chosen for that particular implementation.
[0172] Figure 4 This is a schematic illustration of an example implementation of the hand tracking device 140. In some implementations, the hand tracking device 140 ( Figure 1A Controlled by hand tracking unit 244 Figure 2 To track the location / position of one or more parts of a user's hand, and / or the position of one or more parts of the user's hand relative to... Figure 1A The movement is defined in scenario 105 (e.g., relative to a portion of the user's surrounding physical environment, relative to display generation component 120, or relative to a portion of the user (e.g., the user's face, eyes, or head), and / or relative to a coordinate system (defined relative to the user's hand)). In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).
[0173] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras, etc.) that captures three-dimensional scene information including at least a human user's hand 406. The image sensor 404 captures hand images with sufficient resolution to enable the fingers and their respective positions to be distinguished. The image sensor 404 typically captures images of other portions of the user's body, and can also or possibly capture images of all portions of the body, and can have zoom capabilities or dedicated sensors with increased magnification to capture images of the hand with the desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with other image sensors to capture the physical environment of the scene 105, or serves as the image sensor that captures the physical environment of the scene 105. In some embodiments, the image sensor 404 is positioned relative to the user or the user's environment in a manner that uses the field of view of the image sensor 404 or a portion thereof to define an interaction space in which movement of the hand captured by the image sensor is treated as input to the controller 110.
[0174] In some embodiments, the image sensor 404 outputs a sequence of frames containing 3D map data (and, in addition, possibly color image data) to the controller 110, which extracts high-level information from the map data. This high-level information is typically provided to an application running on the controller via an application program interface (API), which accordingly drives the display generation component 120. For example, a user can interact with software running on the controller 110 by moving his hand 406 and changing his hand pose.
[0175] In some embodiments, the image sensor 404 projects a pattern of dots onto the scene containing the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 computes 3D coordinates of points in the scene (including points on the surface of the user's hand) based on lateral shifts of the dots in the pattern by triangulation. This approach is advantageous because it does not require the user to hold or wear any kind of beacon, sensor, or other marker. The approach gives depth coordinates of points in the scene at a particular distance from the image sensor 404 relative to a predetermined reference plane. In this disclosure, it is assumed that the image sensor 404 defines an orthogonal set of x, y, z axes so that the depth coordinates of points in the scene correspond to the z component measured by the image sensor. Alternatively, the image sensor 404 (e.g., hand tracking device) can use other 3D mapping methods such as stereo imaging or time-of-flight measurement based on a single or multiple cameras or other types of sensors.
[0176] In some embodiments, the hand tracking device 140 captures and processes a time series of depth maps containing the user's hand as he moves his hand (e.g., the entire hand or one or more fingers). Software running on the processor in the image sensor 404 and / or the controller 110 processes the 3D map data to extract image patch descriptors of the hand in these depth maps. The software can match these descriptors to image patch descriptors stored in the database 408 based on a previous learning process in order to estimate the pose of the hand in each frame. The pose typically includes the 3D positions of the user's hand joints and finger tips.
[0177] The software can also analyze the trajectories of the hand and / or fingers over multiple frames in the sequence to identify a gesture. The pose estimation functionality described herein can alternate with motion tracking functionality such that image patch based pose estimation is performed only once every two (or more) frames, while tracking is used to find changes in pose that occur over the remaining frames. The pose, motion, and gesture information is provided to applications running on the controller 110 via the API described above. The program can move and modify images presented on the display generation component 120, for example, in response to the pose and / or gesture information, or perform other functions.
[0178] In some embodiments, a gesture includes an air gesture. An air gesture is a gesture that is detected without the user touching an input element that is part of a device (e.g., computer system 101, one or more input devices 125, and / or hand tracking device 140) (or independent of an input element that is part of a device) and is based on detected motion of a portion of the user's body (e.g., head, one or more arms, one or more hands, one or more fingers, and / or one or more legs) through the air (including motion of the user's body relative to an absolute reference (e.g., an angle of the user's arm relative to the ground or a distance of the user's hand from the ground), relative to another portion of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the other of the user's hands, and / or movement of a finger of the user's hand relative to another finger of the user's hand or a portion of the user's hand), and / or absolute motion of a portion of the user's body (e.g., a tap gesture including a hand moving a predetermined amount and / or speed in a predetermined pose, or a shake gesture including a predetermined speed or amount of rotation of a portion of the user's body)).
[0179] In some embodiments, input gestures used in the various examples and embodiments described herein, in accordance with some embodiments, include air gestures performed by movement of a user’s fingers relative to other fingers (or a portion of the user’s hand) for interacting with an XR environment (e.g., a virtual or mixed reality environment). In some embodiments, air gestures are gestures that are detected without the user touching an input element that is part of the device (or independent of an input element that is part of the device) and are based on detected motion of a portion of the user’s body through the air (including motion of the user’s body relative to an absolute reference (e.g., an angle of the user’s arm relative to the ground or a distance of the user’s hand from the ground), motion relative to another portion of the user’s body (e.g., movement of the user’s hand relative to the user’s shoulder, movement of one of the user’s hands relative to the other of the user’s hands, and / or movement of a user’s finger relative to another finger or portion of the hand of the user), and / or absolute motion of a portion of the user’s body (e.g., including a tap gesture that includes a hand moving a predetermined amount and / or speed in a predetermined gesture, or a shake gesture that includes a predetermined speed or amount of rotation of a portion of the user’s body)).
[0180] In some embodiments where the input gesture is an air gesture (e.g., where there is no physical contact with an input device that provides information to the computer system about which user interface element is the target of the user input, such as contact with a user interface element displayed on a touchscreen, or contact with a mouse or trackpad to move a cursor to a user interface element), the gesture takes into account the user’s attention (e.g., gaze) to determine the target of the user input (e.g., for direct input, as described below). Thus, in embodiments involving air gestures, for example, the input gesture is combined with (e.g., simultaneously with) detection of attention (e.g., gaze) toward a user interface element with movement of a user’s finger and / or hand to perform a pinch and / or tap input, as described in more detail below.
[0181] In some embodiments, input gestures directed to user interface objects are performed directly or indirectly with reference to the user interface objects. For example, user input is performed directly on a user interface object in accordance with performing an input at a location corresponding to the location of the user interface object in the three-dimensional environment (e.g., as determined based on the user’s current viewpoint). In some embodiments, input gestures are performed indirectly on a user interface object in accordance with the location of the user’s hand while the user performs the input gesture not being at the location corresponding to the location of the user interface object in the three-dimensional environment upon detecting the user’s attention (e.g., gaze) to the user interface object. For example, for direct input gestures, the user is enabled to direct the user’s input to a user interface object by initiating a gesture at or near a location corresponding to the displayed location of the user interface object (e.g., within 0.5 cm, 1 cm, 5 cm, or a distance between 0 and 5 cm, measured from the outer edge of the option or the center portion of the option). For indirect input gestures, the user is enabled to direct the user’s input to a user interface object by focusing on the user interface object (e.g., by gazing at the user interface object), and while focusing on the option, the user initiates an input gesture (e.g., at any location detectable by the computer system) (e.g., at a location that does not correspond to the displayed location of the user interface object).
[0182] In some embodiments, input gestures (e.g., air gestures) used in various examples and embodiments described herein include pinch inputs and tap inputs for interacting with virtual or mixed reality environments in accordance with some embodiments. For example, the pinch inputs and tap inputs described below are performed as air gestures.
[0183] In some embodiments, the pinch input is part of an air gesture that includes one or more of: a pinch gesture, a long pinch gesture, a pinch-and-drag gesture, or a double pinch gesture. For example, a pinch gesture as an air gesture includes movement of two or more fingers of a hand to contact each other, that is, optionally followed by an immediate (e.g., within 0-1 seconds) break in contact with each other. A long pinch gesture as an air gesture includes movement of two or more fingers of a hand to contact each other for at least a threshold amount of time (e.g., at least 1 second) before a break in contact with each other is detected. For example, a long pinch gesture includes a user holding a pinch gesture (e.g., with two or more fingers in contact), and the long pinch gesture continues until a break in contact between the two or more fingers is detected. In some embodiments, a double pinch gesture as an air gesture includes two (e.g., or more) pinch inputs (e.g., performed by the same hand) that are detected consecutively immediately (e.g., within a predefined time period). For example, a user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., breaks contact between the two or more fingers), and performs a second pinch input within a predefined time period (e.g., within 1 second or within 2 seconds) after releasing the first pinch input.
[0184] In some embodiments, a pinch-and-drag gesture performed as an in-air gesture includes a pinch gesture (e.g., a pinch gesture or a long pinch gesture) performed in conjunction with (e.g., following) a drag input that changes a position of a user's hand from a first position (e.g., a start position of the drag) to a second position (e.g., an end position of the drag). In some embodiments, the user maintains the pinch gesture while performing the drag input, and releases the pinch gesture (e.g., opens their two or more fingers) to end the drag gesture (e.g., at the second position). In some embodiments, the pinch input and the drag input are performed by the same hand (e.g., the user pinches two or more fingers to contact each other and moves the same hand in the air to a second position with a drag gesture). In some embodiments, the pinch input is performed by a first hand of the user, and the drag input is performed by a second hand of the user (e.g., the second hand of the user moves in the air from the first position to the second position while the user continues the pinch input with the first hand of the user). In some embodiments, an input gesture performed as an in-air gesture includes an input (e.g., a pinch input and / or a tap input) performed using both hands of the user. For example, the input gesture includes two (e.g., or more) pinch inputs performed in conjunction with each other (e.g., simultaneously or within a predefined time period). For example, a first pinch gesture (e.g., a pinch input, a long pinch input, or a pinch-and-drag input) is performed using a first hand of the user, and in conjunction with the pinch input performed using the first hand, a second pinch input is performed using another hand (e.g., a second hand of the user). In some embodiments, movement between the two hands of the user (e.g., increasing and / or decreasing a distance or a relative orientation between the two hands of the user).
[0185] In some embodiments, a tap input performed as an in-air gesture (e.g., pointing to a user interface element) includes movement of a finger of the user toward the user interface element, movement of a hand of the user toward the user interface element (optionally, a finger of the user extending toward the user interface element), a downward motion of a finger of the user (e.g., mimicking a mouse click motion or a tap on a touch screen), or other predefined movement of a hand of the user. In some embodiments, a tap input performed as an in-air gesture is detected based on movement characteristics of a finger or a hand performing a tap gesture movement of the finger or the hand away from a viewpoint of the user and / or toward an object targeted as the tap input, followed by an end of the movement. In some embodiments, the end of the movement is detected based on a change in movement characteristics of the finger or the hand performing the tap gesture (e.g., an end of movement away from a viewpoint of the user and / or toward an object targeted as the tap input, a reversal of a direction of movement of the finger or the hand, and / or a reversal of an acceleration direction of movement of the finger or the hand).
[0186] In some embodiments, determining that the attention of the user is directed to the portion of the three-dimensional environment based on detection of a gaze directed to the portion of the three-dimensional environment (optionally, without requiring other conditions). In some embodiments, determining that the attention of the user is directed to the portion of the three-dimensional environment based on detection of a gaze directed to the portion of the three-dimensional environment with one or more additional conditions, such as requiring the gaze to be directed to the portion of the three- dimensional environment for at least a threshold duration (e.g., dwell duration) and / or requiring the gaze to be directed to the portion of the three-dimensional environment while the user’s viewpoint is within a distance threshold from the portion of the three- dimensional environment for the device to determine that the attention of the user is directed to the portion of the three-dimensional environment, where if one of these additional conditions is not met, the device determines that the attention is not directed to the portion of the three-dimensional environment that the gaze is directed to (e.g., until the one or more additional conditions are met).
[0187] In some embodiments, detection of the readiness state configuration of the user or portion of the user is detected by the computer system. Detection of the readiness state configuration of the hand is used by the computer system as an indication that the user can be preparing to use one or more mid-air hand gesture inputs performed by the hand (e.g., pinch, tap, pinch-and-drag, double pinch, long pinch, or other mid-air hand gestures described herein) to interact with the computer system. For example, the readiness state of the hand is determined based on whether the hand has a predetermined hand shape (e.g., a pre-pinch shape in which the thumb and one or more fingers are extended and spaced apart in preparation to make a pinch or grasp gesture, or a pre-tap in which one or more fingers are extended and the back of the hand is facing the user), based on whether the hand is in a predetermined positioning relative to the viewpoint of the user (e.g., below the head of the user and above the waist of the user and extending at least 15 cm, 20 cm, 25 cm, 30 cm, or 50 cm from the body), and / or based on whether the hand has moved in a particular manner (e.g., moving towards an area in front of the user above the waist of the user and below the head of the user or moving away from the body or legs of the user). In some embodiments, the readiness state is used to determine whether an interactive element of a user interface is responsive to attention (e.g., gaze) inputs.
[0188] In scenarios where inputs are described with reference to in-air gestures, it will be understood that similar gestures can be detected using a hardware input device attached to or held by one or more hands of a user, where the position of the hardware input device in space can be tracked using optical tracking, one or more accelerometers, one or more gyroscopes, one or more magnetometers, and / or one or more inertial measurement units, and the position and / or movement of the hardware input device is used in place of the position and / or movement of the one or more hands in the corresponding in-air gesture. In scenarios where inputs are described with reference to in-air poses, it will be understood that similar poses can be detected using a hardware input device attached to or held by one or more hands of a user. User inputs can be detected with controls contained in the hardware input device, such as one or more touch-sensitive input elements, one or more pressure-sensitive input elements, one or more buttons, one or more knobs, one or more dials, one or more joysticks, one or more hands or finger covers that can detect the position or change in position of the hands and / or fingers relative to each other, relative to the user’s body, and / or relative to the user’s physical environment, and / or other hardware input device controls, where user inputs made with controls contained in the hardware input device are used in place of hand and / or finger gestures such as in-air taps or in-air pinches in the corresponding in-air gesture. For example, a selection input described as being performed with an in-air tap or in-air pinch can alternatively be detected with a button press, a tap on a touch-sensitive surface, a press on a pressure-sensitive surface, or other hardware input. As another example, a movement input described as being performed with an in-air pinch and drag can alternatively be detected based on interaction with a hardware input control such as a button press and hold, a touch on a touch-sensitive surface, a press on a pressure-sensitive surface, or other hardware input following movement of the hardware input device (e.g., along with the hand associated with the hardware input device) through space. Similarly, two-handed inputs that include movement of the hands relative to each other can be performed with one in-air gesture and one hardware input device held in a hand that is not performing the in-air gesture, two hardware input devices held in different hands, or two in-air gestures performed by different hands and / or various combinations of inputs detected by one or more of the above hardware input devices.
[0189] In some embodiments, the software can be downloaded to the controller 110 in electronic form, over a network, for example, or it can alternatively be provided on a tangible non-transitory medium, such as optical, magnetic or electronic memory media. In some embodiments, the database 408 is likewise stored in memory associated with the controller 110. Alternatively or additionally, some or all of the described functionality of the computer can be implemented in dedicated hardware, such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP), for example. Although the described functionality is implemented in one or more computers in the example of FIG. 4, it will be appreciated that the functionality can be implemented in other ways, such as in one or more special-purpose hardware components, for example. Figure 4The controller 110 is shown in FIG. 4, but by way of example, as a separate unit from the image sensor 404, some or all of the processing functions of the controller can be performed by a suitable microprocessor and software or by dedicated circuitry within the housing of the image sensor 404 (e.g., a hand tracking device) or other device associated with the image sensor 404. In some embodiments, at least some of these processing functions can be performed by a suitable processor integrated with the display generation component 120 (e.g., in a television receiver, a hand-held device, or a head-mounted device) or with any other suitable computerized device, such as a game console or a media player. The sensing functions of the image sensor 404 likewise can be integrated into a computer or other computerized apparatus that will be controlled by the sensor output.
[0190] Figure 4 Also included is a schematic of a depth map 410 captured by the image sensor 404, in accordance with some embodiments. As described above, the depth map includes a matrix of pixels with corresponding depth values. Pixels 412 corresponding to the hand 406 have been segmented from the background and wrist in this map. The intensity of each pixel within the depth map 410 is inversely proportional to its depth value (i.e., the measured z-distance from the image sensor 404), with the gray shading becoming darker as the depth increases. The controller 110 processes these depth values in order to identify and segment the constituent parts of the image (i.e., groups of adjacent pixels) that have the characteristics of a human hand. These characteristics can include, for example, overall size, shape, and motion from frame to frame in a sequence of depth maps.
[0191] Figure 4 Also schematically illustrated is a hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406, in accordance with some embodiments. In Figure 4 In FIG. 4, the hand skeleton 414 is superimposed on a hand background 416 that has been segmented from the original depth map. In some embodiments, key feature points of the hand, and optionally on a wrist or arm connected to the hand (e.g., points corresponding to finger joints, finger tips, center of palm, end of the hand connected to the wrist, etc.) are identified and located on the hand skeleton 414. In some embodiments, the controller 110 uses the locations and movements of these key feature points across multiple image frames to determine a gesture performed by the hand or a current state of the hand, in accordance with some embodiments.
[0192] Figure 5 An example embodiment of the eye tracking device 130 ( Figure 1A ) is illustrated. In some embodiments, the eye tracking device 130 is implemented by an eye tracking unit 243 ( Figure 2) control to track the position and movement of the user’s gaze relative to the scene 105 or relative to XR content displayed via the display generation component 120. In some embodiments, the eye tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, when the display generation component 120 is a head-mounted device (such as headphones, a helmet, eyewear, or glasses) or a handheld device that is placed in a wearable frame, the head-mounted device includes both components that generate XR content for the user to view and components to track the user’s gaze relative to the XR content. In some embodiments, the eye tracking device 130 is separate from the display generation component 120. For example, when the display generation component is a handheld device or an XR room, the eye tracking device 130 is optionally a device that is separate from the handheld device or the XR room. In some embodiments, the eye tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, the head-mounted eye tracking device 130 is used in combination with a display generation component that is also head-mounted or a display generation component that is not head-mounted. In some embodiments, the eye tracking device 130 is not a head-mounted device and is used in combination with a head-mounted display generation component. In some embodiments, the eye tracking device 130 is not a head-mounted device and is part of a non-head-mounted display generation component.
[0193] In some embodiments, the display generation component 120 uses display mechanisms (e.g., left and right near-eye display panels) to display frames including left and right images in front of a user’s eyes to provide a 3D virtual view to the user. For example, a head-mounted display generation component can include left and right optical lenses (referred to herein as eye lenses) positioned between the display and the user’s eyes. In some embodiments, the display generation component can include or be coupled to one or more external cameras that capture video of the user’s environment for display. In some embodiments, a head-mounted display generation component can have a transparent or semi-transparent display and display virtual objects on the transparent or semi-transparent display through which the user can directly view the physical environment. In some embodiments, the display generation component projects virtual objects into the physical environment. The virtual objects may, for example, be projected on a physical surface or as a hologram so that an individual using the system observes the virtual objects superimposed over the physical environment. In this case, separate display panels and image frames for the left and right eyes can not be needed.
[0194] As Figure 5As shown in FIG. 1, in some embodiments, eye tracking device 130 (e.g., a gaze tracking device) includes at least one eye tracking camera (e.g., an infrared (IR) or near-IR (NIR) camera), and an illumination source (e.g., an IR or NIR light source, such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user’s eyes. The eye tracking camera can be pointed at the user’s eyes to receive IR or NIR light that is directly reflected from the eyes by the light source, or alternatively can be pointed at “hot” mirrors that are positioned between the user’s eyes and the display panel, which reflect IR or NIR light from the eyes to the eye tracking camera while allowing visible light to pass through. Eye tracking device 130 optionally captures images of the user’s eyes (e.g., as a video stream captured at 60-120 frames per second (fps)), analyzes the images to generate gaze tracking information, and communicates the gaze tracking information to controller 110. In some embodiments, both of the user’s eyes are tracked separately by respective eye tracking cameras and illumination sources. In some embodiments, only one of the user’s eyes is tracked by a respective eye tracking camera and illumination source.
[0195] In some embodiments, eye tracking device 130 is calibrated using a device-specific calibration process to determine parameters of the eye tracking device for a particular operating environment 100, such as 3D geometry and parameters of the LEDs, camera, hot mirrors (if present), eye lenses, and display screen. The device-specific calibration process can be performed at a factory or another facility prior to delivery of the AR / VR equipment to an end user. The device-specific calibration process can be an automatic calibration process or a manual calibration process. According to some embodiments, a user-specific calibration process can include estimation of eye parameters for a particular user, such as pupil position, fovea position, optical axis, visual axis, interpupillary distance, etc. According to some embodiments, once the device-specific parameters and user-specific parameters are determined for eye tracking device 130, a glint-assisted method can be used to process images captured by the eye tracking camera to determine a current visual axis and a gaze point of the user relative to the display.
[0196] As Figure 5As shown in FIG. 5A, eye tracking device 130 (e.g., 130A or 130B) includes an eye lens 520 and a gaze tracking system that includes at least one eye tracking camera 540 (e.g., an infrared (IR) or near-IR (NIR) camera) positioned on the side of the user’s face on which eye tracking is performed, and an illumination source 530 (e.g., an IR or NIR light source such as an array or ring of NIR light-emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user’s eye 592. Eye tracking camera 540 can be directed at a mirror 550 (which mirrors IR or NIR light from eye 592 while allowing visible light to pass) located between user’s eye 592 and display 510 (e.g., a left display panel or a right display panel of a head-mounted display, or a display of a handheld device, a projector, etc.) (e.g., as shown in the top portion of FIG. 5A), or alternatively can be directed at user’s eye 592 to receive reflected IR or NIR light from eye 592 (e.g., as shown in the bottom portion of FIG. 5A). Figure 5 Figure 5
[0197] In some embodiments, controller 110 renders AR or VR frames 562 (e.g., left and right frames for left and right display panels) and provides frames 562 to display 510. Controller 110 uses gaze tracking input 542 from eye tracking camera 540 for various purposes, e.g., for processing frames 562 for display. Controller 110 optionally estimates a gaze point of the user on display 510 based on gaze tracking input 542 acquired from eye tracking camera 540 using a glint-assisted method or other suitable method. The gaze point estimated from gaze tracking input 542 is optionally used to determine a direction in which the user is currently looking.
[0198] Several possible use cases of the user’s current gaze direction are described below and are not intended to be limiting. As an example use case, the controller 110 can render virtual content differently based on the determined direction of the user’s gaze. For example, the controller 110 can generate virtual content in a foveal region determined from the user’s current gaze direction with a higher resolution than in a peripheral region. As another example, the controller can position or move virtual content in a view based at least in part on the user’s current gaze direction. As another example, the controller can display particular virtual content in a view based at least in part on the user’s current gaze direction. As another example use case in an AR application, the controller 110 can direct an external camera used to capture the physical environment of the XR experience to focus in the determined direction. The autofocus mechanism of the external camera can then focus on an object or surface in the environment that the user is currently looking at on the display 510. As another example use case, the eye lenses 520 can be focusable lenses, and the controller uses gaze tracking information to adjust the focal point of the eye lenses 520 so that the virtual object the user is currently looking at has the proper vergence to match the convergence of the user’s eyes 592. The controller 110 can utilize gaze tracking information to direct the eye lenses 520 to adjust the focal point so that objects the user is looking at that are close appear at the correct distance.
[0199] In some embodiments, the eye tracking device is part of a head-mounted device that includes a display (e.g., display 510) mounted in a wearable housing, two eye lenses (e.g., eye lenses 520), an eye tracking camera (e.g., eye tracking camera 540), and a light source (e.g., illumination source 530 (e.g., IR or NIR LED)). The light source emits light (e.g., IR or NIR light) toward the user’s eyes 592. In some embodiments, the light source can be arranged in a ring or circle around each of the lenses, as shown in FIG. 6B. In some embodiments, as an example, eight illumination sources 530 (e.g., LEDs) are arranged around each lens 520. However, more or fewer illumination sources 530 can be used, and other arrangements and positions of the illumination sources 530 can be used. Figure 5
[0200] In some embodiments, the display 510 emits light in the visible light range and does not emit light in the IR or NIR range, and therefore does not introduce noise into the gaze tracking system. It should be noted that the position and angle of the eye-tracking camera 540 are given by way of example and are not intended to be limiting. In some embodiments, a single eye-tracking camera 540 is located on each side of the user's face. In some embodiments, two or more NIR cameras 540 may be used on each side of the user's face. In some embodiments, cameras 540 with a wider field of view (FOV) and cameras 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, cameras 540 operating at one wavelength (e.g., 850 nm) and cameras 540 operating at different wavelengths (e.g., 940 nm) may be used on each side of the user's face.
[0201] like Figure 5 The illustrated gaze tracking system implementation can be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide users with computer-generated reality, virtual reality, augmented reality, and / or augmented virtual experiences.
[0202] Figure 6 Examples of flash-assisted gaze tracking pipelines according to some embodiments are illustrated. In some embodiments, the gaze tracking pipeline uses a flash-assisted gaze tracking system (e.g., such as...) Figure 1A and Figure 5 The illustrated eye-tracking device 130) is used to implement this. The flash-assisted gaze tracking system can maintain a tracking state. Initially, the tracking state is off or "no". When in tracking state, the flash-assisted gaze tracking system uses previous information from previous frames when analyzing the current frame to track the pupil outline and flash in the current frame. When not in tracking state, the flash-assisted gaze tracking system attempts to detect the pupil and flash in the current frame, and if successful, initializes the tracking state to "yes" and continues to the next frame in tracking state.
[0203] like Figure 6 As shown, the gaze-tracking camera captures left and right images of the user's left and right eyes. The captured images are then fed into a gaze-tracking pipeline for processing to begin at 610. As indicated by the arrow returning to element 600, the gaze-tracking system can continue capturing images of the user's eyes, for example, at a rate of 60 to 120 frames per second. In some embodiments, each set of captured images can be fed into the pipeline for processing. However, in some embodiments or under certain conditions, not all captured frames are processed by the pipeline.
[0204] At 610, for the current captured image, if the tracking status is yes, the method proceeds to element 640. At 610, if the tracking status is no, the image is analyzed to detect the user’s pupils and glints in the image as indicated at 620. At 630, if the pupils and glints are successfully detected, the method proceeds to element 640. Otherwise, the method returns to element 610 to process the next image of the user’s eye.
[0205] At 640, if proceeding from element 610, the current frame is analyzed to track the pupils and glints based in part on previous information from previous frames. At 640, if proceeding from element 630, the tracking status is initialized based on the pupils and glints detected in the current frame. The results of the processing at element 640 are checked to verify that the results of the tracking or detection can be trusted. For example, the results can be checked to determine whether the pupils and a sufficient number of glints for performing gaze estimation were successfully tracked or detected in the current frame. At 650, if the results can not be trusted, at element 660, the tracking status is set to no and the method returns to element 610 to process the next image of the user’s eye. At 650, if the results are trusted, the method proceeds to element 670. At 670, the tracking status is set to yes (if it is not already) and the pupil and glint information is passed to element 680 to estimate the user’s gaze point.
[0206] Figure 6 It is intended to serve as one example of an eye tracking technique that can be used in particular implementations. As will be recognized by one of ordinary skill in the art, in accordance with various implementations, other eye tracking techniques that are currently existing or are developed in the future can be used in place of or in combination with the glint-assisted eye tracking technique described herein in computer systems 101 for providing XR experiences to users.
[0207] In some implementations, portions of the captured real-world environment 602 are used to provide an XR experience to the user, such as a mixed reality environment in which one or more virtual objects are overlaid over a representation of the real-world environment 602.
[0208] Accordingly, the description herein describes some embodiments of a three-dimensional environment (e.g., an XR environment) that includes representations of real-world objects and representations of virtual objects. For example, the three-dimensional environment optionally includes a representation of a table that is present in a physical environment, which is captured and displayed in the three-dimensional environment (e.g., actively displayed via a camera and display of the computer system or passively displayed via a transparent or semi-transparent display of the computer system). As previously described, the three-dimensional environment is optionally a mixed reality system in which the three-dimensional environment is based on a physical environment that is captured by one or more sensors of the computer system and displayed via the display generation component. As a mixed reality system, the computer system is optionally able to selectively display portions and / or objects of the physical environment such that the respective portions and / or objects of the physical environment appear as if they are present in the three-dimensional environment displayed by the computer system. Similarly, the computer system is optionally able to display virtual objects in the three-dimensional environment at respective locations that have corresponding locations in the real world (e.g., the physical environment) by placing the virtual objects in the three-dimensional environment at the respective locations to appear as if the virtual objects are present in the real world. For example, the computer system optionally displays a vase such that the vase appears as if a real vase is placed on top of a table in the physical environment. In some embodiments, the respective locations in the three-dimensional environment have corresponding locations in the physical environment. Accordingly, when the computer system is described as displaying a virtual object at a respective location relative to a physical object (e.g., a location such as at or near a user’s hand or a location at or near a physical table), the computer system displays the virtual object at a particular location in the three-dimensional environment such that it appears as if the virtual object is at or near the physical object in the physical environment (e.g., the virtual object is displayed in the three-dimensional environment at a location that corresponds to a location in the physical environment where the virtual object would be displayed if the virtual object were a real object at the particular location).
[0209] In some embodiments, real-world objects that are present in a physical environment that are displayed in the three-dimensional environment (e.g., and / or are visible via the display generation component) can interact with virtual objects that are only present in the three-dimensional environment. For example, the three-dimensional environment can include a table and a vase placed on top of the table, where the table is a view (or representation) of a physical table in a physical environment, and the vase is a virtual object.
[0210] In three-dimensional environments (e.g., real environments, virtual environments, or environments that include a mix of real and virtual objects), objects are sometimes referred to as having a depth or simulated depth, or objects are referred to as being visible, displayed, or placed at different depths. In this context, depth refers to a dimension that is different from height or width. In some embodiments, depth is defined relative to a fixed set of coordinates (e.g., where a room or object has a height, depth, and width defined relative to a fixed set of coordinates). In some embodiments, depth is defined relative to a user’s position or viewpoint, in which case the depth dimension varies based on the user’s position and / or the position and angle of the user’s viewpoint. In some embodiments where depth is defined relative to a user’s position that is located relative to a surface of the environment (e.g., a surface of the floor or ground of the environment), objects that are farther away from the user along a line that extends parallel to the surface are considered to have greater depth in the environment, and / or the depth of an object is measured along an axis that extends outward from the user’s position and parallel to the surface of the environment (e.g., depth is defined in a cylindrical or substantially cylindrical coordinate system, where the user’s position is at the center of the cylinder that extends from the user’s head toward the user’s feet). In some embodiments where depth is defined relative to a user’s viewpoint (e.g., relative to a direction of a point in space that determines which portion of the environment is visible via a head-mounted device or other display), objects that are farther away from the user’s viewpoint along a line that extends parallel to the direction of the user’s viewpoint are considered to have greater depth in the environment, and / or the depth of an object is measured along an axis that extends outward from the user’s viewpoint and parallel to the direction of the user’s viewpoint (e.g., depth is defined in a spherical or substantially spherical coordinate system, where the origin of the viewpoint is at the center of the sphere that extends outward from the user’s head). In some embodiments, depth is defined relative to a user interface container (e.g., a window or application in which application and / or system content is displayed), where the user interface container has a height and / or width, and the depth is a dimension that is orthogonal to the height and / or width of the user interface container. In some embodiments where depth is defined relative to a user interface container, when the container is placed in a three-dimensional environment or initially displayed (e.g., such that the depth dimension of the container extends outward away from the user or the user’s viewpoint), the height and / or width of the container is typically orthogonal or substantially orthogonal to a straight line that extends from the user’s position (e.g., the user’s viewpoint or the user’s position) to the user interface container (e.g., the center of the user interface container or another characteristic point of the user interface container). In some embodiments where depth is defined relative to a user interface container, the depth of an object relative to the user interface container refers to the position of the object along the depth dimension of the user interface container. In some embodiments, multiple different containers can have different depth dimensions (e.g., different depth dimensions that extend in different directions and / or from different starting points away from the user or the user’s viewpoint).In some embodiments, when defining depth relative to a user interface container, the direction of the depth dimension remains constant for the user interface container as the location of the user interface container, the user, and / or the user’s point of view changes (e.g., or when multiple different viewers are viewing the same container in a three-dimensional environment, such as during a physical collaboration session and / or when multiple participants are in a live communication session with shared virtual content including the container). In some embodiments, for curved containers (e.g., containers that include a curved surface or a curved content area), the depth dimension optionally extends into the surface of the curved container. In some cases, a z-separation (e.g., the separation of two objects in the depth dimension), a z-height (e.g., the distance of one object from another object in the depth dimension), a z-position (e.g., the positioning of one object in the depth dimension), a z-depth (e.g., the positioning of one object in the depth dimension), or a simulated z-dimension (e.g., a depth used as a dimension of an object, a dimension of an environment, a direction in space, and / or a simulated direction in space) is used to refer to the concept of depth as described above.
[0211] In some embodiments, the user optionally is able to interact with virtual objects in the three-dimensional environment using one or both hands as if the virtual objects were real objects in the physical environment. For example, as described above, one or more sensors of the computer system optionally capture the user’s hand(s) and display a representation of the user’s hand(s) in the three-dimensional environment (e.g., in a manner similar to displaying real-world objects in the three-dimensional environment as described above), or in some embodiments, the user’s hand(s) are viewable via the display generation component via the ability to see the physical environment through the user interface due to the transparency / translucency of the portion of the user interface being displayed by the display generation component, or due to the projection of the user interface onto a transparent / translucent surface or onto or into the field of view of the user’s eyes. Thus, in some embodiments, the user’s hands are displayed at their respective locations in the three-dimensional environment and are treated as if they were objects in the three-dimensional environment that are able to interact with virtual objects in the three-dimensional environment as if the virtual objects were physical objects in the physical environment. In some embodiments, the computer system is able to update the display of the representation of the user’s hands in the three-dimensional environment in conjunction with the movement of the user’s hands in the physical environment.
[0212] In some of the embodiments described below, the computer system is optionally able to determine an "effective" distance between a physical object in the physical world and a virtual object in the three-dimensional environment, e.g., for determining whether a physical object is directly interacting with a virtual object (e.g., whether a hand is touching, grasping, holding, etc. a virtual object or within a threshold distance of a virtual object). For example, a hand directly interacting with a virtual object optionally includes one or more of: a finger of the hand pressing a virtual button, a hand of the user grasping a virtual vase, two fingers of the hand of the user coming together and pinching / holding an application's user interface, and any of the other types of interactions described herein. For example, when determining whether a user is interacting with a virtual object and / or how a user is interacting with a virtual object, the computer system optionally determines a distance between a hand of the user and the virtual object. In some embodiments, the computer system determines a distance between a hand of the user and a virtual object by determining a distance between a location of the hand in the three-dimensional environment and a location of the virtual object of interest in the three-dimensional environment. For example, the hand or hands of the user are at a particular location in the physical world, the computer system optionally captures the hand or hands and displays the hand or hands at a particular corresponding location in the three-dimensional environment (e.g., if the hand is a virtual hand rather than a physical hand, the hand will be displayed at a location in the three-dimensional environment). The location of the hand in the three-dimensional environment is optionally compared to the location of the virtual object of interest in the three-dimensional environment to determine a distance between the hand or hands of the user and the virtual object. In some embodiments, the computer system optionally determines a distance between a physical object and a virtual object by comparing locations in the physical world (e.g., rather than comparing locations in the three-dimensional environment). For example, when determining a distance between a hand or hands of the user and a virtual object, the computer system optionally determines a corresponding location of the virtual object in the physical world (e.g., if the virtual object is a physical object rather than a virtual object, the location at which the virtual object would be located in the physical world), and then determines a distance between the corresponding physical location and the hand or hands of the user. In some embodiments, the same techniques are optionally used to determine a distance between any physical object and any virtual object. Thus, as described herein, when determining whether a physical object is in contact with a virtual object or whether a physical object is within a threshold distance of a virtual object, the computer system optionally performs any of the techniques described above to map a location of the physical object to the three-dimensional environment and / or to map a location of the virtual object to the physical environment.
[0213] In some embodiments, the same or similar techniques are used to determine where and what a user’s gaze is directed at, and / or where and what a physical stylus held by the user is directed at. For example, if a user’s gaze is directed at a particular location in the physical environment, the computer system optionally determines a corresponding location in the three-dimensional environment (e.g., a virtual location of the gaze), and if a virtual object is located at that corresponding virtual location, the computer system optionally determines that the user’s gaze is directed at that virtual object. Similarly, the computer system optionally is able to determine a direction in which a physical stylus is directed in the physical environment based on the orientation of the stylus. In some embodiments, based on the determination, the computer system determines a corresponding virtual location in the three-dimensional environment that corresponds to the location in the physical environment at which the stylus is directed, and optionally determines that the stylus is directed at the corresponding virtual location in the three-dimensional environment.
[0214] Similarly, the embodiments described herein can refer to a location of a user (e.g., a user of a computer system) in a three-dimensional environment and / or a location of a computer system in a three-dimensional environment. In some embodiments, a user of a computer system is holding, wearing, or otherwise located at or near the computer system. Thus, in some embodiments, the location of the computer system is used as a proxy for the location of the user. In some embodiments, the location of the computer system and / or the user in the physical environment corresponds to a respective location in the three-dimensional environment. For example, the location of the computer system would be a location in the physical environment (and its corresponding location in the three-dimensional environment) from which the user would see the physical environment in the same location, orientation, and / or size (e.g., in absolute terms and / or relative to each other) as the objects in the physical environment if the user were standing in that location facing the respective portion of the physical environment visible via the display generation component. Similarly, if the virtual objects displayed in the three-dimensional environment are physical objects in the physical environment (e.g., physical objects placed in the physical environment at the same locations as the virtual objects in the three-dimensional environment, and physical objects in the physical environment having the same size and orientation as when in the three-dimensional environment), the location of the computer system and / or the user is a location from which the user would see the virtual objects in the physical environment in the same location, orientation, and / or size (e.g., in absolute terms and / or relative to each other and real-world objects) as the virtual objects displayed in the three-dimensional environment by the display generation component of the computer system.
[0215] In this disclosure, various input methods are described with respect to interactions with a computer system. When one input device or input method is used to provide an example, and another input device or input method is used to provide another example, it will be understood that each example can be compatible with and optionally utilize the input device or input method described with respect to the other example. Similarly, various output methods are described with respect to interactions with a computer system. When one output device or output method is used to provide an example, and another output device or output method is used to provide another example, it will be understood that each example can be compatible with and optionally utilize the output device or output method described with respect to the other example. Similarly, various methods are described with respect to interactions with a virtual environment or a mixed reality environment by a computer system. When an interaction with a virtual environment is used to provide an example, and a mixed reality environment is used to provide another example, it will be understood that each example can be compatible with and optionally utilize the methods described with respect to the other example. Thus, this disclosure discloses embodiments that are combinations of features of multiple examples without the need to exhaustively list all features of an embodiment in the description of each example embodiment.
[0216] User interface and related processes
[0217] Attention is now directed to embodiments of user interfaces (“UIs”) and associated processes that can be implemented on a computer system, such as a portable multifunctional device or a head-mounted device, that is in communication with a display generation component and, optionally, one or more sensors.
[0218] Examples described herein illustrate ways in which a user of a computer system (e.g., device 700) can initiate and / or modify a live communication session in which the user communicates with one or more users of other respective computer systems. In some embodiments, the live communication session is an audio communication session (e.g., a voice call or a telephone call). In some embodiments, the live communication session is a video communication session (e.g., a video call and / or a video conference). In some embodiments, the live communication session is an XR communication session, such as a spatial communication session or a non-spatial communication session. During a spatial communication session, the one or more users are respectively represented in an XR environment by three-dimensional (3D) representations (e.g., avatars) that correspond to the users. In some embodiments, the 3D representations have spatial agents such that the 3D representations can move within the XR environment relative to other elements and / or users in the XR environment. During a non-spatial communication session, the one or more users are respectively represented in an XR environment by two-dimensional (2D) representations that correspond to the users. In some embodiments, the 2D representations include video feeds of the users and optionally have fixed positions (e.g., locations) within the XR environment.
[0219] Figures 7A to 7Q An example of managing a live communication session is illustrated. Figure 8 is a flow diagram of an example method 800 for managing a live communication session. Figure 9 is a flow diagram of an example method 900 for providing an avatar in a live communication session. Figures 7A to 7Q The user interfaces in Figure 8 and / or Figure 9 are used to illustrate processes described below, which include
[0220] While Figures 7A to 7Q The device 700 is illustrated as a handheld device (e.g., a tablet, a smartphone, or a laptop) with a display 702, in some embodiments, the device 700 is a head-mounted device (HMD). The HMD is configured to be worn on the head of a user of the device 700 and includes the display 702 on and / or in an interior portion of the HMD. The display 702 is visible to the user when the device 700 is worn on the head of the user. For example, in some embodiments, the HMD at least partially covers the eyes of the user when worn on the head of the user such that the display 702 is positioned over and / or in front of the eyes of the user. In such embodiments, the display 702 is configured to display an XR environment during a live communication session in which the user of the HMD is participating.
[0221] In Figure 7A , the device 700 displays an XR environment 704 including elements (e.g., virtual elements and / or physical elements) such as a table 704a and a sofa 704b on the display 702. While displaying the XR environment 704, the device 700 receives a request to display a communication interface. In some embodiments, the request to display the communication interface is a press of a button 703 of the device 700. As Figure 7B shown, in response to receiving the request, the device 700 displays the communication interface 710. In some embodiments, the communication interface 710 is displayed within the XR environment 704.
[0222] In general, the communication interface 710 can be used to initiate and / or modify live communication sessions (e.g., audio communication sessions, video communication sessions, or XR communication sessions). The communication interface 710 includes pinned contacts 712 (e.g., pinned contacts 712a-712g) and recent contacts 714 (e.g., recent contacts 714a-714i). In some embodiments, the pinned contacts 712 are a set of contacts (e.g., contacts that are favorited or pinned by the user of the device 700) that are selected by the user of the device 700 to be included in the communication interface 710. In some embodiments, the recent contacts 714 are contacts that the user of the device 700 recently used the device 700 and, optionally, one or more other devices associated with the user of the device 700 to communicate with (e.g., via text, phone, and / or live communication sessions). In some embodiments, the recent contacts 714 are arranged (e.g., ordered or ranked) based on recency of communication between the recent contacts 714 and the user of the device 700.
[0223] In some embodiments, one or more of the pinned contacts 712 and / or the recent contacts 714 correspond to a defined group of contacts. As an example, the pinned contact 712d corresponds to the group of contacts “Surfers.” As another example, the recent contact 714d corresponds to the group of contacts “Lake Crew.”
[0224] In some embodiments, the pinned contacts 712 and / or the recent contacts 714 indicate a most recent communication between the user of the device 700 and various contacts. As an example, the pinned contact 712b (“John”) indicates that the contact last transmitted a text message 1 minute ago. Optionally, the communication interface 710 includes a preview 716b that indicates content of the text message transmitted by the pinned user 712b. As another example, the pinned user 712c indicates that the contact last transmitted a text reaction (e.g., a “heart” reaction) at 2:10. As yet another example, the recent contact 714a (“Mom”) indicates that the user of the device 700 most recently communicated with the contact 714a in an XR communication session (e.g., a spatial live communication session or a non-spatial live communication session) at 3:32. As yet another example, the recent contact 714e (“Uncle Bob”) indicates that the user of the device 700 most recently communicated with the contact 714e in an audio communication session (e.g., a phone call) at 9:41.
[0225] In some embodiments, the pinned contacts 712 and / or the recent contacts 714 indicate pending invitations to live communication sessions. As an example, the recent contact 714b (“Dad”) indicates that the user of the device 700 can join a live communication session with the recent contact 714b. As a further example, the recent contact 714d (“Lake Crew”) indicates that three members of the group are currently in an ongoing live communication session to which the user of the device 700 has been invited to join.
[0226] In some embodiments, the contacts 712 and 714 of the communication interface 710 can be used to manage contacts. By way of example, while displaying the communication interface 710, the device 700 detects selection of the contact 712e (“Jo”). In some embodiments, the selection of the contact 712e is a tap gesture 705b on the contact 712e. In some embodiments, the selection of the contact 712e is an air gesture that, for example, indicates selection of the contact 712e. As shown, in response to detecting the selection of the contact 712e, the device 700 displays a contact menu 720 associated with the contact 712e. Figure 7C1 and / or Figure 7C2 As shown, in response to detecting the selection of the contact 712e, the device 700 displays a contact menu 720 associated with the contact 712e.
[0227] The contact menu 720 includes an invite option 720a and an expand option 720b. The invite option 720a, when selected, causes the device 700 to invite the contact 712e to an XR communication session. The expand option 720b, when selected, causes the device 700 to display one or more additional options for managing the contact 712e. For example, while displaying the contact menu 720, the device 700 detects selection of the expand option 720b. In some embodiments, the selection of the expand option 720b is a tap gesture 705c on the expand option 720b. In some embodiments, the selection of the expand option 720b is an air gesture that, for example, indicates selection of the expand option 720b. As shown, in response to detecting the selection of the expand option 720b, the device 700 expands the contact menu 720 to display (e.g., replaces the display of the expand option 720b with) one or more additional options (e.g., options 720c-720f). Figure 7D
[0228] In some embodiments, the contacts menu 720, when expanded, includes an audio option 720c, a message option 720d, an info option 720e, and an edit option 720f. The audio option 720c, when selected, causes the device 700 to initiate an audio communication session with the contact 712e (e.g., without a live video component). In some embodiments, the device 700 is not capable of communicating over a cellular network and / or is configured to use an external device for audio calls. Thus, in some examples, the device 700 uses a nearby device (e.g., a mobile phone and / or a tablet) that is capable of communicating over a cellular network to initiate the audio communication session. The edit option 720f, when selected, allows the user of the device 700 to remove the contact 712e from the pinned contacts 712 (or add the contact 712e to the pinned contacts in embodiments in which the contact 712e is not already a pinned contact). The message option 720d, when selected, allows the user to transmit a message to the contact 712e. For example, while displaying the contacts menu 720, the device 700 detects a selection of the message option 720d. In some embodiments, the selection of the message option 720d is a tap gesture 705d on the message option 720d. In some embodiments, the selection of the message option 720d is an air gesture that indicates a selection of the message option 720d, for example. As shown in Figure 7E FIG. 7B, in response to detecting the selection of the message option 720d, the device 700 displays (e.g., replaces the display of the communication interface 710 with) a message interface 730. Thereafter, the message interface 730 is available for transmitting a message to the contact 712e.
[0229] Referring again to Figure 7D , the info option 720e, when selected, causes the device 700 to display information corresponding to the contact 712e (e.g., without displaying additional information corresponding to other contacts). For example, while displaying the contacts menu 720, the device 700 detects a selection of the info option 720e. In some embodiments, the selection of the info option 720e is a tap gesture 707d on the info option 720e. In some embodiments, the selection of the info option 720e is an air gesture that indicates a selection of the info option 720e, for example. As shown in Figure 7F FIG. 7C, in response to detecting the selection of the info option 720e, the device 700 displays an info interface 740. The info interface 740 includes various details corresponding to the contact 712e, including but not limited to a name and contact information.
[0230] In some embodiments, the device of the contact is not capable of participating in an XR communication session with device 700. Accordingly, in some embodiments, one or more options of the contact menu can be omitted, de-emphasized (e.g., grayed out or darkened), and / or replaced to accurately reflect the capabilities of the device of the contact. For example, again with reference to Figure 7B While displaying communication interface 710, device 700 detects selection of contact 712g (“Sam”). In some embodiments, the selection of contact 712g is a tap gesture 709b on contact 712g. In some embodiments, the selection of contact 712g is an air gesture that indicates selection of contact 712g, for example. As shown, in response to detecting selection of contact 712g, device 700 displays a contact menu 722 associated with contact 712g. Figure 7C1
[0231] Because, in some embodiments, the device of contact 712g is not capable of communicating in an XR communication session with device 700, menu 722 does not include an invite option (e.g., invite option 720a), and instead includes an audio option 722a. Audio option 722a, when selected, causes device 700 to initiate an audio communication session with contact 712g. Menu 722 also includes an expand option 722b that, when selected, causes device 700 to display one or more additional options for contact 712g.
[0232] In some embodiments, a contact menu associated with a contact includes one or more additional options based on a status of device 700. As an example, in some embodiments, in instances in which device 700 is participating in a live communication session (e.g., an XR communication session or an audio communication session), the contact menu includes an option to invite the contact to the live communication session. For example, with reference to Figure 7B While participating in an XR communication session and while displaying communication interface 710, device 700 detects selection of contact 714g (“Dylan”). In some embodiments, the selection of contact 714g is a tap gesture 711b on contact 714g. In some embodiments, the selection of contact 714g is an air gesture that indicates selection of contact 714f, for example. As shown, in response to detecting selection of contact 714g, device 700 displays a contact menu 724 associated with contact 714g. Figure 7C1
[0233] The contacts menu 724 includes an invite option 724a, an invite option 724b, and an expand option 724c. The invite option 724a, when selected (e.g., tap gesture 709c), causes the device 700 to invite the contact 714g to a new live communication session. The invite option 724b, when selected, causes the device 700 to invite the contact 714g to the live communication session in which the device 700 is currently participating. The expand option 720c, when selected, causes the device 700 to display one or more additional options for the contact 714g.
[0234] In some embodiments, inviting a contact to a new live communication session (e.g., in response to a selection of the option 724a) will cause the device 700 to disconnect from and / or terminate the live communication session in which the device 700 is currently participating. In some embodiments, prior to terminating the existing live communication session in this manner, the device 700 confirms that the user wishes to disconnect from the current live communication session prior to initiating a new live communication session. For example, as shown in Figure 7G response to a selection of the invite option 724a, the device 700 displays a confirmation interface 740 that includes a confirmation affordance 742. In response to a selection of the confirmation affordance 742, the device 700 terminates the current live communication session and invites the contact 714f to a new live communication session.
[0235] In some embodiments, the user optionally uses the communication interface 710 to transmit a message to a contact. For example, referring to Figure 7B While displaying the communication interface 710, the device 700 detects a selection of the preview 716b associated with the pinned contact 712b. In some embodiments, the selection of the preview 716b is a tap gesture 707b on the preview 716b. In some embodiments, the selection of the preview 716b is an air gesture that, for example, indicates a selection of the preview 716b. As shown in Figure 7C1 In response to detecting the selection of the preview 716b, the device 700 expands the preview 716b to display a reply option 718.
[0236] The reply option 718, when selected, causes the device 700 to display a reply interface for transmitting a message to the contact 712b. For example, while displaying the reply option 718 in the preview 716b, the device 700 detects a selection of the reply option 718. In some embodiments, the selection of the reply option 718 is a tap gesture 707c on the reply option 718. In some embodiments, the selection of the reply option 718 is an air gesture that, for example, indicates a selection of the reply option 718. As shown in Figure 7H In response to detecting the selection of the reply option 718, the device 700 displays a reply interface 750 that is usable to transmit a message to the contact 712b.
[0237] In some implementation schemes, Figure 7C1 The technologies and user interfaces described in the text are by Figures 1A to 1P One or more of the devices described herein are provided. Figure 7C2 Examples are given (for example, such as...) Figure 7B and Figure 7C1 The communication interface X710 described herein is displayed on the display module X702 of a head-mounted device (HMD) X700. In some embodiments, the device X700 includes a pair of display modules that provide stereoscopic content to different eyes of the same user. For example, the HMD X700 includes display module X702 (which provides content to the user's left eye) and a second display module (which provides content to the user's right eye). In some embodiments, the second display module displays an image that is slightly different from that of display module X702 to generate the illusion of stereoscopic depth.
[0238] like Figure 7C2 As shown, in response to detecting a selection of contact X712e, HMD X700 displays a contact menu X720 associated with contact X712e. In some embodiments, HMD X700 detects the selection of contact X712e based on an air gesture performed by the user of HMD X700. In some embodiments, HMD X700 detects the user's hand X750a and / or X750b and determines whether the movement of hand X750a and / or X750b performs a predetermined air gesture corresponding to the selection of contact X712e. In some embodiments, the predetermined air gesture for selecting contact X712e includes a pinch gesture. In some embodiments, a pinch gesture includes detecting the movement of fingers X750c and thumb X750d toward each other. In some embodiments, HMD X700 detects the selection of contact X712e based on gaze and air gesture input performed by the user of HMD X700. In some implementations, gaze and air gesture input includes detecting that the user of the HMD X700 is looking at the contact X712e (e.g., for a duration greater than a predetermined amount of time) and that the user's hands X750a and / or X750b of the HMD X700 are performing a pinch gesture.
[0239] The contacts menu X720 includes an invite option X720a and an expand option X720b. The invite option X720a, when selected (e.g., via an aerial gesture such as a pinch gesture and / or via a gaze-and-pinch gesture), causes the HMD X700 to invite the contact X712e to the XR communication session. The expand option X720b, when selected, causes the HMD X700 to display one or more additional options for managing the contact X712e. For example, while displaying the contacts menu X720, the HMD X700 detects a selection of the expand option X720b. In some embodiments, the selection of the expand option X720b is, for example, an aerial gesture (e.g., a pinch gesture, and / or a gaze-and-pinch gesture) that indicates selection of the expand option X720b. In response to detecting the selection of the expand option X720b, the HMD X700 expands the contacts menu X720 to display (e.g., replaces the display of the expand option X720b with) the one or more additional options (e.g., options 720c-720f, as Figure 7D illustrated).
[0240] In some embodiments, when expanded, the contacts menu X720 includes an audio option (e.g., 720c), a message option (e.g., 720d), an information option (e.g., 720e), and an edit option (e.g., 720f), for example, as described with respect to Figure 7D FIG. 6. The audio option, when selected, causes the HMD X700 to initiate an audio communication session (e.g., without a live video component) with the contact X712e. In some embodiments, the HMD X700 is not capable of communicating over a cellular network and / or is configured to use an external device for audio calls. Thus, in some examples, the HMD X700 uses a nearby device (e.g., a mobile phone and / or a tablet) that is capable of communicating over a cellular network to initiate the audio communication session. The edit option, when selected, allows a user of the HMD X700 to remove the contact X712e from the pinned contacts X712 (or add the contact X712e to the pinned contacts in embodiments in which the contact X712e is not already a pinned contact), for example, as described with respect to Figure 7D FIG. 6. The message option, when selected, allows the user to transmit a message to the contact X712e, for example, as described with respect to Figure 7D FIG. 6. For example, while displaying the expanded contacts menu X720, the HMD X700 detects a selection of the message option (e.g., 720d). In some embodiments, the selection of the message option is, for example, an aerial gesture (e.g., a pinch gesture, and / or a gaze-and-pinch gesture) that indicates selection of the message option. In some embodiments, as Figure 7EAs shown, in response to detecting selection of the message option (e.g., 720d), HMD X700 displays (e.g., replaces display of communication interface X710 with) a message interface (e.g., 730). Thereafter, the message interface can be used to transmit a message to contact X712e.
[0241] In some embodiments, the contact menu associated with a contact includes one or more additional options based on a state of HMD X700. As an example, in some embodiments, in instances in which HMD X700 is engaged in a live communication session (e.g., an XR communication session or an audio communication session), the contact menu includes an option to invite the contact to the live communication session. For example, while engaged in an XR communication session and while displaying communication interface X710, HMD X700 detects selection of contact X714g (“Dylan”). In some embodiments, the selection of contact X714g is an air gesture (e.g., a pinch gesture, and / or a gaze-and-pinch gesture) that indicates selection of contact X714f, for example. As shown, in response to detecting selection of contact X714g, HMD X700 displays a contact menu X724 associated with contact X714g. Figure 7C2
[0242] Contact menu X724 includes an invite option X724a, an invite option X724b, and an expand option X724c. Invite option X724a, when selected (e.g., via an air gesture such as a pinch gesture, and / or a gaze-and-pinch gesture), causes HMD X700 to invite contact X714g to a new live communication session. Invite option X724b, when selected (e.g., via an air gesture such as a pinch gesture, and / or a gaze-and-pinch gesture), causes HMD X700 to invite contact X714g to the live communication session in which HMD X700 is currently engaged. Expand option X724c, when selected, causes HMD X700 to display one or more additional options for contact X714g.
[0243] In some embodiments, inviting a contact to a new live communication session (e.g., in response to selection of option X724a) will cause HMD X700 to disconnect from and / or terminate the live communication session in which HMD X700 is currently engaged. In some embodiments, prior to terminating an existing live communication session in this manner, HMD X700 confirms that the user wishes to disconnect from the current live communication session prior to initiating a new live communication session. For example, as shown, in response to detecting selection of invite option X724a, HMD X700 displays a confirmation interface X730a. Figure 7G As shown, in response to selection of the invite option X724a, the HMD X700 can display a confirmation interface 740 that includes a confirmation affordance 742. In response to selection of the confirmation affordance 742 (e.g., via an air gesture such as a pinch gesture, and / or a gaze-and-pinch gesture), the HMD X700 terminates the current live communication session and invites the contact X714f to a new live communication session.
[0244] In some embodiments, the user optionally uses the communication interface X710 to transmit a message to a contact. For example, while displaying the communication interface X710, the HMD X700 detects selection of the preview 716b associated with the pinned contact X712b (e.g., as shown in FIG. 7G). In some embodiments, the selection of the preview 716b is, for example, an air gesture (e.g., a pinch gesture, and / or a gaze-and-pinch gesture) that indicates selection of the preview 716b. As shown in FIG. 7H, in response to detecting the selection of the preview 716b, the HMD X700 expands the preview 716b to display a reply option X718. Figure 7B Figure 7C2 As shown, in response to detecting the selection of the preview 716b, the HMD X700 expands the preview 716b to display a reply option X718.
[0245] The reply option X718, when selected, causes the HMD X700 to display a reply interface for transmitting a message to the contact X712b. For example, while displaying the reply option X718 in the preview 716b, the HMD X700 detects selection of the reply option X718. In some embodiments, the selection of the reply option X718 is, for example, an air gesture (e.g., a pinch gesture, and / or a gaze-and-pinch gesture) that indicates selection of the reply option X718. As shown in FIG. 7I, in response to detecting the selection of the reply option X718, the HMD X700 can display a reply interface 750 that can be used to transmit a message to the contact X712b. Figure 7H
[0246] Figures 1B to 1P Any of the features, components, and / or parts shown, including their arrangements and configurations, can be included in the HMD X700, alone or in any combination. For example, in some embodiments, the HMD X700 includes any of the features, components, and / or parts of the HMDs 1-100, 1-200, 3-100, 6-100, 6-200, 6-300, 6-400, 11.1.1-100, and / or 11.1.2-100, alone or in any combination. In some embodiments, the display module X702 includes any of the features, components, and / or parts of the display units 1-102, display units 1-202, display units 1-306, display units 1-406, display generation component 120, display screens 1-122a-b, first rear-facing display screen 1-322a and second rear-facing display screen 1-322b, display 11.3.2-104, first display assembly 1-120a and second display assembly 1-120b, display assembly 1-320, display assembly 1-421, first display sub-assembly 1-420a and second display sub-assembly 1-420b, display assembly 3-108, display assembly 11.3.2-204, first optical module 11.1.1-104a and second optical module 11.1.1-104b, optical module 11.3.2-100, optical module 11.3.2-200, lenticular lens array 3-110, display regions or zones 6-232, and / or display / display regions 6-334, alone or in any combination. In some embodiments, the HMD X700 includes a sensor that includes any of the features, components, and / or parts of any of the sensors 190, sensors 306, image sensors 314, image sensors 404, sensor assembly 1-356, sensor assembly 1-456, sensor system 6-102, sensor system 6-202, sensor 6-203, sensor system 6-302, sensor 6-303, sensor system 6-402, and / or sensors 11.1.2-110a-f, alone or in any combination. In some embodiments, the input device X703 includes any of the features, components, and / or parts of any of the first button 1-128, button 11.1.1-114, second button 1-132, and / or dial or button 1-328, alone or in any combination. In some embodiments, the HMD X700 includes one or more audio output components (e.g., electronic components 1-112) for generating audio feedback (e.g., audio output), which is optionally generated based on detected events and / or user input detected by the HMD X700.
[0247] In some embodiments, the communication interface 710 is used to generate an avatar. In some embodiments, the avatar serves as a representation (e.g., a 3D representation) of the user of the device 700 in an XR communication session. For example, with reference to Figure 7B While displaying the communication interface 710, the device 700 detects selection of the avatar option 715. In some embodiments, the selection of the avatar option 715 is a tap gesture 713b on the avatar option 715. In some embodiments, the selection of the avatar option 715 is an air gesture that indicates selection of the avatar option 715, for example. As shown, in response to detecting the selection of the avatar option 715, the device 700 displays an avatar interface 760. Figure 7I
[0248] At Figure 7I the avatar interface 760 includes a first option 762 (e.g., more realistic than the second option) and a second option 764 (e.g., less realistic than the first option). The first option 762, when enabled, causes the avatar of the user of the device 700 to reflect the appearance of the user. For example, in some embodiments, when the first option 762 is enabled, the avatar includes one or more visual characteristics that correspond to one or more physical characteristics of the user. The second option 764, when enabled, causes the avatar of the user of the device 700 to indicate the motion of the user (e.g., during a live communication session) without reflecting the appearance of the user. For example, in some embodiments, when the second option 764 is enabled, an avatar with a default appearance is used. In some embodiments, when the first option 762 is enabled (as compared to the second option 764), the avatar of the user of the device 700 is represented by a first representation style, and the avatar is displayed at a first level of detail (e.g., a first level of detail with respect to the appearance of the user and / or one or more portions of the user) and indicates positioning and movement of a first user portion of the user relative to positioning and movement of a second user portion of the user in a first manner. In some embodiments, when the second option 764 is enabled (as compared to the first option 762), the avatar of the user of the device 700 is represented by a second representation style that is different from the first representation style, and the avatar is displayed at a second level of detail (e.g., a second level of detail with respect to the appearance of the user and / or one or more portions of the user) that is lower than the first level of detail (e.g., less than the appearance of the user and / or mimics the appearance of the user with less detail and / or a lower amount of detail) and indicates positioning and movement of the first user portion of the user relative to positioning and movement of the second user portion of the user in a second manner that is different from the first manner.
[0249] The avatar interface 760 also includes a menu option 766 that, when selected, causes the device 700 to display an avatar menu, as shown in Figure 7I For example, at Figure 7I When the avatar interface 760 is displayed, device 700 detects a selection of menu option 766. In some embodiments, the selection of menu option 766 is a tap gesture 705i on menu option 766. In some embodiments, the selection of menu option 766 is, for example, an air gesture indicating a selection of menu option 766. Figure 7J As shown, in response to the detection of a selection of menu option 766, device 700 displays avatar menu 768.
[0250] exist Figure 7J In some embodiments, the avatar menu 768 includes an edit option 768a, a create option 768b, and / or a delete option 768c. If an avatar has not yet been created for the user of device 700, the avatar menu 768 includes the create option 768b and does not include the edit option 768a and the delete option 768c. In other embodiments, if an avatar has already been created for the user of device 700, the avatar menu 768 includes the edit option 768a and the delete option 768c, and does not include the create option 768b.
[0251] exist Figure 7J At the location where the avatar menu 768 is displayed, device 700 detects a selection of creation option 768b. In some embodiments, the selection of creation option 768b is a tap gesture 705j on creation option 768b. In some embodiments, the selection of creation option 768b is, for example, an air gesture indicating a selection of creation option 768b. Figure 7K As shown, in response to the detection of a selection of creation option 768b, device 700 displays settings interface 770.
[0252] exist Figure 7K In this context, the settings interface 770 includes a settings option 772, which, when selected, causes the device 700 to display an avatar editing interface. For example, when the settings interface 770 is displayed, the device 700 detects a selection of the settings display 772. In some embodiments, the selection of the settings display 772 is a tap gesture 705k on the settings display 772. In some embodiments, the selection of the settings display 772 is, for example, an air gesture indicating a selection of the settings display 772. Figure 7L1 As shown, in response to the detection of a selection of the setting display 772, the device 700 displays the avatar editing interface 780.
[0253] exist Figure 7L1At the location, the avatar editing interface 780 includes a live view 781 of the avatar of the user of the device 700. In some embodiments, the device 700 displays the live view 781 of the avatar editing interface 780 in accordance with movements and / or behavioral gestures of the user of the device 700 as detected by the device 700, which is updated in real-time. The avatar editing interface 780 also includes various settings and / or parameters to adjust visual characteristics of the avatar. As an example, the avatar editing interface 780 includes settings 782, including a brightness setting 782a and a warmth setting 782b. The brightness setting 782a and the warmth setting 782b are used to adjust the simulated lighting and the temperature of the skin of the avatar, respectively. As another example, the avatar editing interface 780 includes a color palette 783 including a set of one or more colors and / or shades from which to select the color of the skin of the avatar.
[0254] In some embodiments, as Figure 7L1 illustrated, the avatar editing interface 780 includes a set of parameters 784, such as a shirt parameter 784a and a headwear parameter 784b. In some embodiments, selection of a parameter allows selection of a visual characteristic of one or more aspects of the avatar. For example, with reference to Figure 7M selection of the shirt parameter 784a (e.g., a tap input 7051 or an air gesture at a location corresponding to the parameter 784a) causes the device 700 to display a parameter menu 790 from which the user can select from any number of options (e.g., options 790a-790c) for the shirt parameter 784a. With reference to Figure 7N Once an option has been selected (e.g., a tap input 705m or an air gesture at a location corresponding to the option 790b), the user can select from styles 792 (e.g., 792a-792f) of the selected option, and the visual characteristic of the avatar is updated accordingly. Similarly, in some embodiments, selection of the headwear parameter 784b causes the device 700 to display headwear options for the headwear parameter 784b, and selection of an option causes the device 700 to display the type of the selected option.
[0255] While described herein with respect to parameters 784a and 784b that correspond to shirts and headwear, respectively, it will be appreciated that, in some embodiments, the parameters of the avatar editing interface 780 optionally correspond to other / additional visual aspects of the avatar. By way of example, in some embodiments, the parameters 784 are used to select one or more aspects of the eyewear of the avatar (e.g., parameter 784a corresponds to glasses, and parameter 784b corresponds to eye coverings). In examples in which the device 700 receives a user selection of the parameter 784a that corresponds to glasses, the device 700 displays options for various designs of glasses (e.g., rimless, thin frame, thick frame, etc.). Once the device 600 receives a user selection of a design, the device 700 displays various styles of the selected design as the type 792 for the user to select. In examples in which the device 700 receives a user selection of eye coverings, the device 700 displays options for various designs of eye coverings (e.g., left eye covering or right eye covering). Once the device 600 receives a user selection of a design, the device 700 displays various styles of the selected design as the type 792 for the user to select.
[0256] In some embodiments, the avatar editing interface 780 includes a set of parameters 786. As shown, in some embodiments, the parameters 786 are used to select one or more aspects of the hair. By way of example, parameter 786a corresponds to a hairstyle, parameter 786b corresponds to a hair color, and parameter 786c corresponds to a hair highlight. In other embodiments, the parameters 786 are used to select one or more aspects of accessibility features. By way of example, in some embodiments, parameter 786a corresponds to a hand prosthesis, parameter 786b corresponds to a hearing aid, and parameter 786c corresponds to a wheelchair.
[0257] In some embodiments, Figure 7L1 The techniques and user interfaces described in Figures 1A to 1P are provided by one or more of the devices described in Figure 7L2 Embodiments are illustrated in which the avatar editing interface X780 is displayed on a display module X702 of a head-mounted device (HMD) X700 (e.g., as described in Figures 7L1 to 7N In some embodiments, the device X700 includes a pair of display modules that provide stereoscopic content to different eyes of the same user. For example, the HMD X700 includes a display module X702 that provides content to the user’s left eye and a second display module that provides content to the user’s right eye. In some embodiments, the second display module displays slightly different images than the display module X702 to generate the illusion of stereoscopic depth.
[0258] In Figure 7L2At the place, the avatar editing interface X780 includes a live view X781 of the user's avatar of the HMD X700. In some embodiments, the HMD X700 displays the live view X781 of the avatar editing interface X780 in accordance with movements and / or behavioral gestures of the user of the HMD X700 as detected by the HMD X700, which is updated in real-time. The avatar editing interface X780 also includes various settings and / or parameters to adjust visual characteristics of the avatar. As an example, the avatar editing interface X780 includes settings X782, which include a brightness setting X782a and a warmth setting X782b. The brightness setting X782a and the warmth setting X782b are used to adjust the simulated lighting and the temperature of the skin of the avatar, respectively. As another example, the avatar editing interface X780 includes a color palette X783, which includes one or more sets of colors and / or shades from which to select the color of the skin of the avatar.
[0259] In some embodiments, as Figure 7L2 illustrated, the avatar editing interface X780 includes a set of parameters X784, such as a shirt parameter X784a and a headwear parameter X784b. In some embodiments, selecting a parameter allows for selection of a visual characteristic of one or more aspects of the avatar. For example, selection of the shirt parameter X784a (e.g., a gaze and pinch gesture, where the gaze is represented by a gaze indicator X705L corresponding to the location of the parameter X784a) causes the HMD X700 to display a parameter menu 790 from which the user can select from any number of options (e.g., options 790a-c) for the shirt parameter X784a, as Figure 7M illustrated.
[0260] In some embodiments, HMD X700 detects selection of shirt parameter X784a based on an air gesture performed by the user of HMD X700. In some embodiments, HMD X700 detects hands X750a and / or X750b of the user of HMD X700 and determines whether a motion of hands X750a and / or X750b performs a predetermined air gesture that corresponds to selection of shirt parameter X784a. In some embodiments, the predetermined air gesture that selects shirt parameter X784a includes a pinching gesture. In some embodiments, the pinching gesture includes detecting movement of fingers X750c and thumb X750d toward each other. In some embodiments, HMD X700 detects selection of shirt parameter X784a based on gaze and air gesture input performed by the user of HMD X700. In some embodiments, the gaze and air gesture input includes detecting that the user of HMD X700 is looking at shirt parameter X784a (e.g., for greater than a predetermined amount of time) and that hands X750a and / or X750b of the user of HMD X700 perform a pinching gesture.
[0261] Referring to Figure 7N Once an option has been selected (e.g., tap input 705m or an air gesture corresponding to the location of option 790b), the user can select from the styles 792 (e.g., 792a-792f) of the selected option and the visual properties of the avatar are updated accordingly. Similarly, in some embodiments, selection of headwear parameter 784b causes device 700 to display headwear options for headwear parameter 784b, and selection of an option causes device 700 to display the type of the selected option.
[0262] While described herein with respect to parameters X784a and X784b that correspond to shirts and headwear, respectively, it will be appreciated that in some embodiments, the parameters of the avatar editing interface X780 optionally correspond to other / additional visual aspects of the avatar. By way of example, in some embodiments, the parameters X784 are used to select one or more aspects of the eyewear of the avatar (e.g., parameter X784a corresponds to glasses and parameter X784b corresponds to eye patches). In examples in which the HMD X700 receives a user selection of the parameter X784a that corresponds to glasses, the HMD X700 displays options for various designs of glasses (e.g., rimless, thin frame, thick frame, etc.). Once the HMD X700 receives a user selection of a design, the HMD X700 displays various styles of the selected design as type 792 for the user to select. In examples in which the HMD X700 receives a user selection of eye patches, the HMD X700 displays options for various designs of eye patches (e.g., left eye patch or right eye patch). Once the HMD X700 receives a user selection of a design, the HMD X700 displays various styles of the selected design as type 792 for the user to select.
[0263] In some embodiments, the avatar editing interface X780 includes a set of parameters X786. As shown, in some embodiments, the parameters X786 are used to select one or more aspects of the hair. By way of example, parameter X786a corresponds to a hairstyle, parameter X786b corresponds to a hair color, and parameter X786c corresponds to a hair highlight. In other embodiments, the parameters X786 are used to select one or more aspects of accessibility features. By way of example, in some embodiments, parameter X786a corresponds to a hand prosthesis, parameter X786b corresponds to a hearing aid, and parameter X786c corresponds to a wheelchair.
[0264] Figures 1B to 1PAny of the illustrated features, components, and / or parts, including their arrangements and configurations, can be included in the HMD X700, alone or in any combination. For example, in some embodiments, the HMD X700 includes any of the features, components, and / or parts of the HMDs 1-100, 1-200, 3-100, 6-100, 6-200, 6-300, 6-400, 11.1.1-100, and / or 11.1.2-100, alone or in any combination. In some embodiments, the display module X702 includes any of the features, components, and / or parts of the display units 1-102, 1-202, 1-306, 1-406, the display generation component 120, the display screens 1-122a-b, the first and second rear-facing display screens 1-322a and 1-322b, the display 11.3.2-104, the first and second display assemblies 1-120a and 1-120b, the display assembly 1-320, the display assembly 1-421, the first and second display subassemblies 1-420a and 1-420b, the display assembly 3-108, the display assembly 11.3.2-204, the first and second optical modules 11.1.1-104a and 11.1.1-104b, the optical module 11.3.2-100, the optical module 11.3.2-200, the lenticular lens array 3-110, the display regions or zones 6-232, and / or the display / display regions 6-334, alone or in any combination. In some embodiments, the HMD X700 includes a sensor that includes any of the features, components, and / or parts of any of the sensors 190, the sensor 306, the image sensor 314, the image sensor 404, the sensor assembly 1-356, the sensor assembly 1-456, the sensor system 6-102, the sensor system 6-202, the sensor 6-203, the sensor system 6-302, the sensor 6-303, the sensor system 6-402, and / or the sensors 11.1.2-110a-f, alone or in any combination. In some embodiments, the input device X703 includes any of the features, components, and / or parts of any of the first button 1-128, the button 11.1.1-114, the second button 1-132, and / or the dial or button 1-328, alone or in any combination. In some embodiments, the HMD X700 includes one or more audio output components (e.g., the electronic component 1-112) for generating audio feedback (e.g., audio output), which is optionally generated based on detected events and / or user input detected by the HMD X700.
[0265] Reference Figure 7NWhile displaying the avatar editing interface 780, the device 700 detects selection of the save option 788. In some embodiments, the selection of the save option 788 is a tap gesture 705n on the save option 788. In some embodiments, the selection of the save option 788 is an air gesture, e.g., indicating selection of the save option 788. In response to detecting the selection of the save option 788, the device 700 stores, e.g., locally and / or remotely, the configuration of the avatar selected by the user of the device 700 for subsequent use in XR communication sessions. As Figure 7O shown, in further response to detecting selection of the settings affordance 772, the device 700 displays a completion interface 795 indicating that the user has successfully created and / or updated the user’s avatar.
[0266] In Figure 7P the user of the device 700 is engaged in an XR communication session with a contact 712f (“Ann,” Figure 7B ) within the XR environment 704. In some embodiments, the XR communication session is a spatial communication session. Thus, in some embodiments, the contact 712f and / or the user of the device 700 are represented in the XR environment 704 by 3D representations, e.g., avatars. For example, as Figure 7P shown, the user of the device 700 is represented by a representation 700A (as shown in a self-preview 706A) and the contact 712f is represented by a representation 701A.
[0267] In some embodiments, the view of the XR environment 704 for the user of the device 700 is provided from the perspective of the representation 700A within the XR environment 704. As this can prevent the user from otherwise viewing the representation 700A, the device 700 displays a self-preview 706A that includes a live view of the representation 700A in the XR environment 704. While the self-preview 706A is shown as being located in the lower right corner of the display 702, it will be appreciated that the self-preview 706A can optionally be displayed at any location on the display 702. By way of example, in some embodiments, the self-preview 706A is positioned proximate to the representation of the contact in the XR communication session. In some embodiments, the self-preview 706A is located at a positioning 708A, e.g., proximate to the representation 701A.
[0268] In some embodiments, the participants represented by the 3D representation in the XR environment have spatial agents. Therefore, during an XR communication session, the 3D representation optionally moves within the XR environment 704, such that the 3D representation moves relative to elements (e.g., table 704a and sofa 704b) and / or other participants in the XR environment 704. In some embodiments, the 3D representation moves in response to the movement of a corresponding device. By way of example, 3D representation 700A may move within the XR environment 704 in response to the movement of device 700. In some embodiments, the 3D representation moves in a manner corresponding to the movement of the device. For example, if device 700 first moves in a first direction (e.g., to the left) and then subsequently moves in a second direction (e.g., to the right), then 3D representation 700A will move within the XR environment 704 in a similar manner in both the first and second directions.
[0269] In some implementations, when participating in an XR communication session, device 700 displays a set of controls 704A for managing one or more aspects of the XR communication session, such as... Figure 7P As shown. The set of controls 704A includes a message option 704Aa, an information option 704Ab, a microphone option 704Ac, an avatar option 704Ad, a camera option 704Ae, and a termination option 704Af. The message option 704Aa, when selected, causes the device 700 to display a message interface for sending messages to the contact 712f. The information option 704Ab, when selected, causes the device 700 to display an information interface corresponding to the contact 712f. The microphone option 704Ac, when selected, toggles the state of the device 700's microphone (e.g., enabling or disabling the microphone). In some embodiments, disabling the device 700's microphone prevents the device 700 from providing audio during an XR communication session. The camera option 704Ae, when selected, toggles the state of the device 700's camera (e.g., enabling or disabling the camera). In some embodiments, disabling the device 700's camera prevents the device 700 from providing video during an XR communication session (e.g., video feeds from the user of device 700 and / or movement of the user's representation on device 700). The termination option 704Af disconnects device 700 from the XR communication session when selected.
[0270] Avatar option 704Ad, when selected, toggles (e.g., enables or disables) the use of 3D representation in XR environment 704. For example, when XR environment 704 is displayed, device 700 detects the selection of avatar option 704Ad. In some embodiments, the selection of avatar option 704Ad is a tap gesture 705p on avatar option 704Ad. In some embodiments, the selection of avatar option 704Ad is, for example, an air gesture indicating the selection of avatar option 704Ad. Figure 7QAs shown, in response to selection of avatar option 704Ad, device 700 disables use of 3D representations in XR environment 704.
[0271] In some embodiments, when toggling use of 3D representations in XR environment 704, device 700 transitions the XR communication session between a spatial communication session and a non-spatial communication session. Disabling use of 3D representations causes device 700 to transition the XR communication session from a spatial communication session to a non-spatial communication session. Enabling use of 3D representations causes device 700 to transition the XR communication session from a non-spatial communication session to a spatial communication session.
[0272] In some embodiments, in a non-spatial communication session, participants in the XR communication session are represented by 2D representations. By way of example, as shown in FIG. 7, user of device 700 is represented by 2D representation 710A (as shown in self-preview 706A), and contact 712f is represented by 2D representation 712A. Figure 7Q As shown, user of device 700 is represented by 2D representation 710A (as shown in self-preview 706A), and contact 712f is represented by 2D representation 712A.
[0273] In some embodiments, the 2D representation includes a video feed (e.g., a live video feed) of the user. In some embodiments, if the video feed of the user is unavailable (e.g., the camera of the device is disabled), the 2D representation of the user includes, instead, an image (e.g., a thumbnail) associated with the user, a letter combination corresponding to the user, and / or another 2D representation. In some embodiments, a user represented by a 2D representation in an XR environment does not have a spatial agent, and is optionally positioned at one or more predetermined locations in XR environment 704. In some embodiments, device 600 is configured to move a 2D representation of a remote participant in an XR environment based on user input received at device 700 (e.g., input that drags a representation from a first location to a second location). In some embodiments, device 600 is not configured to move a 3D representation of a remote participant in an XR environment based on user input received at device 700.
[0274] The following reference methods 800 and 900 provide additional description regarding Figures 7A to 7Q each of which is described with respect to Figures 7A to 7Q .
[0275] Figure 8 is a flow diagram of an example method 800 for managing a live communication session, in accordance with some embodiments. In some embodiments, method 800 is performed at a computer system (e.g., computer system 101, computer system 700, and / or HMD X700) (e.g., a smartphone, a tablet, and / or a head-mounted device) that includes a display generation component (e.g., Figure 1A , in accordance with some embodiments. In some embodiments, method 900 is performed at a computer system (e.g., computer system 101, computer system 700, and / or HMD X700) (e.g., a smartphone, a tablet, and / or a head-mounted device) that includes a display generation component (e.g., Figure 1A , in accordance with some embodiments. In some embodiments, method 900 is performed at a computer system (e.g., computer system 101, computer system 700, and / or HMD X700) (e.g., a smartphone, a tablet, and / or a head-mounted device) that includes a display generation component (e.g、 Figure 3 and Figure 4 The display generation component 120, the display 702, and / or the display X 702) (e.g., a visual output device, a 3D display, a display with a transparent or semi-transparent at least portion on which images can be projected (e.g., a see-through display), a projector, a heads-up display, and / or a display controller) and one or more sensors (e.g., a touch- sensitive surface, a gyroscope, an accelerometer, a motion sensor, a movement sensor, a microphone, an infrared sensor, a camera sensor, a depth camera, a visible light camera, an eye tracking sensor, a gaze tracking sensor, a physiological sensor, and / or an image sensor) of the computer system (e.g., 700 and / or X 700) are in communication. In some embodiments, the method 800 is governed by instructions that are stored in a non-transitory (or transitory) computer readable storage medium and that are executed by one or more processors of a computer system, such as the one or more processors 202 of the computer system 101 (e.g., Figure 1A The controls 110 of FIG. 1M). Some operations in the method 800 are, optionally, combined and / or the order of some operations is, optionally, changed.
[0276] The computer system (e.g., 700 and / or X 700) displays (802), via the display generation component, representations (e.g., 712a-712g and 714a-714i) (e.g., static avatars, animated avatars, images, and / or letter combinations) of a plurality of users (e.g., users that are not operating the computer system (remote users) and / or users other than the user of the computer system).
[0277] The computer system (e.g., 700 and / or X 700) receives (804), via the one or more sensors, a selection (e.g., 705b and / or 711b) (e.g., via a touch input on a touch- sensitive surface and / or via an air gesture) of a representation (e.g., 712e, X712e, 714g, and / or X714g) (e.g., a static avatar, an animated avatar, an image, and / or a letter combination) of a respective user of the plurality of users.
[0278] In response to (806) receiving the selection of the representation (e.g., 712e, X712e, 714g, and / or X714g) of the respective user and in accordance with a determination that there is an ongoing (e.g., active and / or currently established) communication session (e.g., a video communication session, an audio communication session, an extended reality communication session, a spatial communication session, and / or a non-spatial communication session), the computer system (e.g., 700 and / or X 700) displays (808), via the display generation component (e.g., 702 and / or X 702), an option (e.g., Figure 7C1 at 724b and / or Figure 7C2 at X724b) to invite the respective user to join the ongoing communication session.
[0279] In response to (806) receiving a selection of the corresponding user (e.g., 712e, X712e, 714g and / or X714g) and based on determining that no ongoing communication session exists, the computer system (e.g., 700 and / or X700) abandons displaying (810) the option to invite the corresponding user to join the ongoing communication session (e.g., as...). Figure 7C1 Menu 720 and / or Figure 7C2 (In the menu X720). Conditionally displaying the option to invite the corresponding user to join the ongoing communication session allows the computer system user to invite the corresponding user without navigating to a different user interface, thereby reducing the amount of input required to perform the invitation operation.
[0280] In some implementations, in response to receiving a selection of a representation of the corresponding user (e.g., 712e, X712e, 714g and / or X714g) (e.g., regardless of the determination of whether an ongoing communication session exists), the computer system (e.g., 700 and / or X700) displays options for initiating a new space communication session with the corresponding user (e.g., ...) via a display generation component (e.g., 702 and / or X702). Figure 7C1 720a and / or 724a and / or Figure 7C2 (X720a and / or X724a at the location) and options for additional features (e.g., Figure 7C1 720b and / or 724c and / or Figure 7C2 (e.g., without displaying options for sending text messages to the corresponding user and / or displaying additional information about the corresponding user). In some embodiments, options for displaying additional features (e.g., ...) are provided. Figure 7C1 720a and / or 724a and / or Figure 7C2 When X720a and / or X724a are present, the computer system (e.g., 700 and / or X700) receives options for additional features (e.g., ... Figure 7C1 720a and / or 724a and / or Figure 7C2 The selection of x720a and / or x724a at the location (e.g., 705c and / or 709d). In some embodiments, in response to receiving a selection of an option for an additional feature, the computer system (e.g., 700 and / or X700) displays (e.g., by replacing the display of the option to invite the corresponding user to join an ongoing communication session) one or more options associated with the corresponding user via a display generation component (e.g., 702 and / or X702). Figure 7D(e.g., options for communicating with a corresponding user, such as by sending a text message to the corresponding user and / or displaying additional information about the corresponding user). In some embodiments, a space communication session is a communication session represented by at least some (e.g., fewer than all, multiple, and / or all) of participating users distributed in a 3D environment. Displaying options for initiating a new space communication session with the corresponding user, as well as options for accessing additional features, allows users of the computer system to quickly access the options for initiating a new space communication session without cluttering the user interface, while still providing access to additional (and potentially less frequently used) features, thereby improving the human-computer interface.
[0281] In some implementations, in response to receiving a selection of the corresponding user (e.g., regardless of whether an ongoing communication session is determined), the computer system (e.g., 700 and / or X700) displays options for additional features (e.g., ...) via a display generation component (e.g., 702 and / or X702). Figure 7C1 720b and / or Figure 7C2 (e.g., without displaying options for sending text messages to the corresponding user and / or displaying additional information about the corresponding user). In some embodiments, options for displaying additional features (e.g., ...) are shown. Figure 7C1 720b and / or Figure 7C2 When X720b is in use, the computer system (e.g., 700 and / or X700) receives options for additional features (e.g., ...) via one or more sensors. FIG. 7C1 720b and / or FIG. 7C2 The selection of X720b) (e.g., 705c) (e.g., via touch input on a touch-sensitive surface and / or via air gestures). In some embodiments, in response to receiving an option for additional features (e.g., FIG. 7C1 720b and / or FIG. 7C2 The computer system (e.g., 700 and / or X700) displays (e.g., by replacing the display of options for inviting the corresponding user to join the ongoing communication session) options for initiating an audio communication session with the corresponding user (e.g., excluding live visual representations of participants and / or excluding video components) via a display generation component (e.g., X720b). FIG. 7D (e.g., as part of one or more options associated with the corresponding user). In some implementations, the computer system receives options for initiating an audio communication session with the corresponding user via one or more sensors (e.g., 720c). FIG. 7Dselection of 720c) at the place (e.g., via touch input on the touch- sensitive surface and / or via an over-the-air gesture). In some embodiments, in response to receiving a selection of the option to initiate an audio communication session with the respective user (e.g., FIG. 7D In response to receiving a selection of 720c) at the place, the computer system (e.g., 700 and / or X700) initiates an audio communication session (e.g., that does not include a live visual representation of the participant and / or that does not include a video component) with the respective user (e.g., and not with other users). Displaying the option to initiate the audio communication session enables a user of the computer system to begin a communication session that does not include a live visual representation of the user without having to initiate a video communication session and separately disable the video portion, thereby reducing the number of inputs needed to initiate the audio communication session.
[0282] In some embodiments, the computer system initiating the audio communication session includes using an external electronic device (e.g., a smartphone and / or a cellular phone) that is within a predetermined range (e.g., a distance and / or a wireless range) of the computer system (e.g., 700 and / or X700) to initiate an audio call (e.g., a voice call and / or a telephone call). In some embodiments, the option to initiate the audio communication session is displayed for respective users that do not have an account that includes a particular online service (or that do not have an active account) (e.g., the user of the computer system has an account that includes a particular online service for video and / or extended reality communication, but the respective user does not have that account). Using the external electronic device to initiate the audio communication session via the audio call enables the computer system to use resources of the external computer system (e.g., cellular connectivity and / or CPU processing of the external computer system), thereby improving the functionality of the computer system while reducing the work load of the computer system.
[0283] In some embodiments, in response to receiving selection (e.g., 705c and / or 711b) of a representation of a respective user (e.g., 712e, X712e, 714g, and / or X714g) (e.g., independent of a determination of whether there is an ongoing communication session), the computer system (e.g., 700 and / or X700) displays, via the display generation component (e.g., 702 and / or X702), an option for an additional feature (e.g., 720b, X720b, 724c, and / or X724c) (e.g., without displaying a process for initiating a message transmission to the respective user and / or without displaying an option for additional information about the respective user). In some embodiments, while displaying the option for the additional feature (e.g., 720b, X720b, 724c, and / or X724c), the computer system (e.g., 700 and / or X700) receives, via the one or more sensors, selection (e.g., 705c) of the option for the additional feature (e.g., 720b and / or X720b) (e.g., via a touch input on the touch- sensitive surface and / or via an air-gesture). In some embodiments, in response to receiving selection (e.g., 705c) of the option for the additional feature (e.g., 720b and / or X720b), the computer system (e.g., 700 and / or X700) displays, via the display generation component (e.g., 702 and / or X702), an option (e.g., 720d) to initiate a process (e.g., that does not include a live transmission of audio and / or video) to transmit a message to the respective user (e.g., as part of one or more options associated with the respective user) (e.g., by replacing display of the option to invite the respective user to join an ongoing communication session). In some embodiments, the computer system receives, via the one or more sensors, selection (e.g., 705d) of the option to initiate the process to transmit a message to the respective user (e.g., 720d) (e.g., via a touch input on the touch-sensitive surface and / or via an air-gesture). In some embodiments, in response to receiving selection (e.g., 705d) of the option to initiate the process to transmit a message to the respective user (e.g., 720d), the computer system (e.g., 700 and / or X700) initiates transmission of a message to the respective user (e.g., FIG. 7E at 730 and / or FIG. 7Hthe process of transmitting the message to the respective user includes displaying a user interface that includes a conversation between the user of the computer system and the respective user, displaying a keyboard, and / or displaying a text entry field for entering the message. Providing the option to initiate the process of transmitting the message to the respective user via the additional feature selection enables the user of the computer system to quickly initiate the process without having to specify a recipient, thereby reducing the number of inputs required to transmit the message.
[0284] In some implementations, in response to receiving a selection (e.g., 705c and / or 711b) of a representation to a corresponding user (e.g., 712e, X712e, 714g and / or X714g) (e.g., regardless of whether an ongoing communication session exists), the computer system (e.g., 700 and / or X700) displays options for additional features (e.g., 720b, X720b, 724c and / or X724c) via a display generation component (e.g., 702 and / or X702) (e.g., without displaying options for initiating a process to transmit a message to the corresponding user and / or displaying additional information about the corresponding user). In some embodiments, when displaying options for additional features (e.g., 720b, X720b, 724c, and / or X724C), the computer system receives the selection of the options for the additional features (e.g., 720b, X720b, 724c, and / or X724c) via one or more sensors (e.g., via touch input on a touch-sensitive surface and / or via air gestures). In some embodiments, in response to receiving the selection of the options for the additional features (e.g., 720b, X720b, 724c, and / or X724c), the computer system displays (e.g., by replacing the display of options for inviting the corresponding user to join an ongoing communication session) an option (e.g., 720e) for displaying additional information about the corresponding user (e.g., without displaying additional information about other users) (e.g., as part of one or more options associated with the corresponding user) via a display generation component. In some embodiments, the computer system receives a selection (e.g., 707d) of an option (e.g., 720e) for displaying additional information about the corresponding user via one or more sensors (e.g., via touch input on a touch-sensitive surface and / or via air gestures). In some embodiments, in response to receiving a selection (e.g., 707d) of an option (e.g., 720e) for displaying additional information about the corresponding user via a display generation component (e.g., 702 and / or X702), the computer system (e.g., 700 and / or X700) displays additional information about the corresponding user (e.g., information not displayed when an option for additional features is selected and / or when an option for additional information is selected) via a display generation component (e.g., information not displayed when an option for additional features is selected). FIG. 7F (e.g., previous communication history with the corresponding user, the corresponding user's phone number, and / or the corresponding user's email address) (e.g., without displaying additional information about other users). Providing users with additional information about the corresponding user provides feedback to users of the computer system about the corresponding user and / or the corresponding user's device, thereby providing improved visual feedback.
[0285] In some implementations, in response to receiving a selection (e.g., 705c and / or 711b) of a representation to a corresponding user (e.g., 712e, X712e, 714g and / or X714g) (e.g., regardless of whether an ongoing communication session exists), the computer system (e.g., 700 and / or X700) displays options for additional features (e.g., 720b, X720b, 724c and / or X724c) via a display generation component (e.g., 702 and / or X702) (e.g., without displaying options for initiating a process to transmit a message to the corresponding user and / or displaying additional information about the corresponding user). In some embodiments, when displaying options for additional features (e.g., 720b, X720b, 724c, and / or X724c), the computer system (e.g., 700 and / or X700) receives the selection of the option for the additional feature via one or more sensors (e.g., 705c) (e.g., via touch input on a touch-sensitive surface and / or via air gestures). In some embodiments, in response to receiving a selection (e.g., 705c) for the additional feature (e.g., 720b and / or X720b), the computer system displays (e.g., by replacing the display of options for inviting the corresponding user to join an ongoing communication session) an option (e.g., 720f) for removing the corresponding user from a favorites page (e.g., a list or group of favorite users) (e.g., without removing other users) as part of one or more options associated with the corresponding user. In some implementations, removing a user from the Favorites page includes stopping the display of the user's representation as part of a representation of multiple users (e.g., static avatars, animated avatars, images, and / or letter combinations) (e.g., the representation is optionally displayed as part of the main user interface). In some implementations, the computer system receives a selection (e.g., 720f) of an option (e.g., via touch input on a touch-sensitive surface and / or via air gestures) via one or more sensors. In some implementations, in response to receiving a selection of an option (e.g., 720f) to remove a user from the Favorites page (e.g., 712), the computer system (e.g., 700 and / or X700) initiates a process to remove the user from the Favorites page (e.g., requesting confirmation to remove the user from the Favorites page and / or removing the user from the Favorites page). Initiating a process to remove a user from the Favorites page allows the user to limit the users accessible via the Favorites page and / or the main user interface, thereby reducing visual clutter and enabling the addition of other users to the Favorites page, thus providing improved visual feedback.
[0286] In some implementations, in response to receiving a selection (e.g., 711b) of a representation to the corresponding user (e.g., 714g and / or X714g) and based on the determination that an ongoing (e.g., active and / or currently established) communication session exists (e.g., video communication session, audio communication session, extended reality communication session, space communication session, and / or non-space communication session), the computer system (e.g., 700 and / or X700) displays, via a display generation component (e.g., 702 and / or X702), options (e.g., 724a and / or X724a) for initiating a process to start a new communication session with the corresponding user, along with options for inviting the corresponding user to join the ongoing communication session (e.g., 724b and / or X724b). In some embodiments, the computer system (e.g., 700 and / or X700) receives, via one or more sensors, a selection (e.g., 709c) of an option (e.g., 724a and / or X724a) for initiating a process to begin a new communication session with the corresponding user (e.g., via touch input on a touch-sensitive surface and / or via air gestures). In some embodiments, in response to receiving a selection (e.g., 709c) of an option (e.g., 724a and / or X724a) for initiating a process to begin a new communication session with the corresponding user, the computer system (e.g., 700 and / or X700) initiates an termination of the ongoing communication session (e.g., ...). FIG. 7G The process of initiating a new communication session with the corresponding user (at location 740). In some implementations, in response to receiving a selection of an option for initiating a new communication session with the corresponding user, the computer system automatically (e.g., without requiring and / or receiving additional input from the user and / or without requesting user confirmation) ends the ongoing communication session and begins a new communication session with the corresponding user. Providing the option to begin a new communication session with the corresponding user enables the computer system to end the ongoing communication session and begin a new communication session without requiring separate user input to end the ongoing communication session and begin a new communication session, thereby reducing the amount of input required to perform the operation.
[0287] In some implementations, the process of initiating a new communication session with the corresponding user (e.g., FIG. 7GDuring the period at 740), the computer system (e.g., 700 and / or X700) prompts (e.g., 742) (e.g., via audio using a speaker and / or via a display using a display generation component) to confirm the end of the ongoing communication session. In some embodiments, the computer system (e.g., 700 and / or X700) receives confirmation of the end of the ongoing communication session via one or more sensors (e.g., when the prompt confirming the end of the ongoing communication session is displayed). In some embodiments, in response to receiving confirmation of the end of the ongoing communication session, the computer system ends the ongoing communication session (and optionally, begins a new communication session with the corresponding user). The user request for confirmation of the end of the ongoing communication session from the computer system allows the computer system to prevent the user from unintentionally ending the ongoing communication session, thereby improving the human-computer interface.
[0288] In some implementations, representations of multiple users (e.g., static avatars, animated avatars, images, and / or letter combinations) are displayed as part of a main user interface (e.g., optionally including representations of contacts with whom they have recently communicated). FIG. 7B-7D As shown), and displays options for inviting the corresponding user to join an ongoing communication session (e.g., 724b and / or X724b) and / or displays one or more options associated with the corresponding user (e.g., 720a-720b, X720a-X720b, 724a-724c and / or X724a-X724c), including obscuring the main user interface (e.g., partially blocking the display of the main user interface, blurring the main user interface, and / or otherwise partially obscuring the main user interface). In some embodiments, the main user interface includes multiple user interface objects for displaying the corresponding application (e.g., a first user interface object that displays the user interface of a first application when activated, and a second user interface object that displays the user interface of a second application different from the first application when activated). In some embodiments, in response to detecting corresponding user input (e.g., detecting the pressing of a physical button and / or detecting a corresponding gesture, such as an air gesture), the computer system displays the main user interface (e.g., regardless of what the computer system is displaying when the corresponding user input is received). In some implementations, the computer system automatically displays the main user interface after it exits a low-power mode (e.g., wakes up) and / or receives user input to unlock the computer system. Continuing to display the main user interface (when obscured) provides the user with context about the content they are accessing, including information about the relevant user (e.g., name and / or contact information).
[0289] In some embodiments, one or more options associated with the corresponding user include an option (e.g., 720d) for initiating a process of sending a message to the corresponding user (e.g., excluding live transmission of audio and / or video) (e.g., as part of one or more options associated with the corresponding user). In some embodiments, the computer system (e.g., 700 and / or X700) receives a selection (e.g., 705d) of the option (e.g., 720d) for initiating the process of sending a message to the corresponding user via one or more sensors. In some embodiments, in response to receiving (e.g., 705d) the option (e.g., 720d) for initiating the process of sending a message to the corresponding user (e.g., 705d) (and optionally, based on a determination that no message from the corresponding user is displayed upon receiving a selection of the representation of the corresponding user), the computer system (e.g., 700 and / or X700) displays a messaging user interface (e.g., for messaging with the corresponding user) having a first appearance via a display generation component (e.g., 702 and / or X702). FIG. 7E At point 730 (e.g., a messaging user interface having a first size and / or including a messaging user interface displaying a keyboard), the main user interface is not displayed. Displaying a messaging user interface with a first appearance instead of the main user interface provides the user with feedback that the messaging user interface is in a first state, thereby providing the user with improved visual feedback.
[0290] In some implementations, one or more options associated with a corresponding user include options (e.g., 718 and / or X718) for initiating a process of transmitting a message to the corresponding user (e.g., excluding live transmission of audio and / or video) (e.g., as part of one or more options associated with the corresponding user). In some implementations, a computer system (e.g., 700 and / or X700) receives selections (e.g., 707c) of the options (e.g., 718 and / or X718) for initiating the process of transmitting a message to the corresponding user via one or more sensors. In some implementations, in response to receiving a selection (e.g., 707c) of an option (e.g., 718 and / or X718) for initiating a message transmission to the corresponding user (and optionally, based on a determination that a message from the corresponding user is being displayed upon receiving a selection of the representation of the corresponding user), the computer system (e.g., 700 and / or X700) simultaneously displays, via a display generation component, a messaging user interface (e.g., 750) having a second appearance (e.g., different from the first appearance) for messaging with the corresponding user (e.g., a messaging user interface of a second size smaller than the first size and / or a messaging user interface excluding a display keyboard) and at least a portion of the main user interface (e.g., such as...). FIG. 7Hdisplaying the messaging user interface with the second appearance (e.g., displaying a blocked main user interface). Displaying the messaging user interface with the second appearance and a portion of the main user interface provides feedback to the user that the messaging user interface is in the second state, and thus provides the user with improved visual feedback.
[0291] In some embodiments, the main user interface is not movable by the user, and wherein the messaging user interface (e.g., 750) (e.g., with the first appearance and / or with the second appearance) is movable by the user. Enabling the user to move the messaging user interface without enabling the user to move the main user interface provides feedback to the user about which elements are part of the main user interface and which elements are not part of the main user interface, and thus provides the user with improved visual feedback.
[0292] In some embodiments, displaying, via the display generation component (e.g., 702 and / or X702), the representations of the plurality of users (e.g., 712a-712g and 714a-714i) includes displaying, via the display generation component (e.g., 702 and / or X702), a first grouping of a first plurality of representations of a first plurality of users (e.g., 712), where the first plurality of users (e.g., 712a-712g) are selected to be included as part of the representations of the plurality of users (e.g., included to be displayed based on being manually selected as part of a favorites page contact and / or based on frequency of communication) independent of recency of communication between the user of the computer system and the first plurality of users. In some embodiments, displaying, via the display generation component, the representations of the plurality of users (e.g., 712a-712g and 714a-714i) includes displaying, via the display generation component (e.g., 702 and / or X702), a second grouping of a second plurality of representations of a second plurality of users (e.g., 714), where the second plurality of users (e.g., 714a-714i) are selected to be included as part of the plurality of users based on recency of communication between the user of the computer system and the second plurality of users (e.g., included to be displayed based on most recent users communicated with). In some embodiments, an order of the second plurality of representations of the second plurality of users is based on recency of communication between the user of the computer system and the second plurality of users. In some embodiments, displaying, via the display generation component, the representations of the plurality of users includes: displaying, via the display generation component, a first grouping of a first plurality of representations of a first plurality of users, where the first plurality of users are selected to be included as part of the representations of the plurality of users independent of recency of communication between the user of the computer system and the first plurality of users; and a second grouping of a second plurality of representations of a second plurality of users, where: in accordance with a determination that the recent communications by the user of the computer system include communications with a first set of one or more users and do not include communications with a second set of one or more users, the second plurality of users includes the first set of one or more users and does not include the second set of one or more users, and in accordance with a determination that the recent communications by the user of the computer system include communications with the second set of one or more users and do not include communications with the first set of one or more users, the second plurality of users includes the second set of one or more users and does not include the first set of one or more users. Grouping the first plurality of users together and the second plurality of users together enables the computer system to provide feedback to the user regarding which users are selected independent of recency of communication and which users are included based on recency of communication, thereby providing improved visual feedback.
[0293] In some embodiments, recency of communication between a user of the computer system (e.g., 700 and / or X700) and a second plurality of users is based on a plurality of communication modalities (e.g., text messaging, telephone calling, and / or communication sessions (e.g., video communication sessions, audio communication sessions, extended reality communication sessions, spatial communication sessions, and / or non-spatial communication sessions)). Grouping the second plurality of users together based on recency of communication using a plurality of communication modalities enables the computer system to group recent contacts regardless of how the communication occurred, providing improved visual feedback.
[0294] In some embodiments, the computer system (e.g., 700 and / or X700) displays, via the display generation component (e.g., 702 and / or X702) and contemporaneously with an option (e.g., 724b and / or X724b) to invite a respective user to join an ongoing communication session and / or one or more options associated with the respective user (e.g., 724a, X724a, 724c, and / or X724c), an indication (e.g., 714f and / or X714f) of recent communication activity between a user of the computer system and the respective user (e.g., information about a recent c...
Claims
1. A computer system configured to communicate with a display generation component, the computer system comprising: one or more processors; and memory storing one or more programs configured for execution by the one or more processors, the one or more programs including instructions for: while participating in a communication session as a spatial communication session, wherein the spatial communication session includes displaying representations of a plurality of participants in the communication session in a spatial distribution arrangement in a 3D environment via the display generation component, wherein displaying the plurality of participants in the spatial distribution arrangement includes displaying: the representations of the plurality of participants spaced apart from each other and a user of the computer system in a first non-vertical direction in the 3D environment by at least a threshold amount; and the representations of the plurality of participants spaced apart from each other and the user in a second non-vertical direction different from the first non-vertical direction by at least the threshold amount; while displaying the representations of the plurality of participants distributed in the 3D environment, detecting an event; and in response to detecting the event, transitioning the communication session from the spatial communication session to a non-spatial communication session, the non-spatial communication session including displaying representations of at least a subset of the plurality of participants of the communication session in a grouped arrangement via the display generation component, wherein in the grouped arrangement: the representations of the plurality of participants are spaced apart from each other in the 3D environment in the first non-vertical direction by less than the threshold amount; a representation of a first participant in the grouped arrangement has a different positioning than a representation of the first participant in the spatial distribution arrangement; and a representation of a second participant in the grouped arrangement has a different positioning than a representation of the second participant in the spatial distribution arrangement.
2. The computer system of claim 1, wherein: in the non-spatial communication session: a representation of a first participant of the plurality of participants is in a first window region, and a representation of a second participant of the plurality of participants is in a second window region different from the first window region; and in the spatial communication session: the representation of the first participant of the plurality of participants is not in a window region, and the representation of the second participant of the plurality of participants is not in a window region.
3. The computer system of claim 2, wherein the representation of the first participant is a simulated three-dimensional representation and the representation of the second participant is a two-dimensional representation.
4. The computer system of claim 2, wherein the plurality of participants are two-dimensional representations.
5. The computer system of claim 2, wherein the plurality of participants are three-dimensional representations.
6. The computer system of claim 1, wherein the event is a request received during the communication session to transition a representation of the user of the computer system from a 3D representation to a 2D representation.
7. The computer system of claim 6, wherein the request is based on input in a communication session control region. 8. The computer system of claim 7, wherein the communication session control region includes an option to transition a representation of the user of the computer system from the 3D representation to the 2D representation and one or more options corresponding to other communication session controls.
9. The computer system of claim 1, wherein the event is a request received during the communication session to transition the communication session from the spatial communication session to the non-spatial communication session.
10. The computer system of claim 1, wherein the event is an additional participant joining the communication session.
11. The computer system of claim 10, wherein the additional participant joining the communication session causes a number of participants represented by the simulated three- dimensional representation to exceed a threshold number of participants.
12. The computer system of claim 1, the one or more programs further including instructions to: when the communication session is a non-spatial communication session, shift a positioning of a respective window region corresponding to a respective participant based on movement of the respective participant.
13. The computer system of claim 12, wherein the respective window region moves forward and / or backward in the virtual environment based on a head positioning of the respective participant.
14. The computer system of claim 12, wherein the respective window region tilts based on a head positioning of the respective participant.
15. The computer system of claim 12, wherein a first window shifts in a first direction based on movement of a participant displayed in the first window, and a second window shifts in a second direction different from the first direction based on movement of a participant displayed in the second window.
16. The computer system of claim 1, the one or more programs further including instructions to: while participating in the communication session as a non-spatial communication session, detect a second event; and in response to detecting the second event, transition the communication session from the non-spatial communication session to the spatial communication session.
17. The computer system of claim 16, wherein the second event is a participant leaving the communication session.
18. The computer system of claim 16, wherein the second event is a request received during the communication session to transition a representation of the user of the computer system from a 2D representation to a 3D representation.
19. The computer system of claim 16, wherein the second event is a request received during the communication session to transition the communication session from a non-spatial communication session to a spatial communication session.
20. The computer system of claim 1, the one or more programs further including instructions to: while in a spatial communication session, display, via the display generation component, a first-person view of a representation of the user of the computer system in a first-person view window region.
21. The computer system of claim 20, wherein the self-view window area overlaps a window area that includes a representation of another participant of an ongoing communication session.
22. The computer system of claim 20, wherein the self-view window area is smaller than a window area that includes a representation of another participant.
23. The computer system of claim 1, wherein: during a spatial communication session, causing a first participant of the communication session to be able to move a respective representation of the first participant, and causing a second participant of the communication session to be able to move a respective representation of the second participant; and during a non-spatial communication session, causing a user of the computer system to be able to move respective window areas that include respective representations of the plurality of participants of the communication session.
24. The computer system of claim 23, wherein the computer system is configured to communicate with one or more sensors, the one or more programs further including instructions for: detecting, via the one or more sensors, a user input that repositions a respective window area that includes a respective representation of a participant; and in response to detecting the user input that repositions the respective window area that includes the respective representation of the participant, repositioning a plurality of window areas of the plurality of participants.
25. The computer system of claim 23, wherein the respective representations of the plurality of participants are placed with an initial placement based on predetermined placement rules.
26. The computer system of claim 23, the one or more programs further including instructions for: displaying, via the display generation component, a representation of an invited user that is not currently a participant in the communication session; in accordance with a determination that the communication session is a non-spatial communication session, causing a user of the computer system to be able to reposition the representation of the invited user that is not currently a participant; and in accordance with a determination that the communication session is a spatial communication session, forgoing causing the user of the computer system to be able to reposition the representation of the invited user that is not currently a participant.
27. The computer system of claim 1, the one or more programs further including instructions for: in response to changing between a spatial communication and a non-spatial communication session, displaying, via the display generation component, an indication that a mode of the communication session has changed.
28. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with a display generation component, the one or more programs including instructions for: while participating in a communication session that is a spatial communication session, wherein the spatial communication session includes displaying representations of a plurality of participants in the communication session in a 3D environment with a spatial distribution arrangement via the display generation component, wherein displaying the plurality of participants with the spatial distribution arrangement includes displaying: the representations of the plurality of participants spaced apart from each other and a user of the computer system in the 3D environment in a first non-vertical direction by at least a threshold amount; and the representations of the plurality of participants spaced apart from each other and the user in a second non-vertical direction different from the first non-vertical direction by at least the threshold amount; while displaying the representations of the plurality of participants distributed in the 3D environment, detecting an event; and in response to detecting the event, transitioning the communication session from the spatial communication session to a non-spatial communication session, the non-spatial communication session including displaying, via the display generation component, representations of at least a subset of the plurality of participants of the communication session in a grouped arrangement, wherein in the grouped arrangement: the representations of the plurality of participants are spaced apart from each other in the 3D environment in the first non-vertical direction by less than the threshold amount; a representation of a first participant in the grouped arrangement has a different positioning than a representation of the first participant in the spatial distributed arrangement; and a representation of a second participant in the grouped arrangement has a different positioning than a representation of the second participant in the spatial distributed arrangement.
29. A method, comprising: at a computer system in communication with a display generation component: while participating in a communication session as a spatial communication session, wherein the spatial communication session includes displaying, via the display generation component, representations of a plurality of participants in the communication session in a 3D environment in a spatial distributed arrangement, wherein displaying the plurality of participants in the spatial distributed arrangement includes displaying: the representations of the plurality of participants spaced apart from each other and a user of the computer system in a first non-vertical direction in the 3D environment by at least a threshold amount; and the representations of the plurality of participants spaced apart from each other and the user in a second non-vertical direction different from the first non-vertical direction by at least the threshold amount; while displaying the representations of the plurality of participants distributed in the 3D environment, detecting an event; and in response to detecting the event, transitioning the communication session from the spatial communication session to a non-spatial communication session, the non-spatial communication session including displaying, via the display generation component, representations of at least a subset of the plurality of participants of the communication session in a grouped arrangement, wherein in the grouped arrangement: the representations of the plurality of participants are spaced apart from each other in the 3D environment in the first non-vertical direction by less than the threshold amount; a representation of a first participant in the grouped arrangement has a different positioning than a representation of the first participant in the spatial distributed arrangement; and a representation of a second participant in the grouped arrangement has a different positioning than a representation of the second participant in the spatial distributed arrangement.
30. The non-transitory computer-readable storage medium of claim 28, wherein: in the non-spatial communication session: a representation of a first participant of the plurality of participants is in a first window region, and a representation of a second participant of the plurality of participants is in a second window region different from the first window region; and in the spatial communication session: the representation of the first participant of the plurality of participants is not in a window region, and the representation of the second participant of the plurality of participants is not in a window region. the representation of the second participant of the plurality of participants is not in a window region.
31. The non-transitory computer-readable storage medium of claim 28, wherein the event is a request received during the communication session to transition the representation of the user of the computer system from a 3D representation to a 2D representation.
32. The non-transitory computer-readable storage medium of claim 28, wherein the event is a request received during the communication session to transition the communication session from the spatial communication session to the non-spatial communication session.
33. The non-transitory computer-readable storage medium of claim 28, wherein the event is an additional participant joining the communication session.
34. The non-transitory computer-readable storage medium of claim 28, the one or more programs further including instructions to: when the communication session is a non-spatial communication session, offset positioning of respective window regions corresponding to respective participants based on movement of the respective participants.
35. The non-transitory computer-readable storage medium of claim 28, the one or more programs further including instructions to: while participating in the communication session as a non-spatial communication session, detect a second event; and in response to detecting the second event, transition the communication session from the non-spatial communication session to the spatial communication session.
36. The non-transitory computer-readable storage medium of claim 28, the one or more programs further including instructions to: while in a spatial communication session, display, via the display generation component, a self-view of a representation of the user of the computer system in a self-view window region.
37. The non-transitory computer-readable storage medium of claim 28, wherein: during a spatial communication session, enabling a first participant of the communication session to move a respective representation of the first participant, and enabling a second participant of the communication session to move a respective representation of the second participant; and during a non-spatial communication session, enabling a user of the computer system to move respective window regions that include respective representations of the plurality of participants of the communication session.
38. The non-transitory computer-readable storage medium of claim 28, the one or more programs further including instructions to: in response to changing between a spatial communication and a non-spatial communication session, display, via the display generation component, an indication that the mode of the communication session has changed.
39. The method of claim 29, wherein: in the non-spatial communication session: a representation of a first participant of the plurality of participants is in a first window region, and a representation of a second participant of the plurality of participants is in a second window region different from the first window region; and in the spatial communication session: the representation of the first participant of the plurality of participants is not in a window region, and the representation of the second participant of the plurality of participants is not in a window region.
40. The method of claim 29, wherein the event is a request received during the communication session to transition a representation of the user of the computer system from a 3D representation to a 2D representation.
41. The method of claim 29, wherein the event is a request received during the communication session to transition the communication session from the spatial communication session to the non-spatial communication session.
42. The method of claim 29, wherein the event is an additional participant joining the communication session.
43. The method of claim 29, the one or more programs further including instructions to: when the communication session is a non-spatial communication session, offset positioning of respective window regions corresponding to respective participants based on movement of the respective participants.
44. The method of claim 29, the one or more programs further including instructions to: while participating in the communication session as a non-spatial communication session, detect a second event; and in response to detecting the second event, transition the communication session from the non-spatial communication session to the spatial communication session.
45. The method of claim 29, the one or more programs further including instructions to: while in a spatial communication session, display, via the display generation component, a self-view of a representation of the user of the computer system in a self-view window region.
46. The method of claim 29, wherein: during a spatial communication session, enabling a first participant of the communication session to move a respective representation of the first participant, and enabling a second participant of the communication session to move a respective representation of the second participant; and during a non-spatial communication session, enabling a user of the computer system to move respective window regions including respective representations of the plurality of participants of the communication session.
47. The method of claim 29, the one or more programs further including instructions to: in response to changing between a spatial communication and a non-spatial communication session, display, via the display generation component, an indication that a mode of the communication session has changed.
Citation Information
Patent Citations
Method and device for group video session
CN108513088A
Virtual environments associated with processed video streams
EP4024854A1
Virtual environments associated with processed video streams
US11233974B1
Artificial Reality Spatial Interactions
US20220197403A1