User interface for managing live communication sessions
The system addresses inefficiencies in live communication sessions by using gaze and gesture detection, spatial transitions, and avatar editing to create a more intuitive and energy-efficient interface for augmented and mixed reality environments.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- APPLE INC
- Filing Date
- 2025-09-22
- Publication Date
- 2026-04-13
AI Technical Summary
Existing methods for managing live communication sessions in augmented and mixed reality environments are cumbersome, inefficient, and require excessive user input, leading to a cognitive burden and inefficient energy use, particularly in battery-powered devices.
The system employs intuitive interfaces that utilize gaze and gesture detection, spatial arrangement transitions, and avatar editing to streamline user interactions, reducing the need for explicit inputs and optimizing energy consumption.
The system enhances user interaction efficiency, reduces cognitive load, conserves energy, and extends battery life by minimizing unnecessary inputs and providing intelligent feedback, thus improving the overall user experience and device performance.
Smart Images

Figure 0007844727000001 
Figure 0007844727000002 
Figure 0007844727000003
Abstract
Description
Technical Field
[0001] (Cross - Reference to Related Applications) This application claims priority to U.S. Patent Application No. 18 / 367,418, titled "USER INTERFACES FOR MANAGING LIVE COMMUNICATION SESSIONS", filed on September 12, 2023; U.S. Provisional Patent Application No. 63 / 470,882, titled "USER INTERFACES FOR MANAGING LIVE COMMUNICATION SESSIONS", filed on June 3, 2023; and U.S. Provisional Patent Application No. 63 / 409,583, titled "USER INTERFACES FOR MANAGING LIVE COMMUNICATION SESSIONS", filed on September 23, 2022. The content of each of these applications is hereby incorporated by reference in its entirety.
[0002] The present disclosure generally relates to display - generating components, and optionally, computer systems that communicate with one or more sensors that provide computer - generated experiences, including but not limited to electronic devices that provide virtual reality and mixed - reality experiences via a display.
Background Art
[0003] The development of computer systems for augmented reality has advanced significantly in recent years. Exemplary augmented reality environments include at least some virtual elements that replace or enhance the physical world. Input devices such as cameras, controllers, joysticks, touch - sensitive surfaces, and touch - screen displays for computer systems and other electronic computing devices are used to interact with virtual / augmented reality environments. Exemplary virtual elements include virtual objects such as digital images, videos, text, icons, and control elements such as buttons and other graphics.
Summary of the Invention
[0004] Some methods and interfaces for managing live communication sessions, including those involving at least some virtual elements (e.g., applications, augmented reality environments, mixed reality environments, and virtual reality environments), are cumbersome, inefficient, and limited. For example, systems that provide insufficient control over performing actions associated with virtual objects, systems that require a series of inputs to achieve desired results within an augmented reality environment, and systems where manipulating virtual objects is complex, tedious, and error-prone impose a considerable cognitive burden on the user and detract from the virtual / augmented reality experience. In addition, these methods are unnecessarily time-consuming, thereby wasting the energy of the computer system. This latter consideration is particularly important in battery-powered devices.
[0005] Therefore, there is a need for computer systems with improved methods and interfaces for managing live communication sessions that are more efficient and intuitive for the user. Such methods and interfaces optionally complement or replace conventional methods for managing live communication sessions. Such methods and interfaces reduce the number, extent, and / or types of user input by helping the user understand the connection between the inputs provided and the device responses to those inputs, thereby generating a more efficient human-machine interface.
[0006] The above-mentioned drawbacks and other problems associated with the user interface of a computer system are mitigated or eliminated by the disclosed system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a notebook computer, tablet computer, or handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device such as a wristwatch or a head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also known as a “touchscreen” or “touchscreen display”). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, the computer system has one or more output devices in addition to display-generating components, the output devices include one or more tactile output generators and / or one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, memory, and one or more modules, programs, or instruction sets stored in memory for performing multiple functions. In some embodiments, the user interacts with the GUI (and / or computer system) through stylus and / or finger touch and gestures on a touch-sensitive surface, the movement of the user's eyes and hands in space relative to the user's body, and / or voice input captured by one or more audio input devices.In some embodiments, the functions performed through interaction optionally include image editing, drawing, presentation, word processing, spreadsheet creation, gameplay, making phone calls, video conferencing, sending emails, instant messaging, training support, digital photography, digital videography, web browsing, digital music playback, note-taking, and / or digital video playback. The executable instructions for performing those functions optionally reside in a primary computer-readable storage medium and / or a non-primary computer-readable storage medium, or in other computer program products configured to be executed by one or more processors.
[0007] There is a need for electronic devices with improved methods and interfaces for managing live communication sessions. Such methods and interfaces can complement or replace conventional methods for managing live communication sessions. Such methods and interfaces reduce the number, extent, and / or types of user input, resulting in a more efficient human-machine interface. In the case of battery-operated computing devices, such methods and interfaces conserve power and extend the interval between battery charges.
[0008] In some embodiments, the computer system displays a set of controls associated with controlling the playback of media content (e.g., transport controls and / or other types of controls) in response to detecting the user's gaze and / or gestures. In some embodiments, the computer system first displays a first set of controls in a reduced-sight state (e.g., with reduced visual prominence) in response to detecting a first input, and then displays a second set of controls (optionally including additional controls) in an increased-sight state in response to detecting a second input. In this way, the computer system optionally provides the user with feedback that the user has initiated the display of controls without excessively diverting the user's attention from the content (e.g., by initially displaying the controls in a visually inconspicuous manner), and then, based on the detection of user input indicating that the user wishes to interact with the controls further, displays the controls in a more visually prominent manner to enable easier and more precise interaction with the computer system.
[0009] Exemplary methods are described herein. Exemplary methods include, in a computer system communicating with a display generation component and one or more sensors, displaying representations of multiple users via the display generation component; receiving selections of representations of individual users from among the multiple users via one or more sensors; and, in response to receiving the selection of representations of individual users, displaying an option via the display generation component to invite the individual user to join an ongoing communication session, according to a determination that an ongoing communication session exists, and discontinuing the display of the option to invite the individual user to join an ongoing communication session, according to a determination that no ongoing communication session exists.
[0010] An exemplary method is a computer system communicating with a display generation component and one or more sensors, wherein the display generation component displays a communication user interface for communicating with other users in a real-time communication session, during which a user of the computer system is represented by an avatar that moves according to the movement of the user of the computer system detected by one or more sensors during the real-time communication session; while the communication user interface is being displayed, the display generation component displays a selectable user interface object; one or more inputs, including selection inputs directed to the selectable user interface object, are detected via one or more sensors; and in response to the detection of one or more inputs, including selection inputs directed to the selectable user interface object, the display generation component simultaneously displays an avatar editing user interface, which includes an avatar representing the user of the computer system and one or more options for modifying the appearance of the avatar representing the user of the computer system.
[0011] An exemplary method involves a computer system communicating with a display generation component, and while participating in a communication session which is a spatial communication session, displaying representations of multiple participants in a communication session in a spatially distributed arrangement in a 3D environment via the display generation component, which includes displaying representations of multiple participants spaced at least a threshold amount apart from each other and from the user of the computer system in a first non-vertical direction in the 3D environment, and representations of multiple participants spaced at least a threshold amount apart from each other and from the user in a second non-vertical direction different from the first non-vertical direction, and detecting events while displaying representations of multiple participants distributed in the 3D environment. A method comprising: performing a communication session from a spatial communication session to a non-spatial communication session in response to the detection of an event, wherein the transition includes displaying representations of at least a subset of the participants of the communication session in a grouped arrangement via a display generation component, in which the representations of the participants are spaced less than a threshold amount apart from each other in a first non-vertical direction in a 3D environment, the representation of the first participant in the grouped arrangement is in a different position than the representation of the first participant in a spatially distributed arrangement, and the representation of the second participant in the grouped arrangement is in a different position than the representation of the second participant in a spatially distributed arrangement.
[0012] An exemplary method includes, in a computer system communicating with a display generation component and one or more sensors, detecting eye-gaze input of a user of the computer system via one or more sensors during a communication session with one or more participants in a communication session; and, in response to the detection of eye-gaze input, displaying information about a first participant in the communication session via the display generation component according to a determination that the eye-gaze input satisfies one or more sets of eye-gaze criteria, and ceasing to display information about the first participant in the communication session according to a determination that the eye-gaze input does not satisfy one or more sets of eye-gaze criteria.
[0013] An exemplary non-temporary computer-readable storage medium is described herein. The exemplary non-temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generating component and one or more sensors, and includes instructions to display representations of multiple users via the display generating component, receive selections of representations of individual users among the multiple users via one or more sensors, and, in response to receiving selections of representations of individual users, display options via the display generating component to invite the individual user to join an ongoing communication session, according to a determination that an ongoing communication session exists, and to discontinue displaying options to invite the individual user to join an ongoing communication session, according to a determination that no ongoing communication session exists.
[0014] An exemplary non-temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more sensors, and includes instructions to display a communication user interface for communicating with other users in a real-time communication session, during which a user of the computer system is represented by an avatar that moves according to the movement of the user of the computer system detected by one or more sensors during the real-time communication session; while the communication user interface is being displayed, the display generation component displays selectable user interface objects; one or more inputs, including selection inputs directed to the selectable user interface objects, are detected via one or more sensors; and in response to the detection of one or more inputs, including selection inputs directed to the selectable user interface objects, the display generation component simultaneously displays an avatar representing the user of the computer system and one or more options for modifying the appearance of the avatar representing the user of the computer system.
[0015] An exemplary non-temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component, and while participating in a communication session which is a spatial communication session, the display generation component displays representations of multiple participants in a communication session in a spatially distributed arrangement in a 3D environment, which includes displaying representations of multiple participants separated from each other and from the user of the computer system by at least a threshold amount in a first non-vertical direction in the 3D environment, and representations of multiple participants separated from each other and from the user by at least a threshold amount in a second non-vertical direction different from the first non-vertical direction, and displays representations of multiple participants distributed in a 3D environment A non-temporary computer-readable storage medium includes instructions for detecting an event while displaying and, in response to detecting the event, transitioning a communication session from a spatial communication session to a non-spatial communication session, wherein the transition includes displaying representations of at least a subset of multiple participants of a communication session in a grouped arrangement via a display generating component, in which the representations of the multiple participants are spaced less than a threshold amount apart from each other in a first non-vertical direction in a 3D environment, the representation of the first participant in the grouped arrangement has a different position from the representation of the first participant in a spatially distributed arrangement, and the representation of the second participant in the grouped arrangement has a different position from the representation of the second participant in a spatially distributed arrangement.
[0016] An exemplary non-temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more sensors, and includes instructions for detecting eye-gaze input of a user of the computer system via one or more sensors during a communication session with one or more participants in a communication session, and, in response to detecting eye-gaze input, displaying information about a first participant in the communication session via the display generation component according to a determination that the eye-gaze input satisfies one or more sets of eye-gaze criteria, and ceasing to display information about the first participant in the communication session according to a determination that the eye-gaze input does not satisfy one or more sets of eye-gaze criteria.
[0017] An exemplary temporary computer-readable storage medium is described herein. The exemplary temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generating component and one or more sensors, and includes instructions to display representations of multiple users via the display generating component, receive selections of representations of individual users among the multiple users via one or more sensors, and, in response to receiving selections of representations of individual users, display options via the display generating component to invite the individual user to join an ongoing communication session, according to a determination that an ongoing communication session exists, and to deactivate the display of options to invite the individual user to join an ongoing communication session, according to a determination that no ongoing communication session exists.
[0018] An exemplary temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more sensors, and includes instructions to display a communication user interface for communicating with other users in a real-time communication session, during which a user of the computer system is represented by an avatar that moves according to the movement of the user of the computer system detected by one or more sensors during the real-time communication session; while the communication user interface is being displayed, the display generation component displays selectable user interface objects; one or more inputs, including selection inputs directed to the selectable user interface objects, are detected via one or more sensors; and in response to the detection of one or more inputs, including selection inputs directed to the selectable user interface objects, the display generation component simultaneously displays an avatar representing the user of the computer system and one or more options for modifying the appearance of the avatar representing the user of the computer system.
[0019] An exemplary temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component, and while participating in a communication session which is a spatial communication session, the display generation component displays representations of multiple participants in a communication session in a spatially distributed arrangement in a 3D environment, which includes displaying representations of multiple participants separated from each other and from the user of the computer system by at least a threshold amount in a first non-vertical direction in the 3D environment, and representations of multiple participants separated from each other and from the user by at least a threshold amount in a second non-vertical direction different from the first non-vertical direction, and displays representations of multiple participants distributed in a 3D environment A temporary computer-readable storage medium includes instructions for detecting an event while displaying and, in response to detecting the event, transitioning a communication session from a spatial communication session to a non-spatial communication session, wherein the transition includes displaying representations of at least a subset of multiple participants of a communication session in a grouped arrangement via a display generating component, in which the representations of the multiple participants are spaced less than a threshold amount apart from each other in a first non-vertical direction in a 3D environment, the representation of the first participant in the grouped arrangement has a different position from the representation of the first participant in a spatially distributed arrangement, and the representation of the second participant in the grouped arrangement has a different position from the representation of the second participant in a spatially distributed arrangement.
[0020] An exemplary temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more sensors, and includes instructions to detect eye-gaze input of a user of the computer system via one or more sensors during a communication session with one or more participants in a communication session, and, in response to the detection of eye-gaze input, to display information about a first participant in the communication session via the display generation component according to a determination that the eye-gaze input satisfies one or more sets of eye-gaze criteria, and to stop displaying information about a first participant in the communication session according to a determination that the eye-gaze input does not satisfy one or more sets of eye-gaze criteria.
[0021] An exemplary computer system is described herein. The exemplary computer system is configured to communicate with a display generating component and one or more sensors, and includes one or more processors and a memory for storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for displaying representations of multiple users via the display generating component, receiving selections of representations of individual users from among the multiple users via one or more sensors, and, in response to receiving selections of representations of individual users, displaying options via the display generating component to invite the individual user to join an ongoing communication session, according to a determination that an ongoing communication session exists, and ceasing to display options to invite the individual user to join an ongoing communication session, according to a determination that no ongoing communication session exists.
[0022] An exemplary computer system includes a display generation component and one or more sensors, and comprises one or more processors and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions to display a communication user interface for communicating with other users in a real-time communication session, during which a user of the computer system is represented by an avatar that moves according to the movement of the user of the computer system detected by one or more sensors during the real-time communication session; while the communication user interface is being displayed, the display generation component displays a selectable user interface object; one or more sensors detect one or more inputs, including selection inputs directed to the selectable user interface object; and in response to the detection of one or more inputs, including selection inputs directed to the selectable user interface object, the display generation component simultaneously displays an avatar representing the user of the computer system and one or more options for modifying the appearance of the avatar representing the user of the computer system.
[0023] An exemplary computer system is configured to communicate with a display generation component and includes one or more processors and memory that stores one or more programs configured to run by one or more processors, wherein one or more programs, while participating in a communication session which is a spatial communication session, display representations of multiple participants in a communication session in a spatially distributed arrangement in a 3D environment via the display generation component, which includes representations of multiple participants spaced at least a threshold amount apart from each other and from the user of the computer system in a first non-vertical direction in the 3D environment, and representations of multiple participants spaced at least a threshold amount apart from each other and from the user in a second non-vertical direction different from the first non-vertical direction, and displays the multiple participants in a spatially distributed arrangement, distributed within the 3D environment A computer system including instructions for detecting an event while displaying representations of multiple participants, and, in response to detecting the event, transitioning a communication session from a spatial communication session to a non-spatial communication session, wherein the transition includes displaying representations of at least a subset of the participants of the communication session in a grouped arrangement via a display generation component, in which the representations of the multiple participants are spaced less than a threshold amount apart from each other in a first non-vertical direction in a 3D environment, the representation of the first participant in the grouped arrangement has a different position from the representation of the first participant in a spatially distributed arrangement, and the representation of the second participant in the grouped arrangement has a different position from the representation of the second participant in a spatially distributed arrangement.
[0024] An exemplary computer system is configured to communicate with a display generation component and one or more sensors, and includes one or more processors and memory for storing one or more programs configured to be executed by the one or more processors, wherein one or more programs include instructions for detecting eye-gaze input of a user of the computer system via one or more sensors during a communication session with one or more participants in a communication session, and, in response to detecting eye-gaze input, displaying information about a first participant in the communication session via the display generation component according to a determination that the eye-gaze input satisfies one or more sets of eye-gaze criteria, and ceasing to display information about the first participant in the communication session according to a determination that the eye-gaze input does not satisfy one or more sets of eye-gaze criteria.
[0025] An exemplary computer system is configured to communicate with a display generation component and one or more sensors, and includes means for displaying representations of multiple users via the display generation component; means for receiving selections of representations of individual users from among the multiple users via one or more sensors; and means for displaying an option to invite the individual user to join an ongoing communication session via the display generation component, in accordance with a determination that an ongoing communication session exists, in response to receiving the selection of an individual user's representation, and for discontinuing the display of the option to invite the individual user to join an ongoing communication session, in accordance with a determination that no ongoing communication session exists.
[0026] An exemplary computer system is configured to communicate with a display generation component and one or more sensors, and includes means for displaying a communication user interface for communicating with other users in a real-time communication session, during which a user of the computer system is represented by an avatar that moves according to the movement of the user of the computer system detected by one or more sensors during the real-time communication session; means for displaying selectable user interface objects via the display generation component while the communication user interface is being displayed; means for detecting one or more inputs via one or more sensors, including selection inputs directed to the selectable user interface objects; and means for simultaneously displaying an avatar editing user interface via the display generation component in response to the detection of one or more inputs, including selection inputs directed to the selectable user interface objects, an avatar representing the user of the computer system, and one or more options for modifying the appearance of the avatar representing the user of the computer system.
[0027] An exemplary computer system is configured to communicate with a display generation component and, while participating in a communication session that is a spatial communication session, display, via the display generation component, representations of a plurality of participants in the communication session in a spatially distributed arrangement in a 3D environment, which includes displaying representations of a plurality of participants that are at least a threshold amount apart from each other and from a user of the computer system in a first non-vertical direction in the 3D environment, and representations of a plurality of participants that are at least a threshold amount apart from each other and from the user in a second non-vertical direction different from the first non-vertical direction, means for displaying the plurality of participants in a spatially distributed arrangement, means for detecting an event while displaying representations of the plurality of participants distributed within the 3D environment, and means for transitioning the communication session from a spatial communication session to a non-spatial communication session via the display generation component in response to detecting the event, wherein transitioning includes displaying representations of at least a subset of the plurality of participants in the communication session in a grouped arrangement, in which the representations of the plurality of participants are less than a threshold amount apart from each other in a first non-vertical direction in the 3D environment, the representation of a first participant in the grouped arrangement has a different position from the representation of the first participant in the spatially distributed arrangement, and the representation of a second participant in the grouped arrangement has a different position from the representation of the second participant in the spatially distributed arrangement.
[0028] An exemplary computer system is configured to communicate with a display generation component and one or more sensors and, during a communication session with one or more participants in the communication session, detect, via the one or more sensors, a line-of-sight input of a user of the computer system, and, in response to detecting the line-of-sight input, display, via the display generation component, information regarding a first participant in the communication session according to a determination that the line-of-sight input meets a set of one or more line-of-sight criteria, and cease displaying information regarding the first participant in the communication session according to a determination that the line-of-sight input does not meet a set of one or more line-of-sight criteria.
[0029] An exemplary computer program product is described herein. The exemplary computer program product includes one or more programs configured to run on one or more processors of a computer system communicating with a display generating component and one or more sensors, the one or more programs including instructions to display representations of multiple users via the display generating component, receive selections of representations of individual users from among the multiple users via one or more sensors, and, in response to receiving selections of representations of individual users, display options via the display generating component to invite the individual user to join an ongoing communication session, according to a determination that an ongoing communication session exists, and to discontinue displaying options to invite the individual user to join an ongoing communication session, according to a determination that no ongoing communication session exists.
[0030] An exemplary computer program product includes one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more sensors. The one or more programs, via the display generation component, display a communication user interface for communicating with other users in a real-time communication session, where, during the real-time communication session, the user of the computer system is represented by an avatar that moves according to the movement of the user of the computer system detected by the one or more sensors. While displaying the communication user interface, via the display generation component, display selectable user interface objects, detect one or more inputs including a selection input directed to a selectable user interface object via the one or more sensors, and in response to detecting the one or more inputs including a selection input directed to a selectable user interface object, simultaneously display, via the display generation component, an avatar representing the user of the computer system and one or more options for modifying the appearance of the avatar representing the user of the computer system, including instructions for an avatar editing user interface.
[0031] An exemplary computer program product includes one or more programs configured to run on one or more processors of a computer system that communicate with a display generation component, and while participating in a communication session which is a spatial communication session, the one or more programs, via the display generation component, display representations of multiple participants in a communication session in a spatially distributed arrangement in a 3D environment, which includes displaying representations of multiple participants separated from each other and from the user of the computer system by at least a threshold amount in a first non-vertical direction in the 3D environment, and representations of multiple participants separated from each other and from the user by at least a threshold amount in a second non-vertical direction different from the first non-vertical direction, and displaying multiple participants in a spatially distributed arrangement in the 3D environment A computer program product comprising instructions for detecting an event while displaying representations of participants, and, in response to the detection of the event, transitioning a communication session from a spatial communication session to a non-spatial communication session, wherein the transition includes displaying representations of at least a subset of multiple participants of the communication session in a grouped arrangement via a display generation component, in which the representations of the multiple participants are spaced less than a threshold amount apart from each other in a first non-vertical direction in a 3D environment, the representation of the first participant in the grouped arrangement having a different position from the representation of the first participant in a spatially distributed arrangement, and the representation of the second participant in the grouped arrangement having a different position from the representation of the second participant in a spatially distributed arrangement.
[0032] An exemplary computer program product includes one or more programs configured to run by one or more processors of a computer system communicating with a display generation component and one or more sensors, the one or more programs including instructions to detect eye-gaze input of a user of the computer system via one or more sensors during a communication session with one or more participants in a communication session, and, in response to the detection of eye-gaze input, to display information about a first participant in the communication session via the display generation component according to a determination that the eye-gaze input satisfies one or more sets of eye-gaze criteria, and to stop displaying information about a first participant in the communication session according to a determination that the eye-gaze input does not satisfy one or more sets of eye-gaze criteria.
[0033] It should be noted that the various embodiments described herein can be combined with any other embodiments described herein. The features and advantages described herein are not exhaustive, and many additional features and advantages will become apparent to those skilled in the art, in particular, in light of the drawings, specification and claims. Furthermore, it should be noted that the language used herein has been selected solely for readability and explanatory purposes and not to define or limit the subject matter of the invention.
[0034] To better understand the various embodiments described, the following “Modes for Carrying Out the Invention” should be referenced in conjunction with the following drawings, and similar reference numbers throughout the following drawings refer to the corresponding parts. [Brief explanation of the drawing]
[0035] [Figure 1A] This is a block diagram illustrating the operating environment of a computer system for providing an XR experience, according to several embodiments.
[0036] [Figure 1B] This is an example of a computer system for providing an XR experience in the operating environment shown in Figure 1A. [Figure 1C]This is an example of a computer system for providing an XR experience in the operating environment shown in Figure 1A. [Figure 1D] This is an example of a computer system for providing an XR experience in the operating environment shown in Figure 1A. [Figure 1E] This is an example of a computer system for providing an XR experience in the operating environment shown in Figure 1A. [Figure 1F] This is an example of a computer system for providing an XR experience in the operating environment shown in Figure 1A. [Figure 1G] This is an example of a computer system for providing an XR experience in the operating environment shown in Figure 1A. [Figure 1H] This is an example of a computer system for providing an XR experience in the operating environment shown in Figure 1A. [Figure 1I] This is an example of a computer system for providing an XR experience in the operating environment shown in Figure 1A. [Figure 1J] This is an example of a computer system for providing an XR experience in the operating environment shown in Figure 1A. [Figure 1K] This is an example of a computer system for providing an XR experience in the operating environment shown in Figure 1A. [Figure 1L] This is an example of a computer system for providing an XR experience in the operating environment shown in Figure 1A. [Figure 1M] This is an example of a computer system for providing an XR experience in the operating environment shown in Figure 1A. [Figure 1N] This is an example of a computer system for providing an XR experience in the operating environment shown in Figure 1A. [Figure 10] This is an example of a computer system for providing an XR experience in the operating environment shown in Figure 1A. [Figure 1P] This is an example of a computer system for providing an XR experience in the operating environment shown in Figure 1A.
[0037] [Figure 2]Block diagram showing a controller for a computer system configured to manage and adjust the XR experience for a user, according to several embodiments.
[0038] [Figure 3] This block diagram shows display generation components of a computer system configured to provide users with visual components of an XR experience, according to several embodiments.
[0039] [Figure 4] This is a block diagram showing a hand tracking unit for a computer system configured to capture user gesture input, according to several embodiments.
[0040] [Figure 5] This is a block diagram showing an eye-tracking unit for a computer system configured to capture user eye-gaze input, according to several embodiments.
[0041] [Figure 6] Flowcharts illustrating glint-assisted eye-tracking pipelines in several embodiments.
[0042] [Figure 7A] This document presents exemplary techniques for managing live communication sessions in several embodiments. [Figure 7B] This document presents exemplary techniques for managing live communication sessions in several embodiments. [Figure 7C1] This document presents exemplary techniques for managing live communication sessions in several embodiments. [Figure 7C2] This document presents exemplary techniques for managing live communication sessions in several embodiments. [Figure 7D] This document presents exemplary techniques for managing live communication sessions in several embodiments. [Figure 7E]This document presents exemplary techniques for managing live communication sessions in several embodiments. [Figure 7F] This document presents exemplary techniques for managing live communication sessions in several embodiments. [Figure 7G] This document presents exemplary techniques for managing live communication sessions in several embodiments. [Figure 7H] This document presents exemplary techniques for managing live communication sessions in several embodiments. [Figure 7I] This document presents exemplary techniques for managing live communication sessions in several embodiments. [Figure 7J] This document presents exemplary techniques for managing live communication sessions in several embodiments. [Figure 7K] This document presents exemplary techniques for managing live communication sessions in several embodiments. [Figure 7L1] This document presents exemplary techniques for managing live communication sessions in several embodiments. [Figure 7L2] This document presents exemplary techniques for managing live communication sessions in several embodiments. [Figure 7M] This document presents exemplary techniques for managing live communication sessions in several embodiments. [Figure 7N] This document presents exemplary techniques for managing live communication sessions in several embodiments. [Figure 7O] This document presents exemplary techniques for managing live communication sessions in several embodiments. [Figure 7P] This document presents exemplary techniques for managing live communication sessions in several embodiments. [Figure 7Q] This document presents exemplary techniques for managing live communication sessions in several embodiments.
[0043] [Figure 8]This is a flowchart illustrating methods for managing live communication sessions using various embodiments.
[0044] [Figure 9] This is a flowchart illustrating various methods for providing avatars in live communication sessions.
[0045] [Figure 10A] This document presents exemplary techniques for providing representations in live communication sessions, based on several embodiments. [Figure 10B] This document presents exemplary techniques for providing representations in live communication sessions, based on several embodiments. [Figure 10C] This document presents exemplary techniques for providing representations in live communication sessions, based on several embodiments. [Figure 10D1] This document presents exemplary techniques for providing representations in live communication sessions, based on several embodiments. [Figure 10D2] This document presents exemplary techniques for providing representations in live communication sessions, based on several embodiments. [Figure 10E] This document presents exemplary techniques for providing representations in live communication sessions, based on several embodiments.
[0046] [Figure 11] This is a flowchart illustrating various methods for providing representation in live communication sessions.
[0047] [Figure 12A] This document presents exemplary techniques for providing information in a live communication session, based on several embodiments. [Figure 12B1] This document presents exemplary techniques for providing information in a live communication session, based on several embodiments. [Figure 12B2] This document presents exemplary techniques for providing information in a live communication session, based on several embodiments. [Figure 12C] This document presents exemplary techniques for providing information in a live communication session, based on several embodiments. [Figure 12D] This document presents exemplary techniques for providing information in a live communication session, based on several embodiments. [Figure 12E] This document presents exemplary techniques for providing information in a live communication session, based on several embodiments. [Figure 12F] This document presents exemplary techniques for providing information in a live communication session, based on several embodiments.
[0048] [Figure 13] This is a flowchart illustrating various methods for providing information in a live communication session. [Modes for carrying out the invention]
[0049] This disclosure relates to user interfaces that provide users with Extended Reality (XR) experiences, in several embodiments.
[0050] The systems, methods, and GUIs described herein improve user interface interactions with virtual / augmented reality environments in multiple ways.
[0051] In some embodiments, the computer system enables live communication between users. The computer system displays representations of multiple users and receives selections of representations of individual users from among the multiple users. In response to receiving selections of representations of individual users, and in accordance with the determination that an ongoing communication session exists, the computer system displays an option to invite the individual user to join the ongoing communication session; and, in accordance with the determination that no ongoing communication session exists, the computer system discontinues displaying the option to invite the individual user to join the ongoing communication session.
[0052] In some embodiments, the computer system provides the user with options to modify the appearance of their avatar. The computer system displays a communication user interface for communicating with other users in a real-time communication session. During a real-time communication session, the user is represented by an avatar that moves during the real-time communication session in accordance with the user's movement in the computer system. While displaying the communication user interface, the computer system simultaneously displays selectable user interface objects. While simultaneously displaying the communication user interface and the selectable user interface objects, the computer system detects one or more inputs, including selection inputs directed to the selectable user interface objects. In response to detecting one or more inputs, including selection inputs directed to the selectable user interface objects, the computer system simultaneously displays an avatar editing user interface, which includes an avatar representing the user of the computer system and one or more options for modifying the appearance of the avatar representing the user of the computer system.
[0053] In some embodiments, the computer system switches between spatial communication sessions and non-spatial communication sessions. While participating in a communication session, which is a spatial communication session including the computer system, the computer system displays representations of multiple participants in the communication session in a spatially distributed arrangement in a 3D environment. Displaying multiple participants in a spatially distributed arrangement includes displaying representations of multiple participants spaced at least a threshold amount apart from each other and from the user in a first non-vertical direction in the 3D environment, and representations of multiple participants spaced at least a threshold amount apart from each other and from the user in a second non-vertical direction different from the first non-vertical direction. While displaying representations of multiple participants distributed in the 3D environment, the computer system detects an event, and in response to detecting an event, the computer system transitions the communication session from a spatial communication session to a non-spatial communication session. Transitioning to a non-spatial communication session includes displaying representations of at least a subset of the multiple participants in the communication session in a grouped arrangement. In a grouped arrangement, the representations of multiple participants are spaced less than a threshold amount apart from each other in a first non-vertical direction within the 3D environment, the representation of the first participant in a grouped arrangement has a different position than the representation of the first participant in a spatially dispersed arrangement, and the representation of the second participant in a grouped arrangement has a different position than the representation of the second participant in a spatially dispersed arrangement.
[0054] In some embodiments, the computer system provides information during a live communication session based on the user's gaze. During a communication session with one or more participants in the communication session, the computer system detects the user's gaze input. In response to the detection of gaze input, the computer system displays information about the first participant in the communication session according to the determination that the gaze input satisfies one or more sets of gaze criteria, and stops displaying information about the first participant in the communication session according to the determination that the gaze input does not satisfy one or more sets of gaze criteria.
[0055] In some embodiments, the computer system displays content in a first area of the user interface. In some embodiments, while the computer system is displaying content and a first set of controls is not displayed in a first state, the computer system detects a first input from a first part of the user. In some embodiments, in response to the detection of the first input and according to the determination that the user's gaze was directed towards a second area of the user interface when the first input was detected, the computer system displays a first set of one or more controls in the user interface in a first state, and according to the determination that the user's gaze was not directed towards a second area of the user interface when the first input was detected, the computer system discontinues displaying the first set of one or more controls in the first state.
[0056] In some embodiments, the computer system displays content on a user interface. In some embodiments, while displaying content, the computer system detects a first input based on movement of a first part of the user of the computer system. In some embodiments, in response to detecting the first input, the computer system displays a first set of one or more controls within the user interface, the first set of one or more controls being displayed in a first state and within a first area of the user interface. In some embodiments, while the first set of one or more controls is displayed in a first state, the computer system transitions from displaying the first set of one or more controls in a first state to displaying a second set of one or more controls in a second state, based on a determination that one or more first criteria are met, including criteria that are met when the user's attention is directed to a first area of the user interface, based on movement of a second part of the user different from the first part of the user.
[0057] Figures 1A to 6 provide a description of an exemplary computer system for providing an XR experience to a user. Figures 7A to 7Q show exemplary techniques for managing a live communication session in several embodiments. Figure 8 is a flowchart of a method for managing a live communication session in various embodiments. Figure 9 is a flowchart of a method for providing an avatar in a live communication session in various embodiments. The user interfaces in Figures 7A to 7Q are used to illustrate the processes in Figures 8 and 9. Figures 10A to 10E show exemplary techniques for providing representations in a live communication session in several embodiments. Figure 11 is a flowchart of a method for providing representations in a live communication session in various embodiments. The user interfaces in Figures 10A to 10E are used to illustrate the process in Figure 11. Figures 12A to 12F show exemplary techniques for providing information in a live communication session in several embodiments. Figure 13 is a flowchart of a method for providing information in a live communication session in various embodiments. The user interfaces in Figures 10A to 10E are used to illustrate the process in Figure 11.
[0058] The processes described below enhance the usability of the device and make the user device interface more efficient (for example, by helping the user provide appropriate input and reducing user errors when operating / interacting with the device) through various technologies, including providing the user with improved visual feedback, reducing the number of inputs required to perform actions, providing additional control options without cluttering the user interface with additional displayed controls, performing actions without requiring further user input when a set of conditions is met, improving privacy and / or security, providing a more diverse, detailed, and / or realistic user experience while saving memory space, and / or additional technologies. These technologies also reduce power consumption and improve the battery life of the device by enabling the user to use the device more quickly and efficiently. Saving battery power, and therefore weight, improves the ergonomics of the device. These technologies also enable real-time communication, allow the use of fewer and / or less accurate sensors, resulting in more compact, lighter, and less expensive devices, and enabling the device to be used in a variety of lighting conditions. These technologies reduce energy consumption and thereby reduce the heat emitted by the device, which is especially important for wearable devices that can become uncomfortable for the user to wear if they generate excessive heat, even if the device is well within the operating parameters for its components.
[0059] Furthermore, in any method described herein that is conditional on one or more conditions being met in one or more steps, it should be understood that the method described can be repeated in multiple iterations such that all the conditions that the steps of the method are conditional on are met in different iterations of the method. For example, if a method requires that a first step be performed if a condition is met, and a second step be performed if the condition is not met, a person skilled in the art will understand that the steps described in the claim are repeated in a specific order until the conditions are met and then not met. Thus, a method described in one or more steps that depends on one or more conditions being met can be rewritten as a method that is repeated until each of the conditions described in the method is met. However, this is not required for a claim of a system or computer-readable medium that includes instructions for performing a conditional operation based on the satisfaction of the corresponding one or more conditions, and thus can determine whether a contingency has been met without explicitly repeating the steps of the method until all the conditions that the steps of the method are conditional on are met. Those skilled in the art will also understand that, as with a method having conditional steps, a system or computer-readable storage medium may repeat the steps of the method as many times as necessary to ensure that all of the conditional steps have been performed.
[0060] In some embodiments, as shown in Figure 1A, the XR experience is provided to the user via an operating environment 100 which includes a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or remote server), display generation components 120 (e.g., a head-mounted device (HMD), a display, a projector, a touchscreen, etc.), one or more input devices 125 (e.g., an eye-tracking device 130, a hand-tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a tactile output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a tactile sensor, an orientation sensor, a proximity sensor, a temperature sensor, a location sensor, a motion sensor, a velocity sensor, etc.), and optionally one or more peripheral devices 195 (e.g., home appliances, wearable devices, etc.). In some embodiments, one or more of the input device 125, output device 155, sensor 190, and peripheral device 195 are integrated with the display generation component 120 (for example, within a head-mounted device or handheld device).
[0061] When describing an XR experience, various terms are used to refer individually to several related but distinct environments that the user can perceive and / or interact with (for example, using inputs detected by the computer system 101 that generates the XR experience, causing the computer system generating the XR experience to generate audio, visual, and / or haptic feedback corresponding to various inputs provided to the computer system 101). The following is a subset of these terms.
[0062] Physical Environment: The physical environment refers to the physical world that people can perceive and / or interact with without the help of electronic systems. Examples of physical environments, such as a physical park, include physical objects such as physical trees, physical buildings, and physical people. People can directly perceive and / or interact with the physical environment through their senses of sight, touch, hearing, taste, and smell.
[0063] Extended reality: In contrast, an extended reality (XR) environment refers to a fully or partially simulated environment that people perceive and / or interact with through electronic systems. In XR, a subset of a person's bodily movements or their representation is tracked, and accordingly, one or more properties of one or more virtual objects simulated within the XR environment are adjusted to behave according to at least one law of physics. For example, an XR system may detect a person's head rotation and, accordingly, adjust the graphical content and sound field presented to the person in a similar way to how such views and sounds would change in a physical environment. In some situations (e.g., for reasons of accessibility), adjustments to the properties of one or more virtual objects within an XR environment may be made in response to a representation of physical movement (e.g., a voice command). A person may perceive and / or interact with an XR object using any one of their senses, including sight, sound, touch, taste, and smell. For example, a person can perceive and / or interact with audio objects that create a 3D or spatial audio environment, providing the perception of point audio sources in 3D space. In another example, audio objects may enable audio transparency, selectively incorporating ambient sounds from the physical environment, with or without computer-generated audio. In some XR environments, a person may perceive and / or interact with only audio objects.
[0064] Examples of XR include virtual reality and mixed reality.
[0065] Virtual reality: A virtual reality (VR) environment refers to a simulated environment designed to be entirely based on computer-generated sensory input for one or more senses. A VR environment includes multiple virtual objects that a person can perceive and / or interact with. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can perceive and / or interact with virtual objects in a VR environment through a simulation of their presence within the computer-generated environment and / or through a simulation of a subset of their physical movement within the computer-generated environment.
[0066] Mixed Reality: A mixed reality (MR) environment is a simulated environment designed to incorporate sensory input or its representation from a physical environment, in addition to including computer-generated sensory input (e.g., virtual objects), in contrast to a virtual reality (VR) environment designed to rely entirely on computer-generated sensory input. On a virtual continuum, a mixed reality environment is any location between, but not encompassing, the complete physical environment at one end and the virtual reality environment at the other. In some MR environments, computer-generated sensory input may respond to changes in sensory input from the physical environment. Also, some electronic systems for presenting an MR environment may track location and / or orientation relative to the physical environment to enable virtual objects to interact with real objects (i.e., physical articles or their representations from the physical environment). For example, the system may take movement into account so that a virtual tree appears stationary relative to the physical ground.
[0067] Examples of mixed reality include augmented reality and augmented virtual reality.
[0068] Augmented Reality: An augmented reality (AR) environment refers to a simulated environment in which one or more virtual objects are superimposed on or onto a physical environment. For example, an electronic system for presenting an AR environment may have a transparent or translucent display that allows a person to directly view the physical environment. The system may also be configured to present virtual objects on the transparent or translucent display, thereby allowing a person to use the system to perceive the virtual objects superimposed on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture an image or video of the physical environment, which is a representation of the physical environment. The system composites the image or video with the virtual objects and presents the composite on the opaque display. A person uses this system to perceive the virtual objects superimposed on the physical environment by indirectly viewing the physical environment through the image or video of the physical environment. As used herein, a video of the physical environment displayed on an opaque display is referred to as “pass-through video,” meaning that the system uses one or more image sensors to capture images of the physical environment and uses those images when presenting the AR environment on the opaque display. Alternatively, the system may have a projection system that projects virtual objects, for example, as holograms, into or onto the physical environment, so that a person can use the system to perceive the virtual objects superimposed on the physical environment. An augmented reality environment also refers to a simulated environment in which the representation of the physical environment is transformed by computer-generated sensory information. For example, when providing pass-through video, the system may transform one or more sensor images to plane a selected perspective (e.g., viewpoint) different from the perspective captured by the imaging sensor. As another example, the representation of the physical environment may be transformed by graphically modifying (e.g., enlarging) a portion of it, so that the modified portion is a non-photorealistic altered version of the original captured image. As yet another example, the representation of the physical environment may be transformed by graphically removing or obscuring a portion of it.
[0069] Augmented Virtuality (AV) refers to a simulated environment in which a virtual or computer-generated environment incorporates one or more sensory inputs from a physical environment. These sensory inputs may be representations of one or more characteristics of the physical environment. For example, an AV park might have virtual trees and buildings, but people with faces might be realistically reproduced from images of real people. Another example is that a virtual object might adopt the shape or color of a physical article captured by one or more imaging sensors. A further example is that a virtual object might adopt shadows that correspond to the position of the sun in the physical environment.
[0070] In augmented reality, mixed reality, or virtual reality environments, a view of a three-dimensional environment is visible to the user. Typically, the view of the three-dimensional environment is visible to the user through one or more display-generating components (e.g., a display or a pair of display modules providing stereoscopic content to different eyes of the same user) via a virtual viewport having a viewport boundary that defines the extent of the three-dimensional environment visible to the user through one or more display-generating components. In some embodiments, the area defined by the viewport boundary is smaller than the user's field of view in one or more dimensions (e.g., based on the user's field of view, the size, optical properties, or other physical characteristics of one or more display-generating components, and / or the location and / or orientation of one or more display-generating components relative to the user's eyes). In some embodiments, the area defined by the viewport boundary is larger than the user's field of view in one or more dimensions (e.g., based on the user's field of view, the size, optical properties, or other physical characteristics of one or more display-generating components, and / or the location and / or orientation of one or more display-generating components relative to the user's eyes). Viewports and viewport boundaries typically move as one or more display-generating components move (for example, with the user's head in the case of a head-mounted device, or with the user's hand in the case of a handheld device such as a tablet or smartphone). The user's viewpoint determines which content is visible within the viewport, and the viewpoint generally specifies the location and orientation of the three-dimensional environment, so that as the viewpoint shifts, the view of the three-dimensional environment also shifts within the viewport. In the case of head-mounted devices, the viewpoint is typically based on the location and orientation of the user's head, face, and / or eyes to provide a view of the three-dimensional environment that is perceptually accurate and provides an immersive experience when the user is using the head-mounted device.For handheld or stationary devices, the viewpoint shifts as the handheld or stationary device moves and / or as the user's position relative to the handheld or stationary device changes (for example, as the user moves toward or away from the device, above or below the device, to the right of the device, and / or to the left of the device). In devices that include display-generating components with virtual passthrough, the portion of the physical environment visible (e.g., displayed and / or projected) through one or more display-generating components moves as the user's viewpoint moves as the field of view of one or more cameras moves (and the appearance of one or more virtual objects displayed through one or more display-generating components is updated based on the user's viewpoint (e.g., the displayed position and orientation of the virtual objects are updated based on the user's viewpoint) and therefore typically moves with the display-generating components (e.g., moves with the user's head in a head-mounted device, or moves with the user's hand in a handheld device such as a tablet or smartphone), communicating with the display-generating components. Based on the field of view of one or more cameras. In the case of a display-generating component with optical passthrough, parts of the physical environment that are visible through one or more display-generating components (e.g., optically visible through one or more partially or completely transparent parts of the display-generating component) are based on the user's field of view through the partially or completely transparent parts of the display-generating component (e.g., moving with the user's head in the case of a head-mounted device, or moving with the user's hand in the case of a handheld device such as a tablet or smartphone), because the user's viewpoint moves as the user's field of view moves through the partially or completely transparent parts of the display-generating component (one or more), and the appearance of one or more virtual objects is updated based on the user's viewpoint.
[0071] In some embodiments, a representation of the physical environment (e.g., displayed via virtual or optical passthrough) can be partially or completely obscured by the virtual environment. In some embodiments, the amount of virtual environment displayed (e.g., the amount of physical environment not displayed) is based on the level of immersion of the virtual environment (e.g., relative to the representation of the physical environment). For example, increasing the immersion level optionally displays more virtual environment and replaces and / or obscures more of the physical environment, while decreasing the immersion level optionally displays less virtual environment and reveals portions of the physical environment that were not previously displayed and / or obscured. In some embodiments, at a certain level of immersion, one or more first background objects (e.g., in a representation of the physical environment) are visually less emphasized than one or more second background objects (e.g., dimmed, blurred, and / or displayed with increased transparency), and one or more third background objects are discontinued. In some embodiments, the immersion level includes the relevant degree to which the virtual content displayed by the computer system (e.g., a virtual environment and / or virtual content) obscures the background content surrounding / behind the virtual content (e.g., content other than the virtual environment and / or virtual content), and optionally includes the number of items of the background content displayed and / or the visual characteristics of the background content on which it is displayed (e.g., color, contrast, and / or opacity), the angular range of the virtual content displayed through the display-generating components (e.g., 60-degree content displayed at low immersion, 120-degree content displayed at medium immersion, or 180-degree content displayed at high immersion), and / or the percentage of the field of view displayed through the display-generating components consumed by the virtual content (e.g., 33% of the field of view consumed by the virtual content at low immersion, 66% of the field of view consumed by the virtual content at medium immersion, or 100% of the field of view consumed by the virtual content at high immersion). In some embodiments, the background content is included in the background on which the virtual content is displayed (e.g., background content within a representation of a physical environment).In some embodiments, background content includes a user interface (e.g., a user interface generated by a computer system corresponding to the application), virtual objects not associated with or included in the virtual environment and / or virtual content (e.g., files or representations of other users generated by the computer system), and / or real objects (e.g., pass-through objects representing real objects in the physical environment around the user, which are visible so as to be displayed through the display generation components and / or are visible through transparent or translucent components of the display generation components so as not to obscure / hinder their visibility through the display generation components by the computer system). In some embodiments, at low immersion levels (e.g., a first immersion level), the background, virtual and / or real objects are displayed in a non-obscuring manner. For example, a low-immersion virtual environment is optionally displayed simultaneously with the background content, and the background content is optionally displayed with full brightness, color, and / or translucency. In some embodiments, at higher immersion levels (e.g., a second immersion level higher than a first immersion level), backgrounds, virtual and / or real objects are displayed in an obscured manner (e.g., dimmed, blurred, or removed from the display). For example, a separate virtual environment at a high immersion level is displayed without simultaneously displaying background content (e.g., in full-screen or fully immersive mode). As another example, a virtual environment displayed at an intermediate immersion level is displayed simultaneously with dimmed, blurred, or otherwise de-emphasized background content. In some embodiments, the visual characteristics of background objects differ among them. For example, at a particular immersion level, one or more first background objects are visually de-emphasized more than one or more second background objects (e.g., dimmed, blurred, and / or displayed with increased transparency), and one or more third background objects are not displayed at all.In some embodiments, a null or zero level of immersion corresponds to the discontinuation of the display of the virtual environment, and instead, the representation of the physical environment is displayed (optionally together with one or more virtual objects such as applications, windows, or virtual three-dimensional objects) without the representation of the physical environment being obscured by the virtual environment. Adjusting the level of immersion using physical input elements provides a quick and efficient way to adjust immersion, improving the usability of the computer system and making the user device interface more efficient.
[0072] Viewpoint-locked virtual objects: A virtual object is viewpoint-locked when the computer system displays the virtual object in the same location and / or position within the user's view, even if the user's view shifts (e.g., changes). In embodiments where the computer system is a head-mounted device, the user's view is locked in the forward direction of the user's head (e.g., the user's view is at least a portion of the user's field of vision when the user is looking straight ahead). Thus, the user's view remains fixed even if the user's gaze moves, without moving the user's head. In embodiments where the computer system has a display generation component (e.g., a display screen) that can be repositioned relative to the user's head, the user's view is the augmented reality view presented to the user on the display generation component of the computer system. For example, a viewpoint-locked virtual object displayed in the upper-left corner of the user's view when the user's view is in a first orientation (e.g., the user's head is facing north) will continue to be displayed in the upper-left corner of the user's view even if the user's view changes to a second orientation (e.g., the user's head is facing west). In other words, the location and / or position in which a viewpoint-locked virtual object is displayed from the user's viewpoint is independent of the user's position and / or orientation in the physical environment. In embodiments where the computer system is a head-mounted device, the user's viewpoint is locked to the orientation of the user's head, so that the virtual object is also referred to as a "head-locked virtual object."
[0073] Environment-Locked Virtual Objects: A virtual object is environment-locked (or "world-locked") when a computer system displays it at a location and / or position in the user's viewpoint that is based on (e.g., selected by reference to and / or fixed to) a location and / or object in a three-dimensional environment (e.g., a physical or virtual environment). As the user's viewpoint shifts, the location and / or object in the environment relative to the user's viewpoint changes, and as a result, the environment-locked virtual object will appear at a different location and / or position in the user's viewpoint. For example, an environment-locked virtual object locked to a tree directly in front of the user will appear at the center of the user's viewpoint. If the user's viewpoint shifts to the right (e.g., the user's head is turned to the right) and the tree becomes left-leaning in the user's viewpoint (e.g., the tree's position in the user's viewpoint shifts), the environment-locked virtual object locked to the tree will appear left-leaning in the user's viewpoint. In other words, the location and / or position in which an environment-locked virtual object is displayed in the user's viewpoint depends on the location and / or object's position and / or orientation in the environment to which the virtual object is locked. In some embodiments, the computer system uses a stationary reference frame (e.g., a fixed location in the physical environment and / or a coordinate system fixed to an object) to determine the position in which the environment-locked virtual object is displayed from the user's viewpoint. The environment-locked virtual object can be locked to a stationary part of the environment (e.g., a floor, wall, table, or other stationary object) or to a moving part of the environment (e.g., a vehicle, animal, person, or a representation of a part of the user's body that moves independently of the user's viewpoint, such as the user's hands, wrists, arms, or feet), so that the virtual object moves as the viewpoint or the part of the environment moves in order to maintain a fixed relationship between the virtual object and the part of the environment.
[0074] In some embodiments, an environment-locked or viewpoint-locked virtual object exhibits delayed tracking behavior, reducing or delaying its movement in response to the movement of a reference point that the virtual object is following. In some embodiments, when exhibiting delayed tracking behavior, the computer system detects movement of the reference point that the virtual object is following (e.g., a part of the environment, a viewpoint, or a point fixed to the viewpoint, such as a point between 5 and 300 cm from the viewpoint) and intentionally delays the movement of the virtual object. For example, when the reference point (e.g., a part of the environment or the viewpoint) moves at a first velocity, the virtual object is moved by the device so as to remain locked to the reference point, but at a second velocity slower than the first velocity (e.g., the virtual object begins to catch up to the reference point until the reference point stops or slows down). In some embodiments, when a virtual object exhibits delayed tracking behavior, the device ignores small movements of the reference point (e.g., ignoring movements of the reference point that are below a threshold movement amount, such as a movement of 0 to 5 degrees or a movement of 0 to 50 cm). For example, when the reference point (e.g., the part of the environment or viewpoint from which the virtual object is locked) moves by a first amount, the distance between the reference point and the virtual object increases (e.g., because the virtual object is displayed to maintain a fixed or substantially fixed position relative to a different viewpoint or part of the environment from which the virtual object is locked), and when the reference point (e.g., the part of the environment or viewpoint from which the virtual object is locked) moves by a second amount greater than the first amount, the distance between the reference point and the virtual object first increases (e.g., because the virtual object is displayed to maintain a fixed or substantially fixed position relative to a different viewpoint or part of the environment from which the virtual object is locked), and then decreases as the amount of movement of the reference point increases beyond a threshold (e.g., a "delayed tracking" threshold) as the virtual object is moved by the computer system to maintain a fixed or substantially fixed position relative to the reference point.In some embodiments, a virtual object that maintains a substantially fixed position with respect to a reference point includes the virtual object being displayed within a threshold distance (e.g., 1, 2, 3, 5, 15, 20, 50 cm) of the reference point in one or more dimensions (e.g., above / below, left / right, and / or forward / behind the position of the reference point).
[0075] Hardware: There are many different types of electronic systems that enable a person to perceive and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be positioned over a person's eyes (e.g., contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system may include speakers and / or other audio output devices integrated into the head-mounted system to provide audio output. A head-mounted system may have one or more speakers and an integrated opaque display. Alternatively, a head-mounted system may be configured to accept an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing images or videos of the physical environment and / or one or more microphones for capturing audio of the physical environment. The head-mounted system may have a transparent or translucent display instead of an opaque display. The transparent or translucent display may have a medium through which light representing an image is directed to the person's eye. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium may be an optical waveguide, holographic medium, optical coupler, optical reflector, or any combination thereof. In one embodiment, the transparent or translucent display may be configured to be selectively opaque. The projection-based system may employ retinal projection technology to project a graphical image onto the person's retina.The projection system may also be configured to project virtual objects into the physical environment, for example, as holograms or onto physical surfaces. In some embodiments, the controller 110 is configured to manage and coordinate the XR experience for the user. In some embodiments, the controller 110 includes a preferred combination of software, firmware, and / or hardware. The controller 110 is described in more detail below with respect to Figure 2. In some embodiments, the controller 110 is a computing device that is local or remote to the scene 105 (e.g., the physical environment). For example, the controller 110 is a local server located within the scene 105. In another example, the controller 110 is a remote server located outside the scene 105 (e.g., a cloud server, a central server, etc.). In some embodiments, the controller 110 is communicably coupled to a display generation component 120 (e.g., an HMD, display, projector, touchscreen, etc.) via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). In another example, the controller 110 is housed in a casing (e.g., a physical housing) of one or more of the display generation components 120 (e.g., an HMD, or a portable electronic device including a display and one or more processors, etc.), one or more of the input devices 125, one or more of the output devices 155, one or more of the sensors 190, and / or peripheral devices 195, or shares the same physical housing or support structure as one or more of the above.
[0076] In some embodiments, the display generation component 120 is configured to provide the user with an XR experience (e.g., at least the visual components of the XR experience). In some embodiments, the display generation component 120 includes a preferred combination of software, firmware, and / or hardware. The display generation component 120 is described in more detail below with respect to Figure 3. In some embodiments, the functions of the controller 110 are provided by and / or combined with the display generation component 120.
[0077] According to some embodiments, the display generation component 120 provides the user with an XR experience while the user is virtually and / or physically present in the scene 105.
[0078] In some embodiments, the display generation component is mounted on a part of the user's body (e.g., their head or hand). Thus, the display generation component 120 includes one or more XR displays provided for displaying XR content. For example, in various embodiments, the display generation component 120 surrounds the user's field of view. In some embodiments, the display generation component 120 is a handheld device (such as a smartphone or tablet) configured to present XR content, and the user holds the device, which has a display directed towards the user's field of view and a camera directed towards scene 105. In some embodiments, the handheld device is optionally placed in a housing mounted on the user's head. In some embodiments, the handheld device is optionally placed on a support in front of the user (e.g., a tripod). In some embodiments, the display generation component 120 is an XR chamber, housing, or room configured to present XR content when the user is not mounting or holding the display generation component 120. Many user interfaces described with reference to one type of hardware for displaying XR content (e.g., a handheld device or a device on a tripod) may be implemented on another type of hardware for displaying XR content (e.g., an HMD or other wearable computing device). For example, a user interface showing interaction with XR content triggered based on interaction occurring in the space in front of a handheld or tripod-mounted device may be implemented similarly to an HMD where the interaction occurs in the space in front of the HMD and the XR content response is displayed through the HMD. Similarly, a user interface showing interaction with XR content triggered based on the movement of a handheld or tripod-mounted device relative to a physical environment (e.g., Scene 105 or a part of the user's body (e.g., the user's eyes, head, or hands)) may be implemented similarly to an HMD where the movement is triggered by the movement of the HMD relative to a physical environment (e.g., Scene 105 or a part of the user's body (e.g., the user's eyes, head, or hands)).
[0079] Relevant features of the operating environment 100 are shown in Figure 1A, but those skilled in the art will understand from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the exemplary embodiments disclosed herein.
[0080] Figures 1A to 1P illustrate various examples of computer systems used to perform the method and provide audio, visual, and / or haptic feedback as part of the user interface described herein. In some embodiments, the computer system optionally includes one or more display generating components (e.g., first and second display assemblies 1-120a, 1-120b and / or first and second optical modules 11.1.1-104a and 11.1.1-104b) for displaying to a user of the computer system a representation of a virtual element and / or physical environment generated based on detected events and / or user input detected by the computer system. The user interface generated by the computer system is optionally corrected by one or more corrective lenses 11.3.2 to 216 optionally detachably attached to one or more of the optical modules, enabling the user interface to be more easily viewed by a user who otherwise corrects their vision using glasses or contact lenses. Many of the user interfaces described herein show a single view of the user interface, but the user interface within the HMD may optionally be displayed using two optical modules (e.g., first and second display assemblies 1-120a, 1-120b and / or first and second optical modules 11.1.1-104a and 11.1.1-104b), one for the user's right eye and a different one for the user's left eye, with slightly different images presented to the two different eyes to produce a three-dimensional depth illusion, and the single view of the user interface is typically either the right-eye or left-eye view, and the depth effect is described in text or using other schematic diagrams or views.In some embodiments, the computer system includes one or more external displays (e.g., display assembly 1-108) for displaying status information of the computer system to the user of the computer system (when the computer system is not installed) and / or to other people near the computer system, which is optionally generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more audio output components (e.g., electronic component 1-112) for generating audio feedback, which is optionally generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more input devices for detecting inputs, such as one or more sensors (e.g., sensor assembly 1-356 and / or one or more sensors in Figure 1I) for detecting information about the physical environment of a device that can be used (optionally in conjunction with one or more illuminators, such as the illuminator shown in Figure 1I) to generate a digital passthrough image, capture a visual medium (e.g., photograph and / or video) corresponding to a physical environment, or determine the pose (e.g., position and / or orientation) of physical objects and / or surfaces in the physical environment, so that virtual objects can be positioned based on the detected pose of physical objects and / or surfaces. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting hand position and / or movement (e.g., sensor assembly 1-356 and / or one or more sensors in Figure 1I), which can be used to determine when one or more air gestures were performed (optionally in conjunction with one or more illuminators, such as illuminator 6-124 shown in Figure 1I).In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting eye movements (e.g., the eye-tracking and eye-tracking sensors in Figure 1I), which may be used (optionally, in conjunction with one or more lights, such as lights 11.3.2-110 in Figure 1O) to determine attention or gaze position and / or gaze movement, which may be used to detect gaze-only input based on gaze movement and / or dwell time. Using the various sensor combinations described above, it is possible to determine the user's facial expressions and / or hand movements for use in generating the user's avatar or representation, such as a personified avatar or representation for use in a real-time communication session, the avatar having facial expressions, hand movements and / or body movements that are based on or similar to the detected facial expressions, hand movements and / or body movements of the user of the device. Eye-gaze and / or attention information is optionally combined with hand tracking information to determine interactions between the user and one or more user interfaces based on direct and / or indirect inputs such as air gestures or inputs using one or more hardware input devices, including buttons (e.g., first buttons 1-128, buttons 11.1.1-114, second buttons 1-132, and / or dials or buttons 1-328), knobs (e.g., first buttons 1-128, buttons 11.1.1-114, and / or dials or buttons 1-328), digital crowns (e.g., pressable, twistable, or rotatable first buttons 1-128, buttons 11.1.1-114, and / or dials or buttons 1-328), trackpads, touchscreens, keyboards, mice, and / or other input devices. One or more buttons (for example, the first buttons 1-128, buttons 11.1.1-114, the second button 1-132, and / or the dial or button 1-328) are optionally used to perform system actions such as re-centering content in a three-dimensional environment visible to the device user, displaying a home user interface for launching an application, starting a real-time communication session, or starting to display a virtual three-dimensional background.A knob or digital crown (e.g., a first button 1-128, button 11.1.1-114, and / or dial or button 1-328, which is pressable and twistable or rotatable) is optionally rotatable to adjust parameters of the visual content, such as the level of immersion of the virtual three-dimensional environment (e.g., the extent to which the virtual content occupies the user's viewport into the three-dimensional environment), or other parameters associated with the virtual content displayed via the three-dimensional environment and optical modules (e.g., first and second display assemblies 1-120a, 1-120b, and / or first and second optical modules 11.1.1-104a and 11.1.1-104b).
[0081] Figure 1B shows front, top, and perspective views of an example of a head-mounted display (HMD) device 1-100, which is worn by a user and configured to provide a virtual and augmented / mixed reality (VR / AR) experience. The HMD 1-100 may include a display unit 1-102 or assembly, an electronic strap assembly 1-104 connected to and extending from the display unit 1-102, and a band assembly 1-106 fixed to the electronic strap assembly 1-104 at either end. The electronic strap assembly 1-104 and the band 1-106 may be part of a retaining assembly configured to wrap around the user's head to hold the display unit 1-102 against the user's face.
[0082] In at least one example, the band assembly 1-106 may include a first band 1-116 configured to wrap around the back of the user's head and a second band 1-117 configured to extend over the top of the user's head. The second strap may extend between the first electronic strap 1-105a and the second electronic strap 1-105b of the electronic strap assembly 1-104, as shown in the illustration. The strap assembly 1-104 and the band assembly 1-106 may be part of a fastening mechanism that extends rearward from the display unit 1-102 and is configured to hold the display unit 1-102 against the user's face.
[0083] In at least one example, the fastening mechanism includes a first electronic strap 1-105a, which includes a first proximal end 1-134 coupled to a housing 1-150 of the display unit 1-102, for example, and a first distal end 1-136 opposite the first proximal end 1-134. The fastening mechanism may also include a second electronic strap 1-105b, which includes a second proximal end 1-138 coupled to the housing 1-150 of the display unit 1-102, and a second distal end 1-140 opposite the second proximal end 1-138. The fastening mechanism may also include a first band 1-116 having a first end 1-142 coupled to a first distal end 1-136 and a second end 1-144 coupled to a second distal end 1-140, and a second band 1-117 extending between the first electronic strap 1-105a and the second electronic strap 1-105b. The straps 1-105a and 1-105b and the band 1-116 may be connected via a connecting mechanism or assembly 1-114. In at least one example, the second band 1-117 includes a first end 1-146 coupled to a first electron strap 1-105a between a first proximal end 1-134 and a first distal end 1-136, and a second end 1-148 coupled to a second electron strap 1-105b between a second proximal end 1-138 and a second distal end 1-140.
[0084] In at least one example, the first and second electronic straps 1-105a-b include plastic, metal, or other structural material that forms the shape of substantially rigid straps 1-105a-b. In at least one example, the first and second bands 1-116, 1-117 are formed from an elastic flexible material, including woven fabric, rubber, etc. The first and second bands 1-116, 1-117 may be flexible to conform to the shape of the user's head when the HMD 1-100 is worn.
[0085] In at least one example, one or more of the first and second electronic straps 1-105a to b may define an internal strap volume and include one or more electronic components disposed within that internal strap volume. In one example, as shown in Figure 1B, the first electronic strap 1-105a may include electronic component 1-112. In one example, electronic component 1-112 may include a speaker. In another example, electronic component 1-112 may include a computing component such as a processor.
[0086] In at least one example, the housing 1-150 defines a first forward-facing opening 1-152. The display assembly 1-108 is positioned to block the first opening 1-152 from view when the HMD 1-100 is assembled, so the forward-facing opening is labeled with a dotted line at 1-152 in Figure 1B. The housing 1-150 may also define a second rearward-facing opening 1-154. The housing 1-150 also defines an internal volume between the first opening 1-152 and the second opening 1-154. In at least one example, the HMD 1-100 includes a display assembly 1-108, which may include a front cover and a display screen (shown in other figures) positioned within or across the front opening 1-152 to block the front opening 1-152. In at least one example, the display screen of display assembly 1-108 has a curvature configured to follow the curvature of the user's face, as well as the display assembly 1-108 as a whole. The display screen of display assembly 1-108 can be curved to complement the features of the user's face and the overall curvature from one side of the face to the other, for example, from left to right and / or from top to bottom when the display unit 1-102 is pressed.
[0087] In at least one example, the housing 1-150 may define a first aperture 1-126 between a first opening 1-152 and a second opening 1-154, and a second aperture 1-130 between the first opening 1-152 and the second opening 1-154. The HMD 1-100 may also include a first button 1-128 disposed in the first aperture 1-126 and a second button 1-132 disposed in the second aperture 1-130. The first and second buttons 1-128 and 1-132 may be pressable through their respective apertures 1-126 and 1-130. In at least one example, the first button 1-126 and / or the second button 1-132 may be a twistable dial and a pressable button. In at least one example, the first button 1-128 is a pressable and twistable dial button, and the second button 1-132 is a pressable button.
[0088] Figure 1C shows a rear perspective view of HMD1-100. HMD1-100 may include an optical seal 1-110 that extends rearward from housing 1-150 around the housing 1-150 of display assembly 1-108, as shown. The optical seal 1-110 may be configured to extend from housing 1-150 to the user's face around the user's eyes to block external light from being visible. In one example, HMD1-100 may include first and second display assemblies 1-120a, 1-120b that are disposed in or within a rearward-facing second opening 1-154 defined by housing 1-150 and / or disposed within the internal volume of housing 1-150 and configured to project light through the second opening 1-154. In at least one example, each display assembly 1-120a-b may include respective display screens 1-122a, 1-122b configured to project light backward through a second opening 1-154 toward the user's eyes.
[0089] In at least one example, referring to both Figures 1B and 1C, the display assembly 1-108 may be a forward-facing display assembly including a display screen configured to project light in a first forward direction, and the rear-facing display screens 1-122a-b may be configured to project light in a second rear direction opposite to the first direction. As described above, the light seal 1-110 may be configured to prevent external light from the HMD 1-100, including light projected by the forward-facing display screen of the display assembly 1-108 shown in the front perspective view of Figure 1B, from reaching the user's eyes. In at least one example, the HMD 1-100 may also include a curtain 1-124 that closes a second opening 1-154 between the housing 1-150 and the rear-facing display assemblies 1-120a-b. In at least one example, the curtain 1-124 may be elastic or at least partially elastic.
[0090] Any feature, component, and / or part, including its arrangement and configuration shown in Figures 1B and 1C, may be included, alone or in any combination, in any other example of devices, features, components, and parts shown in Figures 1D to 1F and described herein. Similarly, any feature, component, and / or part shown and described with reference to Figures 1D to 1F may be included, alone or in any combination, in the examples of devices, features, components, and parts shown in Figures 1B and 1C.
[0091] Figure 1D shows an exploded view of an example of HMD1-200, which includes various parts or components separated according to modularity and the selective coupling of their components. For example, HMD1-200 may include a band 1-216 that can be selectively coupled to first and second electronic straps 1-205a, 1-205b. The first fastening strap 1-205a may include a first electronic component 1-212a, and the second fastening strap 1-205b may include a second electronic component 1-212b. In at least one example, the first and second straps 1-205a and 1-205b may be detachably coupled to a display unit 1-202.
[0092] In addition, the HMD1-200 may include an optical seal 1-210 configured to be detachably coupled to a display unit 1-202. The HMD1-200 may also include a lens 1-218 that can be detachably coupled to the display unit 1-202, for example, on first and second display assemblies including a display screen. The lens 1-218 may include a customized prescription lens configured for vision correction. As stated, each component shown in the exploded view of Figure 1D and described above may be detachably coupled, mounted, reattached, and replaced in order to update or replace parts for different users. For example, bands such as band 1-216, optical seals such as optical seal 1-210, lenses such as lens 1-218, and electronic straps such as straps 1-205a~b may be replaced on a user-by-user basis so that these components are customized to fit and correspond to individual users of the HMD1-200.
[0093] Any feature, component, and / or part, including its arrangement and configuration shown in Figure 1D, may be included, alone or in any combination, in any other example of devices, features, components, and parts shown in Figures 1B, 1C, and 1E-1F and described herein. Similarly, any feature, component, and / or part, including its arrangement and configuration shown and described with reference to Figures 1B, 1C, and 1E-1F, may be included, alone or in any combination, in the examples of devices, features, components, and parts shown in Figure 1D.
[0094] Figure 1E shows an exploded view of an example of a display unit 1-306 of an HMD. Display unit 1-306 may include a front display assembly 1-308, a frame / housing assembly 1-350, and a curtain assembly 1-324. Display unit 1-306 may also include a sensor assembly 1-356, a logic board assembly 1-358, and a cooling assembly 1-360, disposed between the frame assembly 1-350 and the front display assembly 1-308. In at least one example, display unit 1-306 may also include a rear-facing display assembly 1-320, which includes first and second rear-facing display screens 1-322a, 1-322b, disposed between the frame 1-350 and the curtain assembly 1-324.
[0095] In at least one example, the display unit 1-306 may also include a motor assembly 1-362 configured as an adjustment mechanism for adjusting the position of the display screens 1-322a-b of the display assembly 1-320 relative to the frame 1-350. In at least one example, the display assembly 1-320 is mechanically coupled to a motor assembly 1-362 with at least one motor for each display screen 1-322a-b, so that the motors can translate the display screens 1-322a-b to match the interpupillary distance of the user's eyes.
[0096] In at least one example, the display unit 1-306 may include a dial or button 1-328 that is pressable relative to the frame 1-350 and accessible to the user outside the frame 1-350. The button 1-328 may be electronically connected to the motor assembly 1-362 via a controller so that the user can operate the button 1-328 to cause the motors of the motor assembly 1-362 to adjust the position of the display screens 1-322a-b.
[0097] Any feature, component, and / or part, including its arrangement and configuration shown in Figure 1E, may be included, alone or in any combination, in any other example of devices, features, components, and parts shown in Figures 1B to 1D and Figure 1F and described herein. Similarly, any feature, component, and / or part, including its arrangement and configuration, as illustrated and described with reference to Figures 1B to 1D and Figure 1F, may be included, alone or in any combination, in the examples of devices, features, components, and parts shown in Figure 1E.
[0098] Figure 1F shows an exploded view of another example of a display unit 1-406 of an HMD device similar to other HMD devices described herein. Display unit 1-406 may include a forward display assembly 1-402, a sensor assembly 1-456, a logic board assembly 1-458, a cooling assembly 1-460, a frame assembly 1-450, a rear-facing display assembly 1-421, and a curtain assembly 1-424. Display unit 1-406 may also include a motor assembly 1-462 for adjusting the positions of the first and second display subassemblies 1-420a, 1-420b of the rear-facing display assembly 1-421, which include the first and second display screens, respectively, for interpupillary adjustment, as described above.
[0099] Various components, systems, and assemblies shown in the exploded view of Figure 1F are described in more detail herein with reference to Figures 1B to 1E and subsequent figures referenced herein. The display unit 1-406 shown in Figure 1F may be assembled and integrated with the fastening mechanisms shown in Figures 1B to 1E, which include other components such as electronic straps, bands, and optical seals, and connecting assemblies.
[0100] Any feature, component, and / or part, including its arrangement and configuration shown in Figure 1F, may be included, individually or in any combination, in any other example of devices, features, components, and parts shown in Figures 1B to 1E and described herein. Similarly, any feature, component, and / or part shown and described with reference to Figures 1B to 1E, including its arrangement and configuration, may be included, individually or in any combination, in the examples of devices, features, components, and parts shown in Figure 1F.
[0101] Figure 1G shows a perspective exploded view of a front cover assembly 3-100 of an HMD device described herein, for example, front cover assembly 3-1 of the HMD 3-100 shown in Figure 1G, or any other HMD device illustrated and described herein. The front cover assembly 3-100 shown in Figure 1G may include a transparent or translucent cover 3-102, a shroud 3-104 (or "canopy"), an adhesive layer 3-106, a display assembly 3-108 including a lenticular lens panel or array 3-110, and a structural trim 3-112. The adhesive layer 3-106 can fasten the shroud 3-104 and / or the transparent cover 3-102 to the display assembly 3-108 and / or the trim 3-112. The trim 3-112 can fasten various components of the front cover assembly 3-100 to the frame or chassis of the HMD device.
[0102] In at least one example, as shown in Figure 1G, a display assembly 3-108 including a transparent cover 3-102, a shroud 3-104, and a lenticular lens array 3-110 can be curved to accommodate the curvature of the user's face. The transparent cover 3-102 and shroud 3-104 can be curved in two or three dimensions, for example, curving perpendicularly in the Z direction inside and outside the ZX plane, and curving horizontally in the X direction inside and outside the ZX plane. In at least one example, the display assembly 3-108 may include a display panel having a lenticular lens array 3-110, as well as pixels configured to project light through the shroud 3-104 and the transparent cover 3-102. The display assembly 3-108 can be curved in at least one direction, for example, horizontally, to adapt to the curvature of the user's face from one side (e.g., the left side) to the other side (e.g., the right side). In at least one example, as shown and described in more detail in subsequent figures, each layer or component of the display assembly 3-108, which may include a lenticular lens array 3-110 and a display layer, can be curved horizontally, similarly or concentrically, to adapt to the curvature of the user's face.
[0103] In at least one example, the shroud 3-104 may include a transparent or translucent material from which the display assembly 3-108 projects light. In one example, the shroud 3-104 may include one or more opaque portions, such as opaque ink-printed portions or other opaque film portions, on the rear surface of the shroud 3-104. The rear surface may be the surface of the shroud 3-104 that faces the user's eyes when the HMD device is worn. In at least one example, the opaque portions may be on the front surface of the shroud 3-104 opposite the rear surface. In at least one example, one or more opaque portions of the shroud 3-104 may include perimeter portions that visually conceal any components around the perimeter of the display screen of the display assembly 3-108. In this way, the opaque portions of the shroud conceal any other components, including electronic components, structural components, etc., of the HMD device that would otherwise be visible through the transparent or translucent cover 3-102 and / or the shroud 3-104.
[0104] In at least one example, the shroud 3-104 can define one or more aperture transparent portions 3-120 through which a sensor can send and receive signals. In one example, portion 3-120 is an aperture through which a sensor can extend or send and receive signals. In one example, portion 3-120 is a transparent portion, or a portion more transparent than the translucent or opaque portion around the shroud, through which the sensor can send and receive signals through the shroud and through the transparent cover 3-102. In one example, the sensor may include a camera, an IR sensor, a LUX sensor, or any other visual or non-visual environment sensor of the HMD device.
[0105] Any feature, component, and / or part, including its arrangement and configuration shown in Figure 1G, may be included, individually or in any combination, in any other example of devices, features, components, and parts described herein. Similarly, any feature, component, and / or part, including its arrangement and configuration illustrated and described herein, may be included, individually or in any combination, in the example of devices, features, components, and parts shown in Figure 1G.
[0106] Figure 1H shows an exploded view of an example of HMD device 6-100. HMD device 6-100 may include a sensor array or system 6-102 which includes one or more sensors, cameras, projectors, etc., attached to one or more components of HMD 6-100. In at least one example, the sensor system 6-102 may include a bracket 1-338 to which one or more sensors of the sensor system 6-102 can be fixed / attached.
[0107] Figure 1I shows a portion of the HMD device 6-100, including the front transparent cover 6-104 and the sensor system 6-102. The sensor system 6-102 may include multiple different sensors, emitters, and receivers, including a camera, IR sensor, and projector. The transparent cover 6-104 is shown in front of the sensor system 6-102 to show the relative positions of the various sensors and emitters and the orientation of each sensor / emitter in the system 6-102. As used herein, “lateral,” “side,” “lateral,” “horizontal,” and other similar terms refer to orientation or direction as indicated by the X-axis shown in Figure 1J. Terms such as “vertical,” “up,” “down,” and similar terms refer to orientation or direction as indicated by the Z-axis shown in Figure 1J. Terms such as “forward,” “backward,” “front,” “rear,” and similar terms refer to orientation or direction as indicated by the Y-axis shown in Figure 1J.
[0108] In at least one example, a transparent cover 6-104 can define the outer front surface of the HMD device 6-100, and a sensor system 6-102, including various sensors and their components, can be positioned behind the cover 6-104 in the Y-axis / direction. The cover 6-104 may be transparent or translucent to allow both the light detected by the sensor system 6-102 and the light emitted thereby to pass through the cover 6-104.
[0109] As described elsewhere in this specification, the HMD device 6-100 may include one or more controllers, including processors, for electrically coupling the various sensors and emitters of the sensor system 6-102 to one or more motherboards, processing units, and other electronic devices such as display screens. Furthermore, as shown in more detail below with reference to other figures, the various sensors, emitters, and other components of the sensor system 6-102 may be coupled to various structural frame members, brackets, etc., of the HMD device 6-100, which are not shown in Figure 1I. Figure 1I shows components of the sensor system 6-102 that are not attached to other components and are not electrically coupled, for the sake of clarity as an example.
[0110] In at least one example, the device may include one or more controllers having processors configured to execute instructions stored on memory components electrically coupled to the processors. The instructions may include, or be executed by, one or more algorithms for self-correcting the angles and positions of various cameras described herein over time with use as the initial position, angle, or orientation of the cameras is impacted or deformed due to an unintended fall event or other event.
[0111] In at least one example, the sensor system 6-102 may include one or more scene cameras 6-106. System 6-102 may include two scene cameras 6-106 positioned on either side of the bridge or arch of the HMD device 6-100, such that each of the two cameras 6-102 roughly corresponds to the positions of the user's left and right eyes behind the cover 6-103. In at least one example, the scene cameras 6-106 are generally oriented forward in the Y direction to capture images in front of the user while the HMD 6-100 is in use. In at least one example, the scene cameras are color cameras and provide images and content for MR video passthrough to a display screen facing the user's eyes when the HMD device 6-100 is in use. The scene cameras 6-106 can also be used for environment and object reconstruction.
[0112] In at least one example, the sensor system 6-102 may include a first depth sensor 6-108 that is generally oriented forward in the Y direction. In at least one example, the first depth sensor 6-108 can be used for reconstructing the environment and objects, as well as tracking the user's hands and body. In at least one example, the sensor system 6-102 may include a second depth sensor 6-110 that is centrally positioned along the width of the HMD device 6-100 (for example, along the X axis). For example, the second depth sensor 6-110 can be positioned to align with the central bridge or feature above the user's nose when the HMD 6-100 is worn. In at least one example, the second depth sensor 6-110 can be used for reconstructing the environment and objects, as well as tracking the hands and body. In at least one example, the second depth sensor may include a LIDAR sensor.
[0113] In at least one example, the sensor system 6-102 may include a generally forward-facing depth projector 6-112 to project electromagnetic waves, for example, in the form of a predetermined pattern of light dots, into and within the field of view of the user and / or scene camera 6-106, or into and within the field of view including and beyond the field of view of the user and / or scene camera 6-106. In at least one example, the depth projector may project electromagnetic waves of light in the form of a dot light pattern that is reflected from objects and returned to the aforementioned depth sensors, including depth sensors 6-108, 6-110. In at least one example, the depth projector 6-112 may be used for environment and object reconstruction and hand and body tracking.
[0114] In at least one example, the sensor system 6-102 may include a downward-facing camera 6-114 having a field of view generally directed downward relative to the HMD device 6-100 in the Z-axis. In at least one example, the downward-facing camera 6-114 may be positioned on the left and right sides of the HMD device 6-100 as shown in the figure and may be used for hand and body tracking, headset tracking, and face avatar detection and creation in order to display a user avatar on the forward-facing display screen of the HMD device 6-100 as described elsewhere in this specification. The downward-facing camera 6-114 may be used to capture the facial expressions and movements of the user below the HMD device 6-100, including, for example, the cheeks, mouth, and chin.
[0115] In at least one example, the sensor system 6-102 may include a jaw camera 6-116. In at least one example, the jaw camera 6-116 may be positioned on the left and right sides of the HMD device 6-100 as shown in the figure and may be used for hand and body tracking, headset tracking, and face avatar detection and creation in order to display a user avatar on the forward-facing display screen of the HMD device 6-100 as described elsewhere in this specification. The jaw camera 6-116 may be used to capture the user's facial expressions and movements below the HMD device 6-100, including, for example, the user's jaw, cheeks, mouth, and chin. Regarding hand and body tracking, headset tracking, and face avatar,
[0116] In at least one example, the sensor system 6-102 may include a side camera 6-118. The side camera 6-118 may be oriented to capture left and right side views in the X-axis or direction relative to the HMD device 6-100. In at least one example, the side camera 6-118 may be used for hand and body tracking, headset tracking, and detection and reproduction of a facial avatar.
[0117] In at least one example, the sensor system 6-102 may include multiple eye-tracking and gaze-tracking sensors for determining the user's eye identification information, status, and gaze direction during and / or before use. In at least one example, the eye / gaze-tracking sensor may include nasal eye cameras 6-120 positioned on either side of the user's nose and adjacent to the user's nose when the HMD device 6-100 is worn. The eye / gaze sensor may also include lower eye cameras 6-122 positioned below each user's eye for capturing images of the eye for face avatar detection and creation, gaze tracking, and iris recognition functions.
[0118] In at least one example, the sensor system 6-102 includes an infrared illuminator 6-124 directed outward from the HMD device 6-100, which can illuminate the external environment and any objects within it with IR light for IR detection by one or more IR sensors of the sensor system 6-102. In at least one example, the sensor system 6-102 may include a flicker sensor 6-126 and an ambient light sensor 6-128. In at least one example, the flicker sensor 6-126 may detect the overhead light refresh rate to avoid display flicker. In one example, the infrared illuminator 6-124 may include a light-emitting diode and can be used in low-light environments, in particular, to illuminate the user's hand and other objects with low light for detection by the infrared sensors of the sensor system 6-102.
[0119] In at least one example, multiple sensors, including a scene camera 6-106, a downward-facing camera 6-114, a jaw camera 6-116, a side camera 6-118, a depth projector 6-112, and depth sensors 6-108 and 6-110, can be used in combination with an electrically coupled controller to combine depth data with camera data for hand tracking and sizing, for better hand tracking and object recognition and tracking capabilities of the HMD device 6-100. In at least one example, the downward-facing camera 6-114, jaw camera 6-116, and side camera 6-118 described above and shown in Figure 1I may be wide-angle cameras capable of operating in the visible and infrared spectra. In at least one example, these cameras 6-114, 6-116, and 6-118 may operate with monochrome light detection only to simplify image processing and increase sensitivity.
[0120] Any feature, component, and / or part, including its arrangement and configuration shown in Figure 1I, may be included, alone or in any combination, in any other example of devices, features, components, and parts shown in Figures 1J to 1L and described herein. Similarly, any feature, component, and / or part illustrated and described with reference to Figures 1J to 1L may be included, alone or in any combination, in the example of devices, features, components, and parts shown in Figure 1I.
[0121] Figure 1J shows a downward perspective view of an example of the HMD6-200, including a cover or shroud 6-204 fixed to the frame 6-230. In at least one example, the sensor 6-203 of the sensor system 6-202 may be positioned around the periphery of the HDM6-200 such that the sensor 6-203 is positioned outward around the periphery of the display area or area 6-232 so as not to obstruct the view of the displayed light. In at least one example, the sensor may be positioned behind the shroud 6-204 and aligned with the transparent portion of the shroud to allow the sensor and projector to pass light back and forth through the shroud 6-204. In at least one example, opaque ink or other opaque material or film / layer can be placed on the shroud 6-204 around the display area 6-232 to conceal components of the HMD 6-200 outside the display area 6-232 other than the transparent portion defined by the opaque portion, through which sensors and projectors transmit and receive light and electromagnetic signals during operation. In at least one example, the shroud 6-204 allows light to pass through from the display (e.g., within the display area 6-232) but not radially outward from the display area around the periphery of the display and the shroud 6-204.
[0122] In some examples, the shroud 6-204 includes a transparent portion 6-205 and an opaque portion 6-207, as described above and elsewhere in this specification. In at least one example, the opaque portion 6-207 of the shroud 6-204 can define one or more transparent regions 6-209 from which sensors 6-203 of the sensor system 6-202 can send and receive signals. In the illustrated example, the sensor 6-203 of the sensor system 6-202, which transmits and receives signals through the shroud 6-204, or more specifically through the transparent area 6-209 of (or defined by) the opaque portion 6-207 of the shroud 6-204, may include the same or similar sensors as those shown in the example in Figure 1I, e.g., depth sensors 6-108 and 6-110, depth projector 6-112, first and second scene cameras 6-106, first and second downward-facing cameras 6-114, first and second side cameras 6-118, and first and second infrared illuminators 6-124. These sensors are also shown in the examples in Figures 1K and 1L. Other sensors, sensor types, number of sensors, and their relative positions may be included in one or more other examples of the HMD.
[0123] Any feature, component, and / or part, including its arrangement and configuration shown in Figure 1J, may be included, alone or in any combination, in any other example of devices, features, components, and parts shown in Figures 1I and 1K-1L and described herein. Similarly, any feature, component, and / or part, including its arrangement and configuration, illustrated and described with reference to Figures 1I and 1K-1L, may be included, alone or in any combination, in the examples of devices, features, components, and parts shown in Figure 1J.
[0124] Figure 1K shows a partial front view of an example of an HMD device 6-300, including a display 6-334, brackets 6-336 and 6-338, and a frame or housing 6-330. The example shown in Figure 1K does not include a front cover or shroud to show brackets 6-336 and 6-338. For example, the shroud 6-204 shown in Figure 1J includes an opaque portion 6-207 that visually covers / blocks the view of anything outside (e.g., radially / peripherally outward) of the display / display area 6-334, including sensors 6-303 and brackets 6-338.
[0125] In at least one example, various sensors of sensor system 6-302 are coupled to brackets 6-336, 6-338. In at least one example, scene cameras 6-306 have tight tolerances for angles relative to each other. For example, the tolerance for the mounting angle between two scene cameras 6-306 may be 0.5 degrees or less, e.g., 0.3 degrees or less. To achieve and maintain such tight tolerances, in one example, scene cameras 6-306 can be mounted to bracket 6-338 rather than to the shroud. The bracket may include a cantilever arm to which scene cameras 6-306 and other sensors of sensor system 6-302 can be mounted, such that their position and orientation remain undeformed in the event of a user-induced drop event resulting in any deformation of the other brackets 6-226, housing 6-330, and / or shroud.
[0126] Any feature, component, and / or part, including its arrangement and configuration shown in Figure 1K, may be included, alone or in any combination, in any other example of devices, features, components, and parts shown in Figures 1I, 1J, and 1L and described herein. Similarly, any feature, component, and / or part, including its arrangement and configuration shown and described with reference to Figures 1I, 1J, and 1L, may be included, alone or in any combination, in the examples of devices, features, components, and parts shown in Figure 1K.
[0127] Figure 1L shows a bottom view of an example of the HMD 6-400, including the front display / cover assembly 6-404 and the sensor system 6-402. The sensor system 6-402 may be similar to other sensor systems described above and elsewhere in this specification, including referring to Figures 1I to 1K. In at least one example, the jaw camera 6-416 may be oriented downward to capture an image of the user's lower facial features. In one example, the jaw camera 6-416 may be directly coupled to the frame or housing 6-430, or to one or more internal brackets directly coupled to the illustrated frame or housing 6-430. The frame or housing 6-430 may include one or more apertures / openings 6-415 from which the jaw camera 6-416 can send and receive signals.
[0128] Any feature, component, and / or part, including its arrangement and configuration shown in Figure 1L, may be included, alone or in any combination, in any other example of devices, features, components, and parts shown in Figures 1I to 1K and described herein. Similarly, examples of devices, features, components, and parts shown in Figure 1L may include, alone or in any combination, any feature, component, and / or part, including its arrangement and configuration shown and described with reference to Figures 1I to 1K.
[0129] Figure 1M shows a rear perspective view of the interpupillary distance (IPD) adjustment system 11.1.1-102, which includes first and second optical modules 11.1.1-104a-b that are slidably engaged / coupled to the respective guide rods 11.1.1-108a-b and motors 11.1.1-110a-b of the left and right adjustment subsystems 11.1.1-106a-b. The IPD adjustment system 11.1.1-102 can be coupled to a bracket 11.1.1-112 and may include buttons 11.1.1-114 that electrically communicate with the motors 11.1.1-110a-b. In at least one example, buttons 11.1.1-114 can electrically communicate with the first and second motors 11.1.1-110a~b via a processor or other circuit component to activate the first and second motors 11.1.1-110a~b and change the positions of the first and second optical modules 11.1.1-104a~b relative to each other.
[0130] In at least one example, the first and second optical modules 11.1.1-104a~b may include respective display screens configured to project light toward the user's eyes when the HMD 11.1.1-100 is worn. In at least one example, the user can operate (e.g., press and / or rotate) the button 11.1.1-114 to activate the position adjustment of the optical modules 11.1.1-104a~b to match the interpupillary distance of the user's eyes. The optical modules 11.1.1-104a~b may also include one or more cameras or other sensors / sensor systems for imaging and measuring the user's IPD so that the optical modules 11.1.1-104a~b can be adjusted to match the IPD.
[0131] In one example, the user can operate buttons 11.1.1-114 to trigger automatic position adjustment of the first and second optical modules 11.1.1-104a~b. In another example, the user can operate buttons 11.1.1-114 to trigger manual adjustment, for example, by rotating buttons 11.1.1-114 in one or the other direction, so that the optical modules 11.1.1-104a~b move further away or closer until the user visually matches their IPD. In another example, the manual adjustment is communicated electronically via one or more circuits, and power for the movement of the optical modules 11.1.1-104a~b via motors 11.1.1-110a~b is provided by a power supply. In yet another example, the adjustment and movement of the optical modules 11.1.1-104a~b via the operation of buttons 11.1.1-114 is mechanically actuated via the movement of buttons 11.1.1-114.
[0132] Any feature, component, and / or part, including its arrangement and configuration shown in Figure 1M, may be included, alone or in any combination, in any other example of devices, features, components, and parts shown in any other figures illustrated and described herein. Similarly, any feature, component, and / or part, including its arrangement and configuration, as illustrated and described with reference to any other figures illustrated and described herein, in any example of devices, features, components, and parts shown in Figure 1M, either alone or in any combination.
[0133] Figure 1N shows a partial front perspective view of the HMD 11.1.2-100, including the outer structural frame 11.1.2-102 and the inner or intermediate structural frame 11.1.2-104 that define the first and second apertures 11.1.2-106a and 11.1.2-106b. The apertures 11.1.2-106a-b are shown as dotted lines in Figure 1N because, as shown, the views of the apertures 11.1.2-106a-b may be obstructed by one or more other components of the HMD 11.1.2-100 coupled to the inner frame 11.1.2-104 and / or the outer frame 11.1.2-102. In at least one example, the HMD 11.1.2-100 may include a first mounting bracket 11.1.2-108 coupled to the inner frame 11.1.2-104. In at least one example, the mounting bracket 11.1.2-108 is coupled to the inner frame 11.1.2-104 between the first and second apertures 11.1.2-106a and 11.1.2-106b.
[0134] The mounting bracket 11.1.2-108 may include an intermediate or central portion 11.1.2-109 coupled to the inner frame 11.1.2-104. In some examples, the intermediate or central portion 11.1.2-109 may not be the geometric middle or center of the bracket 11.1.2-108. Rather, the intermediate / central portion 11.1.2-109 may be positioned between a first cantilever extension arm and a second cantilever extension arm extending away from the intermediate portion 11.1.2-109. In at least one example, the mounting bracket 108 includes a first cantilever arm 11.1.2-112 and a second cantilever arm 11.1.2-114 that extend away from the intermediate portion 11.1.2-109 of the mounting bracket 11.1.2-108 coupled to the inner frame 11.1.2-104.
[0135] As shown in Figure 1N, the outer frame 11.1.2-102 may be defined with a curved shape on its underside to accommodate the user's nose when the user wears the HMD 11.1.2-100. The curved shape may be referred to as the nose bridge 11.1.2-111 and may be located in the center of the underside of the HMD 11.1.2-100 as shown. In at least one example, the mounting bracket 11.1.2-108 may be connected to the inner frame 11.1.2-102 between apertures 11.1.2-106a-b, such that the cantilever arms 11.1.2-112, 11.1.2-114 extend downward and laterally outward away from the intermediate portion 11.1.2-109 to complement the shape of the nose bridge 11.1.2-111 of the outer frame 11.1.2-104. In this way, the mounting bracket 11.1.2-108 is configured to accommodate the user's nose as described above. The shape of the nose bridge 11.1.2-111 accommodates the nose in such a way that the nose bridge 11.1.2-111 provides a curvature that curves above, over, and around the nose, along with the user's nose, for comfort and fit.
[0136] The first cantilever arm 11.1.2-112 may extend away from the intermediate portion 11.1.2-109 of the mounting bracket 11.1.2-108 in a first direction, and the second cantilever arm 11.1.2-114 may extend away from the intermediate portion 11.1.2-109 of the mounting bracket 11.1.2-10 in a second direction opposite to the first direction. The first and second cantilever arms 11.1.2-112 and 11.1.2-114 are referred to as "cantilevered" or "cantilevered" arms because each arm 11.1.2-112 and 11.1.2-114 includes a distal free end 11.1.2-116 and 11.1.2-118 that is not fixed to the inner and outer frames 11.1.2-102 and 11.1.2-104, respectively. In this way, arms 11.1.2-112 and 11.1.2-114 are cantilevered from an intermediate section 11.1.2-109 that can be connected to the inner frame 11.1.2-104, with their distal ends 11.1.2-102 and 11.1.2-104 not attached.
[0137] In at least one example, the HMD 11.1.2-100 may include one or more components coupled to the mounting bracket 11.1.2-108. In one example, the components include a plurality of sensors 11.1.2-110a~f. Each of the plurality of sensors 11.1.2-110a~f may include various types of sensors, such as cameras and IR sensors. In some examples, one or more of the sensors 11.1.2-110a~f may be used for object recognition in three-dimensional space, such that it is important to maintain the precise relative positions of two or more of the plurality of sensors 11.1.2-110a~f. The cantilevered nature of the mounting bracket 11.1.2-108 can protect the sensors 11.1.2-110a~f from damage and displacement in the event of an accidental drop by the user. Since sensors 11.1.2-110a~f are cantilevered on arms 11.1.2-112 and 11.1.2-114 of mounting bracket 11.1.2-108, stresses and deformations in the inner and / or outer frames 11.1.2-104 and 11.1.2-102 are not transmitted to the cantilever arms 11.1.2-112 and 11.1.2-114, and therefore do not affect the relative positioning of sensors 11.1.2-110a~f coupled to / mounted on mounting bracket 11.1.2-108.
[0138] Any feature, component, and / or part, including its arrangement and configuration shown in Figure 1N, may be included, individually or in any combination, in any other example of devices, features, and components described herein. Similarly, any feature, component, and / or part, including its arrangement and configuration illustrated and described herein, may be included, individually or in any combination, in the examples of devices, features, components, and components shown in Figure 1N.
[0139] Figure 10 shows an example of an optical module 11.3.2-100 for use in electronic devices such as HMDs, including the HDM devices described herein. As shown in one or more other examples described herein, the optical module 11.3.2-100 may be one of two optical modules in an HMD, each optical module being positioned to project light toward the user's eye. In this way, the first optical module can project light toward the user's first eye via a display screen, and the second optical module of the same device can project light toward the user's second eye via another display screen.
[0140] In at least one example, the optical module 11.3.2-100 may include an optical frame or housing 11.3.2-102, which may also be referred to as a barrel or optical module barrel. The optical module 11.3.2-100 may also include a display 11.3.2-104, which includes one or more display screens, coupled to the housing 11.3.2-102. The display 11.3.2-104 may be coupled to the housing 11.3.2-102 such that the display 11.3.2-104 is configured to project light toward the user's eyes when the HMD, of which the display module 11.3.2-100 is part, is worn in use. In at least one example, the housing 11.3.2-102 may surround the display 11.3.2-104 and provide a coupling mechanism for coupling other components of the optical module described herein.
[0141] In one example, the optical module 11.3.2-100 may include one or more cameras 11.3.2-106 coupled to the housing 11.3.2-102. The cameras 11.3.2-106 may be positioned relative to the display 11.3.2-104 and the housing 11.3.2-102 so that the cameras 11.3.2-106 are configured to capture one or more images of the user's eyes while in use. In at least one example, the optical module 11.3.2-100 may also include a light strip 11.3.2-108 surrounding the display 11.3.2-104. In one example, the light strip 11.3.2-108 is positioned between the display 11.3.2-104 and the cameras 11.3.2-106. The light strip 11.3.2-108 may include multiple lights 11.3.2-110. Multiple lights may include one or more light-emitting diodes (LEDs) or other lights configured to project light toward the user's eyes when the HMD is worn. Individual lights 11.3.2-110 of the light strip 11.3.2-108 can be spaced apart around the strip 11.3.2-108 and thus can be spaced uniformly or unevenly around the display 11.3.2-104 at various locations on the strip 11.3.2-108 and around the display 11.3.2-104.
[0142] In at least one example, the housing 11.3.2-102 defines a viewing aperture 11.3.2-101 through which the user can see the display 11.3.2-104 when the HMD device is worn. In at least one example, LEDs are configured and positioned to emit light over the user's eyes through the viewing aperture 11.3.2-101. In one example, a camera 11.3.2-106 is configured to capture one or more images of the user's eyes through the viewing aperture 11.3.2-101.
[0143] As described above, each of the components and features of the optical module 11.3.2-100 shown in Figure 1O can be replicated in another (e.g., a second) optical module arranged with the HMD to interact with the user's other eye (e.g., project light and capture images).
[0144] Any feature, component, and / or part, including its arrangement and configuration shown in Figure 1O, may be included alone or in any combination in any other example of devices, features, components, and parts shown in Figure 1P or otherwise described herein. Similarly, any feature, component, and / or part, including its arrangement and configuration, as illustrated and described with reference to Figure 1P or otherwise described herein, may be included alone or in any combination in the examples of devices, features, components, and parts shown in Figure 1O.
[0145] Figure 1P shows a cross-sectional view of an example of an optical module 11.3.2-200, which includes a housing 11.3.2-202, a display assembly 11.3.2-204 coupled to the housing 11.3.2-202, and a lens 11.3.2-216 coupled to the housing 11.3.2-202. In at least one example, the housing 11.3.2-202 defines a first aperture or channel 11.3.2-212 and a second aperture or channel 11.3.2-214. Channels 11.3.2-212 and 11.3.2-214 may be configured to slidably engage with the respective rails or guide rods of the HMD device to allow the optical module 11.3.2-200 to adjust its position relative to the user's eyes to match the user's interpupillary distance (IPD). The housing 11.3.2-202 can slidably engage with the guide rod to fix the optical module 11.3.2-200 in place within the HMD.
[0146] In at least one example, the optical module 11.3.2-200 may also include a lens 11.3.2-216 coupled to the housing 11.3.2-202 and positioned between the display assembly 11.3.2-204 and the user's eyes when the HMD is worn. The lens 11.3.2-216 may be configured to direct light from the display assembly 11.3.2-204 to the user's eyes. In at least one example, the lens 11.3.2-216 may be part of a lens assembly that includes a corrective lens detachably attached to the optical module 11.3.2-200. In at least one example, lens 11.3.2-216 is positioned above light strip 11.3.2-208 and one or more eye-tracking cameras 11.3.2-206, so that the cameras 11.3.2-206 are configured to capture an image of the user's eye through lens 11.3.2-216, and light strip 11.3.2-208 includes a light configured to project light onto the user's eye through lens 11.3.2-216 during use.
[0147] Any feature, component, and / or part, including its arrangement and configuration shown in Figure 1P, may be included, individually or in any combination, in any other example of devices, features, components, and parts described herein. Similarly, any feature, component, and / or part, including its arrangement and configuration shown and described herein, may be included, individually or in any combination, in the example of devices, features, components, and parts shown in Figure 1P.
[0148] Figure 2 is a block diagram of an example of a controller 110 according to several embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that various other features have been omitted for brevity so as not to obscure more suitable embodiments of the embodiments disclosed herein. Therefore, as a non-limiting example, in some embodiments, the controller 110 includes one or more processing units 202 (e.g., a microprocessor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), graphics processing unit (GPU), central processing unit (CPU), processing core, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FireWire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global Mobile Communication System (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), Bluetooth, ZiGBEE, or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these and various other components.
[0149] In some embodiments, one or more communication buses 204 include circuits for interconnecting and controlling communication between system components. In some embodiments, one or more I / O devices 206 include at least one of the following: a keyboard, mouse, touchpad, joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.
[0150] Memory 220 includes high-speed random-access memory such as dynamic random-access memory (DRAM), static random-access memory (SRAM), double-data-rate random-access memory (DDRRAM), or other random-access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. Memory 220 includes a non-temporary computer-readable storage medium. In some embodiments, memory 220, or the non-temporary computer-readable storage medium of memory 220, stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 230 and XR experience module 240.
[0151] The operating system 230 handles various basic system services and includes instructions for performing hardware-dependent tasks. In some embodiments, the XR experience module 240 is configured to manage and coordinate one or more XR experiences for one or more users (e.g., a single XR experience for one or more users, or multiple XR experiences for each group of one or more users). To this end, in various embodiments, the XR experience module 240 includes a data acquisition unit 241, a tracking unit 242, a coordination unit 246, and a data transmission unit 248.
[0152] In some embodiments, the data acquisition unit 241 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the display generation component 120 of Figure 1A, and optionally from one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the data acquisition unit 241 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0153] In some embodiments, the tracking unit 242 is configured to map scene 105 and track the position / location of at least the display generation component 120 relative to scene 105 in Figure 1A, and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the tracking unit 242 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose. In some embodiments, the tracking unit 242 includes a hand tracking unit 244 and / or an eye tracking unit 243. In some embodiments, the hand tracking unit 244 is configured to track the position / location of one or more parts of the user's hand, and / or the movement of one or more parts of the user's hand relative to scene 105 in Figure 1A, relative to the display generation component 120, and / or relative to a coordinate system defined for the user's hand. The hand tracking unit 244 is described in more detail below with respect to Figure 4. In some embodiments, the eye-tracking unit 243 is configured to track the position and movement of the user's gaze (or, more broadly, the user's eyes, face, or head) relative to the scene 105 (e.g., the physical environment and / or the user (e.g., the user's hands)) or to XR content displayed via the display generation component 120. The eye-tracking unit 243 is described in more detail below with reference to Figure 5.
[0154] In some embodiments, the adjustment unit 246 is configured to manage and adjust the XR experience presented to the user by the display generation component 120 and optionally by one or more of the output devices 155 and / or peripheral devices 195. For this purpose, in various embodiments, the adjustment unit 246 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0155] In some embodiments, the data transmission unit 248 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the display generation component 120, and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the data transmission unit 248 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0156] While the data acquisition unit 241, tracking unit 242 (including, for example, eye-tracking unit 243 and hand-tracking unit 244), adjustment unit 246, and data transmission unit 248 are shown as residing on a single device (e.g., controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 241, tracking unit 242 (including, for example, eye-tracking unit 243 and hand-tracking unit 244), adjustment unit 246, and data transmission unit 248 may be located in separate computing devices.
[0157] Furthermore, Figure 2 is intended to illustrate the function of various features that may be present in a particular implementation, in contrast to the structural schematics of the embodiments described herein. As will be recognized by those skilled in the art, the separately shown items can be combined, and some items can be separated. For example, some functional modules shown separately in Figure 2 may be realized in a single module, and the various functions of a single functional block may be realized by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of certain functions and how features are assigned between them, will vary depending on the implementation and, in some embodiments, will partially depend on a particular combination of hardware, software, and / or firmware selected for that particular implementation.
[0158] Figure 3 is a block diagram of an example of a display generation component 120 according to several embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that various other features have been omitted for brevity so as not to obscure more suitable embodiments of the embodiments disclosed herein. For that purpose, in some non-limiting examples, the display generation component 120 (e.g., HMD) may include one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FireWire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, infrared, Bluetooth, ZiGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more XR displays 312, one or more optional in-facing and / or out-facing image sensors 314, memory 320, and one or more communication buses 304 for interconnecting these and various other components.
[0159] In some embodiments, one or more communication buses 304 include circuits for interconnecting and controlling communication between system components. In some embodiments, one or more I / O devices and sensors 306 include at least one of the following: an inertial measuring unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, one or more depth sensors (e.g., structured light, time of flight, etc.).
[0160] In some embodiments, one or more XR displays 312 are configured to provide the user with an XR experience. In some embodiments, one or more XR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface conduction electron emission display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), MEMS, and / or similar display types. In some embodiments, one or more XR displays 312 correspond to waveguide displays such as diffraction, reflection, polarization, and holographic. For example, a display generation component 120 (e.g., HMD) includes a single XR display. In another example, the display generation component 120 includes an XR display for each of the user's eyes. In some embodiments, one or more XR displays 312 can present MR or VR content.
[0161] In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face, including the user's eyes (and may be referred to as an eye-tracking camera). In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hands and optionally a portion of the user's arms (and may be referred to as a hand-tracking camera). In some embodiments, one or more image sensors 314 are configured to face forward to acquire image data corresponding to a scene that the user would view if a display generation component 120 (e.g., an HMD) were not present (and may be referred to as a scene camera). One or more optional image sensors 314 may include one or more RGB cameras (e.g., complementary metal-oxide-semiconductor (CMOS) image sensors or charge-coupled device (CCD) image sensors), one or more infrared (IR) cameras, one or more event-based cameras, and / or similar.
[0162] Memory 320 includes high-speed random-access memory, such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. Memory 320 includes a non-temporary computer-readable storage medium. In some embodiments, memory 320, or the non-temporary computer-readable storage medium of memory 320, stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 330 and XR presentation module 340.
[0163] The operating system 330 includes instructions for handling various basic system services and instructions for performing hardware-dependent tasks. In some embodiments, the XR presentation module 340 is configured to present XR content to the user via one or more XR displays 312. For this purpose, in various embodiments, the XR presentation module 340 includes a data acquisition unit 342, an XR presentation unit 344, an XR map generation unit 346, and a data transmission unit 348.
[0164] In some embodiments, the data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the controller 110 in Figure 1A. To this end, in various embodiments, the data acquisition unit 342 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0165] In some embodiments, the XR presentation unit 344 is configured to present XR content via one or more XR displays 312. For this purpose, in various embodiments, the XR presentation unit 344 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0166] In some embodiments, the XR map generation unit 346 is configured to generate an XR map (for example, a 3D map of a mixed reality scene or a map of a physical environment in which computer-generated objects can be placed to generate extended reality) based on media content data. For this purpose, in various embodiments, the XR map generation unit 346 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0167] In some embodiments, the data transmission unit 348 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the controller 110 and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the data transmission unit 348 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0168] Although the data acquisition unit 342, XR presentation unit 344, XR map generation unit 346, and data transmission unit 348 are shown as existing on a single device (e.g., the display generation component 120 in Figure 1A), it should be understood that in other embodiments, any combination of the data acquisition unit 342, XR presentation unit 344, XR map generation unit 346, and data transmission unit 348 may be located in separate computing devices.
[0169] Furthermore, Figure 3 is intended to illustrate the function of various features that may be present in a particular implementation, in contrast to the structural schematics of the embodiments described herein. As will be recognized by those skilled in the art, the separately shown items can be combined, and some items can be separated. For example, several functional modules shown separately in Figure 3 can be realized within a single module, and the various functions of a single functional block can be realized by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of certain functions and how features are assigned between them, will vary depending on the implementation and, in some embodiments, will partially depend on a particular combination of hardware, software, and / or firmware selected for a particular implementation.
[0170] Figure 4 is a schematic diagram of an exemplary embodiment of the hand tracking device 140. In some embodiments, the hand tracking device 140 (Figure 1A) is controlled by a hand tracking unit 244 (Figure 2) to track the position / location of one or more parts of the user's hand and / or the movement of one or more parts of the user's hand relative to the scene 105 in Figure 1A (e.g., relative to a part of the physical environment surrounding the user, relative to the display generation component 120, or relative to a part of the user (e.g., the user's face, eyes, or head), and / or relative to a coordinate system defined for the user's hand). In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).
[0171] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras) that captures three-dimensional scene information including at least the hand 406 of a human user. The image sensor 404 captures a hand image with sufficient resolution to allow for the distinction of fingers and their respective positions. The image sensor 404 can typically capture images of other parts of the user's body, or images of the entire body, and may have either a zoom function or a dedicated sensor with high magnification to capture an image of the hand at a desired resolution. In some embodiments, the image sensor 404 also captures a 2D color video image of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with other image sensors that capture the physical environment of the scene 105, or functions as an image sensor that captures the physical environment of the scene 105. In some embodiments, the image sensor 404 is positioned relative to the user or the user's environment such that the field of view of the image sensor or a portion thereof is used to define an interaction space in which hand movements captured by the image sensor are processed as input to the controller 110.
[0172] In some embodiments, the image sensor 404 outputs a sequence of frames containing 3D map data (and possibly color image data) to the controller 110, thereby extracting high-level information from the map data. This high-level information is typically provided to an application running on the controller via an application programming interface (API), which drives the display generation components 120 accordingly. For example, a user can interact with the software running on the controller 110 by moving their hand 406 to change the orientation of their hand.
[0173] In some embodiments, the image sensor 404 projects a spot pattern onto a scene including the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral shift of the spot in the pattern. This approach is advantageous in that the user does not need to hold or wear any kind of beacon, sensor, or other marker. This gives the depth coordinates of points in the scene relative to a given reference plane at a specific distance from the image sensor 404. In this disclosure, it is assumed that the image sensor 404 defines an orthogonal set of x, y, and z axes such that the depth coordinates of points in the scene correspond to a z component measured by the image sensor. Alternatively, the image sensor 404 (e.g., a hand tracking device) may use other 3D mapping methods such as stereoscopic imaging or time-of-flight measurement based on one or more cameras or other types of sensors.
[0174] In some embodiments, the hand tracking device 140 captures and processes a time sequence of depth maps containing the user's hand while the user moves their hand (e.g., the entire hand or one or more fingers). Software running on the processor in the image sensor 404 and / or controller 110 processes the 3D map data to extract patch descriptors of the hand within these depth maps. Based on previous training, the software matches these descriptors against patch descriptors stored in the database 408 to estimate the hand pose in each frame. The pose typically includes the 3D location of the user's wrist and fingertips.
[0175] The software can also analyze the trajectory of the hand and / or fingers across multiple frames in a sequence to identify gestures. The posture estimation function described herein may be interleaved with the motion tracking function, so that patch-based posture estimation is performed only once every two (or more) frames, while tracking is used to detect changes in posture that occur over the remaining frames. Posture, motion, and gesture information is provided to an application program running on the controller 110 via the API described above. This program can, for example, move and modify the image presented on the display generation component 120, or perform other functions, depending on the posture and / or gesture information.
[0176] In some embodiments, the gesture includes an air gesture. An air gesture is a gesture detected by the user without (or independently of) touching an input element that is part of a device (e.g., a computer system 101, one or more input devices 125, and / or a hand tracking device 140), and is based on detected movement of a part of the user's body in the air (e.g., head, one or more arms, one or more hands, one or more fingers, and / or one or more legs), including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground, or the distance of the user's hand relative to the ground), movement of the user's body relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the user's other hand, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tap gesture including movement of the hand in a predetermined posture by a predetermined amount and / or speed, or a shake gesture including a predetermined speed or amount of rotation of a part of the user's body).
[0177] In some embodiments, the input gestures used in the various examples and embodiments described herein include air gestures, as in some embodiments, performed by moving one or more of the user's fingers relative to other fingers or parts of the user's hand for interacting with an XR environment (e.g., a virtual or mixed reality environment). In some embodiments, an air gesture is a gesture detected without the user touching (or independently of) an input element that is part of the device, and is based on detected movement of a part of the user's body in the air, including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground, or the distance of the user's hand relative to the ground), movement of the user's body relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of the user's other hand relative to one hand, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tap gesture involving movement of the hand in a predetermined pose by a predetermined amount and / or speed, or a shake gesture involving rotation of a part of the user's body by a predetermined speed or amount).
[0178] In some embodiments where the input gesture is an air gesture (i.e., without physical contact with an input device that provides the computer system with information about which user interface element is the target of user input, such as contact with a user interface element displayed on a touchscreen or contact with a mouse or trackpad to move a cursor over a user interface element), the gesture takes into account the user's attention (e.g., gaze) to determine the target of user input (e.g., in the case of direct input, as described below). Thus, in implementations involving air gestures, the input gesture is the detected attention (e.g., gaze) to the user interface element in combination (e.g., simultaneously) with the movement of the user's fingers (one or more) and / or hand to perform pinch and / or tap input, as described in more detail below.
[0179] In some embodiments, input gestures directed towards a user interface object are performed directly or indirectly by reference to the user interface object. For example, user input is performed directly towards the user interface object in response to the user performing an input gesture with their hand at a position corresponding to the user interface object's position in a three-dimensional environment (e.g., determined based on the user's current viewpoint). In some embodiments, the input gesture is performed indirectly towards the user interface object according to the user performing the input gesture while the user's hand position is not at a position corresponding to the user interface object's position in a three-dimensional environment, while detecting the user's attention (e.g., gaze) to the user interface object. For example, in the case of a direct input gesture, the user can direct their input towards the user interface object by initiating the gesture at or near a position corresponding to the user interface object's display position (e.g., within a distance of 0.5 cm, 1 cm, 5 cm, or 0-5 cm from the optional outer edge or optional central portion). In the case of indirect input gestures, the user can direct their input towards the user interface object by paying attention to the user interface object (for example, by gazing at the user interface object), and while paying attention to the options, the user initiates the input gesture (for example, at any position detectable by the computer system) (for example, at a position that does not correspond to the display position of the user interface object).
[0180] In some embodiments, the input gestures (e.g., air gestures) used in the various examples and embodiments described herein include pinch and tap inputs for interacting with virtual or mixed reality environments, as in some embodiments. For example, the pinch and tap inputs described later are performed as air gestures.
[0181] In some embodiments, a pinch input is part of an air gesture that includes one or more of the following: a pinch gesture, a long pinch gesture, a pinch-and-drag gesture, or a double pinch gesture. For example, a pinch gesture that is an air gesture involves moving two or more fingers of a hand to touch each other, i.e., including an optional interruption (e.g., within 0 to 1 second) immediately after the touch. A long pinch gesture that is an air gesture involves moving two or more fingers of a hand to touch each other for at least a threshold time amount (e.g., at least 1 second) before detecting an interruption of contact between them. For example, a long pinch gesture includes the user holding a pinch gesture (e.g., if two or more fingers are in contact), and the long pinch gesture continues until an interruption of contact between the two or more fingers is detected. In some embodiments, a double pinch gesture that is an air gesture includes two (e.g., or more) pinch inputs (e.g., performed with the same hand) that are detected directly and consecutively (e.g., within a predetermined period of time) to each other. For example, the user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., breaks contact between two or more fingers), and then performs a second pinch input within a predetermined period (e.g., within 1 second or 2 seconds) after releasing the first pinch input.
[0182] In some embodiments, an air gesture, a pinch-and-drag gesture, includes a pinch gesture (e.g., a pinch gesture or a long pinch gesture) performed in relation to (e.g., after) a drag input that changes the user's hand position from a first position (e.g., a drag initiation position) to a second position (e.g., a resistance termination position). In some embodiments, the user maintains the pinch gesture while performing the drag input and releases the pinch gesture (e.g., spreading two or more fingers) to terminate the drag gesture (e.g., at the second position). In some embodiments, the pinch input and drag input are performed by the same hand (e.g., the user pinches two or more fingers together and touches them to each other, and then moves the same hand to a second position in the air with a drag gesture). In some embodiments, the pinch input is performed by the user's first hand and the drag input is performed by the user's second hand (e.g., the user's second hand moves from the first position to the second position in the air while the user continues the pinch input with the user's first hand). In some embodiments, an input gesture that is an air gesture includes an input (e.g., a pinch input and / or a tap input) performed using both of the user's hands. For example, an input gesture includes two (e.g., or more) pinch inputs performed in relation to each other (e.g., simultaneously or within a predetermined period of time). For example, a first pinch gesture (e.g., a pinch input, a long pinch input, or a pinch and drag input) performed using the user's first hand, and a second pinch input performed using the other hand (e.g., a second hand of the user's hands) in relation to performing the pinch input using the first hand. In some embodiments, movement occurs between the user's hands (e.g., to increase and / or decrease the distance or relative orientation between the user's hands).
[0183] In some embodiments, a tap input performed as an air gesture (e.g., directed towards a user interface element) includes the movement of one or more of the user's fingers toward the user interface element, the movement of the user's hand toward the user interface element with the user's fingers (one or more) optionally extended toward the user interface element, a downward movement of the user's fingers (e.g., mimicking a mouse click or a tap on a touchscreen), or other default movements of the user's hand. In some embodiments, a tap input performed as an air gesture is detected based on the movement characteristics of the finger or hand that performs the tap gesture movement away from the user's viewpoint and / or toward the object that is the target of the tap input, followed by the end of the movement. In some embodiments, the end of the movement is detected based on a change in the movement characteristics of the finger or hand that performs the tap gesture (e.g., away from the user's viewpoint and / or the end of the movement toward the object that is the target of the tap input, a reversal of the direction of the finger or hand movement, and / or a reversal of the direction of acceleration of the finger or hand movement).
[0184] In some embodiments, the user's attention is determined to be directed towards a part of the three-dimensional environment based on the detection of a gaze directed towards that part of the three-dimensional environment (optionally, without requiring any other conditions). In some embodiments, for the device to determine that the user's attention is directed towards a part of the three-dimensional environment, the device determines that the user's attention is directed towards a part of the three-dimensional environment based on the detection of a gaze directed towards a part of the three-dimensional environment, with one or more additional conditions such as the gaze being directed towards the part of the three-dimensional environment for at least a threshold duration (e.g., dwell time) while the user's viewpoint is within a distance threshold from the part of the three-dimensional environment, and / or the gaze being directed towards a part of the three-dimensional environment. If one of the additional conditions is not met, the device determines that the user's attention is not directed towards the part of the three-dimensional environment to which the gaze is directed (e.g., until one or more additional conditions are met).
[0185] In some embodiments, the detection of a ready state configuration of a user or part of a user is detected by the computer system. The detection of a ready state configuration of a hand is used by the computer system as an indication that the user is likely to be preparing to interact with the computer system using one or more air gesture inputs performed by the hand (e.g., pinch, tap, pinch and drag, double pinch, long pinch, or other air gestures described herein). For example, the ready state of a hand is determined based on whether the hand has a predetermined hand shape (e.g., a pre-pinch shape where the thumb and one or more fingers are extended and separated, ready to perform a pinch or grab gesture, or a pre-tap shape where one or more fingers are extended and the palm is facing away from the user), whether the hand is in a predetermined position relative to the user's viewpoint (e.g., below the user's head, above the user's waist, extended at least 15 cm, 20 cm, 25 cm, 30 cm, or 50 cm from the body), and / or whether the hand has moved in a particular manner (e.g., moved towards the area in front of the user above the user's waist, below the user's head, or away from the user's body or legs). In some embodiments, the ready state is used to determine whether an interaction element of the user interface is responsive to attention (e.g., gaze) input.
[0186] In scenarios where the input is described in reference to an air gesture, similar gestures may also be detected using hardware input devices attached to or held by one or more of the user's hands, in which case the position of the hardware input device in space may be tracked using optical tracking, one or more accelerometers, one or more gyroscopes, one or more magnetometers, and / or one or more inertial measurement units, and it should be understood that the position and / or movement of the hardware input device is used instead of the position and / or movement of one or more hands in the corresponding air gesture(s). User input can be detected using controls included in hardware input devices, such as one or more touch-sensitive input elements, one or more pressure-sensitive input elements, one or more buttons, one or more knobs, one or more dials, one or more joysticks, one or more hand or finger covers capable of detecting the position or change in position of parts of the hands and / or fingers relative to each other, relative to the user's body, and / or the user's physical environment, and / or other hardware input device controls. User input using controls included in hardware input devices is used in place of hand and / or finger gestures such as air taps or air pinches in corresponding air gestures(single or multiple). For example, a selection input described as being performed by an air tap or air pinch input can alternatively be detected by a button press, a tap on a touch-sensitive surface, a press on a pressure-sensitive surface, or other hardware input.As another example, a movement input described as being performed by air pinch and drag can alternatively be detected based on interaction with hardware input controls such as button press and hold, touch on a touch-sensitive surface, or press on a pressure-sensitive surface, or based on hardware input that follows the movement of other hardware input devices in space (e.g., accompanying the hand to which the hardware input device is associated). Similarly, two-handed inputs, including movements of both hands relative to each other, can also be performed using various combinations of inputs detected by air gestures and / or one or more of the aforementioned hardware input devices, using one air gesture and one hardware input device held in the hand not performing the air gesture, two hardware input devices held in separate hands, or two air gestures performed by separate hands.
[0187] In some embodiments, the software may be downloaded electronically to the controller 110, for example, over a network, or instead, it may be provided on a tangible non-temporary medium such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is similarly stored in memory associated with the controller 110. Alternatively or additionally, some or all of the computer's described functions may be executed on dedicated hardware such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP). Although the controller 110 is shown in Figure 4, for example, as a separate unit from the image sensor 404, some or all of the controller's processing functions may be associated with the image sensor 404 by a suitable microprocessor and software, or by dedicated circuitry within the housing of the image sensor 404 (e.g., a hand-tracking device), or in other ways. In some embodiments, at least some of these processing functions may be executed by a suitable processor integrated with the display generation component 120 (e.g., in a television set, handheld device, or head-mounted device), or by any other suitable computerized device such as a game console or media player. The sensing function of the image sensor 404 can also be integrated into a computer or other computerized device controlled by the sensor output.
[0188] Figure 4 further includes schematic diagrams of depth maps 410 captured by image sensor 404 according to several embodiments. The depth map includes a matrix of pixels, each having a depth value, as described above. Pixels 412 corresponding to the hand 406 are segmented in this map from the background and the wrist. The brightness of each pixel in the depth map 410 is inversely proportional to the depth value, i.e., the measured z-distance from image sensor 404, with the gradation becoming darker as the depth increases. Controller 110 processes these depth values to identify and segment image components (i.e., groups of adjacent pixels) that have the characteristics of a human hand. These characteristics may include, for example, the overall size, shape, and frame-to-frame movement of the depth map sequence.
[0189] Figure 4 also schematically shows the hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406, according to several embodiments. In Figure 4, the hand skeleton 414 is superimposed on the hand background 416, which has been segmented from the original depth map. In some embodiments, the hand (e.g., knuckles, fingertips, center of the palm, end of the hand connected to the wrist), and optionally major feature points on the wrist or arm connected to the hand, are identified and positioned on the hand skeleton 414. In some embodiments, the location and movement of these major feature points across multiple image frames are used by the controller 110 to determine, according to several embodiments, a hand gesture performed by the hand or the current state of the hand.
[0190] Figure 5 shows an exemplary embodiment of the eye-tracking device 130 (Figure 1A). In some embodiments, the eye-tracking device 130 is controlled by an eye-tracking unit 243 (Figure 2) to track the position and movement of the user's gaze toward the scene 105 or toward the XR content displayed via the display generation component 120. In some embodiments, the eye-tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, if the display generation component 120 is a head-mounted device such as a headset, helmet, goggles, or glasses, or a handheld device positioned in a wearable frame, the head-mounted device includes both a component for generating XR content for user viewing and a component for tracking the user's gaze toward the XR content. In some embodiments, the eye-tracking device 130 is separate from the display generation component 120. For example, if the display generation component is a handheld device or an XR chamber, the eye-tracking device 130 is optionally a separate device from the handheld device or XR chamber. In some embodiments, the eye-tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, the head-mounted eye-tracking device 130 is optionally used with a display generation component that is mounted on the head or a display generation component that is not mounted on the head. In some embodiments, the eye-tracking device 130 is not a head-mounted device, but is optionally used in combination with a head-mounted display generation component. In some embodiments, the eye-tracking device 130 is not a head-mounted device, but is optionally part of a non-head-mounted display generation component.
[0191] In some embodiments, the display generation component 120 uses a display mechanism (e.g., left and right near-eye display panels) that displays frames containing left and right images in front of the user's eyes to provide the user with a 3D virtual view. For example, the head-mounted display generation component may include left and right optical lenses (referred to herein as eyepieces) positioned between the display and the user's eyes. In some embodiments, the display generation component may include, or be coupled to, one or more external video cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display on which the user can directly view the physical environment and display virtual objects on a transparent or translucent display. In some embodiments, the display generation component projects virtual objects onto the physical environment. The virtual objects are projected, for example, onto a physical surface or as holograms, so that the individual can use the system to observe the virtual objects superimposed on the physical environment. In such cases, separate display panels and image frames for the left and right eyes may not be required.
[0192] As shown in Figure 5, in some embodiments, the eye-tracking device 130 (e.g., gaze tracking device) includes at least one eye-tracking camera (e.g., an infrared (IR) camera or a near-IR (NIR) camera) and an illumination source (e.g., an IR or NIR light source such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eye. The eye-tracking camera may be directed toward the user's eye to receive reflected IR or NIR light from the light source directly from the eye, or alternatively, it may be directed toward a "hot" mirror positioned between the user's eye and a display panel that reflects IR or NIR light from the eye to the eye-tracking camera while allowing visible light to pass through. The eye-tracking device 130 optionally captures images of the user's eye (e.g., as a video stream captured at 60-120 frames per second (fps)), analyzes the images to generate gaze tracking information, and communicates the gaze tracking information to the controller 110. In some embodiments, the user's two eyes are tracked separately by separate eye-tracking cameras and illumination sources. In some embodiments, only one of the user's eyes is tracked by a separate eye-tracking camera and light source.
[0193] In some embodiments, the eye-tracking device 130 is calibrated using a device-specific calibration process to determine the parameters of the eye-tracking device for a specific operating environment 100, e.g., the 3D geometric relationships and parameters of the LEDs, camera, hot mirror (if present), eyepiece, and display screen. The device-specific calibration process may be performed at the factory or another facility before delivery of the AR / VR device to the end user. The device-specific calibration process may be an automated calibration process or a manual calibration process. A user-specific calibration process may include estimating the eye parameters of a particular user, e.g., pupil location, central visual location, optical axis, visual axis, interpupillary distance. According to some embodiments, once the device-specific and user-specific parameters for the eye-tracking device 130 are determined, the images captured by the eye-tracking camera can be processed using a Glint-assisted method to determine the user's current visual axis and gaze point relative to the display.
[0194] As shown in Figure 5, the eye-tracking device 130 (e.g., 130A or 130B) includes an eyepiece(s) 520, and a gaze tracking system which includes at least one eye-tracking camera 540 (e.g., an infrared (IR) or near-IR (NIR) camera) positioned on the side of the user's face where eye tracking is performed, and an illumination source 530 (e.g., an IR or NIR light source such as an array or ring of NIR light-emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eyes(s) 592. The eye-tracking camera 540 is positioned between the user's eye(s) 592 and the display 510 (e.g., the left or right display panel of a head-mounted display, or the display or projector of a handheld device) and may be directed towards a mirror 550 that transmits visible light while reflecting IR or NIR light from the eye(s) 592 (e.g., as shown at the top of Figure 5), or may be directed towards the user's eye(s) 592 to receive reflected IR or NIR light from the eye(s) 592 (e.g., as shown at the bottom of Figure 5).
[0195] In some embodiments, the controller 110 renders AR or VR frames 562 (e.g., left and right frames of left and right display panels) and provides the frames 562 to the display 510. For various purposes, for example, when processing the frames 562 for display, the controller 110 uses gaze tracking input 542 from the eye-tracking camera 540. The controller 110 optionally uses a glint-assisted method or other appropriate method to estimate the user's viewpoint on the display 510 based on the gaze tracking input 542 obtained from the eye-tracking camera 540. The viewpoint estimated from the gaze tracking input 542 is optionally used to determine the direction the user is currently looking.
[0196] The following describes, but is not intended to be limiting, several possible use cases of the user's current gaze direction. As an exemplary use case, the controller 110 may render virtual content differently based on the determined user's gaze direction. For example, the controller 110 may generate virtual content at a higher resolution in the central visual region determined from the user's current gaze direction than in the peripheral region. As another example, the controller may position or move virtual content within the view based at least partially on the user's current gaze direction. As yet another example, the controller may display specific virtual content within the view based at least partially on the user's current gaze direction. As another exemplary use case in an AR application, the controller 110 may capture the physical environment of the XR experience and orient an external camera to focus in the determined direction. The external camera's autofocus mechanism can then focus on an object or surface in the environment that the user is currently viewing on the display 510. In another exemplary use case, the eyepiece 520 may be a focusing lens, and the controller uses eye-tracking information to adjust the focus of the eyepiece 520 so that the virtual object currently being viewed by the user has appropriate binocular coordination to match the convergence of the user's eye 592. The controller 110 can use the eye-tracking information to orient and adjust the focus of the eyepiece 520 so that the nearby object being viewed by the user appears at the correct distance.
[0197] In some embodiments, the eye-tracking device is part of a head-mounted device, which is housed within a wearable housing and includes a display (e.g., display 510), two eyepieces (e.g., eyepieces 520), an eye-tracking camera (e.g., one or more eye-tracking cameras 540), and a light source (e.g., an illumination source 530 (e.g., IR or NIR LEDs)). The light source emits light (e.g., IR or NIR light) towards the user's eye(s) 592. In some embodiments, the light sources may be arranged in a ring or circle around each lens, as shown in Figure 5. In some embodiments, as an example, eight illumination sources 530 (e.g., LEDs) are arranged around each lens 520. However, more or fewer illumination sources 530 may be used, and other arrangements and locations of the illumination sources 530 may be used.
[0198] In some embodiments, the display 510 emits light within the visible light range and does not emit light within the IR or NIR range, thus not introducing noise into the eye-tracking system. Note that the location and angle of the eye-tracking camera(s) 540 are given as examples and are not intended to be limiting. In some embodiments, a single eye-tracking camera 540 is positioned on each side of the user's face. In some embodiments, two or more NIR cameras 540 can be used on each side of the user's face. In some embodiments, a camera 540 with a wider field of view (FOV) and a camera 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, a camera 540 operating at one wavelength (e.g., 850 nm) and a camera 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.
[0199] An embodiment of the eye-tracking system shown in Figure 5 can be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide users with computer-generated reality, virtual reality, augmented reality, and / or augmented virtual experiences.
[0200] Figure 6 shows glint-assisted eye-tracking pipelines according to several embodiments. In some embodiments, the eye-tracking pipeline is implemented by a glint-assisted eye-tracking system (e.g., an eye-tracking device 130 as shown in Figures 1A and 5). The glint-assisted eye-tracking system can maintain a tracking state. Initially, the tracking state is off or "no". When in tracking state, the glint-assisted eye-tracking system tracks the pupil contour and glint in the current frame by using prior information from previous frames when analyzing the current frame. When not in tracking state, the glint-assisted eye-tracking system attempts to detect the pupil and glint in the current frame, and if successful, initializes the tracking state to "yes" and continues in tracking state for the next frame.
[0201] As shown in Figure 6, the eye-tracking camera can capture left and right images of the user's left and right eyes. The captured images are then fed into the eye-tracking pipeline for processing, which is initiated at 610. As indicated by the arrow returning to element 600, the eye-tracking system can continue to capture images of the user's eyes at a rate of, for example, 60 to 120 frames per second. In some embodiments, each set of captured images may be fed into the pipeline for processing. However, in some embodiments, or under some conditions, not all captured frames are processed by the pipeline.
[0202] At 610, if the tracking status is yes for the currently captured image, the method proceeds to element 640. At 610, if the tracking status is no, the image is analyzed to detect the user's pupil and glint in the image, as shown in 620. At 630, if the pupil and glint are successfully detected, the method proceeds to element 640. If they are not successfully detected, the method returns to element 610 and processes the next image of the user's eyes.
[0203] At 640, if the process proceeds from element 610, the current frame is analyzed to track the pupil and glint based in part on previous information from the previous frame. At 640, if the process proceeds from element 630, the tracking state is initialized based on the detected pupil and glint in the current frame. The results of the processing at element 640 are checked to ensure that the tracking or detection results are reliable. For example, the results may be checked to determine whether a sufficient number of glints for pupil and gaze estimation are successfully tracked or detected in the current frame. At 650, if the results are unreliable, the tracking state is set to no at element 660, and the method returns to element 610 to process the next image of the user's eyes. At 650, if the results are reliable, the method proceeds to element 670. At 670, the tracking state is set to yes (if not already yes), and the pupil and glint information is passed to element 680 to estimate the user's gaze.
[0204] Figure 6 is intended to serve as an example of eye-tracking technology that may be used in a particular implementation. As will be recognized by those skilled in the art, other eye-tracking technologies that currently exist or may be developed in the future may be used in computer system 101 to provide users with XR experiences in various embodiments, either in place of or in combination with the Glint-assisted eye-tracking technology described herein.
[0205] In some embodiments, the captured portion of the real-world environment 602 is used to provide the user with an XR experience, for example, a mixed reality environment in which one or more virtual objects are superimposed on a representation of the real-world environment 602.
[0206] Accordingly, this description describes several embodiments of three-dimensional environments (e.g., XR environments) that include representations of real-world objects and virtual objects. For example, a three-dimensional environment optionally includes a representation of a table existing in a physical environment, which is captured and displayed within the three-dimensional environment (e.g., actively via a computer system's camera and display, or passively via a computer system's transparent or translucent display). As described above, a three-dimensional environment optionally is a mixed reality system based on a physical environment, in which the three-dimensional environment is captured by one or more sensors of a computer system and displayed via a display generation component. As a mixed reality system, the computer system may optionally selectively display parts and / or objects of the physical environment so that each part and / or object of the physical environment appears to exist in the three-dimensional environment displayed by the computer system. Similarly, the computer system may optionally display virtual objects in a three-dimensional environment so that the virtual objects appear to exist in the real world (e.g., the physical environment) by placing virtual objects in each location within the three-dimensional environment that have corresponding locations in the real world. For example, a computer system may optionally display a vase in such a way that it appears as if a real vase were placed on a table in a physical environment. In some embodiments, individual locations in a three-dimensional environment have corresponding locations in the physical environment.Therefore, when a computer system is described as displaying virtual objects in separate locations relative to physical objects (for example, at or near the location of the user's hand, or on or near a physical table), the computer system displays the virtual objects in specific locations within a three-dimensional environment so that they appear to be at or near physical objects in the physical world (for example, if the virtual object is a real object at that specific location, then the virtual object will be displayed in the location within the three-dimensional environment that corresponds to the location within the physical environment where the virtual object would have been displayed).
[0207] In some embodiments, real-world objects existing in a physical environment displayed within a three-dimensional environment (e.g., real-world objects visible via and / or display-generating components) can interact with virtual objects existing only within the three-dimensional environment. For example, the three-dimensional environment may include a table and a vase placed on the table, where the table is a view (or representation) of a physical table in the physical environment, and the vase is a virtual object.
[0208] In a three-dimensional environment (for example, a real environment, a virtual environment, or an environment including a mixture of real and virtual objects), an object may be said to have depth or simulated depth, or an object may be said to be visible, displayed, or positioned at a different depth. In this context, depth refers to dimensions other than height or width. In some embodiments, depth is defined relative to a fixed set of coordinates (for example, a room or object has height, depth, and width defined relative to a fixed set of coordinates). In some embodiments, depth is defined relative to the user's location or viewpoint, in which case the depth dimension varies based on the user's location and / or the location and angle of the user's viewpoint. In some embodiments where depth is defined relative to the user's location positioned with respect to the surface of the environment (e.g., the floor or ground surface of the environment), objects that are further away from the user along a line extending parallel to the surface are considered to have a greater depth in the environment, and / or the depth of an object is measured along an axis that extends outward from the user's location and is parallel to the surface of the environment (e.g., depth is defined in a coordinate system of a cylinder or substantially a cylinder, with the user's position at the center of a cylinder extending from the user's head to the user's feet). In some embodiments, depth is defined relative to the user's viewpoint (e.g., a direction relative to a point in space that determines which parts of the environment are visible through a head-mounted device or other display). Objects that are further away from the user's viewpoint along a line extending parallel to the user's viewpoint are considered to have greater depth in the environment, and / or the depth of an object is measured along an axis extending outward from a line that extends from the user's viewpoint and is parallel to the user's viewpoint (e.g., depth is defined in a spherical or substantially spherical coordinate system with the origin of the viewpoint at the center of a sphere extending outward from the user's head).In some embodiments, depth is defined relative to a user interface container (e.g., a window or application on which application and / or system content is displayed), where the user interface container has height and / or width, and depth is a dimension orthogonal to the height and / or width of the user interface container. In some embodiments, where depth is defined relative to a user interface container, the height and / or width of the container is typically orthogonal or substantially orthogonal to a line extending from a user-based location (e.g., the user's viewpoint or the user's location) to the user interface container (e.g., the center of the user interface container, or another feature point of the user interface container) when the container is placed in a three-dimensional environment or is first displayed (e.g., consequently, the depth dimension of the container extends outward away from the user or the user's viewpoint). In some embodiments, where depth is defined relative to a user interface container, the depth of an object relative to the user interface container refers to the position of the object along the depth dimension of the user interface container. In some embodiments, multiple different containers may have different depth dimensions (e.g., different depth dimensions extending in different directions from the user or the user's viewpoint and / or away from different starting points). In some embodiments, when depth is defined relative to a user interface container, the direction of the depth dimension remains constant relative to the user interface container when the location of the user interface container, the user, and / or the user's viewpoint changes (e.g., when multiple different viewers are viewing the same container in a three-dimensional environment, such as during a face-to-face collaboration session, and / or when multiple participants are in a real-time communication session with shared virtual content containing the container). In some embodiments, for curved containers (e.g., including containers with curved surfaces or curved content areas), the depth dimension optionally extends within the surface of the curved container.In some contexts, z-separation (e.g., separation of two objects in depth dimensions), z-height (e.g., distance of one object from another object in depth dimensions), z-position (e.g., position of one object in depth dimensions), z-depth (e.g., position of one object in depth dimensions), or simulated z-dimension (e.g., depth used as object dimensions, environment dimensions, orientation in space, and / or orientation in simulated space) are used to refer to the concepts of depth as described above.
[0209] In some embodiments, the user may optionally interact with virtual objects in a three-dimensional environment using one or more hands, as if the virtual objects were real objects in a physical environment. For example, as described above, one or more sensors in the computer system may optionally capture one or more of the user's hands and display a representation of the user's hands in the three-dimensional environment (in a similar manner to, for example, displaying real-world objects in the three-dimensional environment as described above), or, in some embodiments, the user's hands are visible through the display-generating components by the ability to see the physical environment through the user interface, due to the transparency / transparency of some of the display-generating components displaying the user interface, or the projection of the user interface onto a transparent / translucent surface, or the projection of the user interface onto the user's eyes or the user's field of view. Thus, in some embodiments, the user's hands are displayed at separate locations in the three-dimensional environment and are processed as if they were objects in the three-dimensional environment that can interact with virtual objects in the three-dimensional environment as if they were real physical objects in the physical environment. In some embodiments, the computer system may update the display of the user's hands in the three-dimensional environment in conjunction with the movement of the user's hands in the physical environment.
[0210] In some of the embodiments described below, for example, to determine whether a physical object is directly interacting with a virtual object (e.g., whether a hand is touching, grasping, or holding a virtual object, or whether it is within a threshold distance from the virtual object), the computer system may optionally determine the "effective" distance between the physical object in the physical world and the virtual object in the three-dimensional environment. For example, a hand directly interacting with a virtual object may optionally include one or more of the fingers of a hand pressing a virtual button, a user's hand grasping a virtual vase, two fingers of a user's hand pinching / holding an application's user interface together, and other types of interactions described herein. For example, when determining whether a user is interacting with a virtual object and / or how a user is interacting with a virtual object, the computer system may optionally determine the distance between the user's hand and the virtual object. In some embodiments, the computer system determines the distance between the user's hand and the virtual object by determining the distance between the location of the hand in the three-dimensional environment and the location of the virtual object of interest in the three-dimensional environment. For example, one or more of the user's hands are located in a specific position in the physical world, which the computer system optionally captures and displays at a specific corresponding position in a three-dimensional environment (e.g., the position in the three-dimensional environment where the hands are displayed, if the hands are virtual hands rather than physical hands). The position of the hands in the three-dimensional environment is optionally compared to the position of a target virtual object in the three-dimensional environment to determine the distance between the one or more of the user's hands and the virtual object. In some embodiments, the computer system optionally determines the distance between the physical object and the virtual object by comparing the position in the physical world (as opposed to comparing the position in the three-dimensional environment).For example, when determining the distance between one or more of the user's hands and a virtual object, the computer system optionally determines the corresponding location of the virtual object in the physical world (e.g., the position in the physical world where the virtual object would be located if it were a physical object rather than a virtual object), and then determines the distance between the corresponding physical position and one or more of the user's hands. In some embodiments, the same technique is optionally used to determine the distance between any physical object and any virtual object. Thus, when determining whether a physical object is in contact with a virtual object, or whether a physical object is within a threshold distance of a virtual object, as described herein, the computer system optionally performs one of the techniques described above to map the location of the physical object to a three-dimensional environment and / or to map the location of the virtual object to a physical environment.
[0211] In some embodiments, the same or similar techniques are used to determine where and what the user's gaze is directed, and / or where and what the physical stylus held by the user is directed. For example, if the user's gaze is directed to a particular position in the physical environment, the computer system optionally determines the corresponding position in the three-dimensional environment (e.g., the virtual position of the gaze), and if a virtual object is located at that corresponding virtual position, the computer system optionally determines that the user's gaze is directed to that virtual object. Similarly, the computer system optionally determines, based on the orientation of the physical stylus, where in the physical environment the stylus is pointing. In some embodiments, based on this determination, the computer system optionally determines the corresponding virtual position in the three-dimensional environment corresponding to the location in the physical environment that the stylus is pointing to, and optionally determines that the stylus is pointing to the corresponding virtual position in the three-dimensional environment.
[0212] Similarly, embodiments described herein may refer to the location of a user (e.g., a user of a computer system) and / or the location of a computer system in a three-dimensional environment. In some embodiments, the user of a computer system is holding, wearing, or otherwise positioned near the computer system. Thus, in some embodiments, the location of the computer system is used as a proxy for the user's location. In some embodiments, the location of the computer system and / or the user in the physical environment corresponds to individual locations in the three-dimensional environment. For example, if a user stands at a location facing an individual part of the physical environment that is visible through a display-generating component, the location of the computer system is the location in the physical environment (and its corresponding location in the three-dimensional environment) where the user will see objects in the physical environment in the same position, orientation, and / or size (e.g., absolutely and / or relative to each other) as the objects are visible through the display-generating component of the computer system in the three-dimensional environment. Similarly, if a virtual object displayed in a three-dimensional environment is a physical object in a physical environment (for example, the physical object is located in the same physical environment location as the one in the three-dimensional environment and has the same size and orientation as the one in the three-dimensional environment), then the computer system and / or user's location is the position from which the user will view the virtual object in the physical environment in the same position, orientation, and / or size (for example, absolutely, and / or relative to each other, and in relation to real-world objects) as it was displayed by the computer system's display generation components in the three-dimensional environment.
[0213] This disclosure describes various input methods for interaction with computer systems. Where one example is provided using one input device or method, and another example is provided using a different input device or method, each example may be compatible with the input device or method described in the other example, and their use should be considered optional. Similarly, various output methods for interaction with computer systems are described. Where one example is provided using one output device or method, and another example is provided using a different output device or method, each example may be compatible with the output device or method described in the other example, and their use should be considered optional. Similarly, various methods for interaction with virtual or mixed reality environments via computer systems are described. Where one example is provided using interaction with a virtual environment, and another example is provided using a mixed reality environment, each example may be compatible with the method described in the other example, and their use should be considered optional. Therefore, this disclosure discloses embodiments that are combinations of features of multiple examples, without exhaustively listing all features of the embodiments in the description of each exemplary embodiment. User interface and related processes
[0214] Here, we draw our attention to embodiments of a user interface ("UI") and related processes that may be implemented on a computer system such as a portable multifunction device or head-mounted device, which communicates with display generation components and (optionally) one or more sensors.
[0215] The examples described herein illustrate how a user of a computer system (e.g., device 700) may initiate and / or modify a live communication session in which the user communicates with one or more other users of each respective computer system. In some embodiments, the live communication session is an audio communication session (e.g., a voice call or telephone call). In some embodiments, the live communication session is a video communication session (e.g., a video telephony and / or video conference). In some embodiments, the live communication session is an XR communication session, such as a spatial communication session or a non-spatial communication session. In a spatial communication session, one or more users are each represented in the XR environment by a three-dimensional (3D) representation (e.g., an avatar) corresponding to the user(s). In some embodiments, the 3D representation has a spatial mediation such that the 3D representation can move within the XR environment relative to other elements and / or users in the XR environment. In a non-spatial communication session, one or more users are each represented in the XR environment by a two-dimensional (2D) representation corresponding to the user(s). In some embodiments, the 2D representation includes a user video feed and optionally has a fixed position (e.g., location) within the XR environment.
[0216] Figures 7A to 7Q illustrate examples of managing live communication sessions. Figure 8 is a flowchart of exemplary method 800 for managing live communication sessions. Figure 9 is a flowchart of exemplary method 900 for providing avatars in live communication sessions. The user interfaces in Figures 7A to 7Q are used to illustrate processes described later, including the processes in Figure 8 and / or Figure 9.
[0217] Figures 7A to 7Q show device 700 as a handheld device (e.g., a tablet, smartphone, or laptop) having a display 702, but in some embodiments, device 700 is a head-mounted device (HMD). The HMD is configured to be worn on the user's head and includes a display 702 on top of and / or inside the HMD. The display 702 is visible to the user when device 700 is worn on the user's head. For example, in some embodiments, the HMD at least partially covers the user's eyes when worn on the user's head so that the display 702 is positioned above and / or in front of the user's eyes. In such embodiments, the display 702 is configured to display an XR environment during a live communication session in which the user of the HMD is participating.
[0218] In Figure 7A, device 700 displays an XR environment 704 on display 702, which includes elements such as a table 704a and a couch 704b (e.g., virtual and / or physical elements). While displaying the XR environment 704, device 700 receives a request to display a communication interface. In some embodiments, the request to display the communication interface is the pressing of a button 703 on device 700. As shown in Figure 7B, in response to receiving the request, device 700 displays the communication interface 710. In some embodiments, the communication interface 710 is displayed within the XR environment 704.
[0219] In general, the communication interface 710 can be used to initiate and / or modify live communication sessions (e.g., audio communication sessions, video communication sessions, or XR communication sessions). The communication interface 710 includes pinned contacts 712 (e.g., pinned contacts 712a-712g) and recent contacts 714 (e.g., recent contacts 714a-714i). In some embodiments, pinned contacts 712 are a set of contacts selected by the user of device 700 to be included in the communication interface 710 (e.g., favorited or pinned contacts). In some embodiments, recent contacts 714 are contacts that the user of device 700 has recently communicated with (e.g., via text, phone, and / or live communication sessions) using device 700 and optionally, one or more other devices associated with the user of device 700. In some embodiments, recent contacts 714 are arranged (e.g., ordered or ranked) based on the recency of communication between the recent contacts 714 and the user of device 700.
[0220] In some embodiments, one or more pinned contacts 712 and / or recent contacts 714 correspond to a defined group of contacts. For example, pinned contact 712d corresponds to the group of contacts "surfers". In another example, recent contact 714d corresponds to the group of contacts "lakecrew".
[0221] In some embodiments, pinned contacts 712 and / or recent contacts 714 indicate the most recent communication between the user of device 700 and various contacts. For example, pinned contact 712b ("John") indicates that the contact last sent a text message one minute ago. Optionally, the communication interface 710 includes a preview 716b that shows the content of the text message sent by pinned user 712b. In another example, pinned user 712c indicates that the contact last sent a text response (e.g., a "heart" response) at 2:10. In yet another example, recent contact 714a ("Mother") indicates that the user of device 700 most recently communicated with contact 714a in an XR communication session (e.g., a spatial live communication session or a non-spatial live communication session) at 3:32. As yet another example, recent contact 714e ("Uncle Bob") indicates that the user of device 700 most recently communicated with contact 714e in an audio communication session (e.g., a phone call) at 9:41.
[0222] In some embodiments, pinned contacts 712 and / or recent contacts 714 indicate pending invitations to live communication sessions. For example, recent contact 714b ("Father") indicates that a user of device 700 can join a live communication session with recent contact 714b. In yet another example, recent contact 714d ("Lake Crew") indicates that three members of a group are currently in an ongoing live communication session to which a user of device 700 has been invited to join.
[0223] In some embodiments, contacts can be managed using contacts 712 and 714 of the communication interface 710. For example, while the communication interface 710 is being displayed, device 700 detects the selection of contact 712e ("Jo"). In some embodiments, the selection of contact 712e is a tap gesture 705b on contact 712e. In some embodiments, the selection of contact 712e is, for example, an air gesture indicating the selection of contact 712e. As shown in Figure 7C1 and / or Figure 7C2, in response to detecting the selection of contact 712e, device 700 displays the contact menu 720 associated with contact 712e.
[0224] The contact menu 720 includes an invitation option 720a and an extension option 720b. When selected, the invitation option 720a causes device 700 to invite contact 712e to an XR communication session. When selected, the extension option 720b causes device 700 to display one or more additional options for managing contact 712e. For example, while displaying the contact menu 720, device 700 detects the selection of extension option 720b. In some embodiments, the selection of extension option 720b is a tap gesture 705c on extension option 720b. In some embodiments, the selection of extension option 720b is, for example, an air gesture indicating the selection of extension option 720b. As shown in Figure 7D, in response to detecting the selection of extension option 720b, device 700 expands the contact menu 720 to display one or more additional options (e.g., options 720c-720f) (e.g., replacing the display of extension option 720b with one or more additional options).
[0225] In some embodiments, when extended, the contact menu 720 includes an audio option 720c, a message option 720d, an information option 720e, and an edit option 720f. When selected, the audio option 720c causes device 700 to initiate an audio communication session (e.g., without live video components) with contact 712e. In some embodiments, device 700 is not able to communicate over a cellular network and / or is configured to use an external device for audio calls. Therefore, in some examples, device 700 initiates an audio communication session using a nearby device (e.g., a mobile phone and / or tablet) (e.g., one that can communicate over a cellular network). When selected, the edit option 720f allows the user of device 700 to remove contact 712e from pinned contacts 712 (or, in embodiments where contact 712e is not yet a pinned contact, to add contact 712e to pinned contacts). When selected, the message option 720d allows the user to send a message to contact 712e. For example, while displaying the contact menu 720, device 700 detects the selection of a message option 720d. In some embodiments, the selection of message option 720d is a tap gesture 705d on message option 720d. In some embodiments, the selection of message option 720d is, for example, an air gesture indicating the selection of message option 720d. As shown in Figure 7E, in response to detecting the selection of message option 720d, device 700 displays the message interface 730 (for example, replacing the display of the communication interface 710 with it). A message can then be sent to contact 712e using the message interface 730.
[0226] Referring again to Figure 7D, when information option 720e is selected, the device 700 displays information corresponding to contact 712e (without displaying additional information corresponding to other contacts, for example). For example, while displaying the contact menu 720, the device 700 detects the selection of information option 720e. In some embodiments, the selection of information option 720e is a tap gesture 707d on information option 720e. In some embodiments, the selection of information option 720e is, for example, an air gesture indicating the selection of information option 720e. As shown in Figure 7F, in response to detecting the selection of information option 720e, the device 700 displays the information interface 740. The information interface 740 includes various details corresponding to contact 712e, including but not limited to name and contact information.
[0227] In some embodiments, the contact's device cannot participate in an XR communication session with device 700. Therefore, in some embodiments, one or more options in the contact menu may be omitted, de-emphasized (e.g., grayed out or darkened), and / or replaced to accurately reflect the capabilities of the contact's device. For example, referring again to Figure 7B, while displaying the communication interface 710, device 700 detects the selection of contact 712g ("Sam"). In some embodiments, the selection of contact 712g is a tap gesture 709b on contact 712g. In some embodiments, the selection of contact 712g is, for example, an air gesture indicating the selection of contact 712g. As shown in Figure 7C1, in response to detecting the selection of contact 712g, device 700 displays the contact menu 722 associated with contact 712g.
[0228] In some embodiments, since the contact 712g's device cannot communicate with device 700 in an XR communication session, menu 722 does not include an invite option (e.g., invite option 720a) and instead includes an audio option 722a. When selected, audio option 722a causes device 700 to initiate an audio communication session with contact 712g. Menu 722 further includes an extension option 722b, which, when selected, causes device 700 to display one or more additional options for contact 712g.
[0229] In some embodiments, the contact menu associated with a contact includes one or more additional options based on the state of device 700. For example, in some embodiments, if device 700 is participating in a live communication session (e.g., an XR communication session or an audio communication session), the contact menu includes an option to invite the contact to the live communication session. For example, referring to Figure 7B, while participating in an XR communication session and displaying the communication interface 710, device 700 detects the selection of contact 714g ("Dylan"). In some embodiments, the selection of contact 714g is a tap gesture 711b on contact 714g. In some embodiments, the selection of contact 714g is an air gesture indicating the selection of contact 714f, for example. As shown in Figure 7C1, in response to detecting the selection of contact 714g, device 700 displays the contact menu 724 associated with contact 714g.
[0230] The contact menu 724 includes invite option 724a, invite option 724b, and extension option 724c. Invite option 724a, when selected (for example, by a tap gesture 709c), causes device 700 to invite contact 714g to a new live communication session. Invite option 724b, when selected, causes device 700 to invite contact 714g to a live communication session that device 700 is currently participating in. Extension option 720c, when selected, causes device 700 to display one or more additional options for contact 714g.
[0231] In some embodiments, inviting a contact to a new live communication session (for example, depending on the selection of option 724a) causes device 700 to disconnect from and / or terminate a live communication session in which device 700 is currently participating. In some embodiments, before terminating an existing live communication session in this manner, device 700 confirms that the user wishes to disconnect from the current live communication session before starting a new live communication session. For example, as shown in Figure 7G, depending on the selection of invitation option 724a, device 700 displays a confirmation interface 740 including a confirmation affordance 742. Depending on the selection of confirmation affordance 742, device 700 terminates the current live communication session and invites contact 714f to a new live communication session.
[0232] In some embodiments, the user optionally sends a message to a contact using the communication interface 710. For example, referring to Figure 7B, while the communication interface 710 is displayed, the device 700 detects a selection of preview 716b associated with the pinned contact 712b. In some embodiments, the selection of preview 716b is a tap gesture 707b on preview 716b. In some embodiments, the selection of preview 716b is, for example, an air gesture indicating the selection of preview 716b. As shown in Figure 7C1, in response to detecting a selection of preview 716b, the device 700 expands preview 716b to display reply options 718.
[0233] When the reply option 718 is selected, the device 700 displays a reply interface for sending a message to contact 712b. For example, while displaying the reply option 718 within preview 716b, the device 700 detects the selection of the reply option 718. In some embodiments, the selection of the reply option 718 is a tap gesture 707c on the reply option 718. In some embodiments, the selection of the reply option 718 is, for example, an air gesture indicating the selection of the reply option 718. As shown in Figure 7H, in response to detecting the selection of the reply option 718, the device 700 displays a reply interface 750 that can be used to send a message to contact 712b.
[0234] In some embodiments, the technology and user interface(s) described in Figure 7C1 are provided by one or more of the devices described in Figures 1A to 1P. Figure 7C2 shows one embodiment in which a communication interface X710 (for example, as described in Figures 7B and 7C1) is displayed on the display module X702 of a head-mounted device (HMD) X700. In some embodiments, the device X700 includes a pair of display modules that provide stereoscopic content to different eyes of the same user. For example, the HMD X700 includes a display module X702 (which provides content to the user's left eye) and a second display module (which provides content to the user's right eye). In some embodiments, the second display module displays a slightly different image from the display module X702 in order to create the illusion of three-dimensional depth.
[0235] As shown in Figure 7C2, upon detecting the selection of contact X712e, the HMD X700 displays the contact menu X720 associated with contact X712e. In some embodiments, the HMD X700 detects the selection of contact X712e based on air gestures performed by the user of the HMD X700. In some embodiments, the HMD X700 detects the user's hands X750a and / or X750b and determines whether the movement of hands X750a and / or X750b performs a predetermined air gesture corresponding to the selection of contact X712e. In some embodiments, the predetermined air gesture for selecting contact X712e includes a pinch gesture. In some embodiments, the pinch gesture includes detecting the movement of fingers X750c and thumb X750d toward each other. In some embodiments, the HMD X700 detects the selection of contact X712e based on gaze and air gesture inputs performed by the user of the HMD X700. In some embodiments, gaze and air gesture inputs include detecting that the user of the HMD X700 is looking at the contact X712e (for example, for a certain amount of time) and that the user's hands X750a and / or X750b of the HMD X700 are performing a pinch gesture.
[0236] The contact menu X720 includes an invite option X720a and an extension option X720b. When the invite option X720a is selected (e.g., via an air gesture such as a pinch gesture, and / or via gaze and pinch gestures), it causes the HMD X700 to invite the contact X712e to an XR communication session. When the extension option X720b is selected, it causes the HMD X700 to display one or more additional options for managing the contact X712e. For example, while the contact menu X720 is displayed, the HMD X700 detects the selection of the extension option X720b. In some embodiments, the selection of the extension option X720b is, for example, an air gesture indicating the selection of the extension option X720b (e.g., a pinch gesture and / or gaze and pinch gestures). Upon detecting the selection of extended option X720b, the HMD X700 extends the contact menu X720 to display one or more additional options (for example, options 720c-720f, as shown in Figure 7D) (for example, replacing the display of extended option X720b with one or more additional options).
[0237] In some embodiments, when extended, the contact menu X720 includes, for example, an audio option (e.g., 720c), a message option (e.g., 720d), an information option (e.g., 720e), and an edit option (e.g., 720f), as described with respect to Figure 7D. When the audio option is selected, it causes the HMD X700 to initiate an audio communication session (e.g., without live video components) with contact X712e. In some embodiments, the HMD X700 is not able to communicate over a cellular network and / or is configured to use an external device for audio calls. Therefore, in some examples, the HMD X700 initiates an audio communication session using a nearby device (e.g., a mobile phone and / or tablet) (e.g., one that can communicate over a cellular network). The Edit option, when selected, allows the user of the HMD X700 to remove contact X712e from pinned contact X712 (or, in embodiments where contact X712e is not yet a pinned contact, to add contact X712e to pinned contacts), for example, as described with respect to Figure 7D. The Message option, when selected, allows the user to send a message to contact X712e, for example, as described with respect to Figure 7D. For example, while displaying the extended contact menu X720, the HMD X700 detects the selection of the Message option (e.g., 720d). In some embodiments, the selection of the Message option is, for example, an air gesture indicating the selection of the Message option (e.g., a pinch gesture and / or gaze and pinch gesture). In some embodiments, as shown in Figure 7E, in response to detecting the selection of the Message option (e.g., 720d), the HMD X700 displays the Message interface (e.g., 730) (e.g., replacing the display of the Communication Interface X710 with the Message interface). Afterward, you can use the messaging interface to send a message to the contact X712e.
[0238] In some embodiments, the contact menu associated with a contact includes one or more additional options based on the state of the HMD X700. For example, in some embodiments, if the HMD X700 is participating in a live communication session (e.g., an XR communication session or an audio communication session), the contact menu includes an option to invite the contact to the live communication session. For example, while participating in an XR communication session and displaying the communication interface X710, the HMD X700 detects the selection of contact X714g ("Dylan"). In some embodiments, the selection of contact X714g is an air gesture (e.g., a pinch gesture and / or gaze and pinch gesture) indicating the selection of contact X714f, for example. As shown in Figure 7C2, in response to detecting the selection of contact X714g, the HMD X700 displays the contact menu X724 associated with contact X714g.
[0239] The contact menu X724 includes the invitation options X724a, X724b, and the extended options X724c. When selected, invitation option X724a causes the HMD X700 to invite contact X714g to a new live communication session (for example, via a pinch gesture and / or an air gesture such as gaze and pinch gestures). When selected, invitation option X724b causes the HMD X700 to invite contact X714g to a live communication session that the HMD X700 is currently participating in (for example, via a pinch gesture and / or an air gesture such as gaze and pinch gestures). When selected, the extended options X724c causes the HMD X700 to display one or more additional options for contact X714g.
[0240] In some embodiments, inviting a contact to a new live communication session (for example, depending on the selection of option X724a) causes the HMD X700 to disconnect from and / or terminate the live communication session in which the HMD X700 is currently participating. In some embodiments, before terminating an existing live communication session in this manner, the HMD X700 confirms that the user wishes to disconnect from the current live communication session before starting a new live communication session. For example, as shown in Figure 7G, depending on the selection of the invitation option X724a, the HMD X700 may display a confirmation interface 740 including a confirmation affordance 742. Depending on the selection of the confirmation affordance 742 (for example, via an air gesture such as a pinch gesture and / or gaze and pinch gesture), the HMD X700 terminates the current live communication session and invites contact X714f to a new live communication session.
[0241] In some embodiments, the user optionally sends a message to a contact using the communication interface X710. For example, while displaying the communication interface X710, the HMD X700 detects a selection of preview 716b associated with a pinned contact X712b (e.g., as shown in Figure 7B). In some embodiments, the selection of preview 716b is, for example, an air gesture indicating the selection of preview 716b (e.g., a pinch gesture and / or gaze and pinch gesture). As shown in Figure 7C2, in response to detecting the selection of preview 716b, the HMD X700 expands preview 716b to display reply options X718.
[0242] When the reply option X718 is selected, the HMD X700 displays a reply interface for sending a message to contact X712b. For example, while displaying the reply option X718 within preview 716b, the HMD X700 detects the selection of the reply option X718. In some embodiments, the selection of the reply option X718 is, for example, an air gesture (e.g., a pinch gesture and / or gaze and pinch gesture) indicating the selection of the reply option X718. As shown in Figure 7H, in response to detecting the selection of the reply option X718, the HMD X700 may display a reply interface 750 that can be used to send a message to contact X712b.
[0243] Any of the features, components, and / or parts, including their arrangement and configuration shown in Figures 1B to 1P, may be included in the HMD X700 individually or in any combination. For example, in some embodiments, the HMD X700 includes any of the features, components, and / or parts of HMD1-100, 1-200, 3-100, 6-100, 6-200, 6-300, 6-400, 11.1.1-100, and / or 11.1.2-100, individually or in any combination. In some embodiments, the display module X702 includes display units 1-102, 1-202, 1-306, 1-406, display generation component 120, display screens 1-122a-b, first and second rear-facing display screens 1-322a, 1-322b, display 11.3.2-104, first and second display assemblies 1-120a, 1-120b, display assembly 1-320, display assembly 1-421, The first and second display subassemblies 1-420a, 1-420b, display assembly 3-108, display assembly 11.3.2-204, first and second optical modules 11.1.1-104a and 11.1.1-104b, optical module 11.3.2-100, optical module 11.3.2-200, lenticular lens array 3-110, display area or area 6-232, and / or features, components, and / or parts of the display / display area 6-334, either individually or in any combination. In some embodiments, the HMD X700 includes a sensor, either alone or in any combination, which includes sensor 190, sensor 306, image sensor 314, image sensor 404, sensor assembly 1-356, sensor assembly 1-456, sensor system 6-102, sensor system 6-202, sensor 6-203, sensor system 6-302, sensor 6-303, sensor system 6-402, and / or any feature, component, and / or part of any of sensors 11.1.2-110a to f.In some embodiments, the input device X703 includes any of the features, components, and / or parts of the first buttons 1-128, buttons 11.1.1-114, second buttons 1-132, and / or dials or buttons 1-328, either individually or in any combination. In some embodiments, the HMD X700 optionally includes one or more audio output components (e.g., electronic components 1-112) for generating audio feedback (e.g., audio output) generated based on detected events and / or user input detected by the HMD X700.
[0244] In some embodiments, the communication interface 710 is used to generate an avatar. In some embodiments, the avatar serves as a representation (e.g., a 3D representation) of the user of device 700 in an XR communication session. For example, referring to Figure 7B, while displaying the communication interface 710, device 700 detects the selection of an avatar option 715. In some embodiments, the selection of an avatar option 715 is a tap gesture 713b on the avatar option 715. In some embodiments, the selection of an avatar option 715 is, for example, an air gesture indicating the selection of an avatar option 715. As shown in Figure 7I, in response to detecting the selection of an avatar option 715, device 700 displays the avatar interface 760.
[0245] Figure 7I shows a first option 762 (e.g., more realistic than the second option) and a second option 764 (e.g., less realistic than the first option) of the avatar interface 760. When enabled, the first option 762 causes the user avatar on device 700 to reflect the user's appearance. For example, in some embodiments, when the first option 762 is enabled, the avatar includes one or more visual characteristics corresponding to one or more physical characteristics of the user. When enabled, the second option 764 causes the user avatar on device 700 to indicate the user's movement (e.g., during a live communication session) without reflecting the user's appearance. For example, in some embodiments, when the second option 764 is enabled, an avatar with a default appearance is used. In some embodiments, when the first option 762 is enabled (compared to the second option 764), the user avatar of device 700 is represented by a first representation style, the avatar is displayed at a first level of detail (e.g., a first level of detail relating to the appearance of the user and / or one or more parts of the user), and the position and movement of the user's first user part relative to the position and movement of the user's second user part is shown in a first manner. In some embodiments, when the second option 764 is enabled (compared to the first option 762), the user avatar of device 700 is represented by a second representation style different from the first representation style, the avatar is displayed at a second level of detail lower than the first level of detail (e.g., a second level of detail relating to the user's appearance and / or one or more parts of the user) (e.g., lower detail and / or a lower amount of detail than and / or mimicking the user's appearance), and the position and movement of the user's first user part relative to the position and movement of the user's second user part is shown in a second manner different from the first.
[0246] The avatar interface 760 further includes a menu option 766 that, when selected, causes the device 700 to display an avatar menu, as shown in Figure 7I. For example, in Figure 7I, while the avatar interface 760 is displayed, the device 700 detects the selection of menu option 766. In some embodiments, the selection of menu option 766 is a tap gesture 705i on menu option 766. In some embodiments, the selection of menu option 766 is, for example, an air gesture indicating the selection of menu option 766. As shown in Figure 7J, in response to detecting the selection of menu option 766, the device 700 displays the avatar menu 768.
[0247] In Figure 7J, the avatar menu 768 includes an edit option 768a, a create option 768b, and / or a delete option 768c. In some embodiments, if the avatar has not yet been created for the user of device 700, the avatar menu 768 includes the create option 768b but does not include the edit option 768a and the delete option 768c. In some embodiments, if the avatar has already been created for the user of device 700, the avatar menu 768 includes the edit option 768a and the delete option 768c but does not include the create option 768b.
[0248] In Figure 7J, while the avatar menu 768 is displayed, device 700 detects the selection of creation option 768b. In some embodiments, the selection of creation option 768b is a tap gesture 705j on creation option 768b. In some embodiments, the selection of creation option 768b is, for example, an air gesture indicating the selection of creation option 768b. As shown in Figure 7K, in response to detecting the selection of creation option 768b, device 700 displays the setup interface 770.
[0249] In Figure 7K, the setup interface 770 includes a setup option 772 that, when selected, causes the device 700 to display the avatar editing interface. For example, while the setup interface 770 is displayed, the device 700 detects the selection of the setup affordance 772. In some embodiments, the selection of the setup affordance 772 is a tap gesture 705k on the setup affordance 772. In some embodiments, the selection of the setup affordance 772 is, for example, an air gesture indicating the selection of the setup affordance 772. As shown in Figure 7L1, in response to detecting the selection of the setup affordance 772, the device 700 displays the avatar editing interface 780.
[0250] In Figure 7L1, the avatar editing interface 780 includes a live view 781 of the user's avatar on device 700. In some embodiments, device 700 displays a real-time updated live view 781 of the avatar editing interface 780 according to the user's movements and / or habits on device 700, as detected by device 700. The avatar editing interface 780 further includes various settings and / or parameters from which the avatar's visual characteristics are adjusted. As an example, the avatar editing interface 780 includes a setting 782, which includes a brightness setting 782a and a warmth setting 782b. The brightness setting 782a and the warmth setting 782b are used to adjust the simulated lighting and temperature of the avatar's skin, respectively. As another example, the avatar editing interface 780 includes a color palette 783, which includes a set of one or more colors and / or shades from which the avatar's skin color is selected.
[0251] In some embodiments, as shown in Figure 7L1, the avatar editing interface 780 includes a set of parameters 784, such as shirt parameters 784a and hat parameters 784b. In some embodiments, selecting parameters allows the selection of visual characteristics for one or more aspects of the avatar. Referring to Figure 7M, for example, the selection of shirt parameter 784a (e.g., a tap input 705l or air gesture corresponding to the location of parameter 784a) causes device 700 to display a parameter menu 790 from which the user can select any number of options (e.g., options 790a to 790c) for shirt parameter 784a. Referring to Figure 7N, once an option is selected (e.g., a tap input 705m or air gesture corresponding to the location of option 790b), the user can select from styles 792 (e.g., 792a to 792f) of the selected option, and the visual characteristics of the avatar are updated accordingly. Similarly, in some embodiments, the selection of headwear parameter 784b causes device 700 to display headwear options for headwear parameter 784b, and the selection of an option causes device 700 to display the type of the selected option.
[0252] This specification describes parameters 784a and 784b corresponding to shirts and hats, respectively, but it will be understood that in some embodiments, parameters of the avatar editing interface 780 may optionally correspond to other / additional visual aspects of the avatar. For example, in some embodiments, parameter 784 is used to select one or more aspects of the avatar's eyewear (e.g., parameter 784a corresponds to glasses and parameter 784b corresponds to an eye patch). In an example where device 700 receives a user selection for parameter 784a corresponding to glasses, device 700 displays options for various designs of glasses (e.g., frameless, thin frame, thick frame, etc.). When device 600 receives a user selection for a design, device 700 displays various styles of the selected design as type 792 for user selection. In an example where device 700 receives a user selection for an eye patch, device 700 displays options for various designs of eye patches (e.g., left eye patch or right eye patch). When device 600 receives a user selection of a design, device 700 displays various styles of the selected design as type 792 for the user to choose from.
[0253] In some embodiments, the avatar editing interface 780 includes a set of parameters 786. As shown, in some embodiments, the parameters 786 are used to select one or more aspects of hair. For example, parameter 786a corresponds to hairstyle, parameter 786b corresponds to hair color, and parameter 786c corresponds to hair highlights. In other embodiments, the parameters 786 are used to select one or more aspects of accessibility features. For example, in some embodiments, parameter 786a corresponds to prosthetic arm, parameter 786b corresponds to hearing aid, and parameter 786c corresponds to wheelchair.
[0254] In some embodiments, the technology and user interface(s) described in Figure 7L1 are provided by one or more of the devices described in Figures 1A to 1P. Figure 7L2 shows one embodiment in which an avatar editing interface X780 (for example, as described in Figures 7L1 to 7N) is displayed on the display module X702 of a head-mounted device (HMD) X700. In some embodiments, the device X700 includes a pair of display modules that provide stereoscopic content to different eyes of the same user. For example, the HMD X700 includes a display module X702 (which provides content to the user's left eye) and a second display module (which provides content to the user's right eye). In some embodiments, the second display module displays a slightly different image from the display module X702 in order to create the illusion of three-dimensional depth.
[0255] In Figure 7L2, the avatar editing interface X780 includes a live view X781 of the user's avatar on the HMD X700. In some embodiments, the HMD X700 displays a real-time updated live view X781 on the avatar editing interface X780 according to the user's movements and / or habits detected by the HMD X700. The avatar editing interface X780 further includes various settings and / or parameters from which the avatar's visual characteristics are adjusted. As an example, the avatar editing interface X780 includes a setting X782, which includes a brightness setting X782a and a warmth setting X782b. The brightness setting X782a and the warmth setting X782b are used to adjust the simulated lighting and temperature of the avatar's skin, respectively. As another example, the avatar editing interface X780 includes a color palette X783, which includes a set of one or more colors and / or shades from which the avatar's skin color is selected.
[0256] In some embodiments, as shown in Figure 7L2, the avatar editing interface X780 includes a set of parameters X784, such as shirt parameters X784a and hat parameters X784b. In some embodiments, selecting parameters allows for the selection of visual characteristics of one or more aspects of the avatar. For example, the selection of shirt parameter X784a (e.g., gaze and pinch gesture, where gaze is represented by gaze indicator X705L, corresponding to the location of parameter X784a) causes the HMD X700 to display a parameter menu 790, as shown in Figure 7M, from which the user can select any number of options (e.g., options 790a-790c) for shirt parameter X784a.
[0257] In some embodiments, the HMD X700 detects the selection of shirt parameter X784a based on air gestures performed by the user of the HMD X700. In some embodiments, the HMD X700 detects the user's hands X750a and / or X750b and determines whether the movement of hands X750a and / or X750b performs a predetermined air gesture corresponding to the selection of shirt parameter X784a. In some embodiments, the predetermined air gesture for selecting shirt parameter X784a includes a pinch gesture. In some embodiments, the pinch gesture includes detecting the movement of fingers X750c and thumb X750d toward each other. In some embodiments, the HMD X700 detects the selection of shirt parameter X784a based on gaze and air gesture inputs performed by the user of the HMD X700. In some embodiments, gaze and air gesture inputs include detecting that the user of the HMD X700 is looking at the shirt parameter X784a (for example, for a predetermined amount of time) and that the user's hands X750a and / or X750b of the HMD X700 are performing a pinch gesture.
[0258] Referring to FIG. 7N, when an option is selected (e.g., a tap input 705m or an air gesture corresponding to the location of option 790b), the user can select from the styles 792 (e.g., 792a - 792f) of the selected option, and the visual characteristics of the avatar are updated accordingly. Similarly, in some embodiments, the selection of the headgear parameter 784b causes the device 700 to display the headgear options of the headgear parameter 784b, and the selection of the option causes the device 700 to display the type of the selected option.
[0259] In this specification, parameters X784a and X784b corresponding to a shirt and a hat respectively are described. However, in some embodiments, it will be understood that the parameters of the avatar editing interface X780 can optionally correspond to other / additional visual aspects of the avatar. As an example, in some embodiments, the parameter X784 is used to select one or more aspects of the avatar's eyewear (e.g., parameter X784a corresponds to glasses and parameter X784b corresponds to an eye patch). In an example where the HMD X700 receives a user selection of the parameter X784a corresponding to glasses, the HMD X700 displays options for various designs of glasses (e.g., frameless, thin frame, thick frame, etc.). When the HMD X700 receives a user selection of a design, the HMD X700 displays various styles of the selected design as type 792 for the user to select. In an example where the HMD X700 receives a user selection of an eye patch, the HMD X700 displays options for various designs of the eye patch (e.g., left eye patch or right eye patch). When the HMD X700 receives a user selection of a design, the HMD X700 displays various styles of the selected design as type 792 for the user to select.
[0260] In some embodiments, the avatar editing interface X780 includes a set of parameters X786. As shown, in some embodiments, the parameters X786 are used to select one or more aspects of the hair. As an example, parameter X786a corresponds to a hairstyle, parameter X786b corresponds to a hair color, and parameter X786c corresponds to a hair highlight. In other embodiments, the parameters X786 are used to select one or more aspects of the accessibility features. As an example, in some embodiments, parameter X786a corresponds to a prosthetic hand, parameter X786b corresponds to a hearing aid, and parameter X786c corresponds to a wheelchair.
[0261] Any of the features, components, and / or parts, including their arrangement and configuration shown in Figures 1B to 1P, may be included in the HMD X700 individually or in any combination. For example, in some embodiments, the HMD X700 includes any of the features, components, and / or parts of HMD1-100, 1-200, 3-100, 6-100, 6-200, 6-300, 6-400, 11.1.1-100, and / or 11.1.2-100, individually or in any combination. In some embodiments, the display module X702 includes display units 1-102, 1-202, 1-306, 1-406, display generation component 120, display screens 1-122a-b, first and second rear-facing display screens 1-322a, 1-322b, display 11.3.2-104, first and second display assemblies 1-120a, 1-120b, display assembly 1-320, display assembly 1-421, The first and second display subassemblies 1-420a, 1-420b, display assembly 3-108, display assembly 11.3.2-204, first and second optical modules 11.1.1-104a and 11.1.1-104b, optical module 11.3.2-100, optical module 11.3.2-200, lenticular lens array 3-110, display area or area 6-232, and / or features, components, and / or parts of the display / display area 6-334, either individually or in any combination. In some embodiments, the HMD X700 includes a sensor, either alone or in any combination, which includes sensor 190, sensor 306, image sensor 314, image sensor 404, sensor assembly 1-356, sensor assembly 1-456, sensor system 6-102, sensor system 6-202, sensor 6-203, sensor system 6-302, sensor 6-303, sensor system 6-402, and / or any feature, component, and / or part of any of sensors 11.1.2-110a to f.In some embodiments, the input device X703 includes any of the features, components, and / or parts of the first buttons 1-128, buttons 11.1.1-114, second buttons 1-132, and / or dials or buttons 1-328, either individually or in any combination. In some embodiments, the HMD X700 optionally includes one or more audio output components (e.g., electronic components 1-112) for generating audio feedback (e.g., audio output) generated based on detected events and / or user input detected by the HMD X700.
[0262] Referring to Figure 7N, while the avatar editing interface 780 is displayed, device 700 detects the selection of save option 788. In some embodiments, the selection of save option 788 is a tap gesture 705n on save option 788. In some embodiments, the selection of save option 788 is, for example, an air gesture indicating the selection of save option 788. In response to detecting the selection of save option 788, device 700 stores the selected configuration of the avatar for the user of device 700 for subsequent use in the XR communication session (e.g., stored locally and / or stored remotely). Further in response to detecting the selection of setup affordance 772, as shown in Figure 7O, device 700 displays a completion interface 795 indicating that the user has successfully created and / or updated the user's avatar.
[0263] In Figure 7P, the user of device 700 is participating in an XR communication session with contact 712f ("Ann", Figure 7B) in the XR environment 704. In some embodiments, the XR communication session is a spatial communication session. Therefore, in some embodiments, contact 712f and / or the user of device 700 are represented in the XR environment 704 by a 3D representation (e.g., an avatar). For example, as shown in Figure 7P, the user of device 700 is represented by representation 700A (as shown in self-preview 706A), and contact 712f is represented by representation 701A.
[0264] In some embodiments, the view of the XR environment 704 for the user of device 700 is provided from the viewpoint of the representation 700A within the XR environment 704. Since this may make it impossible for the user to view the representation 700A in any other way, device 700 displays a self-preview 706A that includes a live view of the representation 700A within the XR environment 704. Although the self-preview 706A is shown as being located in the lower right corner of display 702, it will be understood that the self-preview 706A can optionally be displayed in any location on display 702. For example, in some embodiments, the self-preview 706A is located adjacent to the representation of a contact in an XR communication session. In some embodiments, the self-preview 706A is located, for example, in position 708A adjacent to representation 701A.
[0265] In some embodiments, participants represented by 3D representations within the XR environment have spatial agency. Therefore, during an XR communication session, the 3D representation optionally moves within the XR environment 704 so that the 3D representation moves relative to elements within the XR environment 704 (e.g., table 704a and couch 704b) and / or other participants. In some embodiments, the 3D representation moves in accordance with the movement of a corresponding device. For example, 3D representation 700A may move within the XR environment 704 in accordance with the movement of device 700. In some embodiments, the 3D representation moves in a manner corresponding to the movement of a device. For example, if device 700 first moves in a first direction (e.g., left) and then moves in a second direction (e.g., right), 3D representation 700A will move in the first and second directions within the XR environment 704 in a similar manner.
[0266] In some embodiments, while participating in an XR communication session, device 700 displays a set of controls 704A for managing one or more aspects of the XR communication session, as shown in Figure 7P. The set of controls 704A includes a message option 704Aa, an information option 704Ab, a microphone option 704Ac, an avatar option 704Ad, a camera option 704Ae, and an exit option 704Af. When selected, the message option 704Aa causes device 700 to display a message interface for sending a message to contact 712f. When selected, the information option 704Ab causes device 700 to display an information interface corresponding to contact 712f. When selected, the microphone option 704Ac toggles the state of the device 700's microphone (e.g., enable or disable). In some embodiments, disabling the device 700's microphone prevents device 700 from providing audio during the XR communication session. The camera option 704Ae, when selected, toggles the state of the camera on device 700 (e.g., enable or disable). In some embodiments, disabling the camera on device 700 prevents device 700 from providing video (e.g., a video feed of the user on device 700 and / or a moving representation of the user on device 700) during an XR communication session. The termination option 704Af, when selected, causes device 700 to disconnect from the XR communication session.
[0267] When selected, Avatar Option 704Ad toggles the use of 3D representation in the XR environment 704 (e.g., enables or disables it). For example, while displaying the XR environment 704, Device 700 detects the selection of Avatar Option 704Ad. In some embodiments, the selection of Avatar Option 704Ad is a tap gesture 705p on Avatar Option 704Ad. In some embodiments, the selection of Avatar Option 704Ad is, for example, an air gesture indicating the selection of Avatar Option 704Ad. As shown in Figure 7Q, depending on the selection of Avatar Option 704Ad, Device 700 disables the use of 3D representation in the XR environment 704.
[0268] In some embodiments, when toggling the use of 3D representation within the XR environment 704, device 700 transitions the XR communication session between a spatial communication session and a non-spatial communication session. Disabling the use of 3D representation causes device 700 to transition the XR communication session from a spatial communication session to a non-spatial communication session. Enabling the use of 3D representation causes device 700 to transition the XR communication session from a non-spatial communication session to a spatial communication session.
[0269] In some embodiments, in a non-spatial communication session, participants in the XR communication session are represented by a 2D representation. For example, as shown in Figure 7Q, the user of device 700 is represented by a 2D representation 710A (as shown in the self-preview 706A), and contact 712f is represented by a 2D representation 712A.
[0270] In some embodiments, the 2D representation includes a user video feed (e.g., a live video feed). In some embodiments, if the user video feed is unavailable (e.g., the device's camera is disabled), the user's 2D representation instead includes an image associated with the user (e.g., a thumbnail), a monogram corresponding to the user, and / or another 2D representation. In some embodiments, the user represented by the 2D representation in the XR environment has no spatial agency and is optionally positioned at one or more predetermined locations within the XR environment 704. In some embodiments, device 600 is configured to move the 2D representation of the remote participant in the XR environment based on user input received in device 700 (e.g., input to drag the representation from a first location to a second location). In some embodiments, device 600 is not configured to move the 3D representation of the remote participant in the XR environment based on user input received in device 700.
[0271] Further explanations regarding Figures 7A to 7Q are provided below with reference to Methods 800 and 900, each of which is described in relation to Figures 7A to 7Q.
[0272] Figure 8 is a flowchart of an exemplary method 800 for managing a live communication session according to several embodiments. In some embodiments, the method 800 is performed on a computer system (e.g., computer system 101 in Figure 1A, computer system 700, and / or HMD X700) (e.g., smartphone, tablet, and / or head-mounted device) that communicates with display generating components (e.g., display generating component 120 in Figures 1A, 3, and 4, display 702, and / or display X702) (e.g., visual output device, 3D display, display having at least a transparent or translucent portion capable of projecting an image (e.g., see-through display), projector, head-up display, and / or display controller) and one or more sensors (e.g., touch-sensing surface, gyroscope, accelerometer, motion sensor, movement sensor, microphone, infrared sensor, camera sensor, depth camera, visible light camera, eye-tracking sensor, gaze-tracking sensor, physiological sensor, and / or image sensor). In some embodiments, Method 800 is stored in a non-temporary (or temporary) computer-readable storage medium and managed by instructions executed by one or more processors of the computer system, such as one or more processors 202 of the computer system 101 (e.g., control 110 in Figure 1A). Some operations of Method 800 are optionally combined, and / or the order of some operations is optionally changed.
[0273] A computer system (e.g., 700 and / or X700) displays representations (e.g., 712a-712g and 714a-714i) (e.g., static avatars, animated avatars, images, and / or monograms) of multiple users (e.g., users not operating the computer system (remote users) and / or users other than users of the computer system) via display generation components (802).
[0274] A computer system (e.g., 700 and / or X700) receives, via one or more sensors, a selection (e.g., 705b and / or 711b) of a representation of an individual user among multiple users (e.g., 712e, X712e, 714g, and / or X714g) (e.g., static avatars, animated avatars, images, and / or monograms) (e.g., via touch input on a touch-sensitive surface and / or via air gestures) (804).
[0275] In response to receiving a selection of individual user representations (e.g., 712e, X712e, 714g, and / or X714g) (806), and in accordance with the determination that there are ongoing (e.g., active and / or currently established) communication sessions (e.g., video communication sessions, audio communication sessions, extended reality communication sessions, spatial communication sessions, and / or non-spatial communication sessions), the computer system (e.g., 700 and / or X700) displays options (e.g., 724b in Figure 7C1 and / or X724b in Figure 7C2) via display generation components (e.g., 702 and / or X702) for inviting individual users to join the ongoing communication sessions (808).
[0276] In response to receiving a selection from an individual user (e.g., 712e, X712e, 714g, and / or X714g) (806), and in accordance with the determination that no ongoing communication session exists, the computer system (e.g., 700 and / or X700) discontinues displaying the option to invite the individual user to join an ongoing communication session (e.g., menu 720 in Figure 7C1 and / or menu X720 in Figure 7C2) (810). Conditionally displaying the option to invite an individual user to join an ongoing communication session allows a user of the computer system to invite an individual user, which requires navigating to a different user interface, thereby reducing the number of inputs required to perform the invitation action.
[0277] In some embodiments, in response to receiving a selection of an individual user's representation (e.g., 712e, X712e, 714g, and / or X714g) (for example, regardless of whether there is an ongoing communication session), the computer system (e.g., 700 and / or X700) displays, via display generation components (e.g., 702 and / or X702), options for initiating a new spatial communication session with the individual user (e.g., 720a and / or 724a in Figure 7C1, and / or X720a and / or X724a in Figure 7C2), and options for additional features (e.g., 720b and / or 724c in Figure 7C1, and / or X720b and / or X724c in Figure 7C2) (for example, without displaying an option for sending a text message to the individual user, and / or without displaying additional information about the individual user). In some embodiments, while displaying options for additional features (e.g., 720a and / or 724a in Figure 7C1, and / or X720a and / or X724a in Figure 7C2), the computer system (e.g., 700 and / or X700) receives a selection (e.g., 705c and / or 709d) of the options for additional features (e.g., 720a and / or 724a in Figure 7C1, and / or X720a and / or X724a in Figure 7C2). In some embodiments, in response to receiving an option selection for additional features, the computer system (e.g., 700 and / or X700) displays one or more options associated with an individual user (e.g., 720c-720f in Figure 7D) (e.g., one or more options for communicating with the individual user, such as sending a text message to the individual user and / or displaying additional information about the individual user) via a display generating component (e.g., 702 and / or X702) (e.g., by replacing the display of an option to invite the individual user to join an ongoing communication session).In some embodiments, a spatial communication session is a communication session having representations of at least some (e.g., fewer than all, more, and / or all) of the users participating in the communication session, distributed within a 3D environment. By displaying options for starting a new spatial communication session with individual users and options for accessing additional features, the user of the computer system can quickly access the option to start a new spatial communication session without cluttering the user interface, while still providing access to additional (and potentially less frequently used) features, thereby improving the human-machine interface.
[0278] In some embodiments, in response to receiving a selection of an individual user's representation (for example, independently of determining whether there is an ongoing communication session), the computer system (e.g., 700 and / or X700) displays additional feature options (e.g., 720b in Figure 7C1 and / or X720b in Figure 7C2) via display generation components (e.g., 702 and / or X702) (e.g., without displaying an option to send a text message to the individual user and / or without displaying additional information about the individual user). In some embodiments, while displaying the additional feature options (e.g., 720b in Figure 7C1 and / or X720b in Figure 7C2), the computer system (e.g., 700 and / or X700) receives a selection (e.g., 705c) of the additional feature options (e.g., 720b in Figure 7C1 and / or X720b in Figure 7C2) via one or more sensors (e.g., via touch input on a touch-sensitive surface and / or via an air gesture). In some embodiments, in response to receiving a selection of an additional feature option (e.g., 720b in Figure 7C1 and / or X720b in Figure 7C2), the computer system (e.g., 700 and / or X700) displays an option (e.g., 720c in Figure 7D) via a display generating component (e.g., 702 and / or X702) to initiate an audio communication session with an individual user (e.g., without a live visual representation of the participant and / or without a video component) (e.g., as part of one or more options associated with the individual user) (e.g., by replacing a display of an option to invite the individual user to join an ongoing communication session). In some embodiments, the computer system receives a selection of an option (e.g., 720c in Figure 7D) via one or more sensors (e.g., via touch input on a touch-sensing surface and / or via an air gesture).In some embodiments, upon receiving a selection of an option to initiate an audio communication session with an individual user (e.g., 720c in Figure 7D), the computer system (e.g., 700 and / or X700) initiates an audio communication session with that individual user (e.g., one that does not include a live visual representation of the participant and / or video components) (e.g., one that does not initiate with other users). Displaying an option to initiate an audio communication session allows a user of the computer system to initiate a communication session that does not include a live visual representation of the user without having to initiate a video communication session and separately disable the video portion, thereby reducing the number of inputs required to initiate an audio communication session.
[0279] In some embodiments, the computer system initiates an audio communication session by initiating an audio call (e.g., a voice call and / or telephone call) using an external electronic device (e.g., a smartphone and / or mobile phone) within a predetermined range (e.g., distance and / or wireless range) of the computer system (e.g., 700 and / or X700). In some embodiments, the option to initiate an audio communication session is presented to each user who does not have an account (or does not have an active account) with a particular online service (e.g., a user of the computer system has an account with a particular online service used for video and / or extended reality communication, but individual users do not). Initiating an audio communication session via an audio call using an external electronic device allows the computer system to use the resources of the external computer system (e.g., cellular connectivity and / or CPU processing of the external computer system), thereby improving the functionality of the computer system while reducing the workload on the computer system.
[0280] In some embodiments, in response to receiving a selection (e.g., 705c and / or 711b) of an individual user representation (e.g., 712e, X712e, 714g, and / or X714g) (e.g., independently of determining whether there is an ongoing communication session), the computer system (e.g., 700 and / or X700) displays options for additional features (e.g., 720b, X720b, 724c, and / or X724C) via display generation components (e.g., 702 and / or X702) (e.g., without displaying an option to initiate a process for sending a message to the individual user, and / or without displaying additional information about the individual user). In some embodiments, while options for additional features (e.g., 720b, X720b, 724c, and / or X724c) are displayed, a computer system (e.g., 700 and / or X700) receives a selection (e.g., 705c) of an option for an additional feature (e.g., 720b and / or X720b) via one or more sensors (e.g., via touch input on a touch-sensitive surface and / or via air gestures). In some embodiments, upon receiving a selection (e.g., 705c) of an additional feature option (e.g., 720b and / or X720b), the computer system (e.g., 700 and / or X700) displays an option (e.g., 720d) via a display generating component (e.g., 702 and / or X702) to initiate a process of sending a message (e.g., not including live transmission of audio and / or video) to an individual user (e.g., as part of one or more options associated with an individual user) (e.g., by replacing a display of an option to invite an individual user to join an ongoing communication session). In some embodiments, the computer system receives a selection (e.g., 705d) of an option (e.g., 720d) to initiate a process of sending a message to an individual user via one or more sensors (e.g., via touch input on a touch-sensing surface and / or via an air gesture).In some embodiments, upon receiving a selection (e.g., 705d) of an option (e.g., 720d) for initiating a process to send a message to an individual user, the computer system (e.g., 700 and / or X700) initiates a process to send a message (e.g., 730 in Figure 7E and / or 750 in Figure 7H) to the individual user (e.g., without sending the message to other users). In some embodiments, initiating a process to send a message to an individual user includes displaying a user interface that includes a conversation between the computer system user and the individual user, displaying a keyboard, and / or displaying a text entry field for entering a message. By providing an option to initiate a process to send a message to an individual user via an additional feature selection, the computer system user can quickly initiate the process without having to specify a recipient, thereby reducing the number of inputs required to send a message.
[0281] In some embodiments, in response to receiving a selection (e.g., 705c and / or 711b) of an individual user representation (e.g., 712e, X712e, 714g, and / or X714g) (e.g., independently of determining whether there is an ongoing communication session), the computer system (e.g., 700 and / or X700) displays options for additional features (e.g., 720b, X720b, 724c, and / or X724c) via display generation components (e.g., 702 and / or X702) (e.g., without displaying an option to initiate a process for sending a message to the individual user, and / or without displaying additional information about the individual user). In some embodiments, while displaying options for additional features (e.g., 720b, X720b, 724c, and / or X724C), the computer system receives a selection of the options for additional features (e.g., 720b, X720b, 724c, and / or X724c) via one or more sensors (e.g., via touch input on a touch-sensing surface and / or via air gestures). In some embodiments, in response to receiving a selection of the options for additional features (e.g., 720b, X720b, 724c, and / or X724c), the computer system displays an option (e.g., 720e) via a display generation component to display additional information about an individual user (e.g., as part of one or more options associated with an individual user) (e.g., without displaying additional information about other users), (e.g., by replacing the display of an option to invite an individual user to join an ongoing communication session). In some embodiments, the computer system receives, via one or more sensors, a selection (e.g., 707d) of an option (e.g., 720e) for displaying additional information about an individual user (e.g., via touch input on a touch-sensitive surface and / or via air gestures).In some embodiments, in response to receiving a selection (e.g., 707d) of an option (e.g., 720e) to display additional information about an individual user, a computer system (e.g., 700 and / or X700) causes, via a display generation component (e.g., 702 and / or X702), additional information about the individual user (e.g., 740 of FIG. 7F) (e.g., previous communication history with the individual user, the individual user's phone number, and / or the individual user's email address) (e.g., without displaying additional information about other users) to be displayed that was not displayed (e.g., when an option for an additional feature was not selected and / or when an option for additional information was not selected). Providing the user with additional information about an individual user provides feedback to the user of the computer system about the individual user and / or the individual user's device, thereby providing improved visual feedback.
[0282] In some embodiments, in response to receiving a selection (e.g., 705c and / or 711b) of an individual user representation (e.g., 712e, X712e, 714g, and / or X714g) (e.g., independently of determining whether there is an ongoing communication session), the computer system (e.g., 700 and / or X700) displays options for additional features (e.g., 720b, X720b, 724c, and / or X724c) via display generation components (e.g., 702 and / or X702) (e.g., without displaying an option to initiate a process for sending a message to the individual user, and / or without displaying additional information about the individual user). In some embodiments, while displaying options for additional features (e.g., 720b, X720b, 724c, and / or X724c), the computer system (e.g., 700 and / or X700) receives a selection of options for additional features (e.g., 705c) via one or more sensors (e.g., via touch input on a touch-sensing surface and / or via air gestures). In some embodiments, in response to receiving a selection (e.g., 705c) of options for additional features (e.g., 720b and / or X720b), the computer system displays an option (e.g., 720f) via a display generating component to remove an individual user from favorites (e.g., a list or group of favorite users) (e.g., without removing other users) (e.g., as part of one or more options associated with an individual user). In some embodiments, removing an individual user from favorites includes ceasing to display the individual user's representation as part of a group of user representations (e.g., static avatars, animated avatars, images, and / or monograms) that are optionally displayed as part of the home user interface.In some embodiments, the computer system receives a selection of an option (e.g., 720f) to remove an individual user from favorites (e.g., 712) via one or more sensors (e.g., via touch input on a touch-sensitive surface and / or via air gestures). In some embodiments, upon receiving a selection of an option (e.g., 720f) to remove an individual user from favorites (e.g., 712), the computer system (e.g., 700 and / or X700) initiates a process to remove the individual user from favorites (e.g., requesting confirmation to remove the individual user from favorites and / or removing the individual user from favorites). By initiating a process to remove an individual user from favorites, the user is able to restrict which users are accessible through favorites and / or the home user interface, thereby reducing visual clutter and enabling the user to add other users to favorites, thereby providing improved visual feedback.
[0283] In some embodiments, upon receiving a selection (e.g., 711b) of an individual user's representation (e.g., 714g and / or X714g), and in accordance with the determination that there is an ongoing (e.g., active and / or currently established) communication session (e.g., a video communication session, an audio communication session, an extended reality communication session, a spatial communication session, and / or a non-spatial communication session), the computer system (e.g., 700 and / or X700) displays, via a display generation component (e.g., 702 and / or X702), an option (e.g., 724b and / or X724b) for inviting the individual user to join the ongoing communication session, and simultaneously an option (e.g., 724a and / or X724a) for initiating a process to start a new communication session with the individual user. In some embodiments, a computer system (e.g., 700 and / or X700) receives a selection (e.g., 709c) of an option (e.g., 724a and / or X724a) for initiating a process to start a new communication session with an individual user (e.g., via touch input on a touch-sensing surface and / or via air gesture) via one or more sensors. In some embodiments, upon receiving the selection (e.g., 709c) of an option (e.g., 724a and / or X724a) for initiating a process to start a new communication session with an individual user, the computer system (e.g., 700 and / or X700) terminates the ongoing communication session (e.g., 740 in Figure 7G) and initiates a process to start a new communication session with an individual user. In some embodiments, upon receiving the selection of an option to start a process to start a new communication session with an individual user, the computer system automatically terminates the ongoing communication session and initiates a new communication session with an individual user (e.g., without requesting and / or receiving additional input from the user and / or requiring user confirmation).Providing an option to start a new communication session with an individual user allows a computer system to both terminate an ongoing communication session and start a new one without requiring separate user input for each, thereby reducing the number of inputs required to perform the operation.
[0284] In some embodiments, during the process of initiating a new communication session with an individual user (e.g., 740 in Figure 7G), the computer system (e.g., 700 and / or X700) prompts the user to confirm termination of the ongoing communication session (e.g., via audio using a speaker and / or via display using a display generating component) (e.g., 742). In some embodiments, the computer system (e.g., 700 and / or X700) receives confirmation to terminate the ongoing communication session via one or more sensors (e.g., while displaying a prompt to confirm termination of the ongoing communication session). In some embodiments, upon receiving confirmation to terminate the ongoing communication session, the computer system terminates the ongoing communication session (and optionally initiates a new communication session with the individual user). Requiring confirmation from the user to terminate an ongoing communication session allows the computer system to prevent the user from unintentionally terminating an ongoing communication session, thereby improving the human-machine interface.
[0285] In some embodiments, representations of multiple users (e.g., static avatars, animated avatars, images, and / or monograms) are displayed as part of the home user interface (e.g., optionally including representations of recently communicated contacts, as shown in Figures 7B-7D), and displaying options (e.g., 724b and / or X724b) for inviting individual users to join an ongoing communication session, and / or displaying one or more options associated with individual users (e.g., 720a-720b, X720a-X720b, 724a-724c, and / or X724a-X724c) includes obscuring the home user interface (e.g., partially blocking, blurring, and / or otherwise partially obscuring the display). In some embodiments, the home user interface includes multiple user interface objects for displaying respective applications (e.g., a first user interface object that, when activated, causes the display of the user interface of a first application, and a second user interface object that, when activated, causes the display of the user interface of a second application distinct from the first application). In some embodiments, in response to detecting individual user input (e.g., detecting a physical button press and / or detecting a gesture such as an air gesture), the computer system displays the home user interface (e.g., regardless of what the computer system is displaying when the individual user input is received). In some embodiments, following the computer system exiting low-power mode (e.g., waking up) and / or receiving user input to unlock the computer system, the computer system automatically displays the home user interface. Continuing to display the home user interface (while hidden) provides the user with context about the content they are accessing, including information about the individual user (e.g., name and / or contact information).
[0286] In some embodiments, one or more options associated with an individual user include an option (e.g., 720d) for initiating a process to send a message (e.g., not including live transmission of audio and / or video) to the individual user (e.g., as part of one or more options associated with the individual user). In some embodiments, a computer system (e.g., 700 and / or X700) receives a selection (e.g., 705d) of an option (e.g., 720d) for initiating a process to send a message to an individual user via one or more sensors. In some embodiments, in response to the reception (e.g., 705d) of an option (e.g., 720d) for initiating the process of sending a message to an individual user (and optionally, according to the determination that no message from the individual user was displayed when the selection of the individual user's representation was received), the computer system (e.g., 700 and / or X700) displays a messaging user interface (e.g., 730 in Figure 7E) having a first appearance (e.g., a messaging user interface of a first size and / or a messaging user interface including a displayed keyboard) (for messaging with an individual user) via a display generating component (e.g., 702 and / or X702) without displaying the home user interface. Displaying a messaging user interface having a first appearance and not displaying the home user interface provides the user with feedback that the messaging user interface is in a first state, thereby providing the user with improved visual feedback.
[0287] In some embodiments, one or more options associated with an individual user include options (e.g., 718 and / or X718) for initiating a process to send a message (e.g., not including live transmission of audio and / or video) to the individual user (e.g., as part of one or more options associated with the individual user). In some embodiments, a computer system (e.g., 700 and / or X700) receives a selection (e.g., 707c) of options (e.g., 718 and / or X718) for initiating a process to send a message to an individual user via one or more sensors. In some embodiments, upon receiving a selection (e.g., 707c) of an option (e.g., 718 and / or X718) for initiating the process of sending a message to an individual user (and optionally, according to a determination that a message from an individual user was displayed when the selection of the individual user representation was received), the computer system (e.g., 700 and / or X700) simultaneously displays, via a display generation component, a messaging user interface (e.g., 750) having a second appearance (e.g., a messaging user interface of a second size smaller than the first size and / or a messaging user interface that does not include a displayed keyboard) (for messaging with an individual user), and at least a portion of the home user interface (e.g., as shown in Figure 7H) (e.g., displaying an obscured home user interface). Displaying the messaging user interface having the second appearance and a portion of the home user interface provides the user with feedback that the messaging user interface is in a second state, thereby providing the user with improved visual feedback.
[0288] In some embodiments, the home user interface is not user-movable, while the messaging user interface (e.g., 750) (e.g., having a first appearance and / or a second appearance) is user-movable. Allowing the user to move the messaging user interface without allowing the user to move the home user interface provides the user with feedback on which elements are part of the home user interface and which are not, thereby providing the user with improved visual feedback.
[0289] In some embodiments, displaying representations of multiple users (e.g., 712a-712g and 714a-714i) via display generating components (e.g., 702 and / or X702) includes displaying a first group (e.g., 712) of first multiple representations of a first multiple user via display generating components (e.g., 702 and / or X702), where the first multiple (e.g., 712a-712g) users are selected to be included as part of the representation of the multiple users, regardless of the recency of communication between the computer system user and the first multiple users (e.g., on the basis of being manually selected as part of favorite contacts and / or on the basis of the frequency of communication). In some embodiments, displaying representations of multiple users (e.g., 712a-712g and 714a-714i) via a display generation component includes displaying a second group (e.g., 714) of second multiple representations of a second multiple user via a display generation component (e.g., 702 and / or X702), where the second multiple (e.g., 714a-714i) users are selected to be included as part of the multiple users based on the recency of communication between the computer system user and the second multiple users (e.g., included to be displayed based on being the most recently communicated user). In some embodiments, the order of second multiple representations of second multiple users is based on the recency of communication between the computer system user and the second multiple users.In some embodiments, displaying representations of multiple users via a display generation component includes displaying a first group of first multiple representations of a first multiple user via a display generation component, wherein the first multiple user is selected to be included as part of the representation of multiple users, regardless of the relevance of communications between the computer system user and a second group of second multiple representations of the first multiple user and the second multiple user, and the second multiple user includes the first set of one or more users but not the second set of one or more users, according to a determination that recent communications by the computer system user include communications with the first set of one or more users but not the second set of one or more users, and the second multiple user includes the second set of one or more users but not the first set of one or more users, according to a determination that recent communications by the computer system user include communications with the first set of one or more users but not the second set of one or more users. Grouping a first set of users and a second set of users together allows the computer system to provide users with feedback on which users were selected regardless of the recency of the communication, and which users were included based on the recency of the communication, thereby providing improved visual feedback.
[0290] In some embodiments, the relevance of communications between a user of a computer system (e.g., 700 and / or X700) and a second group of users is based on multiple communication modes (e.g., text messaging, phone calls, and / or communication sessions (e.g., video communication sessions, audio communication sessions, extended reality communication sessions, spatial communication sessions, and / or non-spatial communication sessions)). Grouping the second group of users together based on the relevance of communications using multiple communication modes allows the computer system to group recent contacts regardless of how the communications occurred, thereby providing improved visual feedback.
[0291] In some embodiments, the computer system (e.g., 700 and / or X700) displays, via a display generation component (e.g., 702 and / or X702), an indication of recent communication activity between the computer system user and the individual user (e.g., 714f and / or X714f) (e.g., information about recent calls or communications, information about active calls or communications, and / or information about recent messages) along with an option (e.g., 724b and / or X724b) for inviting an individual user to join an ongoing communication session and / or one or more options associated with the individual user (e.g., 724a, X724a, 724c, and / or X724c). Displaying an indication of recent communication activity along with one or more options allows the user to see which recent communication method has been used and quickly select a communication method for another communication session, thereby reducing the number of inputs required to initiate the appropriate type of communication session.
[0292] In some embodiments, one or more options associated with an individual user (e.g., 720c-720f) include options (e.g., 724a and / or X724a) for initiating a spatial communication session with the individual user (e.g., a communication session with at least some (e.g., less than all, multiple, and / or all) representations of users participating in a distributed communication session within a 3D environment). In some embodiments, the computer system receives a selection (e.g., 709c) of an option (e.g., 724a and / or X724a) for initiating a spatial communication session with an individual user (e.g., via touch input on a touch-sensing surface and / or via air gestures) via one or more sensors. In some embodiments, in response to receiving a selection (e.g., 709c) of an option (e.g., 724a and / or X724a) for initiating a spatial communication session with an individual user, the computer system (e.g., 700 and / or X700) initiates a spatial communication session with the individual user (e.g., Figure 7P, etc.). By providing an option to initiate a spatial communication session with individual users, the computer system can initiate a communication session without requiring separate user input directed at selecting users to participate in the session, thereby reducing the number of inputs required to perform the operation.
[0293] In some embodiments, a computer system (e.g., 700 and / or X700) displays, via display generation components (e.g., 702 and / or X702), multiple user representations (e.g., 712a-712g and 714a-714i) simultaneously with options (e.g., 715) for previewing and / or editing the computer system's user avatar. In some embodiments, a computer system (e.g., 700 and / or X700) receives (e.g., 713b) a selection (e.g., 713b) of the options (e.g., 715) for previewing and / or editing the computer system's user avatar via one or more sensors (e.g., via touch input on a touch-sensitive surface and / or via air gestures). In some embodiments, upon receiving a selection (e.g., 713b) of an option (e.g., 715) for previewing and / or editing the computer system user's avatar, the computer system (e.g., 700 and / or X700) displays a user interface (e.g., 760) for previewing and / or editing the computer system user's avatar (e.g., via a self-view of the avatar moving with the user's movement) through a display generation component (e.g., 702 and / or X702). Displaying a user interface for previewing and / or editing the computer system user's avatar provides the user with visual feedback regarding the avatar's visual characteristics and the edits that can be made, thereby providing improved visual feedback.
[0294] In some embodiments, a computer system (e.g., 700 and / or X700) simultaneously displays representations of multiple users (e.g., 712a-712g and 714a-714i) via display generation components (e.g., 702 and / or X702) and first indications (e.g., 9:41 of 714e in Figure 7B) (e.g., the first) of recent communications (e.g., text messages, established calls, and / or missed calls) between a user of the computer system and a ...
Claims
1. In a computer system that communicates with display generation components, While participating in a communication session, which is a spatial communication session, the representations of multiple participants in the communication session are displayed in a spatially distributed arrangement in a 3D environment via the display generation component, In the 3D environment, the representations of the multiple participants separated from each other and from the user of the computer system by at least a threshold amount in the first non-vertical direction, Displaying the plurality of participants in the spatially dispersed arrangement includes displaying the representations of the plurality of participants spaced at least by the threshold amount apart from each other and from the user in a second non-vertical direction different from the first non-vertical direction, While displaying the representations of the multiple participants distributed within the 3D environment, an event is detected. A method comprising, in response to detecting the aforementioned event, transitioning the communication session from a spatial communication session to a non-spatial communication session, wherein the transition includes, via the display generation component, displaying representations of at least a subset of the plurality of participants in the communication session in a grouped arrangement, in the grouped arrangement, The representations of at least the subset of the plurality of participants are not dispersed in 3D space and are spaced less than the threshold amount apart in the first non-vertical direction in the 3D environment. The representation of the first participant in the grouped arrangement has a different position from the representation of the first participant in the spatially dispersed arrangement. A method wherein the representation of the second participant in the grouped arrangement has a different position from the representation of the second participant in the spatially dispersed arrangement.
2. In the aforementioned non-spatial communication session, The representation of the first participant among the aforementioned multiple participants is located within the first window area. The representation of the second participant among the aforementioned multiple participants is located in a second window area different from the first window area. In the aforementioned spatial communication session, The representation of the first participant among the aforementioned multiple participants is not within the window area. The method according to claim 1, wherein the representation of the second participant among the plurality of participants is not within the window area.
3. The method according to claim 2, wherein the representation of the first participant is a simulated three-dimensional representation, and the representation of the second participant is a two-dimensional representation.
4. The method according to claim 2, wherein the plurality of participants are represented in two dimensions.
5. The method according to claim 2, wherein the plurality of participants are represented in three dimensions.
6. The method according to claim 1, wherein the event is a request received during the communication session to transition the representation of the user of the computer system from a 3D representation to a 2D representation.
7. The method according to claim 6, wherein the above requirement is based on an input in the communication session control area.
8. The method according to claim 7, wherein the communication session control area includes an option for transitioning the representation of the user of the computer system from the 3D representation to the 2D representation, and one or more options corresponding to other communication session controls.
9. The method according to claim 1, wherein the event is a request received during the communication session to transition the communication session from the spatial communication session to the non-spatial communication session.
10. The method according to claim 1, wherein the event is an additional participant joining the communication session.
11. The method according to claim 10, wherein the addition of the aforementioned participants to the communication session causes the number of participants represented by the simulated three-dimensional representation to exceed a threshold number of participants.
12. The method according to claim 1, further comprising shifting the position of an individual window area corresponding to an individual participant based on the movement of the individual participant while the communication session is a non-spatial communication session.
13. The method according to claim 12, wherein the individual window regions move forward and / or backward within the virtual environment based on the head position of the individual participant.
14. The method according to claim 12, wherein the individual window areas are tilted based on the head position of the individual participant.
15. The method according to claim 12, wherein a first window shifts in a first direction based on the movement of a participant displayed in the first window, and a second window shifts in a second direction different from the first direction based on the movement of a participant displayed in the second window.
16. While participating in the aforementioned communication session, which is a non-spatial communication session, a second event is detected, The method according to claim 1, further comprising: transitioning the communication session from the non-spatial communication session to the spatial communication session in response to the detection of the second event.
17. The method according to claim 16, wherein the second event is that a participant leaves the communication session.
18. The method according to claim 16, wherein the second event is a request received during the communication session to transition the computer system's representation of the user from a 2D representation to a 3D representation.
19. The method according to claim 16, wherein the second event is a request received during the communication session to transition the communication session from a non-spatial communication session to a spatial communication session.
20. The method according to claim 1, further comprising displaying a self-view of the user's representation of the computer system in a self-view window area via the display generation component during a spatial communication session.
21. The method according to claim 20, wherein the self-view window area overlaps with a window area that includes a representation of another participant in the ongoing communication session.
22. The method according to claim 20, wherein the self-view window area is smaller than the window area that includes the representation of another participant.
23. During a spatial communication session, the first participant in the communication session can move their individual representation, and the second participant in the communication session can move their individual representation. The method according to claim 1, wherein during a non-spatial communication session, the user of the computer system can move each window area containing each of the representations of the plurality of participants in the communication session.
24. The computer system is configured to communicate with one or more sensors. The aforementioned method, The system detects user input via one or more sensors communicating with the computer system to rearrange individual window areas containing individual representations of participants. The method of claim 23, further comprising rearranging a plurality of window regions for a plurality of participants in response to detecting user input for rearranging the individual window regions containing the individual representations of the participant.
25. The method according to claim 23, wherein each of the aforementioned representations of the plurality of participants is arranged in an initial arrangement based on a predetermined arrangement rule.
26. The display generation component displays a representation of an invited user who is not currently a participant in the communication session, In accordance with the determination that the communication session is a non-spatial communication session, the user of the computer system is made able to rearrange the representation of the invited user who is not currently a participant. The method of claim 23, further comprising, in accordance with the determination that the communication session is a spatial communication session, ceasing to allow the user of the computer system to rearrange the representation of the invited user who is not currently a participant.
27. The method according to claim 1, further comprising displaying an indication via the display generation component that the mode of the communication session has changed in response to a change between a spatial communication session and a non-spatial communication session.
28. A computer program, It is configured to communicate with the display generation component, A computer program that performs the method described in any one of claims 1 to 27.
29. A computer system, wherein the computer system is A memory for storing the computer program described in claim 28, The system comprises one or more processors capable of executing the computer program stored in the memory, The computer system is configured to communicate with the display generation component. Computer system.
30. A computer system configured to communicate with a display generation component, A computer system comprising means for performing the method described in any one of claims 1 to 27.
Citation Information
Patent Citations
Information processing device and program
JP2022109048A
Shared virtual area communication environment based apparatus and methods
US20140237393A1
Controls and Interfaces for User Interactions in Virtual Spaces
US20180095635A1
3D object annotation
US20210256261A1
Interfaces for presenting avatars in three-dimensional environments
US20220262080A1