DEVICE, METHOD, AND GRAPHICAL USER INTERFACE FOR GENERATING AND DISPLAYING A USER'S REPRESENTATION - Patent application

By integrating display generation components in a computer system, the method of generating and displaying user representatives solves the problems of inefficiency and unsatisfactory energy efficiency in the prior art, achieving a more efficient and intuitive user experience and reducing battery consumption.

JP7672580B2Active Publication Date: 2025-05-07APPLE INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024531175
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-11-16
Filing Date
2022-11-22
Publication Date
2025-05-07
Estimated Expiration
2042-11-22

AI Technical Summary

Technical Problem

In the prior art, the method of generating and displaying user representatives in an environment containing virtual elements is inefficient, complex and error-prone, resulting in poor user experience and poor energy efficiency of battery-operated devices.

Method used

By integrating display generation components in a computer system, a system is provided that when placed on the user's body, the user is prompted to remove the device and capture user information to generate a representative. The system detects that the device removes and uses captured information to generate user representatives and optimizes processing to reduce computational load.

Benefits of technology

Improves efficiency and intuitiveness in generating and displaying user representatives, reduces the number and complexity of user input, reduces battery consumption, and extends the battery life of the device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007672580000001
    Figure 0007672580000001
  • Figure 0007672580000002
    Figure 0007672580000002
  • Figure 0007672580000003
    Figure 0007672580000003
Patent Text Reader

Abstract

In some examples, the computer system displays a prompt to remove the computer system from the user's body while the computer system is disposed on the user's body so that the computer system can capture information about the user to generate a representation of the user. In some examples, the computer system captures information about the user, generates a representation of the user, and displays the representation of the user. In some examples, the computer system displays the representation of the user having a different appearance based on direct information about the state of the user's body that has not been received for one or more predetermined amounts of time. In some examples, the computer system displays different mouth representations of the user's representation based on captured information about the physical state of the user's mouth. In some examples, the computer system displays different portions of the hair representation of the user's representation based on the distance the portions of the hair representation are positioned from the individual portions of the user's representation. In some examples, the computer system displays connected portions of the user's representation that are visually emphasized compared to inner portions of the user's representation.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. patent application Ser. No. 17 / 988,532, filed November 16, 2022, entitled "DEVICES, METHODS, AND GRAPHICAL USER INTERFACES FOR GENERATING AND DISPLAYING A REPRESENTATION OF A USER," and U.S. provisional patent application Ser. No. 63 / 283,969, filed November 29, 2021, entitled "DEVICES, METHODS, AND GRAPHICAL USER INTERFACES FOR GENERATING AND DISPLAYING A REPRESENTATION OF A USER," the contents of each of which are incorporated herein by reference in their entirety.

[0002] Technical Field The present disclosure relates generally to computer systems that provide computer-generated experiences, including but not limited to electronic devices that provide virtual reality and mixed reality experiences via a display. [Background technology]

[0003] The development of computer systems for augmented reality has progressed significantly in recent years. Exemplary augmented reality environments include at least some virtual elements that replace or augment the physical world. Input devices such as cameras, controllers, joysticks, touch-sensitive surfaces, and touchscreen displays for computer systems and other electronic computing devices are used to interact with the virtual / augmented reality environment. Exemplary virtual elements include virtual objects such as digital images, video, text, icons, and control elements such as buttons and other graphics. Summary of the Invention

[0004] Some methods and interfaces for generating and / or displaying a user's representation in an environment that includes at least some virtual elements (e.g., applications, augmented reality environments, mixed reality environments, and virtual reality environments) are cumbersome, inefficient, and limited. For example, systems that capture data to generate a user's representation, display the user's representation, and / or receive insufficient feedback while displaying the user's representation are complex, tedious, error-prone, and impose significant cognitive burden on the user, detracting from the experience in the virtual / augmented reality environment. Additionally, these methods are unnecessarily time-consuming, thereby wasting computer system energy. This latter consideration is particularly important in battery-operated devices.

[0005] Thus, there is a need for a computer system having improved methods and interfaces for providing a user with a computer-generated experience that allows for the creation and / or display of a representation of the user in a more efficient and intuitive manner. Such methods and interfaces optionally complement or replace conventional methods for generating and / or displaying a representation of the user in an environment that includes at least some virtual elements. Such methods and interfaces reduce the number, extent, and / or type of inputs from the user by helping the user understand the connection between the input provided and the device response to that input, thereby creating a more efficient human-machine interface.

[0006] The above-mentioned drawbacks and other problems associated with user interfaces of computer systems are reduced or eliminated by the disclosed system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a notebook computer, a tablet computer, or a handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device such as a wristwatch or a head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also known as a "touch screen" or "touchscreen display"). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, the computer system has one or more output devices in addition to the display generating components, the output devices including one or more tactile output generators and / or one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, memory, and one or more modules, programs, or instruction sets stored in the memory for performing a plurality of functions. In some embodiments, a user interacts with the GUI through stylus and / or finger contacts and gestures on the touch-sensitive surface, the movement of the user's eyes and hands in space relative to the GUI (and / or computer system) or the user's body as captured by cameras and other movement sensors, and voice input as captured by one or more audio input devices.In some embodiments, the functions performed through the interactions optionally include image editing, drawing, presenting, word processing, creating spreadsheets, playing games, making phone calls, video conferencing, emailing, instant messaging, training support, digital photography, digital videography, web browsing, playing digital music, note taking, and / or playing digital videos, and executable instructions to perform those functions are optionally contained in a transient and / or non-transitory computer-readable storage medium or other computer program product configured to be executed by one or more processors.

[0007] There is a need for electronic devices with improved methods and interfaces for generating and / or displaying user expressions. Such methods and interfaces can complement or replace conventional methods for generating and / or displaying user expressions. Such methods and interfaces reduce the number, extent, and / or type of input from the user, creating a more efficient human-machine interface. Such methods and interfaces also display relevant portions of the user's expressions such that the processing power of the computer system is reduced, thereby creating a more efficient human-machine interface. For battery-operated computing devices, such methods and interfaces conserve power and increase the time between battery charges.

[0008] According to some embodiments, a method is described. The method is executed in a computer system in communication with one or more display generating components. The method includes displaying, via the one or more display generating components, a prompt while the computer system is disposed on a user's body, instructing the user to remove the computer system from the user's body and use the computer system to capture information related to the user; detecting that the computer system has been removed from the user's body after displaying the prompt instructing the user to remove the computer system from the user's body; and capturing the information related to the user after detecting that the computer system has been removed from the user's body, the computer system configured to generate a representation of the user using the information.

[0009] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generating components, the one or more programs including instructions for displaying, via the one or more display generating components, a prompt while the computer system is disposed on a user's body, instructing the user to remove the computer system from the user's body and to use the computer system to capture information related to the user, detecting that the computer system has been removed from the user's body after displaying the prompt instructing the user to remove the computer system from the user's body, and a computer system configured to generate a representation of the user using the information, capturing the information related to the user after detecting that the computer system has been removed from the user's body.

[0010] According to some embodiments, a temporary computer-readable storage medium is described, the temporary computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generating components, the one or more programs including instructions for displaying, via the one or more display generating components, a prompt while the computer system is disposed on a user's body, instructing the user to remove the computer system from the user's body and to use the computer system to capture information related to the user, detecting that the computer system has been removed from the user's body after displaying the prompt instructing the user to remove the computer system from the user's body, and a computer system configured to generate a representation of the user using the information, capturing the information related to the user after detecting that the computer system has been removed from the user's body.

[0011] According to some embodiments, a computer system is described. The computer system is in communication with one or more display generating components. The computer system includes one or more processors and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions to: display, via the one or more display generating components, a prompt while the computer system is disposed on a user's body, instructing the user to remove the computer system from the user's body and use the computer system to capture information related to the user; detect that the computer system has been removed from the user's body after displaying the prompt instructing the user to remove the computer system from the user's body; and capture information related to the user after detecting that the computer system has been removed from the user's body, the computer system configured to use the information to generate a representation of the user.

[0012] According to some embodiments, a computer system is described. The computer system is in communication with one or more display generating components. The computer system includes: means for displaying, via the one or more display generating components, a prompt while the computer system is disposed on the user's body, instructing the user to remove the computer system from the user's body and use the computer system to capture information related to the user; means for detecting that the computer system has been removed from the user's body after displaying the prompt instructing the user to remove the computer system from the user's body; and means for capturing information related to the user after detecting that the computer system has been removed from the user's body, the computer system being configured to use the information to generate a representation of the user.

[0013] According to some embodiments, a method is described. The method is executed on a computer system in communication with one or more display generation components. The method includes capturing information regarding one or more physical characteristics of a user of the computer system during an enrollment process for generating a representation of the user, generating a representation of the user based on the information regarding the one or more physical characteristics of the user, including selecting one or more physical characteristics of the representation based on the one or more captured physical characteristics of the user after capturing the information regarding the one or more physical characteristics of the user of the computer system, and displaying at least a portion of the representation of the user within an extended reality environment via the one or more display generation components after generating the representation of the user.

[0014] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generation components, the one or more programs including instructions for capturing information regarding one or more physical characteristics of a user of the computer system during an enrollment process for generating a representation of the user, generating a representation of the user based on the information regarding the one or more physical characteristics of the user, including selecting one or more physical characteristics of the representation based on the one or more captured physical characteristics of the user, and displaying at least a portion of the representation of the user within an extended reality environment via the one or more display generation components after generating the representation of the user.

[0015] According to some embodiments, a temporary computer-readable storage medium is described, the temporary computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generation components, the one or more programs including instructions for capturing information regarding one or more physical characteristics of a user of the computer system during an enrollment process for generating a representation of the user, generating a representation of the user based on the information regarding the one or more physical characteristics of the user, including selecting one or more physical characteristics of the representation based on the one or more captured physical characteristics of the user, and displaying at least a portion of the representation of the user within an extended reality environment via the one or more display generation components after generating the representation of the user.

[0016] According to some embodiments, a computer system is described. The computer system is in communication with one or more display generation components. The computer system includes one or more processors and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for capturing information regarding one or more physical characteristics of a user of the computer system during an enrollment process for generating a representation of the user, generating a representation of the user based on the information regarding the one or more physical characteristics of the user, including selecting one or more physical characteristics of the representation based on the one or more captured physical characteristics of the user, and displaying at least a portion of the representation of the user within an extended reality environment via the one or more display generation components after generating the representation of the user.

[0017] According to some embodiments, a computer system is described. The computer system is in communication with one or more display generation components. The computer system includes: means for capturing information regarding one or more physical characteristics of a user of the computer system during an enrollment process for generating a representation of the user; means for generating a representation of the user based on the information regarding the one or more physical characteristics of the user, including selecting one or more physical characteristics of the representation based on the one or more captured physical characteristics of the user after capturing the information regarding the one or more physical characteristics of the user of the computer system; and means for displaying at least a portion of the representation of the user within an extended reality environment via the one or more display generation components after generating the representation of the user.

[0018] According to some embodiments, a method is described that is executed on a first computer system in communication with one or more display generation components, the method including: displaying, via the one or more display generation components, a representation of a second user within an extended reality environment at a first fidelity while the first computer system is being used by a first user of the first computer system, the representation of the second user moving based on detected movements of the second user detected by the second computer system during a live communication session with the first computer system; varying an amount of direct information regarding a bodily state of the second user while displaying the representation of the second user within the extended reality environment; and in response to the varying amount of direct information regarding the bodily state of the second user, the first computer system adjusting a display of the second user's bodily state by ... and initiating displaying a representation of the second user at different fidelity via the above display generation components, wherein the displaying includes: displaying the representation of the second user at a second fidelity lower than the first fidelity via one or more display generation components in accordance with a determination that direct information regarding the second user's physical state is not received for a first amount of time that is longer than a first time threshold and shorter than a second time threshold; and displaying the representation of the second user at a third fidelity lower than the second fidelity via one or more display generation components in accordance with a determination that direct information regarding the second user's physical state is not received for a second amount of time that is longer than the first time threshold and longer than the second time threshold.

[0019] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a first computer system in communication with one or more display generating components, the one or more programs displaying, via the one or more display generating components, a representation of a second user within an extended reality environment at a first fidelity while the first computer system is being used by a first user of the first computer system, the representation of the second user moving based on detected movements of the second user detected by the second computer system during a live communication session with the first computer system, and while displaying the representation of the second user within the extended reality environment, an amount of direct information regarding a bodily state of the second user changes, and the amount of direct information regarding the bodily state of the second user changes. a first computer system beginning to display, via one or more display generation components, a representation of the second user at a different fidelity in response to a change in an amount of direct information regarding the second user's physical state, wherein the displaying includes: displaying, via the one or more display generation components, the representation of the second user at a second fidelity lower than the first fidelity in accordance with a determination that direct information regarding the second user's physical state is not received for a first amount of time that is longer than a first time threshold and shorter than a second time threshold; and displaying, via the one or more display generation components, the representation of the second user at a third fidelity lower than the second fidelity in accordance with a determination that direct information regarding the second user's physical state is not received for a second amount of time that is longer than the first time threshold and longer than the second time threshold.

[0020] According to some embodiments, a temporary computer-readable storage medium is described. The temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a first computer system in communication with one or more display generating components, the one or more programs displaying, via the one or more display generating components, a representation of a second user within an extended reality environment at a first fidelity while the first computer system is being used by a first user of the first computer system, the representation of the second user moving based on detected movements of the second user detected by the second computer system during a live communication session with the first computer system, and while displaying the representation of the second user within the extended reality environment, an amount of direct information regarding a bodily state of the second user changes, and the amount of direct information regarding the bodily state of the second user changes. a first computer system beginning to display, via one or more display generation components, a representation of the second user at a different fidelity in response to a change in an amount of direct information regarding the second user's physical state, wherein the displaying includes: displaying, via the one or more display generation components, the representation of the second user at a second fidelity lower than the first fidelity in accordance with a determination that no direct information regarding the second user's physical state is received for a first amount of time that is longer than a first time threshold and shorter than a second time threshold; and displaying, via the one or more display generation components, the representation of the second user at a third fidelity lower than the second fidelity in accordance with a determination that no direct information regarding the second user's physical state is received for a second amount of time that is longer than the first time threshold and longer than the second time threshold.

[0021] According to some embodiments, a first computer system is described that is in communication with one or more display generation components, the first computer system comprising one or more processors and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs displaying, via the one or more display generation components, a representation of a second user within an extended reality environment at a first fidelity while the first computer system is being used by a first user of the first computer system, the representation of the second user moving based on detected movements of the second user detected by the second computer system during a live communication session with the first computer system, and varying an amount of direct information regarding the second user's bodily state while displaying the representation of the second user within the extended reality environment, the amount of direct information regarding the second user's bodily state varying while displaying the representation of the second user within the extended reality environment. A first computer system including instructions that, in response to a change in the amount of information, the first computer system begins displaying a representation of the second user at a different fidelity via one or more display generation components, wherein the displaying includes: displaying the representation of the second user at a second fidelity lower than the first fidelity via the one or more display generation components in accordance with a determination that direct information regarding the second user's physical state is not received for a first amount of time that is longer than a first time threshold and shorter than a second time threshold; and displaying the representation of the second user at a third fidelity lower than the second fidelity via the one or more display generation components in accordance with a determination that direct information regarding the second user's physical state is not received for a second amount of time that is longer than the first time threshold and longer than the second time threshold.

[0022] According to some embodiments, a first computer system is described, the first computer system being in communication with one or more display generation components, the first computer system including: means for displaying, via the one or more display generation components, a representation of a second user within an extended reality environment at a first fidelity while the first computer system is being used by a first user of the first computer system, the representation of the second user moving based on detected movements of the second user detected by the second computer system during a live communication session with the first computer system; means for varying an amount of direct information regarding a bodily state of the second user while displaying the representation of the second user within the extended reality environment; and means for varying an amount of direct information regarding a bodily state of the second user while the first computer system is being used by a first user of the first computer system, the amount of direct information regarding the bodily state of the second user, the first computer system being configured to generate one or more display generation components. and means for initiating displaying a representation of the second user at different fidelity via the above display generation components, wherein the displaying includes: displaying the representation of the second user at a second fidelity lower than the first fidelity via one or more display generation components in accordance with a determination that direct information regarding the second user's physical state is not received for a first amount of time that is longer than a first time threshold and shorter than a second time threshold; and displaying the representation of the second user at a third fidelity lower than the second fidelity via the one or more display generation components in accordance with a determination that direct information regarding the second user's physical state is not received for a second amount of time that is longer than the first time threshold and longer than the second time threshold.

[0023] According to some embodiments, a method is described that is executed on a first computer system in communication with one or more display generation components, the method including: displaying, via the one or more display generation components, a representation of a second user within an extended reality environment while the first computer system is being used by a first user of the first computer system, the representation of the second user moving based on detected movements of the second user detected by the second computer system during a live communication session with the first computer system; receiving, while displaying the representation of the second user within the extended reality environment, information corresponding to speech of the second user; and, in response to receiving the information corresponding to the speech of the second user, updating an appearance of the representation of the second user based on the information corresponding to the speech of the second user. The displaying includes displaying, via one or more display generation components, a first mouth expression of the representation of the second user in accordance with a determination that the information regarding the detected physical state of the second user's mouth does not satisfy a set of one or more criteria, the first mouth expression being generated based on audio information corresponding to an utterance of the second user; and displaying, via one or more display generation components, a second mouth expression of the representation of the second user in accordance with a determination that the information regarding the detected physical state of the second user's mouth satisfies the set of one or more criteria, the second mouth expression being generated based on information regarding the detected physical state of the second user's mouth without using audio information corresponding to the utterance of the second user to generate the second mouth expression.

[0024] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a first computer system in communication with one or more display generating components, the one or more programs displaying, via the one or more display generating components, within an extended reality environment while the first computer system is being used by a first user of the first computer system, a representation of a second user, the representation of the second user moving based on detected movements of the second user detected by the second computer system during a live communication session with the first computer system; receiving, while displaying the representation of the second user within the extended reality environment, information corresponding to an utterance of the second user; and, in response to receiving the information corresponding to the utterance of the second user, displaying information corresponding to the utterance of the second user. a non-transitory computer-readable storage medium comprising instructions for updating an appearance of a representation of a second user based on: displaying, via one or more display generation components, a first mouth expression of the representation of the second user in accordance with a determination that information regarding the detected physical state of the second user's mouth does not satisfy a set of one or more criteria, the first mouth expression being generated based on audio information corresponding to an utterance of the second user; and displaying, via the one or more display generation components, a second mouth expression of the representation of the second user in accordance with a determination that information regarding the detected physical state of the second user's mouth satisfies the set of one or more criteria, the second mouth expression being generated based on information regarding the detected physical state of the second user's mouth without using audio information corresponding to the second user's utterance to generate the second mouth expression.

[0025] According to some embodiments, a temporary computer-readable storage medium is described. The temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a first computer system in communication with one or more display generating components, the one or more programs displaying, via the one or more display generating components, within an extended reality environment while the first computer system is being used by a first user of the first computer system, a representation of a second user, the representation of the second user moving based on detected movements of the second user detected by the second computer system during a live communication session with the first computer system; receiving, while displaying the representation of the second user within the extended reality environment, information corresponding to an utterance of the second user; and, in response to receiving the information corresponding to the utterance of the second user, displaying information corresponding to the utterance of the second user. a transient computer-readable storage medium comprising instructions for updating an appearance of a representation of a second user based on: displaying, via one or more display generation components, a first mouth expression of the representation of the second user in accordance with a determination that information regarding the detected physical state of the second user's mouth does not satisfy a set of one or more criteria, the first mouth expression being generated based on audio information corresponding to an utterance of the second user; and displaying, via the one or more display generation components, a second mouth expression of the representation of the second user in accordance with a determination that information regarding the detected physical state of the second user's mouth satisfies the set of one or more criteria, the second mouth expression being generated based on information regarding the detected physical state of the second user's mouth without using audio information corresponding to the second user's utterance to generate the second mouth expression.

[0026] According to some embodiments, a first computer system is described that is in communication with one or more display generation components. The first computer system includes one or more processors and a memory that stores one or more programs configured to be executed by the one or more processors. The one or more programs display, via the one or more display generation components, a representation of a second user within an extended reality environment while the first computer system is being used by a first user of the first computer system, the representation of the second user moving based on detected movements of the second user detected by the second computer system during a live communication session with the first computer system. While displaying the representation of the second user within the extended reality environment, the first computer system receives information corresponding to speech of the second user. In response to receiving the information corresponding to the speech of the second user, the one or more programs display a second user's utterance based on the information corresponding to the speech of the second user. A first computer system including instructions for updating an appearance of a representation of a second user, the updating including: displaying, via one or more display generation components, a first mouth expression of the representation of the second user in accordance with a determination that information regarding a detected physical state of the second user's mouth does not satisfy a set of one or more criteria, the first mouth expression being generated based on audio information corresponding to an utterance of the second user; and displaying, via the one or more display generation components, a second mouth expression of the representation of the second user in accordance with a determination that information regarding the detected physical state of the second user's mouth satisfies the set of one or more criteria, the second mouth expression being generated based on information regarding the detected physical state of the second user's mouth without using audio information corresponding to the second user's utterance to generate the second mouth expression.

[0027] According to some embodiments, a first computer system is described, the first computer system being in communication with one or more display generation components, the first computer system comprising: means for displaying, via the one or more display generation components, within an extended reality environment while the first computer system is being used by a first user of the first computer system, a representation of a second user, the representation of the second user moving based on detected movements of the second user detected by the second computer system during a live communication session with the first computer system; means for receiving, while displaying the representation of the second user within the extended reality environment, information corresponding to speech of the second user; and means for updating an appearance of the representation of the second user based on the information corresponding to the speech of the second user in response to receiving the information corresponding to the speech of the second user. In the system, the updating includes: displaying, via one or more display generation components, a first mouth expression of the representation of the second user in accordance with a determination that the information regarding the detected physical state of the second user's mouth does not satisfy a set of one or more criteria, the first mouth expression being generated based on audio information corresponding to an utterance of the second user; and displaying, via the one or more display generation components, a second mouth expression of the representation of the second user in accordance with a determination that the information regarding the detected physical state of the second user's mouth satisfies the set of one or more criteria, the second mouth expression being generated based on information regarding the detected physical state of the second user's mouth without using audio information corresponding to the second user's utterance to generate the second mouth expression.

[0028] According to some embodiments, a method is described that is executed on a first computer system in communication with one or more display generating components, the method including displaying, via the one or more display generating components, within an extended reality environment while the first computer system is being used by a first user of the first computer system, a representation of a second user, the representation of the second user moving based on detected movements of the second user detected by the second computer system during a live communication session with the first computer system, the representation of the second user including a visual representation of hair of the second user, The visual representation of the second user's hair includes a first portion of the hair representation positioned a first distance from a portion of the second user's representation corresponding to an individual body part of the second user, the first portion of the hair representation having a first visual fidelity, and a second portion of the hair representation positioned a second distance greater than the first distance from the portion of the second user's representation corresponding to the individual body part of the second user, the second portion of the hair representation having a second visual fidelity less than the first visual fidelity.

[0029] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a first computer system in communication with one or more display generating components, the one or more programs configured to display, via the one or more display generating components, within an extended reality environment while the first computer system is being used by a first user of the first computer system, a representation of a second user, the representation of the second user moving based on detected movements of the second user detected by the second computer system during a live communication session with the first computer system. a non-transitory computer-readable storage medium comprising instructions for: generating a visual representation of hair of the second user, the visual representation of the second user's hair comprising: a first portion of the hair representation positioned at a first distance from a portion of the second user's representation corresponding to a respective body part of the second user, the first portion of the hair representation comprising a first visual fidelity; and a second portion of the hair representation positioned at a second distance greater than the first distance from the portion of the second user's representation corresponding to the respective body part of the second user, the second portion of the hair representation comprising a second visual fidelity less than the first visual fidelity.

[0030] According to some embodiments, a temporary computer-readable storage medium is described. The temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a first computer system in communication with one or more display generating components, the one or more programs configured to display, via the one or more display generating components, within an extended reality environment while the first computer system is being used by a first user of the first computer system, a representation of a second user, the representation of the second user moving based on detected movements of the second user detected by the second computer system during a live communication session with the first computer system. a first portion of the hair representation positioned a first distance from a portion of the second user's representation corresponding to a respective body part of the second user, the first portion of the hair representation comprising a first visual fidelity; and a second portion of the hair representation positioned a second distance greater than the first distance from the portion of the second user's representation corresponding to the respective body part of the second user, the second portion of the hair representation comprising a second visual fidelity less than the first visual fidelity.

[0031] According to some embodiments, a first computer system is described that is in communication with one or more display generating components, the first computer system comprising one or more processors and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for displaying, via the one or more display generating components, within an extended reality environment while the first computer system is being used by a first user of the first computer system, a representation of a second user, the representation of the second user moving based on detected movements of the second user detected by the second computer system during a live communication session with the first computer system. A first computer system, wherein the representation of a second user includes a visual representation of hair of the second user, the visual representation of the second user's hair including: a first portion of the hair representation positioned at a first distance from a portion of the second user's representation corresponding to a respective body part of the second user, the first portion of the hair representation including a first visual fidelity; and a second portion of the hair representation positioned at a second distance greater than the first distance from the portion of the second user's representation corresponding to the respective body part of the second user, the second portion of the hair representation including a second visual fidelity less than the first visual fidelity.

[0032] According to some embodiments, a first computer system is described, the first computer system being in communication with one or more display generation components, the first computer system comprising means for displaying, via the one or more display generation components, within an extended reality environment while the first computer system is being used by a first user of the first computer system, a representation of a second user, the representation of the second user moving based on detected movements of the second user detected by the second computer system during a live communication session with the first computer system, the representation of the second user including hair of the second user. wherein the visual representation of the second user's hair includes a first portion of the hair representation positioned a first distance from a portion of the second user's representation corresponding to the second user's individual body part, the first portion of the hair representation comprising a first visual fidelity, and a second portion of the hair representation positioned a second distance greater than the first distance from the portion of the second user's representation corresponding to the second user's individual body part, the second portion of the hair representation comprising a second visual fidelity less than the first visual fidelity.

[0033] According to some embodiments, a method is described that is performed on a first computer system in communication with one or more display generation components. The method includes displaying, via one or more display generation components, within an extended reality environment while a first computer system is being used by a first user of the first computer system, a representation of a second user, the representation of the second user moving based on detected movements of the second user detected by the second computer system during a live communication session with the first computer system, the representation of the second user including: a first portion of the second user's representation corresponding to a boundary between the second user's representation and another portion of the extended reality environment, the first portion of the second user's representation being displayed using a first visual appearance; and a second portion of the second user's representation not corresponding to a boundary between the second user's representation and another portion of the extended reality environment, the second portion of the second user's representation being displayed using a second visual appearance, the first visual appearance being emphasized compared to the second visual appearance.

[0034] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a first computer system in communication with one or more display generating components, the one or more programs displaying, via the one or more display generating components, within an extended reality environment while the first computer system is being used by a first user of the first computer system, a representation of a second user, the representation of the second user moving based on detected movements of the second user detected by the second computer system during a live communication session with the first computer system. A non-transitory computer-readable storage medium comprising instructions for displaying representations of two users, the second user's representation including: a first portion of the second user's representation that corresponds to a boundary between the second user's representation and other portions of the extended reality environment, the first portion of the second user's representation being displayed using a first visual appearance; and a second portion of the second user's representation that does not correspond to a boundary between the second user's representation and other portions of the extended reality environment, the second portion of the second user's representation being displayed using a second visual appearance, the first visual appearance being emphasized compared to the second visual appearance.

[0035] According to some embodiments, a temporary computer-readable storage medium is described. The temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a first computer system in communication with one or more display generating components, the one or more programs displaying, via the one or more display generating components, within an extended reality environment while the first computer system is being used by a first user of the first computer system, a representation of a second user, the representation of the second user moving based on detected movements of the second user detected by the second computer system during a live communication session with the first computer system. A temporary computer-readable storage medium comprising instructions for displaying representations of two users, the second user's representation including: a first portion of the second user's representation that corresponds to a boundary between the second user's representation and other portions of the extended reality environment, the first portion of the second user's representation being displayed using a first visual appearance; and a second portion of the second user's representation that does not correspond to a boundary between the second user's representation and other portions of the extended reality environment, the second portion of the second user's representation being displayed using a second visual appearance, the first visual appearance being emphasized compared to the second visual appearance.

[0036] According to some embodiments, a first computer system is described that is in communication with one or more display generation components. The first computer system comprises one or more processors and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs being configured to generate, via the one or more display generation components, a representation of a second user within an extended reality environment while the first computer system is being used by a first user of the first computer system, the representation of the second user moving based on detected movements of the second user detected by the second computer system during a live communication session with the first computer system. a first computer system including instructions for displaying a representation of a second user, the representation of the second user including: a first portion of the second user's representation that corresponds to a boundary between the second user's representation and another portion of the extended reality environment, the first portion of the second user's representation being displayed using a first visual appearance; and a second portion of the second user's representation that does not correspond to a boundary between the second user's representation and another portion of the extended reality environment, the second portion of the second user's representation being displayed using a second visual appearance, the first visual appearance being emphasized compared to the second visual appearance.

[0037] According to some embodiments, a first computer system is described, the first computer system being in communication with one or more display generation components. The first computer system comprises means for displaying, via one or more display generation components, within the extended reality environment while the first computer system is being used by a first user of the first computer system, a representation of a second user, the representation of the second user moving based on detected movements of the second user detected by the second computer system during a live communication session with the first computer system, the first computer system comprising: a first portion of the second user's representation corresponding to a boundary between the second user's representation and other portions of the extended reality environment, the first portion of the second user's representation being displayed using a first visual appearance; and a second portion of the second user's representation not corresponding to a boundary between the second user's representation and other portions of the extended reality environment, the second portion of the second user's representation being displayed using a second visual appearance, the first visual appearance being emphasized compared to the second visual appearance.

[0038] It should be noted that the various embodiments described above can be combined with any other embodiment described herein. The features and advantages described herein are not exhaustive, and many additional features and advantages will become apparent to those skilled in the art, particularly in light of the drawings, specification, and claims. Furthermore, it should be noted that the language used in this specification has been selected solely for the purposes of readability and explanation, and not to define or limit the subject matter of the present invention. [Brief explanation of the drawings]

[0039] For a better understanding of the various described embodiments, reference should be made to the following Detailed Description of the Invention in conjunction with the following drawings, in which like reference numerals refer to corresponding parts throughout:

[0040] [Figure 1] FIG. 1 is a block diagram illustrating an operating environment for a computer system for providing an XR experience, according to some embodiments.

[0041] [Figure 2] FIG. 1 is a block diagram illustrating a controller of a computer system configured to manage and coordinate a user's XR experience, according to some embodiments.

[0042] [Figure 3] FIG. 1 is a block diagram illustrating display generation components of a computer system configured to provide a visual component of an XR experience to a user, according to some embodiments.

[0043] [Figure 4] FIG. 1 is a block diagram illustrating a hand tracking unit of a computer system configured to capture a user's gesture input, according to some embodiments.

[0044] [Figure 5] FIG. 1 is a block diagram illustrating an eye-tracking unit of a computer system configured to capture a user's gaze input, according to some embodiments.

[0045] [Figure 6] FIG. 1 is a flow diagram illustrating a glint-assisted gaze tracking pipeline, according to some embodiments.

[0046] [Figure 7A] 1 illustrates exemplary techniques for generating a representation of a user and displaying a representation of a user, according to some embodiments. [Figure 7B] 1 illustrates exemplary techniques for generating a representation of a user and displaying a representation of a user, according to some embodiments. [Figure 7C]1 illustrates exemplary techniques for generating a representation of a user and displaying a representation of a user, according to some embodiments. [Figure 7D] 1 illustrates exemplary techniques for generating a representation of a user and displaying a representation of a user, according to some embodiments. [Figure 7E] 1 illustrates exemplary techniques for generating a representation of a user and displaying a representation of a user, according to some embodiments. [Figure 7F] 1 illustrates exemplary techniques for generating a representation of a user and displaying a representation of a user, according to some embodiments. [Figure 7G] 1 illustrates exemplary techniques for generating a representation of a user and displaying a representation of a user, according to some embodiments. [Figure 7H] 1 illustrates exemplary techniques for generating a representation of a user and displaying a representation of a user, according to some embodiments. [Figure 7I] 1 illustrates exemplary techniques for generating a representation of a user and displaying a representation of a user, according to some embodiments. [Figure 7J] 1 illustrates exemplary techniques for generating a representation of a user and displaying a representation of a user, according to some embodiments.

[0047] [Figure 8] FIG. 1 is a flow diagram of a method for generating a representation of a user, according to various embodiments.

[0048] [Figure 9] FIG. 1 is a flow diagram of a method for displaying a user's representation, according to various embodiments.

[0049] [Figure 10A] 1 illustrates an exemplary technique for adjusting the appearance of a user's representation, according to some embodiments. [Figure 10B] 1 illustrates an exemplary technique for adjusting the appearance of a user's representation, according to some embodiments. [Figure 10C] 1 illustrates an exemplary technique for adjusting the appearance of a user's representation, according to some embodiments. [Figure 10D] 1 illustrates an exemplary technique for adjusting the appearance of a user's representation, according to some embodiments. [Figure 10E] 1 illustrates an exemplary technique for adjusting the appearance of a user's representation, according to some embodiments. [Figure 10F] 1 illustrates an exemplary technique for adjusting the appearance of a user's representation, according to some embodiments. [Figure 10G] 1 illustrates an exemplary technique for adjusting the appearance of a user's representation, according to some embodiments. [Figure 10H] 1 illustrates an exemplary technique for adjusting the appearance of a user's representation, according to some embodiments. [Figure 10I] 1 illustrates an exemplary technique for adjusting the appearance of a user's representation, according to some embodiments.

[0050] [Figure 11] FIG. 1 is a flow diagram of a method for adjusting the appearance of a user's representation, according to various embodiments.

[0051] [Figure 12] 1 is a flow diagram of a method for displaying mouth expressions of a user's expressions, according to various embodiments.

[0052] [Figure 13] FIG. 1 is a flow diagram of a method for displaying a representation of a user's expression, according to various embodiments.

[0053] [Figure 14] 1 is a flow diagram of a method for displaying a portion of a user's expression with a visual highlight, according to various embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0054] The present disclosure relates to a user interface that provides an extended reality (XR) experience to a user, according to some embodiments.

[0055] The systems, methods, and GUIs described herein improve user interface interaction with virtual / augmented reality environments in several ways.

[0056] In some embodiments, the computer system captures information related to the user and uses the captured information to generate a representation of the user. While the computer system is disposed on the user's body, the computer system prompts the user to remove the computer system from the user's body and use the computer system to capture information related to the user. The computer system detects that the computer system has been removed from the user's body and captures information about the user after detecting that the computer system has been removed from the user's body. In some embodiments, the computer system is a head-mounted display generating component and / or a wearable computer system, such as a watch, that can be worn in a particular orientation and / or position relative to the user's body. In some embodiments, the computer system captures information related to the user's head and / or face while the computer system is removed from the user's body and captures information related to the user's hands while the computer system is disposed on the user's body. In some embodiments, the computer system displays a first prompt on a first display generating component prompting the user to remove the computer system from the user's body and a second prompt on a second display generating component providing instructions to capture information about the user while the computer system is removed from the user's body.

[0057] In some embodiments, the computer system captures information about one or more physical characteristics of the user, generates a representation of the user based on the information about the one or more physical characteristics of the user, and displays the representation of the user within an extended reality environment, such as an augmented reality environment and / or a virtual reality environment. In some embodiments, the computer system displays the representation of the user to include a representative state that mirrors the user's physical state within the physical environment. In some embodiments, the computer system animates and / or displays movement of the representation based on the user's physical movement within the physical environment. In some embodiments, the computer system provides selectable options for editing the representation of the user and / or for recapturing information about one or more physical characteristics of the user while displaying the representation within the extended reality environment.

[0058] In some embodiments, a first computer system used by a first user displays a representation of a second user within an extended reality environment and adjusts the appearance of the representation of the second user based on the amount of direct information about the second user's physical state. For example, the computer system displays the representation of the second user with a first visual fidelity and / or accuracy. When the direct information about the second user's physical state is not received for a first amount of time that is longer than a first time threshold and shorter than a second time threshold, the computer system displays the representation of the second user with a second visual fidelity and / or accuracy that is lower than the first visual fidelity and / or accuracy. When the direct information about the second user's physical state is not received for a second amount of time that is longer than the first time threshold and longer than a second time threshold, the computer system displays the representation of the second user with a third visual fidelity and / or accuracy that is lower than the first visual fidelity and / or accuracy and lower than the second visual fidelity and / or accuracy. In some embodiments, when direct information regarding the second user's physical state is not received for a second amount of time, the computer system displays a representation of the second user in a presentation mode such that the representation of the second user does not have anthropomorphic features and / or is an inanimate object within the extended reality environment.

[0059] In some embodiments, a first computer system used by a first user displays a representation of a second user within an extended reality environment and displays a mouth expression of the second user's representation based on one or more of audio information corresponding to the second user's speech and / or information related to the detected physical state of the second user's mouth. The computer system receives audio information corresponding to the second user's speech and updates the appearance of the second user's representation based on the audio information corresponding to the user's speech. When the information related to the detected physical state of the second user's mouth does not satisfy one or more sets of criteria, such as the information related to the detected physical state of the second user's mouth being less than a confidence level threshold, the computer system displays the representation of the second user having a first mouth expression generated based on the audio information related to the second user's speech. When the information regarding the detected physical state of the second user's mouth meets one or more sets of criteria, such as the information regarding the detected physical state of the second user's mouth being greater than a confidence level threshold, the computer system displays a representation of the second user having a second mouth expression generated based on the information regarding the detected physical state of the second user's mouth without using audio information corresponding to the second user's speech. In some embodiments, the first mouth expression is a combination and / or overlay of a third mouth expression generated based on audio information corresponding to the second user's speech and a fourth mouth expression generated based on the information regarding the detected physical state of the second user's mouth. In some embodiments, the first mouth expression is generated using different amounts of the third mouth expression and the fourth mouth expression based on a confidence level of the information regarding the detected physical state of the second user's mouth.

[0060] In some embodiments, a first computer system used by a first user displays a representation of a second user within an extended reality environment, the representation including a visual representation of the second user's hair. The visual representation of the second user's hair includes a first portion positioned a first distance from a portion of the second user's representation corresponding to a distinct body part of the second user, such as the face and / or neck, and including a first visual fidelity and / or accuracy. The visual representation of the hair includes a second portion positioned a second distance greater than the first distance from the portion of the second user's representation corresponding to the distinct body part of the second user, and including a second visual fidelity and / or accuracy lower than the first visual fidelity and / or accuracy. Thus, the visual representation of the second user's hair becomes more obscured the farther the visual representation of the hair is positioned from the portion of the second user's representation corresponding to the distinct body part of the second user. In some embodiments, the visual representation of the hair corresponds only to the second user's facial hair and / or beard.

[0061] In some embodiments, a first computer system used by a first user displays a representation of a second user within an extended reality environment and displays different portions of the second user's representation with different levels and / or degrees of visual emphasis. For example, a first portion of the second user's representation that corresponds to a boundary between the second user's representation and other portions of the extended reality environment is displayed using a first visual appearance. A second portion of the second user's representation that does not correspond to a boundary between the second user's representation and other portions of the extended reality environment is displayed using a second visual appearance, the first visual appearance being visually emphasized relative to the second visual appearance. In some embodiments, the computer system adjusts the appearance of the second user's representation based on changes in the displayed viewpoint and / or perspective of the second user's representation such that the first and second portions of the second user's representation change based on changes in the displayed perspective and / or viewpoint. In some embodiments, the computer system displays the representation of the second user in a presentation mode when the representation of the second user is displayed in a rearward orientation, the presentation mode including displaying the representation of the second user without anthropomorphic features and / or as an inanimate object.

[0062] FIGS. 1-6 illustrate an exemplary computer system for providing an XR experience to a user. FIGS. 7A-7J illustrate exemplary techniques for generating a user's representation and displaying a user's representation, according to some embodiments. FIG. 8 is a flow diagram of a method for generating a user's representation, according to various embodiments. FIG. 9 is a flow diagram of a method for displaying a user's representation, according to various embodiments. The user interfaces of FIGS. 7A-7J are used to illustrate the processes of FIGS. 8 and 9. FIGS. 10A-10I illustrate an exemplary technique for adjusting the appearance of a user's representation, according to some embodiments. FIG. 11 is a flow diagram of a method for adjusting the appearance of a user's representation, according to various embodiments. FIG. 12 is a flow diagram of a method for displaying a mouth representation of a user's representation, according to various embodiments. FIG. 13 is a flow diagram of a method for displaying a hair representation of a user's representation, according to various embodiments. FIG. 14 is a flow diagram of a method for displaying a portion of a user's representation with visual highlighting, according to various embodiments. The user interfaces of FIGS. 10A-10I are used to illustrate the processes of FIGS. 11-14.

[0063] The processes described below improve device usability (e.g., by helping users provide appropriate inputs and reducing user errors when operating / interacting with the device), making user device interfaces more efficient through various techniques, including providing improved visual feedback to the user, reducing the number of inputs required to perform an action, providing additional control options without cluttering the user interface with additional displayed controls, performing an action when a set of conditions is met without requiring further user input, improving privacy and / or security, providing a more diverse, detailed, and / or realistic user experience while saving storage space, and / or through additional techniques. These techniques also reduce power usage and improve device battery life by allowing users to use the device more quickly and efficiently. Saving battery power, and therefore weight, improves device ergonomics. These techniques also enable real-time communication, allow the use of fewer and / or less accurate sensors, resulting in more compact, lighter, and less expensive devices, and allow devices to be used in a variety of lighting conditions. These techniques reduce energy use and thereby reduce the heat given off by the device, which is particularly important for wearable devices where a device that is well within the operating parameters for the device components may become uncomfortable for the user to wear if it is generating too much heat.

[0064] Furthermore, for methods described herein in which one or more steps are conditioned on one or more conditions being satisfied, it should be understood that the described method can be repeated in multiple iterations, such that over the course of the iterations, all of the conditions on which the method steps are conditioned are satisfied in different iterations of the method. For example, if a method requires performing a first step if a condition is satisfied and a second step if the condition is not satisfied, one skilled in the art will understand that the steps described in step 2 are repeated in a particular order until the conditions are met and are no longer satisfied. Thus, a method described with one or more steps that depend on one or more conditions being satisfied can be rewritten as a method that is repeated until each condition described in the method is satisfied. However, this is not required for system or computer-readable medium claims in which the system or computer-readable medium includes instructions that perform a conditional action based on the satisfaction of the corresponding one or more conditions, and thus can determine whether a contingency is met without explicitly repeating the method steps until all conditions on which the method steps are conditioned are satisfied. Those skilled in the art will also understand that, as with methods having conditional steps, the system or computer-readable storage medium may repeat the steps of the method as many times as necessary to ensure that all of the conditional steps have been performed.

[0065] 1, an XR experience is provided to a user via an operating environment 100 that includes a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted device (HMD), a display, a projector, a touchscreen, etc.), one or more input devices 125 (e.g., an eye-tracking device 130, a hand-tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a tactile output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a tactile sensor, an orientation sensor, a proximity sensor, a temperature sensor, a location sensor, a motion sensor, a velocity sensor, etc.), and optionally one or more peripheral devices 195 (e.g., a consumer electronics device, a wearable device, etc.). In some embodiments, one or more of input device 125, output device 155, sensor 190, and peripheral device 195 are integrated with display generation component 120 (e.g., within a head-mounted or handheld device).

[0066] When describing an XR experience, various terms are used to individually refer to several related, but distinct, environments that a user senses and / or can interact with (e.g., using inputs detected by computer system 101 that cause the computer system generating the XR experience to generate audio, visual, and / or haptic feedback corresponding to various inputs provided to computer system 101 generating the XR experience). The following is a subset of these terms:

[0067] Physical Environment: The physical environment refers to the physical world that people can sense and / or interact with without the aid of electronic systems. A physical environment, such as a physical park, includes physical objects such as physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment through their senses, such as sight, touch, hearing, taste, and smell.

[0068] Augmented reality: In contrast, an extended reality (XR) environment refers to a wholly or partially mimicked environment that people sense and / or interact with through electronic systems. In XR, a subset of a person's body movements or representations thereof are tracked, and one or more properties of one or more virtual objects simulated within the XR environment are adjusted accordingly to behave according to at least one law of physics. For example, an XR system may detect a person's head rotation and adjust the graphical content and sound field presented to the person accordingly, in a manner similar to how such views and sounds change in a physical environment. In some circumstances (e.g., for accessibility reasons), adjustments to property(ies) of virtual object(s) in the XR environment may be made in response to representations of body movements (e.g., voice commands). A person may sense and / or interact with an XR object using any one of these senses, including sight, hearing, touch, taste, and smell. For example, a person may sense and / or interact with audio objects that create a 3D or spatial audio environment that provides the perception of a point audio source in 3D space. In another example, audio objects may enable audio transparency that selectively incorporates ambient sounds from the physical environment, with or without computer-generated audio. In some XR environments, a person may sense and / or interact with only audio objects.

[0069] Examples of XR include virtual reality and mixed reality.

[0070] Virtual Reality: A virtual reality (VR) environment refers to an emulated environment designed to be based entirely on computer-generated sensory input for one or more senses. A VR environment includes multiple virtual objects that a person can sense and / or interact with. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can sense and / or interact with virtual objects in the VR environment through a simulation of the person's presence in the computer-generated environment and / or through a simulation of a subset of the person's physical movement within the computer-generated environment.

[0071] Mixed Reality: A mixed reality (MR) environment refers to a mimicked environment designed to incorporate sensory input from or representations of a physical environment in addition to including computer-generated sensory input (e.g., virtual objects), as opposed to a VR environment designed to be based entirely on computer-generated sensory input. On a virtuality continuum, a mixed reality environment is anywhere between, but not including, a complete physical environment at one end and a virtual reality environment at the other. In some MR environments, computer-generated sensory input may respond to changes in sensory input from the physical environment. Some electronic systems for presenting MR environments may also track location and / or orientation relative to the physical environment to allow virtual objects to interact with real objects (i.e., physical items or representations thereof from the physical environment). For example, the system may take into account movement so that a virtual tree appears stationary relative to the physical ground.

[0072] Examples of mixed reality include augmented reality and augmented virtuality.

[0073] Augmented reality: An augmented reality (AR) environment refers to a simulated environment in which one or more virtual objects are superimposed on a physical environment or a representation thereof. For example, an electronic system for presenting an AR environment may have a transparent or translucent display through which a person can directly view the physical environment. The system may be configured to present virtual objects on the transparent or translucent display, whereby a person using the system perceives the virtual objects superimposed on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture images or videos of the physical environment, which are representations of the physical environment. The system composites the images or videos with virtual objects and presents the composite on the opaque display. The person uses the system to indirectly view the physical environment through the images or videos of the physical environment and perceive the virtual objects superimposed on the physical environment. As used herein, video of a physical environment shown on an opaque display is referred to as "pass-through video," meaning that the system captures images of the physical environment using one or more image sensors and uses those images in presenting the AR environment on the opaque display. Alternatively, the system may include a projection system that projects virtual objects, e.g., as holograms, into a physical environment or onto a physical surface, such that a person using the system perceives the virtual objects superimposed on the physical environment. Augmented reality environments also refer to mimic environments in which a representation of a physical environment is transformed by computer-generated sensory information. For example, when providing pass-through video, a system may distort one or more sensor images to impose a selected perspective (e.g., viewpoint) different from the perspective captured by the imaging sensor. As another example, a representation of a physical environment may be distorted by graphically modifying (e.g., enlarging) portions thereof, thereby rendering the modified portions a non-photorealistic, altered version of the originally captured image. As a further example, a representation of a physical environment may be distorted by graphically removing or obscuring portions thereof.

[0074] Augmented Virtuality: An augmented virtuality (AV) environment refers to a mimicking environment in which a virtual or computer-generated environment incorporates one or more sensory inputs from a physical environment. The sensory inputs may be representations of one or more characteristics of the physical environment. For example, an AV park may have virtual trees and virtual buildings, while people with faces are realistically recreated from images of physical people. As another example, virtual objects may adopt the shape or color of physical items imaged by one or more imaging sensors. As a further example, virtual objects may adopt shadows that match the position of the sun in the physical environment.

[0075] Perspective-Locked Virtual Object: A virtual object is perspective-locked when the computer system displays the virtual object in the same location and / or position within the user's perspective, even as the user's perspective shifts (e.g., changes). In embodiments in which the computer system is a head-mounted device, the user's perspective is locked to the forward-facing orientation of the user's head (e.g., the user's perspective is at least a portion of the user's field of view when the user is looking straight ahead). Thus, the user's perspective remains fixed even as the user's line of sight moves without moving the user's head. In embodiments in which the computer system has a display generating component (e.g., a display screen) that can be repositioned relative to the user's head, the user's perspective is the augmented reality view being presented to the user on the display generating component of the computer system. For example, a perspective-locked virtual object that is displayed in the upper left corner of the user's perspective when the user's perspective is in a first orientation (e.g., the user's head is facing north) continues to be displayed in the upper left corner of the user's perspective when the user's perspective changes to a second orientation (e.g., the user's head is facing west). In other words, the location and / or position at which a viewpoint-locked virtual object is displayed in a user's viewpoint is independent of the user's position and / or orientation in the physical environment. In embodiments in which the computer system is a head-mounted device, the user's viewpoint is locked to the orientation of the user's head, such that the virtual object is also referred to as a "head-locked virtual object."

[0076] Environment-Locked Virtual Object: A virtual object is environment-locked (or "world-locked") when a computer system displays the virtual object at a location and / or position within a user's viewpoint that is based on (e.g., selected with reference to and / or anchored to) locations and / or objects within a three-dimensional environment (e.g., a physical environment or a virtual environment). As the user's viewpoint shifts, the locations and / or objects within the environment relative to the user's viewpoint change, resulting in the environment-locked virtual object appearing at a different location and / or position within the user's viewpoint. For example, an environment-locked virtual object locked to a tree directly in front of the user will appear centered within the user's viewpoint. If the user's viewpoint shifts to the right (e.g., the user's head is turned to the right) and the tree becomes more left-leaning within the user's viewpoint (e.g., the position of the tree within the user's viewpoint shifts), the environment-locked virtual object locked to the tree will appear more left-leaning within the user's viewpoint. In other words, the location and / or position at which the environment-locked virtual object appears within the user's viewpoint depends on the position and / or orientation of the location and / or object in the environment to which the virtual object is locked. In some embodiments, the computer system uses a stationary reference frame (e.g., a coordinate system fixed to a fixed location and / or object in the physical environment) to determine a position at which to display an environment-locked virtual object in the user's viewpoint. The environment-locked virtual object can be locked to a stationary portion of the environment (e.g., a floor, wall, table, or other stationary object) or can be locked to a moving portion of the environment (e.g., a vehicle, an animal, a person, or a representation of a part of the user's body that moves independent of the user's viewpoint, such as the user's hand, wrist, arm, or leg), so that the virtual object moves as the viewpoint or part of the environment moves in order to maintain a fixed relationship between the virtual object and the part of the environment.

[0077] In some embodiments, an environment-locked or viewpoint-locked virtual object exhibits delayed-following behavior, which reduces or delays the movement of the environment-locked or viewpoint-locked virtual object relative to the movement of a reference point that the virtual object is following. In some embodiments, when exhibiting delayed-following behavior, the computer system intentionally delays the movement of the virtual object when it detects movement of the reference point that the virtual object is following (e.g., a part of the environment, the viewpoint, or a point fixed relative to the viewpoint, such as a point between 5 and 300 cm from the viewpoint). For example, when the reference point (e.g., a part of the environment or the viewpoint) moves at a first speed, the virtual object is moved by the device to remain locked to the reference point, but at a second speed that is slower than the first speed (e.g., until the reference point stops or slows down, at which point the virtual object begins to catch up with the reference point). In some embodiments, when the virtual object exhibits delayed-following behavior, the device ignores small amounts of movement of the reference point (e.g., ignores movement of the reference point that is less than a threshold amount of movement, such as movement between 0 and 5 degrees or movement between 0 and 50 cm). For example, when the reference point (e.g., a portion of the environment or a viewpoint to which the virtual object is locked) moves by a first amount, the distance between the reference point and the virtual object increases (e.g., because the virtual object is displayed to maintain a fixed or substantially fixed position relative to a viewpoint or portion of the environment different from the reference point to which the virtual object is locked), and when the reference point (e.g., a portion of the environment or a viewpoint to which the virtual object is locked) moves by a second amount greater than the first amount, the distance between the reference point and the virtual object initially increases (e.g., because the virtual object is displayed to maintain a fixed or substantially fixed position relative to a viewpoint or portion of the environment different from the reference point to which the virtual object is locked), and then decreases as the amount of movement of the reference point increases beyond a threshold (e.g., a “delayed following” threshold) as the virtual object is moved by the computer system to maintain a fixed or substantially fixed position relative to the reference point.In some embodiments, a virtual object maintaining a substantially fixed position relative to a reference point includes the virtual object being displayed within a threshold distance (e.g., 1, 2, 3, 5, 15, 20, 50 cm) of the reference point in one or more dimensions (e.g., above / below, left / right, and / or forward / backward relative to the position of the reference point).

[0078] Hardware: There are many different types of electronic systems that allow a person to sense and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be placed over a person's eyes (e.g., contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system may include speakers and / or other audio output devices integrated into the head-mounted system to provide audio output. A head-mounted system may have one or more speaker(s) and an integrated opaque display. Alternatively, a head-mounted system may be configured to accept an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing images or video of the physical environment and / or one or more microphones for capturing audio of the physical environment. The head-mounted system may have a transparent or translucent display rather than an opaque display. The transparent or translucent display may have a medium through which light representing an image is directed to a person's eyes. The display may utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser-scanned light source, or any combination of these technologies. The medium may be an optical waveguide, a holographic medium, an optical combiner, an optical reflector, or any combination thereof. In one embodiment, the transparent or translucent display may be configured to be selectively opaque. The projection-based system may employ retinal projection technology that projects a graphical image onto a person's retina.The projection system may also be configured to project virtual objects into the physical environment, for example, as holograms or onto physical surfaces. In some embodiments, controller 110 is configured to manage and coordinate the XR experience for the user. In some embodiments, controller 110 includes a suitable combination of software, firmware, and / or hardware. Controller 110 is described in more detail below with reference to FIG. 2. In some embodiments, controller 110 is a computing device that is local or remote to scene 105 (e.g., the physical environment). For example, controller 110 is a local server located within scene 105. In another example, controller 110 is a remote server (e.g., a cloud server, a central server, etc.) located outside scene 105. In some embodiments, controller 110 is communicatively coupled to display generation component 120 (e.g., an HMD, a display, a projector, a touchscreen, etc.) via one or more wired or wireless communication channels 144 (e.g., BLUETOOTH, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). In another example, the controller 110 is contained within the housing (e.g., physical housing) of one or more of the display generating component 120 (e.g., an HMD or a portable electronic device including a display and one or more processors), one or more of the input devices 125, one or more of the output devices 155, one or more of the sensors 190, and / or one or more of the peripheral devices 195, or shares the same physical housing or support structure as one or more of the foregoing.

[0079] In some embodiments, display generation component 120 is configured to provide an XR experience (e.g., at least a visual component of an XR experience) to a user. In some embodiments, display generation component 120 includes a suitable combination of software, firmware, and / or hardware. Display generation component 120 is described in more detail below with reference to FIG. 3. In some embodiments, the functionality of controller 110 is provided by and / or combined with display generation component 120.

[0080] According to some embodiments, the display generation component 120 provides an XR experience to the user while the user is virtually and / or physically present in the scene 105.

[0081] In some embodiments, the display generating component is worn on a part of the user's body (e.g., on their head, their hand, etc.). Thus, display generating component 120 includes one or more XR displays provided for displaying XR content. For example, in various embodiments, display generating component 120 surrounds the user's field of view. In some embodiments, display generating component 120 is a handheld device (e.g., a smartphone or tablet) configured to present XR content, where the user holds the device with a display pointed toward the user's field of view and a camera pointed toward scene 105. In some embodiments, the handheld device is optionally located within a housing worn on the user's head. In some embodiments, the handheld device is optionally located on a support (e.g., a tripod) in front of the user. In some embodiments, display generating component 120 is an XR chamber, housing, or room configured to present XR content without the user wearing or holding display generating component 120. Many user interfaces described with reference to one type of hardware for displaying XR content (e.g., a handheld device or a device on a tripod) may be implemented on another type of hardware for displaying XR content (e.g., an HMD or other wearable computing device). For example, a user interface illustrating interactions with XR content that are triggered based on interactions occurring in the space in front of a handheld or tripod-mounted device may be implemented similarly to an HMD in which the interactions occur in the space in front of the HMD and the XR content responses are displayed via the HMD. Similarly, a user interface illustrating interactions with XR content that are triggered based on movement of a handheld or tripod-mounted device relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eye(s), head, or hands)) may be implemented similarly to an HMD in which the movement is caused by movement of the HMD relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eye(s), head, or hands)).

[0082] While relevant features of operating environment 100 are shown in FIG. 1, those skilled in the art will understand from this disclosure that for the sake of brevity, various other features are not shown so as to not obscure more pertinent aspects of the exemplary embodiments disclosed herein.

[0083] 2 is a block diagram of an example controller 110 according to some embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that various other features are not shown for the sake of brevity so as not to obscure more pertinent aspects of the embodiments disclosed herein. Thus, by way of non-limiting example, in some embodiments, the controller 110 includes one or more processing units 202 (e.g., a microprocessor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a graphics processing unit (GPU), a central processing unit (CPU), a processing core, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), BLUETOOTH, ZIGBEE, or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these and various other components.

[0084] In some embodiments, one or more communication buses 204 include circuitry that interconnects and controls communications between system components. In some embodiments, one or more I / O devices 206 include at least one of a keyboard, a mouse, a touchpad, a joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.

[0085] Memory 220 includes high-speed random-access memory, such as dynamic random-access memory (DRAM), static random-access memory (SRAM), double-data-rate random-access memory (DDRRAM), or other random-access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory, such as one or more magnetic storage devices, optical storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 220 optionally includes one or more storage devices located remotely from the one or more processing units 202. Memory 220 includes a non-transitory computer-readable storage medium. In some embodiments, memory 220, or the non-transitory computer-readable storage medium of memory 220, stores the following programs, modules, and data structures, or a subset thereof, including an optional operating system 230 and an XR experience module 240:

[0086] Operating system 230 includes instructions for handling various basic system services and performing hardware-dependent tasks. In some embodiments, XR experience module 240 is configured to manage and coordinate one or more XR experiences for one or more users (e.g., a single XR experience for one or more users, or multiple XR experiences for respective groups of one or more users). To that end, in various embodiments, XR experience module 240 includes a data acquisition unit 241, a tracking unit 242, an adjustment unit 246, and a data transmission unit 248.

[0087] 1 , and optionally one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. To that end, in various embodiments, data acquisition unit 241 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0088] In some embodiments, tracking unit 242 is configured to map scene 105 and track the position / location of at least display generating component 120 relative to scene 105 of FIG. 1 , and optionally relative to one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. To that end, in various embodiments, tracking unit 242 includes instructions and / or logic therefor, as well as heuristics and metadata therefor. In some embodiments, tracking unit 242 includes hand tracking unit 244 and / or eye tracking unit 243. In some embodiments, hand tracking unit 244 is configured to track the position / location of one or more parts of a user's hand and / or the movement of one or more parts of a user's hand relative to scene 105 of FIG. 1 , relative to display generating component 120, and / or relative to a coordinate system defined relative to the user's hand. Hand tracking unit 244 is described in more detail below with respect to FIG. 4. In some embodiments, eye tracking unit 243 is configured to track the position and movement of the user's gaze (or, more broadly, the user's eyes, face, or head) relative to scene 105 (e.g., relative to the physical environment and / or the user (e.g., the user's hands)), or relative to XR content displayed via display generation component 120. Eye tracking unit 243 is described in more detail below with respect to FIG. 5.

[0089] In some embodiments, coordination unit 246 is configured to manage and coordinate the XR experience presented to the user by display generation component 120 and, optionally, by one or more of output devices 155 and / or peripheral devices 195. To that end, in various embodiments, coordination unit 246 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0090] In some embodiments, data transmission unit 248 is configured to transmit data (e.g., presentation data, location data, etc.) to at least display generation component 120, and optionally to one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. To that end, in various embodiments, data transmission unit 248 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0091] Although the data acquisition unit 241, the tracking unit 242 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the adjustment unit 246, and the data transmission unit 248 are shown as being present on a single device (e.g., the controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 241, the tracking unit 242 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the adjustment unit 246, and the data transmission unit 248 can be located in separate computing devices.

[0092] Furthermore, Figure 2 is intended more to illustrate the functionality of various features that may be present in particular embodiments, as opposed to a structural overview of the embodiments described herein. As will be recognized by those skilled in the art, items shown separately can be combined and some items can be separated. For example, some functional modules shown separately in Figure 2 can be implemented in a single module, and various functions of a single functional block can be implemented by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of specific functions and how functions are allocated among them, will vary depending on implementation and, in some embodiments, will depend in part on the particular combination of hardware, software, and / or firmware selected for a particular implementation.

[0093] 3 is a block diagram of an example of a display generation component 120 according to some embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that for the sake of brevity, various other features are not shown so as to not obscure more pertinent aspects of the embodiments disclosed herein. To that end, by way of non-limiting example, in some embodiments, the display generation component 120 (e.g., an HMD) includes one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, infrared, BLUETOOTH, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more XR displays 312, one or more optional inward-facing and / or outward-facing image sensors 314, memory 320, and one or more communication buses 304 for interconnecting these and various other components.

[0094] In some embodiments, the one or more communication buses 304 include circuitry that interconnects and controls communications between system components. In some embodiments, the one or more I / O devices and sensors 306 include at least one of an inertial measurement unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, one or more depth sensors (e.g., structured light, time of flight, etc.), etc.

[0095] In some embodiments, the one or more XR displays 312 are configured to provide an XR experience to a user. In some embodiments, the one or more XR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conduction electron-emissive element display (SED), field-emission display (FED), quantum dot light-emitting diode (QD-LED), MEMS, and / or similar display types. In some embodiments, the one or more XR displays 312 correspond to waveguide displays, such as diffractive, reflective, polarized, holographic, etc. For example, the display generation component 120 (e.g., an HMD) includes a single XR display. In another example, the display generation component 120 includes an XR display for each eye of the user. In some embodiments, the one or more XR displays 312 are capable of presenting mixed reality (MR) or virtual reality (VR) content. In some embodiments, the one or more XR displays 312 are capable of presenting mixed reality (MR) or virtual reality (VR) content.

[0096] In some embodiments, the one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face, including the user's eyes (and may be referred to as eye-tracking cameras). In some embodiments, the one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hand(s) and optionally the user's arm(s) (and may be referred to as hand-tracking cameras). In some embodiments, the one or more image sensors 314 are configured to face forward to acquire image data corresponding to a scene as the user would view it if the display generating component 120 (e.g., an HMD) were not present (and may be referred to as a scene camera). The one or more optional image sensors 314 may include one or more RGB cameras (e.g., with a complementary metal-oxide semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), one or more infrared (IR) cameras, one or more event-based cameras, and / or the like.

[0097] Memory 320 includes high-speed random-access memory, such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located remotely from the one or more processing units 302. Memory 320 includes a non-transitory computer-readable storage medium. In some embodiments, memory 320, or its non-transitory computer-readable storage medium, stores the following programs, modules, and data structures, or a subset thereof, including an optional operating system 330 and an XR presentation module 340:

[0098] The operating system 330 includes instructions for handling various basic system services and for performing hardware-dependent tasks. In some embodiments, the XR presentation module 340 is configured to present XR content to a user via one or more XR displays 312. To that end, in various embodiments, the XR presentation module 340 includes a data acquisition unit 342, an XR presentation unit 344, an XR map generation unit 346, and a data transmission unit 348.

[0099] In some embodiments, the data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the controller 110 of Figure 1. To that end, in various embodiments, the data acquisition unit 342 includes its instructions and / or logic, as well as heuristics and metadata therefor.

[0100] In some embodiments, the XR presentation unit 344 is configured to present XR content via one or more XR displays 312. To that end, in various embodiments, the XR presentation unit 344 includes its instructions and / or logic, as well as heuristics and metadata therefor.

[0101] In some embodiments, the XR map generation unit 346 is configured to generate an XR map (e.g., a 3D map of a mixed reality scene or a map of a physical environment in which computer-generated objects can be placed to generate an augmented reality) based on the media content data. To that end, in various embodiments, the XR map generation unit 346 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0102] In some embodiments, data transmission unit 348 is configured to transmit data (e.g., presentation data, location data, etc.) to at least controller 110, and optionally to one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. To that end, in various embodiments, data transmission unit 348 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0103] Although the data acquisition unit 342, the XR presentation unit 344, the XR map generation unit 346, and the data transmission unit 348 are shown as residing on a single device (e.g., the display generation component 120 of FIG. 1), it should be understood that in other embodiments, any combination of the data acquisition unit 342, the XR presentation unit 344, the XR map generation unit 346, and the data transmission unit 348 may be located in separate computing devices.

[0104] Furthermore, Figure 3 is intended more to illustrate the functionality of various features that may be present in particular implementations, as opposed to a structural overview of the embodiments described herein. As will be recognized by those skilled in the art, items shown separately can be combined and some items can be separated. For example, some functional modules shown separately in Figure 3 can be implemented within a single module, and various functions of a single functional block can be performed by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of specific functions and how functions are allocated among them, will vary from implementation to implementation and, in some embodiments, will depend in part on the particular combination of hardware, software, and / or firmware selected for a particular implementation.

[0105] 4 is a schematic diagram of an example embodiment of a hand tracking device 140. In some embodiments, hand tracking device 140 (FIG. 1) is controlled by hand tracking unit 244 (FIG. 2) to track the location / position of one or more parts of a user's hand and / or the movement of one or more parts of a user's hand relative to scene 105 of FIG. 1 (e.g., relative to a portion of the physical environment surrounding the user, relative to display generating components 120, or relative to a portion of the user (e.g., the user's face, eyes, or head), and / or relative to the user's hand). In some embodiments, hand tracking device 140 is part of display generating components 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, hand tracking device 140 is separate from display generating components 120 (e.g., located in a separate housing or attached to a separate physical support structure).

[0106] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras) that captures three-dimensional scene information including at least the hand 406 of a human user. The image sensor 404 captures hand images with sufficient resolution to allow for differentiation of the fingers and their respective positions. The image sensor 404 typically captures images of other parts of the user's body, or all of the body, and can have either zoom capabilities or a dedicated sensor with high magnification to capture hand images at a desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with or functions as an image sensor that captures the physical environment of the scene 105. In some embodiments, the image sensor 404 is positioned relative to the user or the user's environment such that the field of view of the image sensor, or a portion thereof, is used to define an interaction space in which hand movements captured by the image sensor are processed as inputs to the controller 110.

[0107] In some embodiments, image sensor 404 outputs a sequence of frames containing 3D map data (and possibly color image data) to controller 110, which extracts high-level information from the map data. This high-level information is provided, typically via an application program interface (API), to an application running on the controller, which drives display generation component 120 accordingly. For example, a user can interact with software running on controller 110 by moving their hand 406 and changing their hand posture.

[0108] In some embodiments, the image sensor 404 projects a spot pattern onto a scene including the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral shift of the pattern's spots. This approach is advantageous in that it does not require the user to hold or wear any type of beacon, sensor, or other marker. This provides depth coordinates of points in the scene relative to a predetermined reference plane at a specific distance from the image sensor 404. In this disclosure, the image sensor 404 is assumed to define a set of orthogonal x, y, and z axes such that the depth coordinate of a point in the scene corresponds to the z component measured by the image sensor. Alternatively, the image sensor 404 (e.g., a hand tracking device) can use other 3D mapping methods, such as stereoscopic imaging or time-of-flight measurement, based on single or multiple cameras or other types of sensors.

[0109] In some embodiments, the hand tracking device 140 captures and processes a time sequence of depth maps containing the user's hand while the user moves the hand (e.g., the entire hand or one or more fingers). Software running on the image sensor 404 and / or a processor in the controller 110 processes the 3D map data to extract patch descriptors of the hand in these depth maps. The software matches these descriptors with patch descriptors stored in the database 408, based on a previous learning process, to estimate the pose of the hand in each frame. The pose typically includes the 3D locations of the user's wrist joints and fingertips.

[0110] The software can also analyze hand and / or finger trajectories across multiple frames in a sequence to identify gestures. The pose estimation functionality described herein may be interleaved with motion tracking functionality, whereby patch-based pose estimation is performed only once every two (or more) frames, while tracking is used to discover pose changes that occur across the remaining frames. The pose, motion, and gesture information is provided to an application program running on controller 110 via the API described above. This program can, for example, move and modify an image presented on display generation component 120 or perform other functions in response to the pose and / or gesture information.

[0111] In some embodiments, the gesture includes an air gesture, which is detected without (or independent of) the user touching an input element that is part of a device (e.g., computer system 101, one or more input devices 125, and / or hand tracking device 140) and is based on detected movement of a part of the user's body in the air (e.g., head, one or more arms, one or more hands, one or more fingers, and / or one or more legs), including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement of the user's body relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to another of the user's hands, and / or movement of a user's finger relative to another finger or part of the user's hand), and / or absolute movement of the user's body part (e.g., a tap gesture involving movement of a hand in a predetermined posture by a predetermined amount and / or speed, or a shake gesture involving a predetermined speed or amount of rotation of the user's body part).

[0112] In some embodiments, input gestures used in various examples and embodiments described herein include air gestures performed by movement of a user's finger(s) relative to other finger(s) or part(s) of the user's hand to interact with an XR environment (e.g., a virtual or mixed reality environment), according to some embodiments. In some embodiments, an air gesture is a gesture that is detected without the user touching an input element that is part of the device (or independent of an input element that is part of the device) and is based on detected movement of a part of the user's body, including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground, or the distance of the user's hand relative to the ground), movement of the user's body relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of the user's other hand relative to one of the user's hands, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tap gesture that includes movement of the hand in a predetermined pose by a predetermined amount and / or speed, or a shake gesture that includes rotation of a part of the user's body at a predetermined speed or amount).

[0113] In some embodiments where the input gesture is an air gesture (e.g., in the absence of physical contact with an input device that provides a computer system with information about which user interface element is the target of the user input, such as contact with a user interface element displayed on a touchscreen or contact with a mouse or trackpad to move a cursor to a user interface element), the gesture takes into account the user's attention (e.g., gaze) to determine the target of the user input (e.g., in the case of direct input, as described below). Thus, in implementations that include air gestures, the input gesture is detected attention (e.g., gaze) to a user interface element in combination with (e.g., simultaneous with) movement of the user's finger(s) and / or hand to perform pinch and / or tap input, as described in more detail below.

[0114] In some embodiments, an input gesture directed at a user interface object is performed directly or indirectly with reference to the user interface object. For example, user input is performed directly at a user interface object when the user performs an input gesture with their hand at a position corresponding to the position of the user interface object in the three-dimensional environment (e.g., as determined based on the user's current viewpoint). In some embodiments, an input gesture is performed indirectly at a user interface object according to the user performing the input gesture while the position of the user's hand is not at a position corresponding to the position of the user interface object in the three-dimensional environment while detecting the user's attention (e.g., gaze) to the user interface object. For example, for a direct input gesture, the user can direct the user's input at a user interface object by initiating the gesture at or near a position corresponding to the displayed position of the user interface object (e.g., within a distance of 0.5 cm, 1 cm, 5 cm, or 0-5 cm, measured from an outer edge of the option or a central portion of the option). For indirect input gestures, a user can direct their input to a user interface object by paying attention to the user interface object (e.g., by gazing at the user interface object), and while paying attention to the option, the user initiates an input gesture (e.g., at any position detectable by the computer system) (e.g., at a position that does not correspond to the displayed position of the user interface object).

[0115] In some embodiments, input gestures (e.g., air gestures) used in various examples and embodiments described herein include pinch inputs and tap inputs for interacting with a virtual or mixed reality environment, in some embodiments. For example, pinch inputs and tap inputs described below are performed as air gestures.

[0116] In some embodiments, the pinch input is part of an air gesture, including one or more of a pinch gesture, a long pinch gesture, a pinch-and-drag gesture, or a double pinch gesture. For example, a pinch gesture that is an air gesture includes moving two or more fingers of a hand to contact each other, i.e., optionally with a short break (e.g., within 0-1 second) after contact with each other. A long pinch gesture that is an air gesture includes moving two or more fingers of a hand to contact each other for at least a threshold amount of time (e.g., at least 1 second) before detecting a break in contact with each other. For example, a long pinch gesture includes a user holding a pinch gesture (e.g., when two or more fingers are in contact), and the long pinch gesture continues until a break in contact between the two or more fingers is detected. In some embodiments, a double pinch gesture that is an air gesture includes two (e.g., or more) pinch inputs (e.g., performed by the same hand) that are detected immediately in succession (e.g., within a predetermined period of time) after each other. For example, a user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., breaking contact between two or more fingers), and performs a second pinch input within a predetermined period of time (e.g., within 1 second or 2 seconds) after releasing the first pinch input.

[0117] In some embodiments, a pinch-and-drag gesture that is an air gesture includes a pinch gesture (e.g., a pinch gesture or a long pinch gesture) performed in conjunction with (e.g., followed by) a drag input that changes the position of a user's hand from a first position (e.g., a start position of the drag) to a second position (e.g., an end position of the drag). In some embodiments, a user maintains the pinch gesture while performing the drag input and releases the pinch gesture (e.g., spreading two or more fingers apart) to end the drag gesture (e.g., at the second position). In some embodiments, the pinch input and the drag input are performed by the same hand (e.g., a user pinches two or more fingers together and moves the same hand to a second position in the air with a drag gesture). In some embodiments, the pinch input is performed by a user's first hand and the drag input is performed by the user's second hand (e.g., the user's second hand moves from a first position to a second position in the air while the user continues the pinch input with the user's first hand). In some embodiments, an input gesture that is an air gesture includes an input (e.g., a pinch input and / or a tap input) performed using both of a user's hands. For example, the input gesture includes two (e.g., or more) pinch inputs performed in conjunction with each other (e.g., simultaneously or within a predetermined period of time). For example, a first pinch gesture (e.g., a pinch input, a long pinch input, or a pinch and drag input) performed using a first hand of the user and a second pinch input performed using the other hand (e.g., a second of the user's hands) in conjunction with performing the pinch input using the first hand. In some embodiments, a movement between a user's hands (e.g., to increase and / or decrease the distance or relative orientation between the user's hands).

[0118] In some embodiments, a tap input (e.g., directed toward a user interface element) performed as an air gesture includes movement(s) of a user's finger(s) toward the user interface element, movement of a user's hand toward a user interface element, optionally with the user's finger(s) extended toward the user interface element, a downward movement of a user's finger (e.g., mimicking a mouse click action or a tap on a touchscreen), or other predefined movement of the user's hand. In some embodiments, a tap input performed as an air gesture is detected based on movement characteristics of the finger or hand performing the tap gesture, moving the finger or hand away from the user's viewpoint and / or toward the object that is the target of the tap input followed by an end of the movement. In some embodiments, an end of the movement is detected based on a change in movement characteristics of the finger or hand performing the tap gesture (e.g., an end of movement away from the user's viewpoint and / or toward the object that is the target of the tap input, a reversal of the direction of movement of the finger or hand, and / or a reversal of the direction of acceleration of the movement of the finger or hand).

[0119] In some embodiments, the user's attention is determined to be directed to a portion of the three-dimensional environment based on detecting a gaze directed to the portion of the three-dimensional environment (optionally, without requiring other conditions). In some embodiments, the device determines that the user's attention is directed to the portion of the three-dimensional environment based on detecting a gaze directed to the portion of the three-dimensional environment with one or more additional conditions, such as requiring the gaze to be directed to the portion of the three-dimensional environment for at least a threshold duration (e.g., dwell time) while the user's viewpoint is within a distance threshold from the portion of the three-dimensional environment, and / or requiring the gaze to be directed to the portion of the three-dimensional environment, and if one of the additional conditions is not met, the device determines that the user's attention is not directed to the portion of the three-dimensional environment to which the gaze is directed (e.g., until one or more additional conditions are met).

[0120] In some embodiments, detection of a ready configuration of a user or a portion of a user is detected by a computer system, and detection of a ready configuration of the hands is used by the computer system as an indication that the user is likely preparing to interact with the computer system using one or more air gesture inputs performed with the hands (e.g., pinch, tap, pinch and drag, double pinch, long pinch, or other air gestures described herein). For example, the ready state of a hand is determined based on whether the hand has a predetermined hand geometry (e.g., a pre-pinch geometry with the thumb and one or more fingers extended and spaced apart, ready to perform a pinch or grab gesture, or a pre-tap geometry with one or more fingers extended and the palm facing away from the user), whether the hand is in a predetermined position relative to the user's viewpoint (e.g., below the user's head, above the user's waist, extended at least 15 cm, 20 cm, 25 cm, 30 cm, or 50 cm from the body), and / or whether the hand has moved in a particular manner (e.g., above the user's waist, moved toward an area in front of the user below the user's head, or away from the user's body or legs). In some embodiments, the ready state is used to determine whether an interactive element of a user interface is responsive to attentional (e.g., gaze) input.

[0121] In some embodiments, the software may be downloaded to the controller 110 in electronic form, for example, over a network, or alternatively may be provided on a tangible, non-transitory medium, such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is similarly stored in memory associated with the controller 110. Alternatively, or additionally, some or all of the described functionality of the computer may be implemented in dedicated hardware, such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP). While the controller 110 is shown in FIG. 4 as, by way of example, a separate unit from the image sensor 404, some or all of the processing functionality of the controller may be implemented by a suitable microprocessor and software, or by dedicated circuitry within the housing of the image sensor 404 (e.g., a hand tracking device), or otherwise associated with the image sensor 404. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with the display generation component 120 (e.g., in a television set, handheld device, or head-mounted device) or using any other suitable computerized device, such as a game console or media player. The sensing function of the image sensor 404 may likewise be integrated into a computer or other computerized device that is controlled by the sensor output.

[0122] FIG. 4 also includes a schematic diagram of a depth map 410 captured by the image sensor 404 in some embodiments. The depth map, as described above, includes a matrix of pixels having respective depth values. A pixel 412 corresponding to the hand 406 is segmented from the background and wrist in this map. The intensity of each pixel in the depth map 410 is inversely proportional to the depth value, i.e., the measured z-distance from the image sensor 404, with increasing intensity as depth increases. The controller 110 processes these depth values ​​to identify and segment components of the image (i.e., groups of adjacent pixels) that have characteristics of a human hand. These characteristics can include, for example, the overall size, shape, and frame-to-frame motion of the depth map sequence.

[0123] 4 also schematically illustrates a hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406 in some embodiments. In FIG. 4, the hand skeleton 414 is overlaid on a hand background 416 that was segmented from the original depth map. In some embodiments, key feature points on the hand (e.g., knuckles, fingertips, center of the palm, end of the hand where it connects to the wrist, etc.), and optionally the wrist or arm connected to the hand, are identified and positioned on the hand skeleton 414. In some embodiments, the location and movement of these key feature points over multiple image frames are used by the controller 110 to determine, in some embodiments, hand gestures performed by the hand or the current state of the hand.

[0124] FIG. 5 shows an exemplary embodiment of eye tracking device 130 ( FIG. 1 ). In some embodiments, eye tracking device 130 is controlled by eye tracking unit 243 ( FIG. 2 ) to track the position and movement of a user's gaze relative to scene 105 or relative to XR content displayed via display generation component 120. In some embodiments, eye tracking device 130 is integrated with display generation component 120. For example, in some embodiments, if display generation component 120 is a head-mounted device such as a headset, helmet, goggles, or glasses, or a handheld device disposed in a wearable frame, the head-mounted device includes both components for generating XR content for viewing by the user and components for tracking the user's gaze relative to the XR content. In some embodiments, eye tracking device 130 is separate from display generation component 120. For example, if the display generation component is a handheld device or an XR chamber, eye tracking device 130 is optionally a device separate from the handheld device or the XR chamber. In some embodiments, eye tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, head-mounted eye tracking device 130 is optionally used in conjunction with head-mounted or non-head-mounted display generating components. In some embodiments, eye tracking device 130 is not a head-mounted device, and is optionally used in combination with head-mounted display generating components. In some embodiments, eye tracking device 130 is not a head-mounted device, and is optionally part of non-head-mounted display generating components.

[0125] In some embodiments, the display generation component 120 uses a display mechanism (e.g., left and right near-eye display panels) that displays frames including left and right images in front of the user's eyes to provide the user with a 3D virtual view. For example, the head-mounted display generation component may include left and right optical lenses (referred to herein as eyepieces) positioned between the display and the user's eyes. In some embodiments, the display generation component may include or be coupled to one or more external video cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display that allows the user to view the physical environment directly and display virtual objects on the transparent or translucent display. In some embodiments, the display generation component projects virtual objects into the physical environment. The virtual objects are projected, for example, onto a physical surface or as a hologram, allowing an individual using the system to observe the virtual objects superimposed on the physical environment. In such cases, separate display panels and image frames for the left and right eyes may not be required.

[0126] As shown in FIG. 5 , in some embodiments, eye tracking device 130 (e.g., gaze tracking device) includes at least one eye tracking camera (e.g., an infrared (IR) camera or near-IR (NIR) camera) and an illumination source (e.g., an IR or NIR light source such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eyes. The eye tracking camera may be aimed at the user's eyes to receive reflected IR or NIR light from the light source directly from the eyes, or alternatively, may be aimed at a “hot” mirror positioned between the user's eyes and a display panel that reflects IR or NIR light from the eyes to the eye tracking camera while allowing visual light to pass through. Eye tracking device 130 optionally captures images of the user's eyes (e.g., as a video stream captured at 60-120 frames per second (fps)), analyzes the images to generate eye tracking information, and communicates the eye tracking information to controller 110. In some embodiments, the user's eyes are tracked separately by their respective eye tracking cameras and illumination sources. In some embodiments, only one eye of the user is tracked by a separate eye-tracking camera and lighting source.

[0127] In some embodiments, the eye tracking device 130 is calibrated using a device-specific calibration process to determine the eye tracking device's parameters for the particular operating environment 100, such as the 3D geometric relationships and parameters of the LEDs, camera, hot mirror (if present), eyepiece, and display screen. The device-specific calibration process may be performed at a factory or another facility before delivery of the AR / VR equipment to the end user. The device-specific calibration process may be an automatic or manual calibration process. The user-specific calibration process may include estimation of a particular user's eye parameters, such as pupil location, central visual location, optical axis, visual axis, eye spacing, etc. In some embodiments, once the device-specific and user-specific parameters for the eye tracking device 130 are determined, images captured by the eye tracking camera can be processed using glint-assisted methods to determine the user's current visual axis and viewpoint relative to the display.

[0128] As shown in FIG. 5, eye tracking device 130 (e.g., 130A or 130B) includes an eyepiece(s) 520 and a gaze tracking system including at least one eye tracking camera 540 (e.g., an infrared (IR) or near-IR (NIR) camera) positioned on the side of the user's face where eye tracking occurs and an illumination source 530 (e.g., an IR or NIR light source such as an array or ring of NIR light emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eye(s) 592. The eye tracking camera 540 may be positioned between the user's eye(s) 592 and the display 510 (e.g., the left or right display panel of a head-mounted display, or the display of a handheld device, a projector, etc.) and may be directed at a mirror 550 that reflects IR or NIR light from the eye(s) 592 while transmitting visible light (e.g., as shown at the top of FIG. 5), or may be directed at the user's eye(s) 592 to receive reflected IR or NIR light from the eye(s) 592 (e.g., as shown at the bottom of FIG. 5).

[0129] In some embodiments, controller 110 renders AR or VR frames 562 (e.g., left and right frames for left and right display panels) and provides frames 562 to display 510. Controller 110 uses gaze tracking input 542 from eye tracking camera 540 for various purposes, such as in processing frames 562 for display. Controller 110 optionally estimates the user's viewpoint on display 510 based on gaze tracking input 542 obtained from eye tracking camera 540, using a glint-assisted method or other suitable method. The viewpoint estimated from gaze tracking input 542 is optionally used to determine the direction the user is currently looking.

[0130] Some possible use cases of the user's current gaze direction are described below, but are not intended to be limiting. As an exemplary use case, the controller 110 can render virtual content differently based on the determined user's gaze direction. For example, the controller 110 may generate virtual content with higher resolution in a central visual area determined from the user's current gaze direction than in a peripheral area. As another example, the controller may position or move virtual content within a view based at least in part on the user's current gaze direction. As another example, the controller may display particular virtual content within a view based at least in part on the user's current gaze direction. As another exemplary use case in an AR application, the controller 110 can orient an external camera to capture the physical environment of the XR experience and focus in the determined direction. The external camera's autofocus mechanism can then focus on an object or surface within the environment the user is currently viewing on the display 510. As another exemplary use case, eyepiece 520 may be a focusable lens, and eye-tracking information is used by the controller to adjust the focus of eyepiece 520 so that the virtual object the user is currently looking at has the proper binocular coordination to match the convergence of the user's eyes 592. Controller 110 can utilize the eye-tracking information to orient and focus eyepiece 520 so that close objects the user is looking at appear at the correct distance.

[0131] In some embodiments, the eye tracking device is part of a head-mounted device that includes a display (e.g., display 510), two eyepieces (e.g., eyepiece(s) 520), an eye tracking camera (e.g., eye tracking camera(s) 540), and a light source (e.g., light source 530 (e.g., IR or NIR LED)) attached to the wearable housing. The light source emits light (e.g., IR or NIR light) toward the user's eye(s) 592. In some embodiments, the light sources may be arranged in a ring or circle around each lens, as shown in FIG. 5. In some embodiments, eight light sources 530 (e.g., LEDs) are arranged around each lens 520, as an example. However, more or fewer light sources 530 may be used, and other arrangements and locations of the light sources 530 may be used.

[0132] In some embodiments, the display 510 emits light in the visible light range and not in the IR or NIR range, and therefore does not introduce noise into the gaze tracking system. Note that the location and angle of the eye tracking camera(s) 540 are given by way of example and are not intended to be limiting. In some embodiments, a single eye tracking camera 540 is located on each side of the user's face. In some embodiments, two or more NIR cameras 540 may be used on each side of the user's face. In some embodiments, a camera 540 with a wider field of view (FOV) and a camera 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, a camera 540 operating at one wavelength (e.g., 850 nm) and a camera 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.

[0133] Embodiments of an eye tracking system such as that shown in FIG. 5 may be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide a user with a computer-generated reality, virtual reality, augmented reality, and / or augmented virtual experience.

[0134] FIG. 6 illustrates a glint-assisted gaze tracking pipeline according to some embodiments. In some embodiments, the gaze tracking pipeline is implemented by a glint-assisted gaze tracking system (e.g., eye tracking device 130 as shown in FIGS. 1 and 5). The glint-assisted gaze tracking system can maintain a tracking state. Initially, the tracking state is off or "no." When in the tracking state, the glint-assisted gaze tracking system tracks the pupil contour and glint in the current frame using prior information from the previous frame when analyzing the current frame. When not in the tracking state, the glint-assisted gaze tracking system attempts to detect the pupil and glint in the current frame, and if successful, initializes the tracking state to "yes" and continues to the next frame in the tracking state.

[0135] As shown in FIG. 6, an eye-tracking camera can capture left and right images of a user's left and right eyes. The captured images are then input into an eye-tracking pipeline for processing beginning at 610. As indicated by the arrow returning to element 600, the eye-tracking system can continue to capture images of the user's eyes at a rate of, for example, 60-120 frames per second. In some embodiments, each set of captured images may be input into the pipeline for processing. However, in some embodiments, or under some conditions, not all captured frames are processed by the pipeline.

[0136] At 610, if the tracking status is yes for the currently captured image, the method proceeds to element 640. If the tracking status is no at 610, the image is analyzed to detect the user's pupil and glint in the image, as shown at 620. If the pupil and glint are successfully detected at 630, the method proceeds to element 640. If not, the method returns to element 610 to process the next image of the user's eyes.

[0137] At 640, proceeding from element 610, the current frame is analyzed to track pupils and glints based in part on previous information from the previous frame. At 640, proceeding from element 630, a tracking state is initialized based on the detected pupils and glints in the current frame. The results of the processing at element 640 are checked to ensure that the tracking or detection results are reliable. For example, the results can be checked to determine whether a sufficient number of glints are successfully tracked or detected in the current frame to perform pupil and gaze estimation. At 650, if the results are not reliable, the tracking state is set to no at element 660 and the method returns to element 610 to process the next image of the user's eyes. At 650, if the results are reliable, the method proceeds to element 670. At 670, the tracking state is set to yes (if not already yes) and the pupil and glint information is passed to element 680 to estimate the user's gaze point.

[0138] 6 is intended to serve as an example of eye-tracking technology that may be used in particular implementations. As will be recognized by those skilled in the art, other eye-tracking technologies, now existing or developed in the future, may be used in place of or in combination with the glint-assisted eye-tracking technology described herein in computer system 101 to provide an XR experience to a user in some embodiments.

[0139] In this disclosure, various input methods are described with respect to interaction with a computer system. Where one example is provided using one input device or input method and another example is provided using a different input device or input method, it should be understood that each example may be compatible with, and optionally utilize, the input device or input method described with respect to the other example. Similarly, various output methods are described with respect to interaction with a computer system. Where one example is provided using one output device or output method and another example is provided using a different output device or output method, it should be understood that each example may be compatible with, and optionally utilize, the output device or output method described with respect to the other example. Similarly, various methods are described with respect to interaction with a virtual environment or a mixed reality environment via a computer system. Where one example is provided using interaction with a virtual environment and another example is provided using a mixed reality environment, it should be understood that each example may be compatible with, and optionally utilize, the method described with respect to the other example. Thus, this disclosure discloses embodiments that are combinations of features of multiple examples, without exhaustively listing all features of the embodiments in the description of each exemplary embodiment. User Interface and Related Processes

[0140] Attention is now directed to embodiments of user interfaces ("UIs") and associated processes that may be implemented on a computer system, such as a portable multifunction device or a head-mounted device, in communication with one or more display generating components.

[0141] Figures 7A-7J show examples of generating a user's representation and displaying a user's representation. Figure 8 is a flow diagram of an example method 800 for generating a user's representation. Figure 9 is a flow diagram of an example method 900 for displaying a user's representation. The user interfaces of Figures 7A-7J are used to illustrate processes described below, including the processes of Figures 8 and 9.

[0142] 7A-7J show examples for capturing information used to generate a user's expression. In some embodiments, the user's expression is displayed and / or otherwise used to communicate during a real-time communication session. In some embodiments, the real-time communication session includes real-time communication between a user of an electronic device and a second user associated with a second electronic device different from the first electronic device, and the real-time communication session includes displaying and / or otherwise communicating, via the electronic device and / or the second electronic device, a facial and / or body expression representation of the user to the second user via the user's expression. In some embodiments, the real-time communication session includes displaying the user's expression and / or outputting audio corresponding to the user's speech in real time. In some embodiments, the first electronic device and the second electronic device are in communication (e.g., wireless communication) with each other to enable information indicative of the user's expression and / or audio corresponding to the user's speech to be transmitted between each other. In some embodiments, the real-time communication session includes displaying a representation of the user (and, optionally, a representation of the second user) within the extended reality environment via display devices of the first electronic device and the second electronic device.

[0143] FIG. 7A illustrates an electronic device 700 (e.g., a watch and / or smartwatch) displaying a prompt 702 on a display 704. Additionally, FIG. 7A illustrates a physical environment 706 of a user 708 using and / or associated with the electronic device 700. In FIG. 7A, the electronic device 700 is worn on the wrist 708a of the user 708 within the physical environment 706. The electronic device 700 is a wearable device configured to be worn on the body of the user 708 (e.g., the wrist 708a of the user 708). In FIG. 7A, the electronic device 700 is a watch (e.g., a smartwatch). In some embodiments, the electronic device 700 is a handheld device disposed within a headset, helmet, goggles, eyeglasses, or wearable frame. In some embodiments, electronic device 700 is configured to be primarily used when worn on the body of user 708, although electronic device 700 may also be used (e.g., interacted with and / or used to capture information via user 708) when electronic device 700 is removed from the body of user 708.

[0144] 7A shows a first portion 710 (e.g., a first face and / or first side, a front side, and / or an interior portion of a head-mounted device (HMD)) of an electronic device 700, including a display 704 and a sensor 712 (e.g., an image sensor such as a camera). When the electronic device 700 is worn on the wrist 708a of a user 708 (or another part of the user's 708's body, such as the head 708d and / or face 708c), the first portion 710 of the device 700 is visible and / or unobstructed from the wrist 708a and / or arm 708b of the user 708. In other words, the first portion 710 of the device is configured to be positioned such that the display 704 is visible to the user 708 (e.g., the display 704 faces away from the wrist 708a and / or the display 704 is positioned above and / or in front of the eyes of the user 708) when the electronic device 700 is positioned on the user's 708's wrist 708a (or another part of the user's 708's body, such as the user's 708's head 708d and / or face 708c). As described below, the electronic device 700 also includes a second portion 714 (e.g., a second face and / or second side, a back side, and / or an outer portion of the HMD) shown in FIG. 7D . When the electronic device 700 is worn on the wrist 708a of the user 708 (or another part of the user's 708's body, such as the user's 708's head 708d and / or face 708c), the second portion 714 of the electronic device 700 is obstructed by (e.g., resting on, touching, and / or otherwise positioned near) the wrist 708a and / or arm 708b of the user 708 (e.g., the second portion 714 of the HMD is not visible to the user when the HMD is positioned on the user's 708's head 708d because the first portion 710 covers and / or is in front of the user's 708's eyes). In other words, the second portion 714 of the electronic device 700 is positioned such that a surface of the second portion 714 faces toward the wrist 708a of the user 708 (e.g., away from the user's face) while the electronic device 700 is worn on the user's wrist 708a.

[0145] 7A-7J depict the electronic device 700 as a wristwatch, in some embodiments, the electronic device 700 is a head-mounted device (HMD). The HMD is configured to be worn on the head 708d of a user 708 and includes a first display on and / or within an interior portion of the HMD. The first display is visible to the user 708 when the user 708 is wearing the HMD on the user's 708's head 708d. For example, the HMD at least partially covers the user's 708's eyes when positioned on the user's 708's head 708d such that the first display is positioned above and / or in front of the user's 708's eyes. In some embodiments, the HMD also includes a second display positioned on and / or within an exterior portion of the HMD. In some embodiments, the second display is not visible to the user 708 when the HMD is positioned on the user's 708's head 708d. Thus, a first display of the HMD displays a prompt 702 instructing the user 708 to remove the HMD from the user's 708 head 708d, and a second display of the HMD displays one or more additional prompts (e.g., content 722a) that provide the user 708 with instructions and / or guidance on using the HMD to capture one or more physical features of the user 708, as described below.

[0146] 7A , electronic device 700 is worn on the body of user 708 (e.g., wrist 708a and / or another part of the body, such as head 708d and / or face 708c) and displays prompt 702 on display 704. Prompt 702 includes an indication (e.g., text and / or image) instructing user 708 to remove electronic device 700 from the body of user 708 (e.g., remove electronic device 700 from wrist 708a of user 708 and / or remove electronic device 700 from another part of the body of user 708, such as head 708d and / or face 708c of user 708) to continue the registration process (e.g., setup process) of electronic device 700. 7A , the electronic device 700 is undergoing an enrollment process, which is a process that includes capturing one or more physical characteristics of the user 708 to generate a representation 726 of the user 708 (e.g., a virtual representation such as an avatar that includes an appearance based on the captured one or more physical characteristics of the user 708). As described below, the electronic device 700 captures at least a portion of the one or more physical characteristics of the user 708 using sensors 720a-720j that are inaccessible, obstructed, and / or otherwise inappropriate for capturing the one or more physical characteristics of the user 708 when the electronic device 700 is worn on the user's 708's body (e.g., when the HMD is worn on the user's 708's head 708d, the HMD's sensors 720a-720j are not aimed at individual body parts of the user 708). Accordingly, the electronic device 700 outputs a prompt 702 instructing the user 708 to remove the electronic device 700 from the body of the user 708 so that one or more of the sensors 720a-720j can be effectively used to capture at least a portion of one or more physical characteristics of the user 708.While FIG. 7A illustrates the prompt 702 as being displayed on the display 704 of the electronic device 700, in some embodiments the prompt 702 includes an audio output (e.g., via a speaker of the electronic device 700) and / or a tactile output (e.g., via one or more haptic output devices of the electronic device 700) instructing the user 708 to remove the electronic device 700 from the user's 708 body (e.g., from another part of the body, such as the wrist 708a and / or head 708d and / or face 708c).

[0147] In some embodiments, electronic device 700 initiates a registration process when electronic device 700 is powered on and / or in response thereto (e.g., when electronic device 700 is initially powered on before user 708 signs in to an account associated with electronic device 700 and / or when electronic device 700 is powered on while in a setup mode of operation). In some embodiments, the registration process is included within the initial setup process of electronic device 700. In some embodiments, the initial setup process of electronic device 700 includes capturing one or more physical characteristics of user 708 (e.g., via sensor 712 and / or sensors 720a-720j), capturing biometric information of user 708 (e.g., facial features, eye features, and / or fingerprints), an input calibration process (e.g., electronic device 700 capturing information that enables electronic device 700 to detect, recognize, and / or respond to user inputs, such as gaze user inputs, air gestures, voice commands, and / or tap gestures), and / or a spatial audio calibration process (e.g., a process that includes electronic device 700 outputting audio to simulate audio generated from a location in physical environment 706 that is not the location of a speaker of electronic device 700, and optionally detecting one or more user inputs that correspond to the perceived location of the output audio). In some embodiments, electronic device 700 initiates the enrollment process based on one or more user inputs that request that the enrollment process be initiated.

[0148] 7B , the electronic device 700 remains positioned on the body of the user 708 (e.g., wrist 708 a and / or another part of the body, such as head 708 d and / or face 708 c) and displays instructions 716 (e.g., directions) via the display 704. In some embodiments, the electronic device 700 displays the instructions 716 after a predetermined amount of time (e.g., 10, 15, 30, and / or 60 seconds) has elapsed since displaying the prompt 702, and when the electronic device 700 has not been removed from the body of the user 708 (e.g., wrist 708 a and / or another part of the body, such as head 708 d and / or face 708 c). The instructions 716 include additional information, suggestions, and / or hints that provide the user 708 with guidance for completing the registration process. 7B, the instructions 716 include text that provides the user 708 with context about the registration process and informs the user 708 about how to use the electronic device 700 to complete at least a portion of the registration process. In FIG. 7B, the instructions 716 prompt the user 708 to point a sensor (e.g., sensors 720a-720j) on a rear portion, such as the second portion 714 (e.g., an outer portion of an HMD), of the electronic device 700 toward the user's 708's face 708c and / or head 708d. While FIG. 7B shows the instructions 716 including text displayed on the display 704 of the electronic device 700, in some embodiments, the instructions 716 include images, symbols, video, animation, audio, and / or text that provide guidance to the user 708 about how to orient and / or use the electronic device 700 to complete at least a portion of the registration process. For example, in some embodiments, the instructions 716 include a video and / or animated series of images that provide the user 708 with visual examples for completing at least a portion of the registration process using the electronic device 700.In some such embodiments, the video and / or animated series of images includes visual indications of the person removing the electronic device 700 from the person's body, orienting the electronic device 700 (e.g., the second portion 714) toward a part of the person's body, and / or the person moving and / or orienting that part of their body so that the user 708 can better understand how to complete at least a portion of the registration process.

[0149] In some embodiments, the instructions 716 include information indicating that one or more physical characteristics of the user 708 captured during at least a portion of the registration process will be used to generate the representation 726. In some embodiments, the instructions 716 include information regarding using the representation 726 in a real-time communication session with another user associated with the external electronic device, which provides the user 708 with context regarding the purpose of capturing the one or more physical characteristics of the user 708.

[0150] 7C , the electronic device 700 remains positioned on the body of the user 708 (e.g., on the wrist 708a and / or another part of the body, such as the head 708d and / or face 708c) and displays a prompt 718 via the display 704. In FIG. 7C , the prompt 718 includes an indication (e.g., text) related to conditions in the physical environment 706 in which the user 708 is located. In some embodiments, the sensor 712 (and / or other sensors) of the electronic device 700 captures information about the physical environment 706, and the electronic device 700 determines whether the captured information indicates one or more conditions that may affect capturing one or more physical characteristics of the user 708. In FIG. 7C , the electronic device 700 determines that the information about the physical environment 706 indicates that the physical environment 706 includes low illumination (e.g., light emitted from one or more light sources, such as a light bulb, a lamp, and / or the sun, is not reaching the user in a sufficient amount to enable the electronic device to effectively capture one or more physical characteristics of the user 708). Accordingly, the electronic device 700 outputs a prompt 718 to warn and / or advise the user 708 that lighting conditions in the physical environment 706 may affect capturing one or more physical characteristics of the user 708. While FIG. 7C illustrates the electronic device 700 providing a prompt 718 related to low lighting conditions in the physical environment 706, in some embodiments, the electronic device 700 is configured to output a prompt related to one or more other conditions in the physical environment 706, such as harsh lighting conditions, an object positioned between the electronic device 700 and the user 708 (e.g., an object obstructing an area from which one or more sensors of the electronic device 700 are configured to capture information), and / or an object and / or accessory (e.g., eyeglasses, a face covering, a head covering, and / or a hat) positioned on a particular part of the body of the user 708. In some embodiments, the electronic device 700 is configured to output a prompt when the electronic device 700 determines that a set of one or more criteria is met, such as when the electronic device 700 includes a power budget and / or battery life that is less than a threshold amount.

[0151] 7C , the prompt 718 includes a first portion 718a (e.g., a first portion of text) that indicates conditions within the physical environment 706 that may affect capturing one or more physical characteristics of the user 708. Additionally, the prompt 718 includes a second portion 718b (e.g., a second portion of text) that provides the user 708 with suggestions and / or guidance about modifying the conditions that may affect capturing one or more physical characteristics of the user 708. In FIG. 7C , the second portion 718b includes a suggestion to the user 708 to move to an area of ​​the physical environment 706 that includes an increased amount of lighting (e.g., a brighter area). In some embodiments, the second portion 718b includes a suggestion to turn on additional light sources and / or increase the amount of power supplied to the light sources. In some embodiments, the second portion 718b includes suggestions to correct and / or adjust for other conditions, such as moving to an area in the physical environment 706 with less harsh lighting (e.g., cooler and / or dimmer lighting), removing and / or moving objects between the electronic device 700 and the user 708, removing and / or moving objects and / or accessories on individual parts of the user's 708's body, and / or charging the electronic device 700.

[0152] In some embodiments, prompt 718 includes a visual prompt, such as a video, image, symbol, emoji, animation, audio prompt, and / or tactile prompt, instead of and / or in addition to text, that notifies and / or alerts user 708 about conditions that affect the capture of one or more physical characteristics of user 708.

[0153] 7D, the user 708 has removed the electronic device 700 from the user's 708 body (e.g., from another part of the body, such as the wrist 708a and / or head 708d and / or face 708c) within the physical environment 706. In addition, FIG. 7D illustrates a second portion 714 of the electronic device 700 (e.g., a back and / or outer portion of an HMD) that is accessible and / or visible after the user 708 has removed the electronic device 700 from the body (e.g., from another part of the body, such as the wrist 708a and / or head 708d and / or face 708c). The second portion 714 of the electronic device 700 includes sensors 720a-720j configured to capture various information about the user 708. In some embodiments, sensors 720a-720j include one or more image sensors (e.g., an IR camera, a 3D camera, a depth camera, a color camera, an RGB camera (e.g., having a complementary metal-oxide semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), one or more infrared (IR) cameras, and / or one or more event-based cameras), an inertial measurement unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, one or more depth sensors (e.g., two or more cameras that determine depth based on structured light, time-of-flight, and / or the difference in perspective of two or more cameras), one or more optical sensors, one or more tactile sensors, one or more orientation sensors, one or more proximity sensors, one or more location sensors, one or more motion sensors, and / or one or more velocity sensors.

[0154] 7D , second portion 714 includes area 722 (e.g., a portion of second portion 714 that does not include sensors 720a-720j and / or an outer portion of the HMD that includes a display different from display 704). Area 722 includes content 722a (e.g., text as shown in FIG. 7D ) that can be viewed and / or perceived by user 708. In some embodiments, area 722 includes one or more display generating components 722b that display content 722a. For example, area 722 is a display. In some embodiments, electronic device 700 causes one or more display generating components 722b to display visual indications that provide instructions and / or otherwise guide user 708 to capture one or more physical characteristics of user 708 using electronic device 700 (e.g., via sensors 720a-720j).

[0155] In some embodiments, the electronic device 700 detects that the electronic device 700 has been removed from the body (e.g., wrist 708a and / or another part of the body, such as head 708d and / or face 708c) of the user 708. In response to detecting that the electronic device 700 has been removed from the body (e.g., wrist 708a and / or another part of the body, such as head 708d and / or face 708c) of the user 708, the electronic device 700 causes one or more display generation components 722b to display one or more prompts that instruct and / or guide the user 708 to use the electronic device 700 to capture one or more physical features of the user 708. In some embodiments, the one or more prompts include text, images, symbols, video, animation, and / or other visual cues that prompt the user 708 to move the electronic device 700 and / or move a part of the user's 708's body in a particular orientation (e.g., move the electronic device 700 in a particular orientation relative to the user's 708's body and / or move a part of the user's 708's body in a particular orientation relative to the electronic device 700). For example, in some embodiments, the one or more prompts instruct the user 708 to adjust the position of the electronic device 700 and / or the user's 708's body within the physical environment 706 so that one or more of the sensors 720a-720j are aimed at a particular part of the user's body, such as the user's 708's face 708c and / or head 708d. In some embodiments, the one or more prompts instruct the user 708 to move a particular body part of the user 708 (e.g., face 708c and / or head 708d) relative to the electronic device 700 so that the sensors 720a-720j capture features of the particular body part of the user 708. In some embodiments, the one or more prompts instruct the user 708 to rotate the head 708d of the user 708 (optionally at a particular speed) relative to the electronic device 700 so that the sensors 720a-720j capture features of the particular body part of the user 708.In some embodiments, the one or more prompts instruct the user 708 to move and / or orient the electronic device 700 so that the sensors 720a-720j are aimed toward the torso 708e (e.g., shoulders and / or chest) of the user 708, thereby causing the electronic device 700 to capture physical characteristics related to the torso 708e and / or clothing worn by the user 708 (e.g., clothing covering and / or positioned on the torso 708e).

[0156] In some embodiments, the electronic device 700 provides one or more prompts (e.g., via one or more display generating components 722b) to instruct the user 708 to make one or more particular facial expressions (e.g., smiling, frowning, opening the mouth, and / or raising and / or lowering eyebrows) to capture one or more physical features of the user's 708's face 708c. In some embodiments, the one or more prompts, similar to prompt 718, include information about conditions in the physical environment 706 that affect the capture of information about the one or more physical features and / or information about adjusting and / or correcting for the conditions. In some embodiments, the one or more prompts instruct the user 708 to move a portion of the user's 708's body and / or move the electronic device 700 so that a discrete portion of the user's 708's body is within a frame (e.g., a frame such as a box and / or outline displayed via one or more display generating components 722b). In some embodiments, the one or more prompts instruct the user 708 to move closer and / or away from the electronic device 700 and / or to move the electronic device 700 closer and / or away from the user 708. In some embodiments, the one or more prompts provided by the electronic device 700 are displayed via one or more display generation components 722b of the area 722. In some embodiments, the one or more prompts are audio prompts (e.g., output via a speaker of the electronic device 700) and / or tactile prompts (e.g., output via one or more tactile output devices of the electronic device 700) that provide instructions and / or guidance to the user 708 about capturing one or more physical characteristics of the user 708.

[0157] 7D, user 708 moves head 708d and / or moves electronic device 700, as indicated by arrows 724a and / or 724b, respectively. While user 708 moves head 708d (and, optionally, other parts of user's 708's body) and / or electronic device 700, sensors 720a-720j capture information about one or more physical characteristics of user 708. The sensors 720a-720j of the electronic device 700 capture information regarding one or more physical features of the user 708, such as one or more facial features, one or more features of the user's 708's hair (e.g., hair on the user's 708's head 708d and / or facial hair), one or more features of the user's 708's torso 708e (e.g., shoulders, chest, and / or clothing), and / or other physical features of the user 708 that are not otherwise accessible while the electronic device 700 is positioned on the user's 708's body (e.g., another part of the body, such as the wrist 708a and / or head 708d and / or face 708c) and / or that are outside the capture area of ​​the sensors 720a-720j. For example, while the user 708 wears the electronic device 700 on their body (e.g., wrist 708a and / or another part of the body, such as head 708d and / or face 708c), the capture areas and / or fields of the sensors 720a-720j may not be directed toward a particular part of the user's 708's body (e.g., face 708c, head 708d, and / or torso 708e) and / or may be prevented from capturing one or more physical features of the user 708. As described below, the electronic device 700 uses the captured information about the one or more physical features of the user 708 to generate a representation 726 of the user 708.

[0158] 7E, the user 708 places the electronic device 700 back on the user's 708 body (e.g., wrist 708 and / or another part of the body, such as head 708d and / or face 708c). In some embodiments, after capturing one or more physical features of the user 708, the electronic device 700 outputs a prompt (e.g., a display via one or more display generating components 722b and / or an audio and / or tactile sensation) instructing the user 708 to return the electronic device 700 to the user's 708 body (e.g., wrist 708a and / or another part of the body, such as head 708d and / or face 708c) to continue the enrollment process. In FIG. 7E, the electronic device 700 determines and / or detects that the electronic device 700 is positioned on the user's 708 body (e.g., wrist 708a and / or another part of the body, such as head 708d and / or face 708c). In response to determining that the electronic device 700 is positioned on the body of the user 708 (e.g., the wrist 708a and / or another part of the body such as the head 708d and / or face 708c) (and, optionally, in response to detecting that one or more physical features of the user 708 have been captured), the electronic device 700 displays a prompt 728 via the display 704.

[0159] In FIG. 7E , the prompt 728 includes an indication (e.g., text) instructing the user 708 to move and / or orient the electronic device 700 (and / or a part of the body of the user 708) so that the sensor 712 is pointing in a direction toward the left hand 708f of the user 708 (e.g., the sensor 712 of the HMD is a camera, and while the HMD is worn on the head 708d of the user 708, the user 708 can orient the head 708d, the left hand 708f, and / or the user's eyes so that the left hand 708f is within the field of view of the camera). In some embodiments, electronic device 700 captures first information regarding one or more first physical characteristics of user 708 via sensors 720a-720j while electronic device 700 is removed from the body of user 708 (e.g., wrist 708a and / or head 708d and / or another part of the body, such as face 708c), and captures second information regarding one or more second physical characteristics of user 708 via sensor 712 while electronic device 700 is positioned on the body of user 708 (e.g., wrist 708a and / or head 708d and / or face 708c). In some embodiments, electronic device 700 uses at least a portion of both the one or more first physical characteristics of user 708 and the one or more second physical characteristics of user 708 to generate representation 726. In some embodiments, electronic device 700 uses only one of the one or more first physical characteristics of user 708 and the one or more second physical characteristics of user 708 to generate representation 726. In some embodiments, the one or more first physical characteristics of user 708 correspond to physical characteristics of parts of user 708's body that are inaccessible while electronic device 700 is on user 708's body (e.g., wrist 708a, and / or another part of the body such as head 708d and / or face 708c) and / or that are outside the capture area and / or field of sensors 720a-720j, e.g., physical characteristics of face 708c, head 708d, and / or torso 708e.In some embodiments, the one or more first physical characteristics of the user 708 correspond to physical characteristics of a part of the user's 708 body that is inaccessible and / or otherwise unsuitable for capturing via the sensor 712 while the electronic device 700 is on the user's 708 body (e.g., another part of the body such as the wrist 708a and / or head 708d and / or face 708c) (e.g., a part of the user's 708 body that is covered by the electronic device when the electronic device 700 is on the user's 708 body (e.g., the face 708c and / or head 708d of the user 708 are covered by the HMD when the HMD is worn on the user's 708's head 708d)). In some embodiments, the one or more second physical characteristics of the user 708 correspond to physical characteristics of a part of the user's 708's body, such as the left hand 708f and / or the right hand 708g, that is accessible and / or suitable for capturing via the sensor 712 while the electronic device 700 is on the user's 708's body (e.g., the wrist 708a and / or another part of the body such as the head 708d and / or the face 708c) (e.g., the left hand 708f and / or the right hand 708g may be captured via a camera (e.g., the sensor 712) of the HMD while the HMD is worn on the user's 708's head 708d). Accordingly, the electronic device 700 outputs one or more prompts instructing the user to remove the electronic device 700 from the body of the user 708 (e.g., another part of the body such as the wrist 708a and / or the head 708d and / or the face 708c) and / or place the electronic device 700 on the body of the user 708 (e.g., another part of the body such as the wrist 708a and / or the head 708d and / or the face 708c) to capture one or more first physical features of the user 708 and / or one or more second physical features of the user 708.

[0160] 7F , the user 708 positions the electronic device 700, left hand 708f, and / or right hand 708g such that the sensor 712 is positioned such that its capture area and / or field is directed toward the left hand 708f (e.g., the sensor 712 of the HMD is a camera, and the user 708 adjusts the position of the head 708d, the position of the left hand 708f, and / or the position of the user's eye gaze so that the left hand 708f is within the field of view of the camera). Additionally, the electronic device 700 displays, via the display 704, a frame 730 indicating a target position of the left hand 708f relative to the electronic device 700 (e.g., the sensor 712 of the electronic device 700). In FIG. 7F , the sensor 712 includes an image sensor such as a camera, and the electronic device 700 displays information captured via the sensor 712 on the display 704. Accordingly, the electronic device 700 displays a hand representation 732 on the display 704 indicating that the sensor 712 has captured and / or otherwise detected the left hand 708f of the user 708. In some embodiments, the user 708 can adjust the position of the left hand 708f and / or the electronic device 700 so that the hand representation 732 is within a frame 730 on the display 704. In some embodiments, when the hand representation 732 is within the frame 730, the left hand 708f of the user 708 is positioned within a target area with respect to the electronic device 700, which enables the sensor 712 to capture information regarding one or more physical characteristics of the left hand 708f. In some embodiments, the electronic device 700 causes the sensor 712 to capture information regarding one or more physical characteristics of the left hand 708f of the user 708 in response to the hand representation 732 being within the frame 730 and / or in response to the hand representation 732 being within the frame 730 for a predetermined amount of time.

[0161] After and / or while the electronic device 700 captures information regarding one or more physical characteristics of the left hand 708f, the electronic device 700 displays a prompt 734 via the display 704. In FIG. 7F , the prompt 734 includes an indication (e.g., text) instructing the user 708 to adjust the position of the left hand 708f and flip the left hand 708f (e.g., rotate the left hand 708f approximately 180 degrees relative to the electronic device 700 and / or the sensor 712). In some embodiments, in response to detecting that the user's left hand 708f has been rotated and / or flipped, the electronic device 700 captures one or more additional physical characteristics related to the left hand 708f of the user 708 via the sensor 712. In some embodiments, the electronic device 700 uses the one or more physical characteristics related to the left hand 708f of the user 708 and / or the additional one or more physical characteristics related to the left hand 708f of the user 708 to generate a portion of the representation 726. In some embodiments, the electronic device 700 uses one or more physical characteristics of the left hand 708f of the user 708 and / or one or more additional physical characteristics of the left hand 708f of the user 708 as part of the input calibration process.

[0162] In some embodiments, electronic device 700 captures information about one or more physical characteristics related to right hand 708g of user 708 while electronic device 700 is positioned on the body of user 708 (e.g., wrist 708a, wrist 708h, and / or another part of the body such as head 708d and / or face 708c). In some embodiments, electronic device 700 captures information about right hand 708g of user 708 (e.g., captures information about right hand 708g of user 708 via sensors 720a-720j) while electronic device 700 is removed from the body of user 708 (e.g., wrist 708a, wrist 708h, and / or another part of the body such as head 708d and / or face 708c).

[0163] After capturing information about the left hand 708f of the user 708 (and, optionally, after completing the capture of information about one or more physical characteristics of the user 708 and / or after detecting that the electronic device 700 has been placed on the body of the user 708 (e.g., another part of the body, such as the wrist 708a and / or head 708d and / or face 708c)), the electronic device 700 displays, via the display 704, a user interface 736 including a representation 726, as shown in FIG. 7G . In FIG. 7G , the representation 726 includes an appearance based on the captured one or more physical characteristics of the user 708, such that the representation 726 resembles and / or otherwise looks like the user 708. For example, the clothing representation 726i of the representation 726 includes an appearance that includes one or more attributes based on one or more physical attributes of the clothing 708i worn by the user 708. In some embodiments, electronic device 700 uses one or more captured physical characteristics of user 708 to generate representation 726 in a stereoscopic manner (e.g., combining and / or overlaying two or more two-dimensional images of user 708 to create the appearance that representation 726 is three-dimensional).

[0164] 7G, electronic device 700 displays representation 726 in a first region 736a of user interface 736 and displays selectable options 738a-738d in a second region 736b of user interface 736. As described below, electronic device 700 is configured to edit the appearance of representation 726 and / or initiate a process to recapture one or more physical features of user 708 in response to detecting user input selecting one or more of selectable options 738a-738d.

[0165] 7G, representation 726 is displayed within environment 740 in first area 726a. In some embodiments, environment 740 is a virtual reality environment. In some embodiments, environment 740 is an augmented reality environment 740. In some embodiments, environment 740 is a static background. In some embodiments, environment 740 includes one or more objects (e.g., virtual objects), such as a frame and / or a mirror.

[0166] In some embodiments, while displaying representation 726, electronic device 700 receives information indicating movement of user 708 within physical environment 706. In response to receiving information indicating movement of user 708 within physical environment 740, electronic device 700 displays movement of representation 726 within environment 706. In some embodiments, electronic device 700 displays movement of representation 726 within environment 740 to reflect the physical movement of user 708 within physical environment 706. In other words, electronic device 700 displays movement of representation 726 as if user 708 were viewing representation 726 in a mirror (e.g., when user 708 moves right hand 708g within physical environment 706, electronic device displays movement of left hand in representation 726). In some embodiments, electronic device 700 displays a frame and / or mirror (e.g., a virtual frame and / or a virtual mirror) within environment 740 to indicate to user 708 that representation 726 is displayed as a mirror image representation of user 708's body.

[0167] In some embodiments, the electronic device 700 displays the representation 726 as a preview of content that will be displayed to another user via an external electronic device while the user 708 is participating in a real-time communication session with the other user. In some embodiments, the electronic device 700 displays the representation 726 and / or at least a portion of the representation 726 via one or more display generating components 722b while the electronic device 700 is removed from the body of the user 708 (e.g., from the wrist 708a and / or another part of the body, such as the head 708d and / or face 708c). In some embodiments, the electronic device 700 displays the representation 726 when the electronic device 700 detects that the electronic device 700 has been placed on the body of the user 708 (e.g., on another part of the body such as the wrist 708a and / or the head 708d and / or the face 708c), and does not display the representation 726 when the electronic device 700 detects that the electronic device 700 has been removed from the body of the user 708 (e.g., on another part of the body such as the wrist 708a and / or the head 708d and / or the face 708c).

[0168] As described above, the electronic device 700 displays selectable options 738a-738d in a second region 736b of the user interface 736 that enable the user 708 to edit the appearance of the representation 726 and / or initiate a process to recapture one or more physical features of the user 708. In FIG. 7G , the first selectable option 738a corresponds to an option for editing the eyewear (e.g., eyeglasses, sunglasses, bifocals, monoculars, goggles, and / or a headset) of the representation 726. In response to detecting user input selecting the first selectable option 738a, the electronic device 700 enables the appearance of the representation 726 to be adjusted and / or changed so that the representation 726 is either wearing (e.g., includes) or not wearing (e.g., does not include) the selected type of eyewear. In some embodiments, the electronic device 700 captures one or more physical characteristics of the user 708 while the user 708 is wearing the eyewear, and therefore the first selectable option 738a allows the user 708 to select whether the representation 726 is wearing eyewear that includes an appearance based on the captured one or more physical characteristics of the user 708 (e.g., the representation 726 is wearing eyewear that includes an appearance having one or more attributes that correspond to the physical eyewear that the user 708 was wearing while the electronic device 700 captured the one or more physical characteristics of the user 708).

[0169] The second selectable option 738b corresponds to an option for editing the accessibility accessory (e.g., an eye patch, prosthetics, and / or hearing aid) of the representation 726. In response to detecting user input selecting the second selectable option 738b, the electronic device 700 allows the appearance of the representation 726 to be adjusted and / or changed so that the representation 726 is wearing (e.g., includes) or is not wearing (e.g., does not include) the selected accessibility accessory. In some embodiments, the electronic device 700 captured one or more physical characteristics of the user 708 while the user 708 was wearing the accessibility accessory, and thus the second selectable option 738b allows the user 708 to select whether the representation 726 is wearing an accessibility accessory that includes an appearance based on the captured one or more physical characteristics of the user 708 (e.g., the representation 726 is wearing an accessibility accessory that includes an appearance having one or more attributes that correspond to a physical accessibility accessory that the user 708 was wearing while the electronic device 700 captured the one or more physical characteristics of the user 708).

[0170] The third selectable option 738c corresponds to an option for editing the skin tone of the representation 726 (e.g., the color, hue, tint, brightness, and / or darkness of the skin representation). In response to detecting user input selecting the third selectable option 738c, the electronic device 700 enables the appearance of the representation 726 to be adjusted and / or changed such that the skin tone of one or more portions of the representation 726 is adjusted. In some embodiments, the captured one or more physical characteristics of the user 708 do not include information regarding the one or more physical skin tones of the user 708, and / or the displayed skin tone of the representation 726 does not otherwise accurately reflect the one or more physical skin tones of the user 708. Accordingly, the third selectable option 738c enables the user 708 to change and / or adjust the skin tone representation of the representation 726 such that the representation 726 includes a skin tone representation that accurately resembles the physical skin tone of the user 708.

[0171] The fourth selectable option 738d corresponds to an option for recapturing one or more physical characteristics of the user 708 such that the electronic device 700 can regenerate and / or update the representation 726 based on the recaptured one or more physical characteristics of the user 708. In some embodiments, in response to detecting user input selecting the fourth selectable option 738d, the electronic device 700 displays the prompt 702 and / or otherwise initiates a process to recapture one or more physical characteristics of the user 708, as shown in FIG.

[0172] In Figure 7G, the electronic device 700 detects user input 750a corresponding to the selection of the third selectable option 738c. In response to detecting user input 750a, the electronic device 700 displays, via the display, a user interface 742 including the representation 726 and selectable skin tone options 742a and 742b, as shown in Figure 7H.

[0173] 7H , the electronic device 700 displays a first skin tone option 742a and a second skin tone option 742b for editing the skin tones of different portions of the representation 726, such that the representation 726 can include different portions having different skin tones (e.g., the color, hue, brightness, and / or darkness of the skin representation on different representations of the body parts of the representation 726). The first skin tone option 742a corresponds to editing the skin tone of a hand representation of the representation 726. In response to detecting user input selecting the first skin tone option 742a, the electronic device 700 allows the appearance of the representation 726 to be adjusted and / or changed such that the skin tone of the hand representation of the representation 726 is adjusted. The second skin tone option 742b corresponds to editing the skin tone of a facial representation 726c of the representation 726. In response to detecting user input selecting second skin tone option 742b, electronic device 700 enables the appearance of representation 726 to be adjusted and / or changed such that the skin tone of facial representation 726c of representation 726 is adjusted. While Figure 7H shows user interface 742 including two selectable skin tone options 742a and 742b, in some embodiments, user interface 742 includes three or more selectable skin tone options corresponding to editing the skin tone of different portions of representation 726.

[0174] In Figure 7H, electronic device 700 detects user input 750b corresponding to selection of completion user interface object 744. After detecting user input 750b, electronic device 700 displays menu user interface 746, as shown in Figure 71. In Figure 71, menu user interface 746 includes menu user interface objects 746a-746f corresponding to various functions, user interfaces, and / or applications configured to be performed and / or displayed by electronic device 700. In some embodiments, menu user interface 746 is a home user interface and / or default user interface of the operating system of electronic device 700.

[0175] In FIG. 7I , the electronic device 700 detects user input 750c corresponding to a selection of a first menu user interface object 746a (e.g., "faces"). In response to detecting the user input 750c, the electronic device 700 displays, via the display 704, a people user interface 748 (e.g., a representation user interface), as shown in FIG. 7J . The people user interface 748 corresponds to different representations of the user 708 (and, optionally, other users of the electronic device 700) generated by the electronic device 700. In FIG. 7J , the people user interface 748 includes a first person user interface object 748a corresponding to a first representation (e.g., representation 726) generated by the electronic device 700 and a second person user interface object 748b corresponding to a second representation (e.g., a different representation from representation 726) generated by the electronic device 700. In some embodiments, the first representation and / or the second representation correspond to representations of the users generated by the electronic device 700 based on one or more captured physical characteristics of the respective users.

[0176] In some embodiments, in response to detecting user input corresponding to a selection of a first person user interface object 748a, the electronic device 700 displays a user interface 736 that includes a first representation (e.g., representation 726) corresponding to the first person user interface object 748a. Similarly, in response to detecting user input corresponding to a selection of a second person user interface object 748b, the electronic device 700 displays a user interface 736 that includes a second representation (e.g., a representation different from representation 726) corresponding to the second person user interface object 748b.

[0177] Additional discussion regarding FIGS. 7A-7J is provided below with reference to methods 800 and 900 described with respect to FIGS. 7A-7J.

[0178] 8 is a flow diagram of an exemplary method 800 for generating a representation of a user, according to some embodiments. In some embodiments, method 800 is performed in a computer system (e.g., 101, 700, and / or 1000) (e.g., a smartphone, a tablet, a head-mounted display generating component) that includes one or more display generating components (e.g., 120, 704, 722, 722b, and / or 1000a) (e.g., a visual output device, a 3D display, and / or a display having at least a portion that is transparent or semi-transparent onto which an image can be projected (e.g., a see-through display), a projector, a head-up display, and / or a display controller) (and, optionally, in communication with one or more cameras (e.g., an infrared camera, a depth camera, and / or a visible light camera)). In some embodiments, method 800 is governed by instructions stored on a non-transitory (or transitory) computer-readable storage medium and executed by one or more processors of a computer system, such as one or more processors 202 of computer system 101 (e.g., control 110 of FIG. 1 ). Some operations of method 800 are, optionally, combined, and / or the order of some operations is, optionally, changed.

[0179] While the computer system (e.g., 101, 700, and / or 1000) is disposed on the body (e.g., 708a) of the user (e.g., 708) (e.g., the computer system is worn in a particular orientation and / or position relative to a particular part of the user's body) (e.g., the computer system is a wearable computer system (e.g., a head-mounted display generating component, eyeglasses, a headset, and / or a watch) configured to be worn on a body part of the computer system user) (in some embodiments, the computer system is a watch configured to be worn on the wrist (e.g., 708a) of the computer system user) (in some embodiments, the computer system is in communication with one or more sensors that capture data indicating whether the computer system is in a wearable position), the computer system The computer system (e.g., 101, 700, and / or 1000), via one or more display generating components (e.g., 120, 704, 722, 722b, and / or 1000a), displays (802) a prompt (e.g., 702) (e.g., text, images, and / or user interface objects containing instructions) instructing the user (e.g., 708) to remove the computer system (e.g., 101, 700, and / or 1000) from the body (e.g., 708a) of the user (e.g., 708) (e.g., remove the wearable computer system so that it is no longer attached to a body portion of the user) and capture information about the user (e.g., 708) (e.g., information about one or more physical characteristics of the user of the computer system) using the computer system (e.g., 101, 700, and / or 1000) (e.g., one or more sensors of the computer system).

[0180] In some embodiments, the computer system (e.g., 101, 700, and / or 1000) displays a prompt (e.g., 702) to remove the computer system from a wearable position during an enrollment process (e.g., a process that includes capturing data (e.g., image data, sensor data, and / or depth data) indicative of the size, shape, position, pose, color, depth, and / or other characteristics of one or more body parts and / or features of the body parts of the user) to generate a representation (e.g., 726) of the user.

[0181] After (e.g., while) displaying a prompt (e.g., 702) instructing a user (e.g., 708) to remove the computer system (e.g., 101, 700, and / or 1000) from the body (e.g., 708a) of the user (e.g., 708), the computer system (e.g., 101, 700, and / or 1000) detects (804) that the computer system (e.g., 101, 700, and / or 1000) has been removed from the body (e.g., 708a) of the user (e.g., 708) (e.g., receives data captured via one or more sensors in communication with the computer system, the data indicating that the computer system is not being worn on the user's body part (e.g., a particular body part)).

[0182] After (e.g., in response to) detecting that a computer system (e.g., 101, 700, and / or 1000) has been removed from the body (e.g., 708a) of a user (e.g., 708), the computer system captures (806) (e.g., via one or more sensors, such as a camera) information associated with (e.g., regarding) the user (e.g., 708) (e.g., image data, sensor data, and / or depth data indicative of the size, shape, position, pose, color, depth, and / or other characteristics of one or more body parts (e.g., head and / or face) of the user) (e.g., information regarding one or more physical characteristics of the user of the computer system). The computer system (e.g., 101, 700, and / or 1000) is configured to use the information to generate a representation (e.g., 726) (e.g., a (2D or 3D) virtual representation, a (2D or 3D) avatar) of the user (e.g., 708) (e.g., the computer system generates a representation (e.g., an avatar) of the user based on information related to the user such that the representation includes visual indications based on (e.g., similar) characteristics of the size, shape, position, pose, color, depth, and / or other characteristics of the user's body, hair, clothing, and / or other features of the user).

[0183] Capturing information related to the user after detecting that the computer system has been removed from the user's body allows the computer system to capture information about parts of the user's body that were not accessible to the computer system while the computer system was on the user's body. Thus, the computer system can capture information about the user without additional and / or external devices and / or sensors. Additionally, the computer system can capture more information related to the user that can be used to generate a more accurate representation of the user.

[0184] In some embodiments, the representation (e.g., 726) of the user (e.g., 708) is configured to be displayed in an augmented reality environment (e.g., 740 and / or 1008) (e.g., a simulated environment in which one or more virtual objects are overlaid on the physical environment or a representation thereof, and / or a simulated environment in which a representation of the physical environment is transformed by computer-generated sensory information) and / or a virtual reality environment (e.g., 740 and / or 1008) (e.g., a simulated environment designed to be entirely based on computer-generated sensory input for one or more senses that includes multiple virtual objects that a person can sense and / or interact with).

[0185] In some embodiments, the representation (e.g., 726) of the user (e.g., 708) is configured to be displayed in an augmented reality environment (e.g., 740 and / or 1008) and / or a virtual reality environment (e.g., 740 and / or 1008) during a real-time communication session between the user (e.g., 708) and a second user (e.g., a user associated with the second representation 1012) associated with a second computer system different from the computer system (e.g., 101, 700, and / or 1000).

[0186] Displaying a user's representation within an augmented reality and / or virtual reality environment allows a user viewing the representation to gain context regarding the state of the device, thereby providing improved feedback regarding the state of the device.

[0187] In some embodiments, a computer system (e.g., 101, 700, and / or 1000) is configured to generate a representation (e.g., 726) of a user (e.g., 708) in a stereoscopic manner (e.g., the representation is a series of two-dimensional images that, when viewed together and / or combined with each other, appear to make the representation exist in three dimensions within the environment in which the representation is displayed). Generating the representation of the user in a stereoscopic manner allows the computer system to generate a more accurate and / or realistic representation of the user.

[0188] In some embodiments, before detecting that the computer system (e.g., 101, 700, and / or 1000) has been removed from the body (e.g., 708a) of the user (e.g., 708) (e.g., while the computer system is disposed on the user's body), the computer system (e.g., 101, 700, and / or 1000) provides (e.g., output and / or display separately from or simultaneously with prompts instructing the user to remove the computer system from the user's body and use the computer system to capture information about the user) instructions (e.g., 716) (e.g., text instructions, image instructions, video instructions, animation instructions, audio instructions, and / or other instructions) to use the computer system (e.g., 101, 700, and / or 1000) to capture information about the user (e.g., 708) (e.g., instructions explaining how a user of the computer system should use, manipulate, and / or otherwise position the computer system and / or the user's body to capture information about the user). In some embodiments, the instructions (e.g., 716) for using the computer system (e.g., 101, 700, and / or 1000) to capture information related to the user (e.g., 708) include a series of images, text instructions, and / or videos that provide examples of how a user is expected to use the computer system (e.g., 101, 700, and / or 1000) to capture information related to the user (e.g., 708). For example, in some embodiments, the instructions (e.g., 716) include examples of using the computer system (e.g., 101, 700, and / or 1000) and / or movements that the user (e.g., 708) should mimic to capture information related to the user (e.g., 708).

[0189] Providing instructions for using a computer system to capture information about a user facilitates a user's ability to use a computer system to capture information about the user, thereby reducing the number of inputs and / or the amount of time required to capture information about the user.

[0190] In some embodiments, providing the instructions (e.g., 716) includes the computer system (e.g., 101, 700, and / or 1000), via one or more display generation components, displaying an animation (e.g., 716) (e.g., a series of visual indications and / or video) that demonstrates (e.g., provides visual examples of) using the computer system (e.g., 101, 700, and / or 1000) to capture information relevant to a user (e.g., 708) (e.g., instructions that explain how a user of the computer system should use, manipulate, and / or otherwise position the computer system and / or the user's body to capture information relevant to the user). Displaying an animation that demonstrates using the computer system to capture information relevant to the user facilitates the user's ability to use the computer system to capture information relevant to the user, thereby reducing the number of inputs and / or amount of time required to capture information relevant to the user.

[0191] In some embodiments, the computer system (e.g., 101, 700, and / or 1000) may detect removal from the body (e.g., 708a) of a user (e.g., 708) (e.g., while the computer system is disposed on the user's body) and a set of criteria are met (e.g., the computer system has low power and / or low battery life (e.g., an amount of power and / or battery life below a threshold amount), an object is occluding one or more parts of the user's body (e.g., eyeglasses, a hat, and / or a face covering), the environment in which the user is located includes harsh lighting (e.g., bright lighting that may affect capturing information about the user), the environment in which the user is located includes low lighting (e.g., lighting that is not sufficient to accurately and / or completely capture information about the user), and / or another condition in the environment in which the user is located prevents the user from viewing the image. Pursuant to a determination that the computer system receives information and / or data from one or more sensors in communication with the computer system indicating that a condition may affect capturing information related to the user, the computer system (e.g., 101, 700, and / or 1000) displays, via one or more display generation components (e.g., 120, 704, 722, 722b, and / or 1000a), an indication (e.g., 718) (e.g., an alert such as a visual and / or audio notification) associated with a condition affecting capturing information related to the user (e.g., the computer system has low power, an object is occluding one or more parts of the user's body, the environment in which the user is located includes harsh lighting, the environment in which the user is located includes low lighting, and / or another condition that may affect capturing information related to the user).Prior to detecting that the computer system (e.g., 101, 700, and / or 1000) has been removed from the body of the user (e.g., 708) (e.g., while the computer system is positioned on the user's body), in accordance with a determination that the set of criteria is not met, the computer system (e.g., 101, 700, and / or 1000) ceases displaying an indication (e.g., 718) associated with a condition affecting the capture of information about the user (e.g., 708) (and, optionally, maintains display of a prompt (e.g., 702) instructing the user to remove the computer system from the user's body and use the computer system to capture information about the user).

[0192] In some embodiments, the set of criteria is a first set of criteria corresponding to a first condition that affects the capture of information related to the user, and the indication is a first indication associated with the first condition that affects the capture of information related to the user. In accordance with the second set of criteria being met and the second set of criteria corresponding to a second condition affecting the capture of information related to the user that is different from the first set of criteria, the computer system (e.g., 101, 700, and / or 1000) displays (e.g., simultaneously with the first indication, before and / or after the first indication, and / or instead of the first indication) a second indication associated with the second condition affecting the capture of information related to the user (in some embodiments, the computer system displays a separate indication that includes a higher priority than one or more other indications, the priority of the separate indication being based on the separate condition affecting the capture of information related to the user (e.g., when the second condition has a higher priority than the first condition because the second condition is associated with a condition that is more likely to affect, or will affect to a greater extent, the capture of information related to the user, the computer system displays the second indication instead of the first indication)). Following a determination that the set of criteria is not met, the computer system (e.g., 101, 700, and / or 1000) ceases displaying a second indication associated with a second condition affecting the capture of information about the user.

[0193] Displaying indications associated with conditions affecting the capture of information relevant to a user allows a user to preemptively address conditions affecting the capture of information, thereby reducing the amount of time required to capture information relevant to the user.

[0194] In some embodiments, an indication (e.g., 718) associated with a condition affecting the capture of information about the user (e.g., 708) includes information (e.g., suggestions and / or instructions) about taking action to help correct the condition (e.g., modify and / or otherwise adjust the condition so that it no longer affects the capture of information about the user) (e.g., one or more steps and / or suggestions that facilitate and / or otherwise improve the capture of information about the user, such as information suggesting that the computer system be charged, information suggesting that the user remove an object that is obstructing one or more parts of the user's body, and / or information suggesting that the user adjust lighting conditions and / or move to a different location and / or environment with improved lighting conditions). Including information about taking action to help correct the condition allows the user to preemptively address a condition affecting the capture of information, thereby reducing the amount of time required to capture information related to the user.

[0195] In some embodiments, in response to detecting that the computer system (e.g., 101, 700, and / or 1000) has been removed from the body (e.g., 708a) of the user (e.g., 708), the computer system (e.g., 101, 700, and / or 1000) initiates a process to capture information about the user (e.g., 708) (e.g., the process of capturing information about the user is triggered, initiated, and / or started in response to detecting that the computer system has been removed from the user's body). Initiating a process to capture information related to the user in response to detecting that the computer system has been removed from the user's body reduces the number of inputs required to capture information related to the user.

[0196] In some embodiments, after detecting that the computer system (e.g., 101, 700, and / or 1000) has been removed from the body (e.g., 708a) of the user (e.g., 708) (e.g., before, simultaneously with, and / or after capturing at least a portion of information about the user), the computer system (e.g., 101, 700, and / or 1000) provides a second prompt (e.g., 722a) including instructions for capturing information about the user (e.g., 708) (e.g., text, images, video, audio, and / or tactile output providing instructions, suggestions, and / or examples of using the computer system to capture information about the user). In some embodiments, the second prompt (e.g., 722a) includes multiple prompts including different instructions for capturing information about the user (e.g., 708). In some embodiments, the second prompt (e.g., 722a) includes a sequence of second prompts including instructions for capturing information about the user (e.g., 708).

[0197] Providing a second prompt that includes instructions to capture information about the user facilitates the user's ability to use the computer system to capture information about the user, thereby reducing the number of inputs and / or the amount of time required to capture information about the user.

[0198] In some embodiments, providing the second prompt (e.g., 722a) includes the computer system (e.g., 101, 700, and / or 1000), via one or more display generation components (e.g., 120, 704, 722, 722b, and / or 1000a), displaying a visual prompt (e.g., text, one or more images, and / or video that provide information, instructions, suggestions, and / or examples that facilitate the user's ability to capture information relevant to the user) along with one or more registration instructions. In some embodiments, the visual prompt (e.g., 722a) includes a plurality of visual prompts that include visual indications of instructions to capture information relevant to the user. In some embodiments, the visual prompt (e.g., 722a) includes a sequence of visual prompts that include visual indications of instructions to capture information relevant to the user. Displaying the visual prompts facilitates the user's ability to use the computer system to capture information relevant to the user, thereby reducing the number of inputs and / or amount of time required to capture information relevant to the user.

[0199] In some embodiments, a prompt (e.g., 702) instructing the user (e.g., 708) to remove the computer system (e.g., 101, 700, and / or 1000) from the user's (e.g., 708) body and use the computer system (e.g., 101, 700, and / or 1000) to capture information about the user (e.g., 708) is displayed via a first display generating component (e.g., 704) of the one or more display generating components (e.g., 120, 704, 722, 722b, and / or 1000a) (e.g., a first display device in communication with the computer system and / or included in a first location on and / or within the housing of the computer system) (in some embodiments, the first display generating component is internal to the computer system when the computer system is placed on the user's body) (in some embodiments, the computer system is a head-mounted device, and the first display generating component displays a message to the user when the head-mounted device is placed on the user's head and / or over the user's eyes) the visual prompt (e.g., 722a) is displayed via a second display generating component (e.g., 722 and / or 722b) that is different from the first display generating component (e.g., 704) of the one or more display generating components (e.g., 120, 704, 722, 722b, and / or 1000a) (e.g., a second display device in communication with the computer system and / or included in a second location on and / or within the housing of the computer system that is different from the first location) (in some embodiments, the second display generating component is external to the computer system when the computer system is positioned on the user's body) (in some embodiments, the computer system is a head-mounted device and the second display generating component is a display generating component configured to be viewed by the user when the head-mounted device is not positioned on the user's head and / or over the user's eyes, and / or the second display generating component is(Not configured to be viewed by a user when the head-mounted device is positioned on the user's head and / or over the user's eyes).

[0200] Displaying a prompt to remove the computer system from the user's body via a first display generating component and displaying the visual prompt via a second display generating component different from the first display generating component displays information to the user on a separate display generating component that is likely to be within the user's point of view, thereby reducing the amount of time required to capture information relevant to the user.

[0201] In some embodiments, providing the second prompt (e.g., 722a) includes outputting an audio prompt (e.g., an audio alert, audio including voice instructions, and / or audio generated to simulate audio generated from a particular location within the environment in which the user is located) along with the one or more registration instructions via an audio device (e.g., speakers and / or headphones) in communication with the computer system (e.g., 101, 700, and / or 1000). Outputting an audio prompt along with the one or more registration instructions facilitates the user's ability to use the computer system to capture information relevant to the user, thereby reducing the number of inputs and / or amount of time required to capture information relevant to the user.

[0202] In some embodiments, providing the second prompt (e.g., 722a) includes providing an indication (e.g., 722a) (e.g., text, image, video, audio, and / or user interface object) instructing the user (e.g., 708) to orient a body part (e.g., face, hand, and / or torso) of the user (e.g., 708) within a target location (e.g., a location relative to one or more sensors in communication with the computer system that facilitates capturing information about the user's body part) relative to the computer system (e.g., 101, 700, and / or 1000). In some embodiments, the indication includes a frame and / or other user interface object displayed via one or more display generating components (e.g., 120, 704, 722, 722b, and / or 1000a) that provides the user (e.g., 708) with a visual indication of a target location to move and / or orient a portion of the user's (e.g., 708) body relative to the computer system (e.g., 101, 700, and / or 1000).

[0203] Providing an indication that instructs the user to orient a portion of the user's body within a target location relative to the computer system encourages the user to capture information relevant to the user using the computer system, thereby reducing the number of inputs and / or time required to capture information relevant to the user.

[0204] In some embodiments, providing the second prompt (e.g., 722a) includes providing an indication (e.g., 718 and / or 722a) (e.g., text, images, video, audio, and / or user interface objects) instructing the user (e.g., 708) to adjust conditions affecting the capture of information about the user (e.g., 708) (e.g., the computer system has low power, an object is obstructing one or more parts of the user's body, the environment in which the user is located includes harsh lighting, the environment in which the user is located includes low lighting, and / or another condition in which the environment in which the user is located includes one or more steps and / or suggestions that facilitate and / or otherwise improve the capturing of information about the user, such as information suggesting that the computer system be charged, information suggesting that the user remove an object obstructing one or more parts of the user's body, and / or information suggesting that the user adjust lighting conditions and / or move to a different location and / or environment with improved lighting conditions).

[0205] Providing an indication that directs a user to adjust conditions that affect the capture of the user's information facilitates the user's ability to use the computer system to capture information relevant to the user, thereby reducing the number of inputs and / or the amount of time required to capture information relevant to the user. Additionally, providing an indication that directs a user to adjust conditions that affect the capture of the user's information allows the computer system to capture more accurate information relevant to the user, which allows the computer system to generate a more accurate representation of the user.

[0206] In some embodiments, providing the second prompt (e.g., 722a) includes providing an indication (e.g., 722a) (e.g., text, image, video, audio, and / or user interface object) instructing the user (e.g., 708) to move the position of the user's (e.g., 708) head (e.g., 708d) (e.g., move the user's head relative to the computer system and / or one or more sensors in communication with the computer system so that the computer system can capture information about the user's head from a particular angle and / or when the user's head is positioned in a particular orientation). Providing an indication instructing the user to move the position of the user's head facilitates the user's ability to capture information relevant to the user using the computer system, thereby reducing the number of inputs and / or amount of time required to capture information relevant to the user.

[0207] In some embodiments, providing the second prompt (e.g., 722a) includes providing an indication (e.g., 722a) (e.g., text, image, video, audio, and / or user interface object) instructing the user (e.g., 708) to position one or more sets of the user's facial features (e.g., 708c) (e.g., eyes, cheeks, forehead, nose, mouth, and / or lips) into a predefined set of one or more facial expressions (e.g., text, image, video, audio, and / or user interface object instructing the user to make a particular facial expression with the user's eyes, cheeks, forehead, nose, mouth, and / or lips). Providing an indication instructing the user to position one or more sets of the user's facial features into a predefined set of one or more facial expressions facilitates the user's ability to use the computer system to capture information relevant to the user, thereby reducing the number of inputs and / or amount of time required to capture information relevant to the user.

[0208] In some embodiments, providing the second prompt (e.g., 722a) includes providing an indication (e.g., 722a) (e.g., text, image, video, audio, and / or user interface object) instructing the user (e.g., 708) to adjust the position of the computer system (e.g., 101, 700, and / or 1000) (e.g., move the computer system relative to the user's body) to orient (e.g., orient one or more sensors in communication with the computer system) the computer system (e.g., 101, 700, and / or 1000) toward a predetermined portion (e.g., 708e) of the user's (e.g., 708) body (e.g., the user's waist and / or torso) (e.g., a predetermined portion of the user's body including clothing such as a shirt, dress, pants, shorts, skirt, jacket, and / or jewelry). Providing an indication that instructs the user to adjust the position of the computer system to orient the computer system toward a predetermined part of the user's body encourages the user to use the computer system to capture information relevant to the user, thereby reducing the number of inputs and / or time required to capture information relevant to the user.

[0209] In some embodiments, a prompt (e.g., 702) instructing a user (e.g., 708) to remove a computer system (e.g., 101, 700, and / or 1000) from the body (e.g., 708a) of the user (e.g., 708) and use the computer system (e.g., 101, 700, and / or 1000) to capture information about the user (e.g., 708) is displayed via a first display generating component (e.g., 704) (e.g., a first display device in communication with the computer system and / or included in a first location on and / or within the housing of the computer system) of one or more display generating components (e.g., 120, 704, 722, 722b, and / or 1000a). After capturing information about the user (e.g., 708), the computer system (e.g., 101, 700, and / or 1000) displays a preview of a representation (e.g., 726) of the user (e.g., 708) (e.g., an image representing the user (e.g., an avatar) including an appearance based on information about the user) via a second display generating component (e.g., 722 and / or 722b) (e.g., a second display device in communication with the computer system and / or included in a second location on and / or within the housing of the computer system that is different from the first location) different from the first display generating component (e.g., 704) of one or more display generating components (e.g., 120, 704, 722, 722b, and / or 1000a). (In some embodiments, the preview of the user's representation is an initial and / or preliminary representation of the user that may be modified and / or regenerated based on one or more user inputs provided by the user.)

[0210] Displaying a preview of the user's representation via a second display generating component allows the user to view the user's generated representation on a display generating component that is likely to be within the user's viewpoint, allowing the user to determine the accuracy of the user's generated representation, thereby providing improved visual feedback.

[0211] In some embodiments, after capturing information about the user (e.g., 708), the computer system (e.g., 101, 700, and / or 1000) detects (e.g., via one or more sensors in communication with the computer system) that the computer system (e.g., 101, 700, and / or 1000) has been placed on the body (e.g., 708a) of the user (e.g., 708) (e.g., the computer system is worn in a particular orientation and / or position relative to a particular part of the user's body) (e.g., the computer system is a wearable computer system (e.g., a head-mounted display generating component, glasses, a headset, and / or a watch) configured to be worn on a part of the body of the user of the computer system). After detecting that a computer system (e.g., 101, 700, and / or 1000) has been placed on the body (e.g., 708a) of a user (e.g., 708), the computer system (e.g., 101, 700, and / or 1000) displays, via one or more display generation components (e.g., 120, 704, 722, and / or 722b), a preview of a representation (e.g., 726) of the user (e.g., 708) (e.g., an image representing the user (e.g., an avatar) that includes an appearance based on information about the user). (In some embodiments, the preview of the user's representation is an initial and / or preliminary representation of the user that may be modified and / or regenerated based on one or more user inputs provided by the user.) In some embodiments, a preview of a representation (e.g., 726) of a user (e.g., 708) is not displayed (e.g., via the first display generating component and / or via the second display generating component) before detecting that a computer system (e.g., 101, 700, and / or 1000) has been placed on the user's (e.g., 708) body (e.g., 708a) (after capturing information about the user).

[0212] Displaying a preview of the user's representation allows the user to determine the accuracy of the user's generated representation, thereby providing improved visual feedback.

[0213] In some embodiments, capturing information about a user (e.g., 708) includes a computer system (e.g., 101, 700, and / or 1000) capturing first information (e.g., one or more facial features of the user's face) about a first portion (e.g., 708c, 708d, and / or 708e) of the user's (e.g., 708) body (e.g., the user's face and / or head). After capturing first information regarding a first portion (e.g., 708c, 708d, and / or 708e) of a user's (e.g., 708) body, the computer system (e.g., 101, 700, and / or 1000) detects (e.g., via one or more sensors in communication with the computer system) that the computer system (e.g., 101, 700, and / or 1000) is placed on the user's (e.g., 708) body (e.g., 708a) (e.g., the computer system is worn in a particular orientation and / or position relative to a particular portion of the user's body) (e.g., the computer system is a wearable computer system (e.g., a head-mounted display generating component, glasses, a headset, and / or a watch) configured to be worn on a body portion of the user of the computer system). After (e.g., in response to) detecting that the computer system (e.g., 101, 700, and / or 1000) has been placed on the body (e.g., 708a) of the user (e.g., 708), the computer system (e.g., 101, 700, and / or 1000) initiates a process to capture second information (e.g., one or more characteristics of the user's hand) regarding a second part (e.g., 708b, 708f, and / or 708g) of the user's (e.g., 708) body (e.g., the user's hand and / or arm), which is different from the first part (e.g., 708c, 708d, and / or 708e) of the user's (e.g., 708) body. In some embodiments, second information regarding a second portion (e.g., 708b, 708f, and / or 708g) of the user's (e.g., 708) body is captured while the system (e.g., 101, 700, and / or 1000) is positioned on the user's (e.g., 708) body (e.g., 708a).

[0214] Initiating the process of capturing second information related to a second part of the user's body after detecting that the computer system has been placed on the user's body reduces the number of inputs required to perform the capture of second information related to the second part of the user's body.

[0215] In some embodiments, initiating the process of capturing second information regarding the second part (e.g., 708b, 708f, and / or 708g) of the user's (e.g., 708) body includes the computer system (e.g., 101, 700, and / or 1000), via one or more display generating components (e.g., 120, 704, 722, 722b, and / or 1000a), displaying a visual indication (e.g., 730) (e.g., a user interface object including an outline, a shape of a human hand) indicating a location (e.g., relative to the computer system and / or relative to one or more sensors in communication with the computer system) for the user (e.g., 708) to position (e.g., relative to the computer system and / or relative to one or more sensors in communication with the computer system) the second part (e.g., 708b, 708f, and / or 708g) of the user's (e.g., 708) body (e.g., 708b, 708f, and / or 708g).

[0216] Displaying a visual indication of a location for a user to position a second part of the user's body facilitates the user's ability to capture second information related to the second part of the user's body using the computer system, thereby reducing the amount of time required to capture second information related to the second part of the user's body.

[0217] In some embodiments, initiating the process of capturing second information regarding a second part (e.g., 708b, 708f, and / or 708g) of the user's (e.g., 708) body includes the computer system (e.g., 101, 700, and / or 1000) providing a prompt (e.g., 734) (e.g., via a visual prompt such as text, image, video, and / or user interface object, and / or via an audio prompt) instructing the user (e.g., 708) to adjust the orientation of the second part (e.g., 708b, 708f, and / or 708g) of the user's (e.g., 708) body (e.g., the position and / or location of the user's hand relative to the computer system and / or one or more sensors in communication with the computer system). In some embodiments, prompting the user (e.g., 708) to adjust the orientation of the second portion (e.g., 708b, 708f, and / or 708g) of the user's (e.g., 708) body includes providing the user with instructions (e.g., 734) to turn over the user's (e.g., 708) hand (e.g., 708f and / or 708g) so that information regarding the palm side and / or back side of the hand (e.g., 708f and / or 708g) can be captured.

[0218] Providing a prompt instructing the user to adjust the orientation of the second part of the user's body facilitates the user's ability to use the computer system to capture second information related to the second part of the user's body, thereby reducing the amount of time required to capture the second information related to the second part of the user's body.

[0219] In some embodiments, after (e.g., in response to) capturing second information related to a second part (e.g., 708b, 708f, and / or 708g) of the user's (e.g., 708), the computer system (e.g., 101, 700, and / or 1000), via one or more display generation components (e.g., 120, 704, 722, 722b, and / or 1000a), displays a representation (e.g., 726) of the user (e.g., 708) within an extended reality environment (e.g., 740 and / or 1008) (e.g., a wholly or partially simulated environment in which people sense and / or interact via electronic systems, in which a person's physical movement or a subset of their representation is tracked and, in response, one or more properties of one or more virtual objects simulated within the extended reality environment are adjusted in a manner consistent with at least one law of physics). In some embodiments, the representation (e.g., 726) of the user (e.g., 708) includes a facial expression (e.g., 726c) and a hand expression based on captured information about the user (e.g., 708) and / or captured second information about a second part of the user's (e.g., 708b, 708g, and / or 708f) body. Displaying the user's representation within the extended reality environment after capturing the second information related to the user's second part of the body allows the user to determine the accuracy of the user's generated representation, thereby providing improved visual feedback.

[0220] In some embodiments, aspects / operations of methods 900, 1100, 1200, 1300, and / or 1400 may be interchanged, substituted, and / or added between these methods. For example, the computer system of method 800 may be used to display a user's expression, adjust the appearance of a user's expression, display a mouth representation of a user's expression, display a hair representation of a user's expression, and / or display a portion of a user's expression with visual emphasis. For the sake of brevity, those details will not be repeated here.

[0221] 9 is a flow diagram of an exemplary method 900 for displaying a user's representation, according to some embodiments. In some embodiments, method 900 is performed in a computer system (e.g., 101, 700, and / or 1000) (e.g., a smartphone, a tablet, a head-mounted display generating component) that includes one or more display generating components (e.g., 120, 704, 722, 722b, and / or 1000a) (e.g., a visual output device, a 3D display, a display having at least a portion that is transparent or translucent onto which an image can be projected (e.g., a see-through display), a projector, a head-up display, and / or a display controller) (and, optionally, in communication with one or more cameras (e.g., an infrared camera, a depth camera, a visible light camera)). In some embodiments, method 900 is governed by instructions stored on a non-transitory (or transitory) computer-readable storage medium and executed by one or more processors of a computer system, such as one or more processors 202 of computer system 101 (e.g., control 110 of FIG. 1 ). Some operations of method 900 are, optionally, combined, and / or the order of some operations is, optionally, changed.

[0222] During an enrollment process (e.g., a process including capturing data (e.g., image data, sensor data, and / or depth data) indicative of the size, shape, position, pose, color, depth, and / or other characteristics of one or more body parts and / or features of the body parts) for generating a representation (e.g., 726) of a user (e.g., 708) (e.g., an avatar and / or virtual representation of at least a portion of the first user), a computer system (e.g., 101, 700, and / or 1000) detects (902) (e.g., via one or more cameras) information regarding one or more physical characteristics (e.g., data (e.g., image data, sensor data, and / or depth data) indicative of the size, shape, position, pose, color, depth, and / or other characteristics of one or more body parts and / or features of the body parts) of the user (e.g., 708) of the computer system (e.g., 101, 700, and / or 1000).

[0223] After capturing information regarding one or more physical characteristics of a user (e.g., 708) of a computer system (e.g., 101, 700, and / or 1000), the computer system (e.g., 101, 700, and / or 1000) generates (904) a representation (e.g., 726) of the user (e.g., 708) based on the information regarding the one or more physical characteristics of the user (e.g., 708), including selecting one or more physical characteristics of the representation (e.g., 726) based on the one or more captured physical characteristics of the user (e.g., 708) (e.g., the computer system uses information about the user of the computer system to generate a representation (e.g., avatar) of the user that includes visual indications of the captured and / or detected size, shape, position, pose, color, depth, and / or other characteristics of the first user's body, clothing, hair, and / or features).

[0224] After generating a representation (e.g., 726) of the user (e.g., 708), the computer system (e.g., 101, 700, and / or 1000), via one or more display generation components (e.g., 120, 704, 722, 722b, and / or 1000a), displays (906) at least a portion of the representation (e.g., 726) of the user (e.g., 708) in an extended reality environment (e.g., 740) (e.g., a fully or partially simulated environment in which people sense and / or interact via an electronic system, where a subset of a person's physical movements or a representation thereof is tracked and one or more properties of one or more simulated virtual objects in the extended reality environment are adjusted accordingly to conform to at least one law of physics). In some embodiments, the computer system (e.g., 101, 700, and / or 1000) displays the representation (e.g., 726) of the user (e.g., 708) within the extended reality environment (e.g., 740) after the registration process is complete so that the user (e.g., 708) can view the representation (e.g., 726) and, in some embodiments, edit and / or modify the representation (e.g., 726).

[0225] Displaying at least a portion of the user's representation within the extended reality environment after generating the user's representation allows the user to determine the accuracy of the user's generated representation and whether to request that the computer system recapture information relevant to the user, thereby providing improved visual feedback.

[0226] In some embodiments, the extended reality environment (e.g., 740) includes an augmented reality environment (e.g., a simulated environment in which one or more virtual objects are superimposed on a physical environment or a representation thereof, and / or a simulated environment in which a representation of a physical environment is transformed by computer-generated perceptual information). Displaying a representation of a user within an augmented reality environment allows a user viewing the representation to gain context regarding the state of the device, thereby providing improved feedback regarding the state of the device.

[0227] In some embodiments, the extended reality environment (e.g., 740) includes a virtual reality environment (e.g., a simulated environment designed to be entirely based on computer-generated sensory input for one or more senses, including multiple virtual objects that a person can sense and / or interact with). Displaying a representation of a user within the virtual reality environment allows a user viewing the representation to gain context regarding the state of the device, thereby providing improved feedback regarding the state of the device.

[0228] In some embodiments, capturing information regarding one or more physical characteristics of a user (e.g., 708) of a computer system (e.g., 101, 700, and / or 1000) includes the computer system (e.g., 101, 700, and / or 1000) capturing information regarding one or more physical characteristics of a user (e.g., 708) of the computer system (e.g., 101, 700, and / or 1000) while the computer system (e.g., 101, 700, and / or 1000) is removed from the body (e.g., 708a) of the user (e.g., 708). The computer system receives data captured via one or more sensors in communication with the computer system, the data indicating that the computer system is not attached to a body part (e.g., a particular body part) of the user (e.g., the computer system is a wearable computer system (e.g., a head-mounted display generating component, glasses, a headset, and / or a watch) configured to be attached to a body part of the user of the computer system) (in some embodiments, the computer system is a watch configured to be worn on the wrist of the user of the computer system). Displaying at least a portion of the representation (e.g., 726) of the user (e.g., 708) within the extended reality environment (e.g., 740) includes the computer system (e.g., 101, 700, and / or 1000) displaying at least a portion of the representation (e.g., 726) of the user (e.g., 708) within the extended reality environment (e.g., 740) after (e.g., in response to) the computer system (e.g., 101, 700, and / or 1000) detecting that the computer system (e.g., 101, 700, and / or 1000) is positioned on the body (e.g., 708a) of the user (e.g., 708) (e.g., the computer system is attached in a particular orientation and / or position relative to a particular part of the user's body).

[0229] Capturing information about one or more physical characteristics of a user while the computer system is removed from the user's body enables the computer system to capture information about parts of the user's body that are not accessible to the computer system while the computer system is located on the user's body. Thus, the computer system can capture information about the user without additional and / or external devices and / or sensors. Additionally, the computer system can capture more information related to the user that can be used to generate a more accurate representation of the user.

[0230] In some embodiments, displaying at least a portion of the representation (e.g., 726) of the user (e.g., 708) within the extended reality environment (e.g., 740) includes the computer system (e.g., 101, 700, and / or 1000) animating the representation (e.g., 726) (e.g., displaying movement of the representation that mirrors and / or mimics the user's movements) based on movement of the user (e.g., 708) relative to at least a portion of the computer system (e.g., 101, 700, and / or 1000) (e.g., within the physical environment in which the user is located) (e.g., the computer system receives information regarding the user's physical state, including the user's movements, and displays at least a portion of the user's representation within the extended reality environment based on the received information). In some embodiments, the animation of the representation (e.g., 726) is displayed in conjunction with (e.g., coincides with) the detected movement of the user (e.g., 708). Animating the representation based on the user's movement relative to at least a portion of the computer system allows the user to understand that the representation is associated with the user, thereby providing improved feedback regarding the state of the device.

[0231] In some embodiments, animating the representation (e.g., 726) includes the computer system (e.g., 101, 700, and / or 1000) displaying movement of the representation (e.g., 726) that is a mirror image of the movement of the user (e.g., 708) relative to at least a portion of the computer system (e.g., 101, 700, and / or 1000) (e.g., in a physical environment) (e.g., the movement of the representation is displayed to the user as if the user were looking at its reflection in a mirror). In some embodiments, the animation of the representation (e.g., 726) is displayed in conjunction with (e.g., coincides with) the detected movement of the user (e.g., 708). Displaying the movement of the representation as a mirror image of the user's movement relative to at least a portion of the computer system allows the user to understand that the representation is associated with the user, thereby providing improved feedback regarding the state of the device.

[0232] In some embodiments, displaying at least a portion of a representation (e.g., 726) of a user (e.g., 708) in an extended reality environment (e.g., 740) includes a computer system (e.g., 101, 700, and / or 1000) displaying a representation (e.g., 726) having a first orientation (e.g., posture, position, pose, and / or stance) that is a mirror image of a second orientation (e.g., posture, position, pose, and / or stance) of the user (e.g., 708) in a physical environment (e.g., 706) in which the user (e.g., 708) is located (e.g., displayed as if the user were viewing the representation as a reflection of the user in a mirror and / or displayed as if the user's representation were flipped on a vertical axis without being flipped on a horizontal axis) (e.g., the computer system receives information regarding the user's physical state and displays at least a portion of the user's representation within the extended reality environment based on the received information). Displaying a representation with a first orientation that is a mirror image of the user's second orientation within the physical environment in which the user is located allows the user to understand that the representation is associated with the user, thereby providing improved feedback regarding the state of the device.

[0233] In some embodiments, displaying at least a portion of a representation (e.g., 726) of a user (e.g., 708) within the extended reality environment (e.g., 740) includes a computer system (e.g., 101, 700, and / or 1000) displaying a frame (e.g., a user interface object resembling a frame surrounding a mirror and / or reflective surface) around the representation (e.g., 726) within the extended reality environment (e.g., 740). The frame indicates (e.g., indicates to the user) that the representation (e.g., 726) of the user (e.g., 708) in the extended reality environment (e.g., 740) has an orientation (e.g., posture, position, pose, and / or stance) that is a mirror image (e.g., displayed as if the user were viewing the representation as a mirror reflection surrounded by the frame) of the user's (e.g., 708) orientation (e.g., physical and / or actual posture, position, pose, and / or stance) in the physical environment (e.g., 706) in which the user (e.g., 708) is located. Displaying the frame around the representation in the extended reality environment allows the user to understand that the representation is associated with the user, thereby providing improved feedback regarding the state of the device.

[0234] In some embodiments, while displaying at least a portion of a representation (e.g., 726) of a user (e.g., 708) in the extended reality environment (e.g., 740), the computer system (e.g., 101, 700, and / or 1000), via one or more display generation components (e.g., 120, 704, 722, 722b, and / or 1000a), edits visual characteristics of the representation (e.g., 726) (e.g., modifies, adjusts, and / or changes the visual appearance of the representation to display accessories (e.g., headwear, head coverings, eyewear, and / or clothing)). 738a-738d) (e.g., selectable user interface objects such as virtual buttons and / or text) for editing the representation's visual characteristics (e.g., adding and / or removing a facial hair color and / or style, adding and / or removing a prosthetic device, an eye patch, and / or hearing aid, adjusting the skin tone of one or more parts of the representation's body, adjusting the representation's hair color and / or hair style, adjusting the representation's facial hair characteristics, recapturing information about one or more physical characteristics of the user, and / or resuming capture of information about one or more physical characteristics of the user). Displaying one or more selectable options for editing the representation's visual characteristics allows the representation to be edited without requiring additional user input to navigate to a separate editing user interface, thereby reducing the number of inputs required to edit the representation's visual characteristics.

[0235] In some embodiments, one or more selectable options (e.g., 738a-738d) include an eyewear selectable option (e.g., 738a) (e.g., a selectable user interface object such as a virtual button and / or text) for editing the eyewear of the representation (e.g., 726) (e.g., selecting whether the representation of the user is wearing glasses (and optionally the type of glasses), a headset, a monocle, and / or sunglasses, and / or the type, shape (e.g., frame shape), color, and / or size of eyewear included in the representation). Including an eyewear selectable option allows the eyewear of the representation to be edited without requiring additional user input to navigate to a separate editing user interface, thereby reducing the number of inputs required to edit the eyewear of the representation.

[0236] In some embodiments, one or more of the selectable options (e.g., 738a-738d) include an accessory selectable option (e.g., 738b) (e.g., a selectable user interface object such as a virtual button and / or text) for editing the representation's (e.g., 726) accessories (e.g., whether the user's representation includes an eye patch, prosthetics, and / or hearing aids). Including an accessory selectable option allows the representation's accessories to be edited without requiring additional user input to navigate to a separate editing user interface, thereby reducing the number of inputs required to edit the representation's accessories.

[0237] In some embodiments, the one or more selectable options (e.g., 738a-738d) include one or more skin tone selectable options (e.g., 738c, 742a, and / or 742b) (e.g., selectable user interface objects such as virtual buttons and / or text) for editing the skin tone of the representation (e.g., 726) (e.g., adjusting, modifying, and / or changing the hue and / or color of a skin representation included in one or more portions of the user's representation). The inclusion of one or more skin tone selectable options allows the skin tone of the representation to be edited without requiring additional user input to navigate to a separate editing user interface, thereby reducing the number of inputs required to edit the skin tone of the representation.

[0238] In some embodiments, the one or more skin tone selectable options (e.g., 738c, 742a, and / or 742b) (e.g., selectable user interface objects such as virtual buttons and / or text) include a first skin tone selectable option (e.g., 742a) (e.g., selectable user interface objects such as virtual buttons and / or text) for editing the skin tone of the face (e.g., 726c) of the representation (e.g., 726) and a second skin tone selectable option (e.g., 724b) (e.g., selectable user interface objects such as virtual buttons and / or text) for editing the skin tone of the hands of the representation (e.g., 726). In some embodiments, a user's skin tone may vary on different parts of the user's body; therefore, providing multiple skin tone selectable options allows a user (e.g., 708) to modify the appearance of the user's (e.g., 708) representation (e.g., 726) to more accurately reflect the user's (e.g., 708) actual appearance. Including a first skin tone selectable option for editing the skin tone of the face of the expression and a second skin tone selectable option for editing the skin tone of the hands of the expression allows the skin tones of different parts of the expression to be edited without requiring additional user input to navigate to a separate editing user interface, thereby reducing the number of inputs required to edit the skin tone of the expression.

[0239] In some embodiments, one or more of the selectable options (e.g., 738a-738d) include a recapture selectable option (e.g., 738d) (e.g., a selectable user interface object such as a virtual button and / or text) that, when selected, initiates a process of recapturing information about one or more physical characteristics of the user (e.g., 708). (E.g., selecting the recapture selectable option causes the computer system to display a user interface and / or initiate a process of recapturing information about one or more of the user's (e.g., 708) physical characteristics.) In some embodiments, the initial capture of information about one or more physical characteristics of the user (e.g., 708) may be inaccurate and / or incomplete; therefore, providing the user (e.g., 708) with the ability to recapture at least a portion of the information about one or more physical characteristics of the user (e.g., 708) enables the computer system (e.g., 101, 700, and / or 1000) to generate a representation (e.g., 726) to more accurately reflect the user's (e.g., 708) actual appearance. Including a recapture selectable option allows information regarding one or more physical characteristics of a user to be recaptured without requiring additional user input to navigate to a separate user interface, thereby reducing the number of inputs required to recapture information regarding one or more physical characteristics of a user.

[0240] In some embodiments, one or more of the selectable options (e.g., 738a-738d) include a resume selectable option (e.g., 738d) (e.g., a selectable user interface object such as a virtual button and / or text) that, when selected, initiates a step of the enrollment process that includes capturing second information regarding one or more physical characteristics of a user (e.g., 708) of the computer system (e.g., 101, 700, and / or 1000). (E.g., selecting the resume selectable option causes the computer system to resume capturing information regarding one or more physical characteristics of a user of the computer system using one or more sensors of the computer system, and optionally causes the computer system to delete and / or otherwise not use the initially captured information regarding the one or more physical characteristics of the user to generate a representation of the user.) Including the resume selectable option allows the second information regarding the one or more physical characteristics of the user to be captured without requiring additional user input to navigate to a separate user interface, thereby reducing the number of inputs required to capture the second information regarding the one or more physical characteristics of the user.

[0241] In some embodiments, the one or more physical characteristics of the user (e.g., 708) include one or more first features (e.g., facial features) of the user's (e.g., 708) face (e.g., 708c) and one or more second features (e.g., size, shape, skin tone, and / or contour) of the user's (e.g., 708) hands (e.g., 708f and / or 708g). During an enrollment process to generate a representation (e.g., 726) of the user (e.g., 708), and while the computer system (e.g., 101, 700, and / or 1000) is removed from the user's (e.g., 708a) body (e.g., 708a) (e.g., the computer system receives data captured via one or more sensors in communication with the computer system, and the data indicates that the computer system is not being worn on a body part (e.g., a particular body part) of the user) (e.g., the computer system is a wearable computer system (e.g., For example, the computer system may be a head-mounted display generating component, glasses, a headset, and / or a watch) (in some embodiments, the computer system is a watch configured to be worn on the wrist of a user of the computer system) (in some embodiments, the computer system is in communication with one or more sensors that capture data indicating whether the computer system is in a wearable position), the computer system (e.g., 101, 700, and / or 1000) captures one or more first features of a face (e.g., 708c) of a user (e.g., 708) (e.g., without capturing features of the user's hands). After capturing the one or more first features of a face (e.g., 708c) of a user (e.g., 708), the computer system (e.g., 101, 700, and / or 1000) detects that the computer system (e.g., 101, 700, and / or 1000) is disposed on a body (e.g., 708a) of the user (e.g., 708) (e.g., the computer system is worn in a particular orientation and / or position relative to a particular part of the user's body).After (e.g., in response to and / or during) detecting that a computer system (e.g., 101, 700, and / or 1000) has been placed on the body (e.g., 708a) of a user (e.g., 708), the computer system (e.g., 101, 700, and / or 1000) captures one or more second features of the user's (e.g., 708) hands (e.g., 708f and / or 708g) (e.g., without capturing the user's facial features) (e.g., capturing one or more second features of the user's hands via one or more sensors (e.g., cameras) in communication with the computer system).

[0242] Capturing one or more first features of the user's face while the computer system is removed from the user's body and capturing one or more second features of the user's hands while the computer system is positioned on the user's body facilitates the computer system's ability to capture information about different parts of the user's body, thereby reducing the amount of time required to capture information about one or more physical features of the user.

[0243] In some embodiments, the registration process is part of a setup process for the computer system (e.g., 101, 700, and / or 1000) (e.g., a setup process that begins when the computer system is first turned on and / or a setup process that begins when a user of the computer system is creating an account for using the computer system and / or signing in to it for the first time). During the setup process for the computer system (e.g., 101, 700, and / or 1000), the computer system (e.g., 101, 700, and / or 1000) captures one or more biometric features of the user (e.g., one or more features of the user's face, one or more features of the user's eyes, one or more features of the user's hands and / or fingers, and / or one or more features of the user's voice). Capturing the one or more biometric features of the user during the setup process for the computer system allows the computer system to obtain additional information without requiring additional user input, thereby reducing the number of inputs required to capture the user's one or more biometric features.

[0244] In some embodiments, the registration process is part of a setup process for a computer system (e.g., 101, 700, and / or 1000) (e.g., a setup process that begins when the computer system is first turned on and / or a setup process that begins when a user of the computer system creates an account for using the computer system and / or signs in to it for the first time). During a setup process of a computer system (e.g., 101, 700, and / or 1000), the computer system (e.g., 101, 700, and / or 1000) may include an input calibration process that enables the computer system (e.g., 101, 700, and / or 1000) to calibrate the detection of one or more input technologies (e.g., detecting, observing, and / or capturing information about a user's gaze (e.g., a user attempts to provide a known and / or predetermined sequence of gaze inputs, and the detected, observed, and / or captured information about the user's gaze is compared to the known and / or predetermined sequence of gaze inputs, and the comparison adjusts how the computer system interprets the gaze inputs so that the detected, observed, and / or captured information about the user's gaze matches the known and / or predetermined sequence of gaze inputs). and / or performing a process including detecting and / or capturing information about a user's hands, movements of the user's hands, and / or gestures made by the user's hands (e.g., a user is attempting to provide a known and / or predetermined series of hand gesture inputs, and detected, observed, and / or captured information about the user's hands is compared to the known and / or predetermined series of hand gesture inputs, the comparison is used to adjust how the computer system interprets the hand gesture inputs so that the detected, observed, and / or captured information about the user's hands matches the known and / or predetermined series of hand gesture inputs, and the computer system can detect and perform one or more functions based on the inputs, and / or the computer system can detect the inputs more accurately).

[0245] By performing the input calibration process during the computer system setup process, the computer system can obtain additional information without requiring additional user input, thereby reducing the number of inputs required to perform the input calibration process.

[0246] In some embodiments, the registration process is part of a setup process for the computer system (e.g., 101, 700, and / or 1000) (e.g., a setup process that begins when the computer system is first turned on and / or a setup process that begins when a user of the computer system creates an account for using the computer system and / or signs in to it for the first time). During the setup process of the computer system (e.g., 101, 700, and / or 1000), the computer system (e.g., 101, 700, and / or 1000) performs a spatial audio calibration process (e.g., a process that includes outputting audio via an audio output device (e.g., speakers and / or headphones) in communication with the computer system, where the output audio is generated to simulate audio generated from at least one location different from the actual location of the audio output device). (In some embodiments, the spatial audio calibration includes outputting audio, detecting feedback and / or one or more user inputs corresponding to a perceived location of the output audio, and calibrating the perceived location to cause the output audio to simulate audio generated from a target location.)

[0247] Performing the spatial audio calibration process during the setup process for the computer system allows the computer system to obtain additional information without requiring additional user input, thereby reducing the number of inputs required to perform the spatial audio calibration.

[0248] In some embodiments, the registration process is part of a setup process for a computer system (e.g., 101, 700, and / or 1000) (e.g., a setup process that begins when the computer system is first turned on and / or a setup process that begins when a user of the computer system creates an account for using the computer system and / or signs in to it for the first time). During a setup process of the computer system (e.g., 101, 700, and / or 1000), the computer system (e.g., 101, 700, and / or 1000) provides an indication (e.g., via one or more display generation components) of instructions for using the representation (e.g., 726) during a real-time communication session (e.g., instructions describing how a user of the computer system can use the representation to communicate with one or more additional users (e.g., additional users associated with an external computer system) during a real-time communication session (e.g., a real-time communication session between the user of the computer system and a second user associated with a second computer system different from the first computer system, the real-time communication session including displaying and / or otherwise communicating a representation of the user's facial and / or body expression to the second user via the computer system and / or the second computer system).

[0249] Providing an indication of instructions to use expressions during a real-time communication session during the computer system setup process causes the device to automatically perform actions that provide the user with additional context regarding how the expressions may be used.

[0250] In some embodiments, the one or more physical characteristics of the user (e.g., 708) include clothing (e.g., 708i) of the user (e.g., 708) (e.g., physical clothing worn by the user in the physical environment in which the user is located), and the representation (e.g., 726) includes clothing representation (726i) (e.g., visual images and / or indications of clothing that resemble and / or include one or more similar attributes of the user's physical clothing) based on the user's (e.g., 708) clothing (e.g., 708i) detected during the enrollment process to generate the representation (e.g., 726) of the user (e.g., 708). Representations that include clothing representations based on the user's clothing enable the appearance of the representation to more closely resemble the user's actual appearance, thereby providing improved visual feedback.

[0251] In some embodiments, after displaying at least a portion of a representation (e.g., 726) of a user (e.g., 708) within an extended reality environment (e.g., 740), the computer system (e.g., 101, 700, and / or 1000), via one or more display generation components (e.g., 120, 704, 722, 722b, and / or 1000a), displays a menu user interface (e.g., 746 and / or 748) (e.g., a user interface including one or more selectable options for performing computer system functions such as initiating a real-time communication session, editing and / or modifying the representation, and / or starting a game). A menu user interface (e.g., 746 and / or 748), when selected, causes a computer system (e.g., 101, 700, and / or 1000) to display at least a portion of a representation (e.g., 726) of a user (e.g., 708) (e.g., the representation of the user within the extended reality environment is displayed in a manner consistent with the orientation (e.g., the user's physical and / or actual posture, position, pose, and / or stance within the physical environment in which the user is located) within the extended reality environment (e.g., 740) (e.g., as if the user were viewing the extended reality environment from a mirrored perspective). The representation (e.g., 726) of the user (e.g., 708) includes mirrored and / or in-frame) selectable options (e.g., 746a, 748a, and / or 748b) (e.g., selectable user-interface objects such as virtual buttons and / or text) that indicate (e.g., to the user) that the representation has a 'mirror-image' orientation (e.g., posture, position, pose, and / or stance) (e.g., displayed as if the user were viewing a mirror image of the representation). In some embodiments, the representation (e.g., 726) of the user (e.g., 708) is animated and displayed in conjunction with (e.g., matches) the detected movement of the user (e.g., 708).In some embodiments, physical movement of a user (e.g., 708) relative to a portion of a computer system (e.g., 101, 700, and / or 1000) within the physical environment (e.g., 706) in which the user (e.g., 708) is located is displayed via movement of a representation (e.g., 726) within the extended reality environment (e.g., 740). In some embodiments, displaying movement of the representation (e.g., 726) within the extended reality environment (e.g., 740) includes displaying movement of the representation (e.g., 726) that mirrors the physical movement of the user (e.g., 708) relative to a portion of a computer system (e.g., 101, 700, and / or 1000) within the physical environment (e.g., 706) in which the user (e.g., 708) is located.

[0252] Displaying a menu user interface including selectable options allows the computer system to quickly and easily display at least a portion of a user's representation within the extended reality environment, thereby reducing the number of inputs required to display the user's representation within the extended reality environment.

[0253] 10A-10I illustrate example techniques for adjusting the appearance of a user's expression. FIG. 11 is a flow diagram of an example method 1100 for adjusting the appearance of a user's expression. FIG. 12 is a flow diagram of an example method 1200 for displaying a mouth representation of a user's expression. FIG. 13 is a flow diagram of an example method 1300 for displaying a hair representation of a user's expression. FIG. 14 is a flow diagram of an example method 1400 for displaying a portion of a user's expression with a visual highlight. The user interfaces of FIGS. 10A-10I are used to illustrate processes described below, including the processes of FIGS. 11-14.

[0254] 10A-10I show an example of an electronic device 1000 displaying a representation 1002 of one or more parts of a body of a user 1004 having different appearances based on information received by the electronic device 1000. FIGS. 10A-10I also show an example of an electronic device 1000 displaying, via a display 1000a, a communication interface 1006 that includes a first participant area 1006a corresponding to the user 1004 and a second participant area 1006b corresponding to a second user (e.g., a second user associated with and / or using the electronic device 1000). In FIG. 10A, the first participant area 1006a includes an extended reality environment 1008, as well as a representation 1002 of the user 1004 within the extended reality environment 1008 and a table representation 1010 (e.g., an image representing a virtual table and / or an image representing a table 1016 in a physical environment 1014). Additionally, second participant area 1006b includes a second representation 1012 of the second user (e.g., an avatar and / or image representing the second user). In some embodiments, user 1004 is participating in a real-time communication session, such as a videoconference and / or virtual videoconference, with the second user (e.g., electronic device 1000 communicates with an external electronic device of user 1004, allowing user 1004 and / or the second user to communicate with each other via audio, video, and / or images displayed on electronic device 1000 and / or the external electronic device).

[0255] 10A-10I also show user 1004 within a physical environment 1014 (e.g., an actual environment in which user 1004 is physically located), where physical environment 1014 includes user 1004 and a table 1016 (e.g., a physical table). Electronic device 1000 communicates with sensors 1018a and 1018b positioned within physical environment 1014 (e.g., wireless communication via an external electronic device associated with and / or used by user 1004). In some embodiments, sensors 1018a and 1018b include cameras, image sensors, light sensors, depth sensors, tactile sensors, orientation sensors, proximity sensors, temperature sensors, location sensors, motion sensors, and / or velocity sensors. Sensors 1018a and 1018b are configured to capture data and / or information regarding the state (e.g., position, orientation, posture, and / or pose) of user 1004 within physical environment 1014. For example, sensors 1018a and 1018b are configured to detect and capture information related to the position and / or movement of various body parts of user 1004 within physical environment 1014. Although Figures 10A-10I show electronic device 1000 in communication with two sensors (e.g., sensor 1018a and sensor 1018b), in some embodiments, electronic device 1000 is in communication with any suitable number of sensors (e.g., via an external electronic device associated with user 1004).

[0256] While FIGS. 10A-10I show electronic device 1000 displaying representation 1002 of user 1004, in some embodiments, electronic device 700 displays communication interface 1006 including representation 1002 of user 1004 via display 704. In some embodiments, electronic device 1000 is configured to capture one or more physical characteristics of user 708 (and / or user 1004) to generate representation 726 (and / or representation 1002), and / or display representation 726 (and / or representation 1002) on display 1000a of electronic device 1000, as described above with reference to FIGS. 7A-7J. In some embodiments, the same electronic device (e.g., electronic device 700 and / or electronic device 1000) is used to generate and display representation 726, as described above with reference to FIGS. 7A-7J, and is used to display representation 1002, as described below with reference to FIGS. 10A-10I.

[0257] 10A , electronic device 1000 receives information (e.g., via sensors 1018a and / or 1018b and / or via an external device) indicative of the status of one or more body parts of user 1004 within physical environment 1014. In response to receiving the information, electronic device 1000 displays representation 1002 within extended reality environment 1008 in first participant area 1006a. As shown in FIG. 10A , representation 1002 includes an appearance that mimics the physical appearance of user 1004 within physical environment 1014. For example, first representation 1002 includes hips 1004a, hands 1004b, hands 1004c, legs 1004d, legs 1004e, head 1004f, and face 1004g, which correspond to hips 1002a, hands 1002b, hands 1002c, legs 1002d, legs 1002e, head 1002f, and face 1002g of user 900. Specifically, hand 1002b of representation 1002 is elevated above hips 1002a in extended reality environment 1008, as is hand 1004b of user 1004 in physical environment 1014. The hands 1002c of the representation 1002 are positioned at and / or near the waist 1002a of the representation 1002 in the extended reality environment 1008, similar to the hands 1004c of the user 1004 positioned at and / or near the waist 1004a of the user 1004 in the physical environment 1014.

[0258] 10A , electronic device 1000 receives information indicating a state (e.g., position, orientation, posture, and / or pose) of a body of user 1004. Based on the received information, electronic device 1000 displays representation 1002 as having a first appearance within extended reality environment 1008 (e.g., as indicated by the solid line shown in FIG. 10A ). In FIG. 10A , electronic device 1000 displays representation 1002 with a first amount of visual fidelity and / or without blurring applied to at least a portion of representation 1002. In some embodiments, electronic device 1000 displays representation 1002 as an anatomically accurate representation of user 1004 without applying any amount of blurring to representation 1002. In some embodiments, the received information indicates a state of a body part of user 1004. In some such embodiments, the electronic device 1000 displays a first portion of the representation 1002 that corresponds to a body part of the user 1004 having a first appearance, and a second portion of the representation 1002 that does not correspond to a body part of the user 1004 having a second appearance that is different from the first appearance.

[0259] 10A , electronic device 1000 displays movement of representation 1002 within extended reality environment 1008, as indicated by arrow 1019. If electronic device 1000 receives information indicative of a state of user 1004, and the information indicative of the state of user 1004 includes direct information regarding a state of at least a portion of the user's 1004's body (e.g., information directly captured via sensors 1018a and / or 1018b indicating a position of a portion of the user's 1004's body within physical environment 1014) received within a first predetermined amount of time, electronic device 1000 maintains display of representation 1002 having a first appearance. In some embodiments, the first appearance of representation 1002 does not include transparency (e.g., zero amount of transparency applied to representation 1002), such that portion 1010a of table representation 1010 is obscured and / or otherwise blocked by representation 1002. In some embodiments, as the electronic device 1000 displays the movement of the representation 1002 within the extended reality environment 1008, the electronic device 1000 displays the representation 1002 as obscuring and / or otherwise blocking other portions of the table representation 1010.

[0260] 10B , electronic device 1000 receives information indicative of a state of user 1004 within physical environment 1014. However, the information indicative of the state of user 1004 within physical environment 1014 does not include direct information regarding a state of at least a portion of the body of user 1004 (e.g., information directly captured via sensors 1018a and / or 1018b indicating the position of the portion of the body of user 1004 within physical environment 1014). In some embodiments, electronic device 1000 determines that the information indicative of the state of user 1004 within physical environment 1014 does not include direct information regarding a state of at least a portion of the body of user 1004 and / or has not included direct information regarding a state of at least a portion of the body of user 1004 for a first predetermined amount of time. In some embodiments, the information indicative of the status of the user 1004 in the physical environment 1014 includes instructions and / or additional information that indicate to the electronic device 1000 that direct information regarding the status of at least a portion of the body of the user 1004 is not available and / or has not been available for a first predetermined amount of time.

[0261] 10B , electronic device 1000 displays representation 1002 as having a second appearance different from the first appearance based on direct information regarding the status of at least a portion of a body of user 1004 that has not been received and / or is unavailable for a first predetermined amount of time. In some embodiments, the first predetermined amount of time is greater than a first time threshold, such as 1 second, 5 seconds, 10 seconds, and / or 30 seconds, but less than a second time threshold, such as 45 seconds, 60 seconds, 90 seconds, and / or 120 seconds. In some embodiments, when direct information regarding the status of at least a portion of a body of user 1004 is received (e.g., received by electronic device 1000 and / or received by another electronic device in communication with electronic device 1000) before the first predetermined amount of time has elapsed (e.g., within an amount of time less than the first predetermined amount of time), electronic device 1000 maintains the display of representation 1002 having the first appearance, as shown in FIG. 10A .

[0262] If direct information regarding the state of at least a portion of the user's 1004 body has not been received for a first predetermined amount of time (e.g., not received by the electronic device 1000 and / or not received by another electronic device in communication with the electronic device 1000), the electronic device 1000 displays the representation 1002 having a second appearance, as shown in FIG. 10B . For example, in FIG. 10B , the representation 1002 is shown as being displayed by the electronic device 1000 with a first dashed line to indicate that the electronic device is displaying the representation 1002 in the second appearance. In some embodiments, the second appearance includes displaying the representation 1002 with a second amount of visual fidelity (e.g., precision and / or clarity) and / or with an increased amount of blur compared to the first amount of visual fidelity. In some embodiments, the second appearance includes displaying the representation 1002 with a grain size that is larger than the grain size of the first appearance. Thus, in some embodiments, electronic device 1000 displays a less accurate version of representation 1002 when direct information regarding the status of at least a portion of the body of user 1004 has not been received (e.g., not received by electronic device 1000 and / or not received by another electronic device in communication with electronic device 1000) for a first predetermined amount of time. While FIG. 10B shows the entire representation 1002 as having the second appearance, in some embodiments, electronic device 1000 displays a first portion of representation 1002 (e.g., a portion of representation 1002 corresponding to a portion of the body of user 1004 for which no direct information is received) using the second appearance and a second portion of representation 1002 using the first appearance.

[0263] In some embodiments, if direct information regarding the state of at least a portion of the body of the user 1004 has not been received (e.g., not received by the electronic device 1000 and / or not received by another electronic device in communication with the electronic device 1000) for a first predetermined amount of time, the electronic device 1000 displays the representation 1002 as static and / or stationary within the extended reality environment 1008. For example, in some embodiments, the information regarding the state of the user 1004 and / or the direct information regarding the state of at least a portion of the body of the user 1004 includes information indicative of movement of the user 1004 within the physical environment 1014 (e.g., movement of one or more body parts of the user 1004). In some embodiments, when the electronic device 1000 receives direct information regarding the state of at least a portion of the body of the user 1004 within a time amount that is less than a first predetermined amount of time, the electronic device 1000 displays movement of the representation 1002 based on the direct information regarding the state of at least a portion of the body of the user 1004, which is indicative of physical movement of the user 1004 within the physical environment 1014. However, in some embodiments, when the electronic device does not receive direct information regarding the state of at least a portion of the body of the user 1004 for the first predetermined amount of time, the electronic device 1000 maintains display of the representation 1002 at a position within the extended reality environment 1008 and does not otherwise display movement of the representation 1002 (e.g., even as the user 1004 moves within the physical environment 1014).

[0264] In some embodiments, electronic device 1000 displays representation 1002 with a second appearance while maintaining the general shape of representation 1002. In other words, electronic device 1000 displays representation 1002 with a first appearance and representation 1002 with a second appearance, each as having the same shape (e.g., a shape that resembles the shape and / or silhouette of user 1004 and / or includes an otherwise similar appearance).

[0265] In some embodiments, electronic device 1000 displays movement of representation 1002 having a second appearance within extended reality environment 1008, as indicated by arrow 1021. In some embodiments, representation 1002's second appearance includes a first amount of transparency (e.g., a non-zero amount of transparency applied to representation 1002) such that portion 1010a of table representation 1010 is at least partially visible and / or discernible through representation 1002. In some embodiments, when electronic device 1000 displays movement of representation 1002 within extended reality environment 1008, electronic device 1000 displays other portions of representation 1002 through table representation 1010 while electronic device 1000 is displaying representation 1002 having the second appearance.

[0266] 10C , electronic device 1000 receives information indicative of a state of user 1004 within physical environment 1014. However, the information indicative of the state of user 1004 within physical environment 1014 does not include direct information regarding a state of at least a portion of the body of user 1004 (e.g., information directly captured via sensors 1018a and / or 1018b indicating the position of a portion of the body of user 1004 within physical environment 1014). In some embodiments, electronic device 1000 determines that the information indicative of the state of user 1004 within physical environment 1014 does not include direct information regarding a state of at least a portion of the body of user 1004 and / or has not included direct information regarding a state of at least a portion of the body of user 1004 for a second predetermined amount of time that is longer than the first predetermined amount of time. In some embodiments, the information indicative of the status of the user 1004 in the physical environment 1014 includes instructions and / or additional information that indicates to the electronic device 1000 that direct information regarding the status of at least a portion of the body of the user 1004 is not available and / or has not been available for a second predetermined amount of time.

[0267] 10C , electronic device 1000 displays representation 1002 as having a third appearance, different from the first appearance and the second appearance, based on direct information regarding the status of at least a portion of a body of user 1004 not being received and / or not being available for a second predetermined amount of time. In some embodiments, the second predetermined amount of time is greater than a first time threshold, such as 1 second, 5 seconds, 10 seconds, and / or 30 seconds, and greater than a second time threshold, such as 45 seconds, 60 seconds, 90 seconds, and / or 120 seconds. In some embodiments, if direct information regarding the status of at least a portion of a body of user 1004 is received within a time period prior to the second predetermined amount of time (e.g., within a time period less than the second predetermined amount of time), electronic device 1000 maintains display of representation 1002 having the first appearance, as shown in FIG. 10A , and / or maintains display of representation 1002 having the second appearance, as shown in FIG. 10B . In some embodiments, the electronic device 1000 displays the representation 1002 having the second appearance when direct information regarding at least a portion of the body of the user 1004 has not been received for a first predetermined amount of time, and the electronic device 1000 displays the representation 1002 having the third appearance when direct information regarding at least a portion of the body of the user 1004 has not been received for a second predetermined amount of time (e.g., transitioning from displaying the representation 1002 having the second appearance to displaying the representation 1002 having the third appearance). In some embodiments, the electronic device 1000 displays the representation 1002 having the second appearance when direct information regarding at least a portion of the body of the user 1004 has not been received for a first predetermined amount of time, and the electronic device 1000 displays the representation 1002 having the first appearance when direct information regarding at least a portion of the body of the user 1004 is received before the second predetermined amount of time but within an amount of time after the first predetermined amount of time has already elapsed (e.g., transitioning from displaying the representation 1002 having the second appearance to displaying the representation having the first appearance).

[0268] If direct information regarding the state of at least a portion of the user's 1004 body has not been received for a second predetermined amount of time (e.g., not received by the electronic device 1000 and / or not received by another electronic device in communication with the electronic device 1000), the electronic device 1000 displays the representation 1002 having a third appearance, as shown in FIG. 10C . For example, in FIG. 10C , the representation 1002 is shown as being displayed by the electronic device 1000 with a second dashed line to indicate that the electronic device is displaying the representation 1002 having the third appearance. In some embodiments, the third appearance includes displaying the representation 1002 with a third amount of visual fidelity (e.g., precision and / or clarity) and / or with an increased amount of blur compared to the first amount of visual fidelity and / or the second amount of visual fidelity. In some embodiments, the third appearance includes displaying the representation 1002 with a grain size larger than the grain size of the first appearance and / or larger than the grain size of the second appearance. Thus, in some embodiments, electronic device 1000 displays a less accurate version of representation 1002 when direct information regarding the state of at least a portion of the body of user 1004 has not been received by electronic device 1000 for a second predetermined amount of time. While FIG. 10C shows the entire representation 1002 as having the third appearance, in some embodiments, electronic device 1000 displays a first portion of representation 1002 having the third appearance (e.g., a portion of representation 1002 corresponding to a portion of the body of user 1004 for which direct information is not received) and a second portion of representation 1002 having the first appearance and / or the second appearance.

[0269] In some embodiments, if direct information regarding the state of at least a portion of the body of the user 1004 has not been received (e.g., not received by the electronic device 1000 and / or not received by another electronic device in communication with the electronic device 1000) for a second predetermined amount of time, the electronic device 1000 displays the representation 1002 in a presentation mode. In some embodiments, the presentation mode includes displaying the representation 1002 as a blurred circle and / or other non-anatomically accurate representation of the user 1004. In some embodiments, the presentation mode includes displaying the representation 1002 in an audio presence mode, where the representation 1002 includes an icon and / or monogram having an appearance based on detected utterances of the user 1004 in the physical environment 1014. In some embodiments, the presentation mode includes displaying the representation 1002 as having a shape that is not visually responsive to changes in the movement of the user 1004. In some embodiments, the presentation mode includes displaying the representation 1002 at a size that is smaller than the size of the representation 1002 when displayed using the first appearance and / or the second appearance.

[0270] In some embodiments, the electronic device 1000 maintains the display of the representation 1002 having the third appearance when direct information regarding the status of at least a portion of the body of the user 1004 is not received for a time amount that is longer than a second predetermined amount of time. In other words, the electronic device 1000 maintains the display of the representation 1002 having the third appearance unless direct information regarding the status of at least a portion of the body of the user 1004 is received after the second predetermined amount of time has elapsed. In some embodiments, the electronic device 1000 transitions from displaying the representation 1002 having the third appearance to displaying the representation 1002 having the first appearance upon receiving direct information regarding the status of at least a portion of the body of the user 1004.

[0271] In some embodiments, electronic device 1000 displays movement of representation 1002 having a third appearance within extended reality environment 1008, as indicated by arrow 1023. In some embodiments, representation 1002's third appearance includes a second amount of transparency (e.g., a non-zero amount of transparency applied to representation 1002 that is greater than the first amount of transparency) such that portion 1010a of table representation 1010 is at least partially visible and / or discernible through representation 1002. In some embodiments, portion 1010a of table representation 1010 is more visible and / or discernible through representation 1002 when electronic device 1000 displays representation 1002 having the third appearance compared to displaying representation 1002 having the second appearance. In some embodiments, when electronic device 1000 displays movement of representation 1002 within extended reality environment 1008, electronic device 1000 displays table representation 1010 to other portions of representation 1002 while electronic device 1000 is displaying representation 1002 having a third appearance.

[0272] 10D , the electronic device 1000 displays a zoomed-in view of the representation 1002 in the extended reality environment 1008 in the first participant region 1006a. In some embodiments, the electronic device 1000 zooms the first participant region 1006a to a particular portion of the extended reality environment 1008 in response to user input (e.g., a tap gesture, a voice command, and / or an air gesture). In some embodiments, the electronic device 1000 zooms the first participant region 1006a to a particular portion of the extended reality environment 1008 when a condition is met, such as when the user 1004 is outputting (e.g., speaking and / or producing) an utterance (e.g., speech, humming, vocalizing, and / or other sounds otherwise verbally produced).

[0273] 10D , user 1004 is outputting speech 1020 ("Hi Jane, how are you today?") within physical environment 1014. Accordingly, user's mouth 1004h is open, indicating user 1004 is speaking speech 1020. In FIG. 10D , electronic device 1000 receives information indicative of the state of user 1004's mouth 1004h within physical environment 1014. Additionally, electronic device 1000 receives audio information (e.g., via a speaker of electronic device 1000 associated with user 1004 and / or via sensors 1018a and / or 1018b) indicating user 1004 is outputting speech 1020. Based on received information indicating the state of the user's 1004 mouth 1004h and / or based on received audio information indicating that the user 1004 is outputting speech 1020, the electronic device 1000 displays a representation 1002 having a first appearance and with the mouth 1002h in an open position, as shown in FIG. 10D.

[0274] 10D , electronic device 1000 displays mouth 1002h of representation 1002 with a first amount of visual fidelity (e.g., precision and / or clarity). In some embodiments, electronic device 1000 displays mouth 1002h of representation 1002 as an anatomically accurate representation of mouth 1004h of user 1004 without applying any amount of blurring to mouth 1002h.

[0275] In some embodiments, electronic device 1000 displays mouth 1002h of representation 1002 in an open position based on received information indicating the state of mouth 1004h of user 1004 and not based on received audio information indicating that user 1004 is outputting speech 1020. In some embodiments, electronic device 1000 outputs audio corresponding to speech 1020 via a speaker while displaying mouth 1002h of representation 1002 in the open position.

[0276] 10E, electronic device 1000 receives information indicative of the state of user 1004's mouth 1004h, but the information indicative of the state of user 1004's mouth 1004h does not satisfy a set of one or more criteria. For example, in some embodiments, the set of one or more criteria may include a first criterion that is satisfied when information indicative of the state of user 1004's mouth 1004h is received within a predetermined time period (e.g., within a recurring predetermined time interval, such as every 1 second, every 5 seconds, and / or every 10 seconds); a second criterion that is met when the information indicative of the state of the user's 1004's mouth 1004h includes information indicative of movement of the user's 1004's mouth 1004h, and / or a third criterion that is met when the information indicative of the state of the user's 1004's mouth 1004h includes an amount of accuracy greater than a threshold amount of accuracy (e.g., the information includes data indicative of the position, pose, orientation, and / or facial expression of the mouth 1004h that is greater than a confidence level threshold determined at least in part based on the amount of information, the amount of information received over time, and / or the accuracy and / or precision of the information in detecting and / or estimating the actual state of the mouth 1004h).

[0277] When the information indicative of the state of the mouth 1004h of the user 1004 does not meet one or more sets of criteria, the electronic device 1000 displays the mouth 1002p of the representation 1002 having a second appearance. For example, in FIG. 10E , the mouth 1002p of the representation 1002 is shown as being displayed by the electronic device 1000 with a dashed line to indicate that the electronic device is displaying the mouth 1002p of the representation 1002 with the second appearance. In some embodiments, the second appearance includes displaying the mouth 1002p of the representation 1002 with a second amount of visual fidelity (e.g., precision and / or clarity) and / or with an increased amount of blur compared to the first amount of visual fidelity.

[0278] In some embodiments, the second appearance includes displaying mouth 1002p of representation 1002 based at least in part on audio information corresponding to speech 1020. For example, electronic device 1000 displays mouth 1002p of representation 1002 to include a particular pose, orientation, facial expression, and / or position based at least in part on audio information corresponding to speech 1020 (e.g., an estimated, extrapolated, and / or predicted pose, orientation, facial expression, and / or position based on audio information corresponding to speech 1020). In some embodiments, electronic device 1000 displays mouth 1002p of representation 1002 based on both audio information corresponding to speech 1020 and information indicative of a state of mouth 1004h of user 1004 when the information indicative of a state of mouth 1004h does not satisfy one or more sets of criteria. In some embodiments, when the information indicative of the state of mouth 1004h does not satisfy one or more sets of criteria, electronic device 1000 displays mouth 1002p of representation 1002 as a combination of a first portion (e.g., mouth 1002h having a first appearance) generated based on the information indicative of the state of user 1004's mouth 1004h and a second portion generated based on audio information corresponding to speech 1020. For example, in some embodiments, the first portion and second portion are combined, overlaid on each other, and / or otherwise used to generate mouth 1002p of representation 1002 displayed by electronic device 1000. In some embodiments, the first portion is a static representation and the second portion is a dynamic representation. In some embodiments, both the first portion and the second portion are dynamic representations.

[0279] In some embodiments, mouth 1002p of representation 1002 includes different amounts and / or degrees of emphasis of a first portion and a second portion based on information indicative of a state of mouth 1004h of user 1004. For example, in some embodiments, if one or more sets of criteria are not met and it is determined (e.g., via electronic device 1000 and / or via another electronic device associated with user 1004) that the information indicative of the state of mouth 1004h of user 1004 includes a confidence level less than a confidence level threshold, mouth 1002p of representation 1002 is generated using a first portion having a first amount of emphasis (e.g., a first visual emphasis amount and / or a first weighting) and a second portion having a second amount of emphasis (e.g., a second visual emphasis amount and / or a second weighting) that is greater than the first amount of emphasis. Similarly, in some embodiments, if one or more sets of criteria are not met and the information indicative of the state of user's 1004's mouth 1004h is determined (e.g., via electronic device 1000 and / or via another electronic device associated with user 1004) to include a confidence level greater than the confidence level threshold, mouth 1002p of representation 1002 is generated using a first portion having a third amount of emphasis (e.g., a third amount of visual emphasis and / or a third weight) and a second portion having a fourth amount of emphasis (e.g., a fourth amount of visual emphasis and / or a fourth weight) that is less than the third amount of emphasis. In some embodiments, electronic device 1000 changes and / or updates the display of mouth 1002p of representation 1002 as the confidence level of the information indicative of the state of mouth 1004h changes.

[0280] In some embodiments, the confidence level of the information indicative of the state of the mouth 1004h is determined based on audio information corresponding to the speech 1020 (e.g., via the electronic device 1000 and / or via another electronic device associated with the user 1004). For example, in some embodiments, the information indicative of the state of the mouth 100...

Claims

1. 1. A method comprising: A computer system in communication with one or more display generation components, comprising: while the computer system is disposed on the user's body, displaying, via the one or more display generating components, a prompt instructing the user to remove the computer system from the user's body and use the computer system to capture information relevant to the user; detecting that the computer system has been removed from the user's body after displaying the prompt instructing the user to remove the computer system from the user's body; capturing information associated with the user after detecting that the computer system has been removed from the body of the user, the computer system being configured to use the information to generate a representation of the user; and A method comprising:

2. The method of claim 1 , wherein the representation of the user is configured to be displayed in an augmented reality and / or virtual reality environment.

3. The method of claim 1 , wherein the computer system is configured to generate the representation of the user in three dimensions.

4. 10. The method of claim 1, further comprising providing instructions for capturing the information associated with the user using the computer system before detecting that the computer system has been removed from the body of the user.

5. The method of claim 4 , wherein providing the instructions includes displaying, via the one or more display generating components, an animation demonstrating capturing the information about the user using the computer system.

6. before detecting that the computer system has been removed from the body of the user; displaying, via the one or more display generation components, an indication associated with a condition affecting the capture of information associated with the user in accordance with a determination that a set of criteria corresponding to the condition affecting the capture of information associated with the user is satisfied; 2. The method of claim 1, further comprising: discontinuing displaying the indication associated with the condition affecting the capture of information associated with the user in accordance with a determination that the set of criteria is not satisfied.

7. The method of claim 6 , wherein the indication associated with the condition affecting the capture of information related to the user includes information regarding taking an action to help remedy the condition.

8. 10. The method of claim 1, further comprising initiating a process of capturing the information associated with the user in response to detecting that the computer system has been removed from the body of the user.

9. 10. The method of claim 1, further comprising providing a second prompt including instructions to capture the information associated with the user after detecting that the computer system has been removed from the body of the user.

10. The method of claim 9 , wherein providing the second prompt comprises displaying, via the one or more display generating components, a visual prompt along with one or more enrollment instructions.

11. the prompt to remove the computer system from the body of the user and to use the computer system to capture information relevant to the user is displayed via a first display generating component of the one or more display generating components; The method of claim 10 , wherein the visual prompt is displayed via a second one of the one or more display generating components that is different from the first display generating component.

12. 10. The method of claim 9, wherein providing the second prompt comprises outputting an audio prompt along with one or more enrollment instructions via an audio device in communication with the computer system.

13. 10. The method of claim 9, wherein providing the second prompt includes providing an indication to the computer system instructing the user to orient the body portion of the user within a target location.

14. The method of claim 9 , wherein providing the second prompt comprises providing an indication instructing the user to adjust a condition affecting the capture of the user's information.

15. The method of claim 9 , wherein providing the second prompt comprises providing an indication to the user to shift a position of the user's head.

16. 10. The method of claim 9, wherein providing the second prompt includes providing an indication instructing the user to position one or more sets of the user's facial features within a predefined set of one or more facial expressions.

17. 10. The method of claim 9, wherein providing the second prompt includes providing an indication to the user to adjust a position of the computer system to orient the computer system toward a predetermined part of the user's body.

18. the prompt to remove the computer system from the body of the user and to capture information relevant to the user using the computer system is displayed via a first display generating component of the one or more display generating components, the method comprising:

2. The method of claim 1, further comprising, after capturing the information related to the user, displaying a preview of the representation of the user via a second display generating component different from the first display generating component of the one or more display generating components.

19. detecting when the computer system is placed on the body of the user after capturing the information related to the user; 10. The method of claim 1, further comprising: displaying a preview of the representation of the user via the one or more display generation components after detecting that the computer system has been placed on the body of the user.

20. Capturing the information related to the user includes capturing first information related to a first part of a body of the user, the method comprising: detecting that the computer system has been placed on the body of the user after capturing the first information related to the first portion of the body of the user; 10. The method of claim 1, further comprising: after detecting that the computer system has been placed on the body of the user, initiating a process of capturing second information related to a second part of the body of the user, different from the first part of the body of the user.

21. 21. The method of claim 20, wherein initiating the process of capturing the second information related to the second part of the body of the user includes displaying, via the one or more display generation components, a visual indication of a location for the user to position the second part of the body of the user.

22. 21. The method of claim 20, wherein initiating the process of capturing the second information related to the second part of the body of the user includes providing a prompt instructing the user to adjust an orientation of the second part of the body of the user.

23. The method of claim 20 , further comprising displaying the representation of the user within an extended reality environment via the one or more display generation components after capturing the second information related to the second part of the body of the user.

24. The method of claim 1, wherein capturing the information related to the user includes capturing the information related to the user via one or more sensors in communication with the computer system, and the computer system is configured to generate the representation of the user using the information captured via the one or more sensors.

25. The information captured via the one or more sensors includes information regarding the appearance of individual parts of the user; the representation of the user generated based on the information captured via the one or more sensors has an appearance that is automatically determined based on the appearance of the distinct parts of the user represented by the information captured via the one or more sensors.

25. The method of claim 24.

26. A computer program product causing a computer to carry out a method according to any one of claims 1 to 25.

27. A memory storing a computer program according to claim 26; one or more processors capable of executing the computer programs stored in the memory; A computer system comprising: the computer system is configured to communicate with one or more display generation components; Computer system.

28. Means for carrying out the method according to any one of claims 1 to 25. A computer system comprising:

Citation Information

Patent Citations

  • Vapor source

    JP1977019184A

  • Augmented reality display with frame modulation functionality

    JP2021527998A

  • JPP6535699B

  • Information processing device, information processing method, information processing program, terminal device, terminal device control method, and control program

    WO2019220742A1

  • Information processing apparatus, information processing method, computer program, and augmented reality sense system

    WO2021145067A1