Presenting avatars in a three-dimensional environment
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-12-08
- Publication Date
- 2026-03-25
Smart Images

Figure 2026053391000001_ABST
Abstract
Description
Technical Field
[0001] (Cross - Reference to Related Applications) This application claims priority to U.S. Provisional Patent Application No. 63 / 036,411, filed on June 8, 2020, entitled "Presenting Avatars in Three - Dimensional Environments", the content of which is incorporated herein by reference.
[0002] The present disclosure generally relates to a display generation component, and optionally, one or more input devices that provide a computer - generated experience, which communicate with a computer system including, but not limited to, an electronic device that provides virtual reality and mixed reality experiences via a display.
Background Art
[0003] The development of computer systems for augmented reality has advanced significantly in recent years. Exemplary augmented reality environments include at least some virtual elements that replace or enhance the physical world. Input devices such as cameras, controllers, joysticks, touch - sensitive surfaces, and touch - screen displays for computer systems and other electronic computing devices are used to interact with virtual / augmented reality environments. Exemplary virtual elements include virtual objects such as digital images, videos, text, icons, and control elements such as buttons and other graphics.
Summary of the Invention
[0004] Some methods and interfaces for interacting with environments that include at least some virtual elements (e.g., applications, augmented reality environments, mixed reality environments, and virtual reality environments) are cumbersome, inefficient, and restrictive. For example, systems that provide insufficient feedback for performing actions associated with virtual objects, systems that require a series of inputs to achieve desired results in augmented reality environments, and systems where manipulating virtual objects is complex and error-prone impose a significant cognitive burden on the user and detract from the virtual / augmented reality experience. In addition, these methods are unnecessarily time-consuming, thereby wasting the energy of the computer system. This latter consideration is particularly important in battery-powered devices.
[0005] Therefore, there is a need for computer systems with improved methods and interfaces to provide users with computer-generated experiences that make interaction with the computer system more efficient and intuitive for the user. Such methods and interfaces optionally complement or replace conventional methods of providing users with computer-generated reality experiences. Such methods and interfaces reduce the number, extent, and / or types of user input by helping the user understand the connection between the inputs provided and the device response to those inputs, thereby generating a more efficient human-machine interface.
[0006] The above-mentioned drawbacks and other problems associated with the user interface of a computer system are mitigated or eliminated by the disclosed system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a notebook computer, tablet computer, or handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device such as a wristwatch or head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also known as a “touchscreen” or “touchscreen display”). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, the computer system has one or more output devices in addition to a display generation component, the output devices include one or more tactile output generators and / or one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, memory, and one or more modules, programs, or instruction sets stored in memory for performing multiple functions. In some embodiments, the user interacts with the GUI (and / or computer system) through stylus and / or finger touch and gestures on a touch-sensitive surface, the user's eye and hand movements in space relative to the user's body as captured by a camera and other motion sensors, and voice input as captured by one or more audio input devices.In some embodiments, the functions performed through interaction optionally include image editing, drawing, presentation, word processing, spreadsheet creation, gameplay, making phone calls, video conferencing, sending emails, instant messaging, training support, digital photography, digital videography, web browsing, digital music playback, note-taking, and / or digital video playback. The executable instructions for performing those functions optionally reside in a primary computer-readable storage medium and / or a non-primary computer-readable storage medium, or in other computer program products configured to be executed by one or more processors.
[0007] There is a need for electronic devices with improved methods and interfaces for interacting with three-dimensional environments. Such methods and interfaces can complement or replace conventional methods for interacting with three-dimensional environments. Such methods and interfaces reduce the number, degree, and / or type of user input, resulting in a more efficient human-machine interface. In the case of battery-operated computing devices, such methods and interfaces conserve power and extend the interval between battery charges.
[0008] There is a need for electronic devices with improved methods and interfaces for interacting with other users in a three-dimensional environment using avatars. Such methods and interfaces map, complement, or replace conventional methods for interacting with other users in a three-dimensional environment using avatars. Such methods and interfaces reduce the number, extent, and / or types of user input, resulting in a more efficient human-machine interface.
[0009] It should be noted that the various embodiments described herein can be combined with any other embodiments described herein. The functions and advantages described herein are not exhaustive, and many additional functions and advantages will become apparent to those skilled in the art, in particular, in light of the drawings, specification and claims. Furthermore, it should be noted that the language used herein has been selected solely for readability and explanatory purposes and not to define or limit the subject matter of the invention. [Brief explanation of the drawing]
[0010] To better understand the various embodiments described, the following “Modes for Carrying Out the Invention” should be referenced in conjunction with the following drawings, and similar reference numbers throughout the following drawings refer to the corresponding parts.
[0011] [Figure 1] This document describes the operating environment of a computer system for providing a CGR (Continuous Grooving) experience in several embodiments.
[0012] [Figure 2] Block diagram showing a controller for a computer system configured to manage and adjust the user's CGR experience, according to several embodiments.
[0013] [Figure 3] This block diagram shows a display generation component of a computer system configured to provide a user with visual components of a CGR experience, according to several embodiments.
[0014] [Figure 4] The following are hand-tracking units for computer systems configured to capture user gesture input, according to several embodiments.
[0015] [Figure 5]An eye-tracking unit of a computer system configured to capture a user's gaze input according to some embodiments is shown.
[0016] [Figure 6] A flowchart showing a glint-assisted gaze-tracking pipeline according to some embodiments.
[0017] [Figure 7A] A virtual avatar having display characteristics whose appearance varies based on the certainty of a user's pose according to some embodiments is shown. [Figure 7B] A virtual avatar having display characteristics whose appearance varies based on the certainty of a user's pose according to some embodiments is shown. [Figure 7C] A virtual avatar having display characteristics whose appearance varies based on the certainty of a user's pose according to some embodiments is shown.
[0018] [Figure 8A] A virtual avatar having display characteristics whose appearance varies based on the certainty of a user's pose according to some embodiments is shown. [Figure 8B] A virtual avatar having display characteristics whose appearance varies based on the certainty of a user's pose according to some embodiments is shown. [Figure 8C] A virtual avatar having display characteristics whose appearance varies based on the certainty of a user's pose according to some embodiments is shown.
[0019] [Figure 9] A flowchart showing an exemplary method of presenting a virtual avatar character using display characteristics whose appearance varies based on the certainty of a user's pose according to some embodiments.
[0020] [Figure 10A] A virtual avatar having an appearance based on different appearance templates according to some embodiments is shown. [Figure 10B]This document shows virtual avatars having appearances based on different appearance templates, according to several embodiments.
[0021] [Figure 11A] This document shows virtual avatars having appearances based on different appearance templates, according to several embodiments. [Figure 11B] This document shows virtual avatars having appearances based on different appearance templates, according to several embodiments.
[0022] [Figure 12] This flowchart shows an exemplary method for presenting avatar characters having appearances based on different appearance templates, according to several embodiments. [Modes for carrying out the invention]
[0023] This disclosure relates to user interfaces that provide a computer-generated reality (CGR) experience to a user, in several embodiments.
[0024] The systems, methods, and GUIs described herein improve user interface interactions with virtual / augmented reality environments in multiple ways.
[0025] In some embodiments, the computer system presents the user with an avatar (e.g., in a CGR environment) having display characteristics whose appearance varies based on the certainty of the user's body part's posture. The computer system receives posture data representing the user's body part's posture (e.g., from a sensor) and presents an avatar (e.g., via a display generation component) that includes avatar features corresponding to the user's body part and having variable display characteristics indicating the certainty of the user's body part's posture. The computer system causes the avatar features to have different values of variable display characteristics that depend on the certainty of the user's body part's posture, and an estimated visual fidelity index of the avatar's posture with respect to the user's body part's posture is provided to the user of the computer system.
[0026] In some embodiments, the computer system presents the user with an avatar having an appearance based on different appearance templates that change based on the activity the user is performing. The computer system receives data (e.g., from sensors) indicating that the current activity of one or more users is a first type of activity (e.g., interactive activity). In response, the computer system updates (e.g., via a display generation component) a representation (e.g., an avatar) of the first user having a first appearance based on a first appearance template (e.g., a character template). While the representation of the first user having the first appearance is being presented, the computer system receives second data (e.g., from sensors) indicating the current activity of one or more users. In response, the computer system presents (e.g., via a display generation component) a representation of the first user having a second appearance based on a first appearance template, or a third appearance based on a second appearance template (e.g., an abstract template), depending on whether the current activity is of a first type or a second type, thereby providing the computer system user with an indicator of whether the user is performing a first type or second type activity.
[0027] Figures 1-6 illustrate exemplary computer systems for providing a CGR experience to a user. Figures 7A-7C and 8A-8C show virtual avatar characters having display characteristics whose appearance changes based on the certainty of the user's posture, according to several embodiments. Figure 9 is a flowchart illustrating exemplary methods for presenting virtual avatar characters using display characteristics whose appearance changes based on the certainty of the user's posture, according to various embodiments. Figures 7A-7C and 8A-8C are used to illustrate the process in Figure 9. Figures 10A and 10B, and 11A and 11B, show virtual avatars having appearances based on different appearance templates, according to several embodiments. Figure 12 is a flowchart illustrating exemplary methods for presenting avatar characters having appearances based on different appearance templates, according to several embodiments. Figures 10A-10B and 11A-11B are used to illustrate the process in Figure 12.
[0028] Furthermore, in methods described herein where one or more steps presuppose that one or more conditions are met, it should be understood that the described method can be repeated multiple times in such a way that all the conditions presupposed by the steps of the method are met in different iterations of the method. For example, if the method requires performing a first step if a condition is met, and a second step if the condition is not met, a person skilled in the art will understand that the claimed steps are repeated in any order, both before and after the condition is met. Thus, a method described in one or more steps presupposing that one or more conditions are met may be rewritten as a method that is repeated until each of the conditions described in the method is met. However, this does not require a claim for a system or computer-readable medium in which the system or computer-readable medium includes instructions to perform a presupposed action based on the meeting of the corresponding one or more conditions, and thus it is possible to determine whether the presuppositions are met without explicitly repeating the steps of the method until all the conditions presupposed by the steps of the method are met. Furthermore, those skilled in the art will understand that, as with methods having dependent steps, a system or computer-readable storage medium may repeat the steps of the method as many times as necessary to ensure that all prerequisite steps have been performed.
[0029] In some embodiments, as shown in Figure 1, the CGR experience is provided to the user via an operating environment 100 including a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or remote server), a display generation component 120 (e.g., a head-mounted device (HMD), a display, a projector, a touchscreen, etc.), one or more input devices 125 (e.g., an eye-tracking device 130, a hand-tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a tactile output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a tactile sensor, an orientation sensor, a proximity sensor, a temperature sensor, a position sensor, a motion sensor, a velocity sensor, etc.), and optionally one or more peripheral devices 195 (e.g., a home appliance, a wearable device, etc.). In some embodiments, one or more of the input device 125, output device 155, sensor 190, and peripheral device 195 are integrated with the display generation component 120 (for example, in a head-mounted device or handheld device).
[0030] When describing a CGR experience, various terms are used to refer individually to several related but distinct environments that the user perceives and / or interacts with (for example, using inputs detected by the computer system 101, which causes the computer system generating the CGR experience to generate audio, visual, and / or haptic feedback corresponding to various inputs provided to the computer system 101 that generates the CGR experience). The following is a subset of these terms.
[0031] Physical Environment: The physical environment refers to the physical world that people can perceive and / or interact with without the help of electronic systems. Examples of physical environments, such as a physical park, include physical objects such as physical trees, physical buildings, and physical people. People can directly perceive and / or interact with the physical environment through their senses of sight, touch, hearing, taste, and smell.
[0032] Computer-Generated Reality: In contrast, a computer-generated reality (CGR) environment refers to a fully or partially simulated environment that people perceive and / or interact with through an electronic system. In CGR, a subset of a person's bodily movements or their representations are tracked, and in response, one or more properties of one or more virtual objects simulated within the CGR environment are adjusted to behave according to at least one law of physics. For example, a CGR system may detect a person's head rotation and, in response, adjust the graphic content and sound field presented to the person in a similar manner to how such views and sounds would change in a physical environment. Depending on the circumstances (e.g., for reasons of accessibility), adjustments to the properties(s) of virtual objects(s) in the CGR environment may be made in response to representations of bodily movements (e.g., voice commands). A person may perceive and / or interact with CGR objects using any one of these senses, including sight, hearing, touch, taste, and smell. For example, a person may perceive and / or interact with an audio object that creates a 3D or spatially expansive audio environment, providing the perception of a point source in 3D space. In another example, an audio object may enable audio transparency, selectively incorporating ambient sounds from the physical environment, with or without computer-generated audio. In some CGR environments, a person may perceive and / or interact with only audio objects.
[0033] Examples of CGR include virtual reality and mixed reality.
[0034] Virtual reality: A virtual reality (VR) environment refers to a simulated environment designed to be entirely based on computer-generated sensory input for one or more senses. A VR environment includes multiple virtual objects that a person can perceive and / or interact with. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can perceive and / or interact with virtual objects in a VR environment through a simulation of their presence within the computer-generated environment and / or through a simulation of a subset of their bodily movements within the computer-generated environment.
[0035] Mixed Reality: A mixed reality (MR) environment is a simulated environment designed to incorporate sensory input or its representation from a physical environment, in addition to including computer-generated sensory input (e.g., virtual objects), in contrast to a virtual reality (VR) environment designed to rely entirely on computer-generated sensory input. On a virtual continuum, a mixed reality environment is any place between, but not including, the complete physical environment at one end and the virtual reality environment at the other end. In some MR environments, computer-generated sensory input may respond to changes in sensory input from the physical environment. Also, some electronic systems for presenting an MR environment may track the position and / or orientation relative to the physical environment to enable virtual objects to interact with real objects (i.e., physical articles or their representations from the physical environment). For example, the system may take motion into account so that a virtual tree appears stationary relative to the physical ground.
[0036] Examples of mixed reality include augmented reality and augmented virtual reality.
[0037] Augmented Reality: An augmented reality (AR) environment refers to a simulated environment in which one or more virtual objects are superimposed on or onto a physical environment. For example, an electronic system for presenting an AR environment may have a transparent or translucent display that allows a person to directly view the physical environment. The system may also be configured to present virtual objects on the transparent or translucent display, thereby allowing a person to use the system to perceive the virtual objects superimposed on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture images or videos of the physical environment, which are representations of the physical environment. The system composites the images or videos with the virtual objects and presents the composite on the opaque display. A person uses this system to perceive the virtual objects superimposed on the physical environment by indirectly viewing the physical environment through images or videos of the physical environment. As used herein, videos of the physical environment shown on an opaque display are referred to as “pass-through videos,” and it means that the system uses one or more image sensors to capture images of the physical environment and uses those images when presenting the AR environment on the opaque display. Alternatively, the system may have a projection system that projects virtual objects, for example, as holograms, into or onto the physical environment, thereby allowing a person to perceive the virtual objects superimposed on the physical environment. An augmented reality environment also refers to a simulated environment in which the representation of the physical environment is transformed by computer-generated sensory information. For example, when providing pass-through video, the system may transform one or more sensor images to plane a selected perspective (e.g., viewpoint) different from the perspective captured by the image sensor. As another example, the representation of the physical environment may be transformed by graphically altering (e.g., enlarging) a portion of it, thereby making the altered portion a modified version that represents the original captured image but is not photorealistic. As yet another example, the representation of the physical environment may be transformed by graphically removing or obscuring a portion of it.
[0038] Augmented Virtual: An Augmented Virtual (AV) environment refers to a simulated environment in which a virtual or computer-generated environment incorporates one or more sensory inputs from a physical environment. These sensory inputs may be representations of one or more characteristics of the physical environment. For example, an AV park might have virtual trees and virtual buildings, but people with faces might be realistically reproduced from images of real people. Another example is that a virtual object might adopt the shape or color of a physical article captured by one or more imaging sensors. A further example is that a virtual object might adopt shadows that correspond to the position of the sun in the physical environment.
[0039] Hardware: There are many different types of electronic systems that enable people to perceive and / or interact with various CGR environments. Examples include head-mounted systems, projection-based systems, heads-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be placed over a person's eyes (similar to contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system may have one or more speakers and an integrated opaque display. Alternatively, a head-mounted system may be configured to accept an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing images or videos of the physical environment and / or one or more microphones for capturing sounds of the physical environment. A head-mounted system may have a transparent or translucent display instead of an opaque display. A transparent or translucent display may have a medium through which light representing an image is directed to the human eye. The display may utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination thereof. The medium may be an optical waveguide, a holographic medium, an optical coupler, an optical reflector, or any combination thereof. In one embodiment, the transparent or translucent display may be configured to be selectively opaque. The projection-based system may employ retinal projection technology to project a graphical image onto the human retina. The projection system may also be configured to project virtual objects into the physical environment, for example, as holograms or as physical surfaces. In some embodiments, the controller 110 is configured to manage and adjust the user's CGR experience.In some embodiments, the controller 110 includes a preferred combination of software, firmware, and / or hardware. The controller 110 is described in more detail below with reference to Figure 2. In some embodiments, the controller 110 is a computing device that is local or remote to the scene 105 (e.g., the physical environment). For example, the controller 110 is a local server located within the scene 105. In another example, the controller 110 is a remote server located outside the scene 105 (e.g., a cloud server, a central server, etc.). In some embodiments, the controller 110 is coupled to a display generation component 120 (e.g., an HMD, display, projector, touchscreen, etc.) via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). In another example, the controller 110 is contained within a housing (e.g., a physical housing) of one or more of the display generation components 120 (e.g., a portable electronic device including a display and one or more processors), one or more of the input devices 125, one or more of the output devices 155, one or more of the sensors 190, and / or peripheral devices 195, or shares the same physical housing or support structure as one or more of the above.
[0040] In some embodiments, the display generation component 120 is configured to provide the user with a CGR experience (e.g., at least the visual component of the CGR experience). In some embodiments, the display generation component 120 includes a preferred combination of software, firmware, and / or hardware. The display generation component 120 is described in more detail below with reference to Figure 3. In some embodiments, the functions of the controller 110 are provided by and / or combined with the display generation component 120.
[0041] According to some embodiments, the display generation component 120 provides the user with a CGR experience while the user is virtually and / or physically present in scene 105.
[0042] In some embodiments, the display generation component is mounted on a part of the user's body (e.g., their head or hand). Thus, the display generation component 120 includes one or more CGR displays provided for displaying CGR content. For example, in various embodiments, the display generation component 120 surrounds the user's field of view. In some embodiments, the display generation component 120 is a handheld device (such as a smartphone or tablet) configured to present CGR content, and the user holds the device, which has a display directed towards the user's field of view and a camera directed towards scene 105. In some embodiments, the handheld device is optionally placed in a housing mounted on the user's head. In some embodiments, the handheld device is optionally placed on a support in front of the user (e.g., a tripod). In some embodiments, the display generation component 120 is a CGR chamber, housing, or room configured to present CGR content when the user is not wearing or holding the display generation component 120. Many user interfaces described with reference to one type of hardware for displaying CGR content (e.g., a handheld device or a device on a tripod) may be implemented on another type of hardware for displaying CGR content (e.g., an HMD or other wearable computing device). For example, a user interface showing interaction with CGR content triggered based on interaction occurring in the space in front of a handheld or tripod-mounted device may be implemented similarly to an HMD where the interaction occurs in the space in front of the HMD and the CGR content response is displayed via the HMD. Similarly, a user interface showing interaction with CGR content triggered based on the movement of a handheld or tripod-mounted device relative to the physical environment (e.g., Scene 105 or a part of the user's body (e.g., the user's eyes, head, or hands)) may be implemented similarly to an HMD where the interaction is triggered by the movement of the HMD relative to the physical environment (e.g., Scene 105 or a part of the user's body (e.g., the user's eyes, head, or hands)).
[0043] While relevant features of the operating environment 100 are shown in Figure 1, those skilled in the art will understand from this disclosure that various other features have been omitted for brevity so as not to obscure more appropriate embodiments of the exemplary embodiments disclosed herein.
[0044] Figure 2 is a block diagram of an example of the controller 110 according to several embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that various other features have been omitted for brevity so as not to obscure more suitable embodiments of the embodiments disclosed herein. Therefore, as a non-limiting example, in some embodiments, the controller 110 includes one or more processing units 202 (e.g., a microprocessor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), graphics processing unit (GPU), central processing unit (CPU), processing core, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FireWire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global Mobile Communication System (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), Bluetooth, ZiGBEE, or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these and various other components.
[0045] In some embodiments, one or more communication buses 204 include circuits that interconnect system components and control communication between system components. In some embodiments, one or more I / O devices 206 include at least one of the following: a keyboard, mouse, touchpad, joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.
[0046] Memory 220 includes high-speed random-access memory such as dynamic random-access memory (DRAM), static random-access memory (SRAM), double-data-rate random-access memory (DDRRAM), or other random-access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. Memory 220 includes a non-temporary computer-readable storage medium. In some embodiments, memory 220, or the non-temporary computer-readable storage medium of memory 220, stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 230 and CGR experience module 240.
[0047] The operating system 230 includes instructions for handling various basic system services and instructions for performing hardware-dependent tasks. In some embodiments, the CGR experience module 240 is configured to manage and coordinate one or more CGR experiences for one or more users (e.g., a single CGR experience for one or more users, or multiple CGR experiences for each group of one or more users). For this purpose, in various embodiments, the CGR experience module 240 includes a data acquisition unit 241, a tracking unit 242, a coordination unit 246, and a data transmission unit 248.
[0048] In some embodiments, the data acquisition unit 241 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the display generation component 120 in Figure 1, and optionally one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. For this purpose, in various embodiments, the data acquisition unit 241 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0049] In some embodiments, the tracking unit 242 is configured to map scene 105 and track the position of at least the display generation component 120 relative to scene 105 in Figure 1, and optionally one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. For this purpose, in various embodiments, the tracking unit 242 includes instructions and / or logic therefor, as well as heuristics and metadata therefor. In some embodiments, the tracking unit 242 includes a hand tracking unit 244 and / or an eye tracking unit 243. In some embodiments, the hand tracking unit 244 is configured to track the position of one or more parts of the user's hand and / or the movement of one or more parts of the user's hand relative to scene 105 in Figure 1, relative to the display generation component 120, and / or relative to a coordinate system defined for the user's hand. The hand tracking unit 244 is described in more detail below with reference to Figure 4. In some embodiments, the eye-tracking unit 243 is configured to track the position and movement of the user's gaze (or, more broadly, the user's eyes, face, or head) relative to the scene 105 (e.g., the physical environment and / or the user (e.g., the user's hands)) or to the CGR content displayed via the display generation component 120. The eye-tracking unit 243 is described in more detail below with reference to Figure 5.
[0050] In some embodiments, the adjustment unit 246 is configured to manage and adjust the CGR experience presented to the user by the display generation component 120 and optionally by one or more of the output devices 155 and / or peripheral devices 195. For this purpose, in various embodiments, the adjustment unit 246 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0051] In some embodiments, the data transmission unit 248 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the display generation component 120, and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral devices 195. For this purpose, in various embodiments, the data transmission unit 248 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0052] While the data acquisition unit 241, tracking unit 242 (including, for example, an eye-tracking unit 243 and a hand-tracking unit 244), adjustment unit 246, and data transmission unit 248 are shown as residing on a single device (e.g., a controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 241, tracking unit 242 (including, for example, an eye-tracking unit 243 and a hand-tracking unit 244), adjustment unit 246, and data transmission unit 248 may be located in separate computing devices.
[0053] Furthermore, Figure 2 is intended to illustrate the functionality of various features that may be present in a particular embodiment, in contrast to the structural schematics of the embodiments described herein. As will be recognized by those skilled in the art, the separately shown items can be combined, and some items can be separated. For example, several functional modules shown separately in Figure 2 can be implemented within a single module, and the various functions of a single functional block can be performed by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of certain functions and how functions are assigned between them, will vary depending on the implementation and, in some embodiments, will partially depend on a particular combination of hardware, software, and / or firmware selected for a particular implementation.
[0054] Figure 3 is a block diagram of an example of a display generation component 120 according to several embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that various other features have been omitted for brevity so as not to obscure more suitable embodiments of the embodiments disclosed herein. For that purpose, in some non-limiting examples, the display generation component 120 (e.g., HMD) may include one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FireWire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, infrared, Bluetooth, ZiGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more CGR displays 312, one or more optional in-facing and / or out-facing image sensors 314, memory 320, and one or more communication buses 304 for interconnecting these and various other components.
[0055] In some embodiments, one or more communication buses 304 include circuits that interconnect system components and control communication between system components. In some embodiments, one or more I / O devices and sensors 306 include at least one of the following: an inertial measuring unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, one or more depth sensors (e.g., structured light, time of flight, etc.).
[0056] In some embodiments, one or more CGR displays 312 are configured to provide a CGR experience to the user. In some embodiments, one or more CGR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface conduction electron emission display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), MEMS, and / or similar display types. In some embodiments, one or more CGR displays 312 correspond to waveguide displays such as diffraction, reflection, polarization, and holographic. For example, a display generation component 120 (e.g., HMD) includes a single CGR display. In another embodiment, the display generation component 120 includes a CGR display for each of the user's eyes. In some embodiments, one or more CGR displays 312 can present MR or VR content.
[0057] In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face, including the user's eyes (and may be referred to as an eye-tracking camera). In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hands and optionally a portion of the user's arms (and may be referred to as a hand-tracking camera). In some embodiments, one or more image sensors 314 are configured to face forward to acquire image data corresponding to a scene that the user would view if a display generation component 120 (e.g., an HMD) were not present (and may be referred to as a scene camera). One or more optional image sensors 314 may include one or more RGB cameras (e.g., complementary metal-oxide-semiconductor (CMOS) image sensors or charge-coupled device (CCD) image sensors), one or more infrared (IR) cameras, one or more event-based cameras, and / or similar.
[0058] Memory 320 includes high-speed random-access memory, such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. Memory 320 includes a non-temporary computer-readable storage medium. In some embodiments, memory 320, or the non-temporary computer-readable storage medium of memory 320, stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 330 and CGR presentation module 340.
[0059] The operating system 330 includes instructions for handling various basic system services and instructions for performing hardware-dependent tasks. In some embodiments, the CGR presentation module 340 is configured to present CGR content to the user via one or more CGR displays 312. For this purpose, in various embodiments, the CGR presentation module 340 includes a data acquisition unit 342, a CGR presentation unit 344, a CGR map generation unit 346, and a data transmission unit 348.
[0060] In some embodiments, the data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the controller 110 in Figure 1. To this end, in various embodiments, the data acquisition unit 342 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0061] In some embodiments, the CGR presentation unit 344 is configured to present CGR content via one or more CGR displays 312. For this purpose, in various embodiments, the CGR presentation unit 344 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0062] In some embodiments, the CGR map generation unit 346 is configured to generate a CGR map (e.g., a 3D map of a mixed reality scene or a map of a physical environment on which computer-generated objects can be placed) based on media content data. For this purpose, in various embodiments, the CGR map generation unit 346 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0063] In some embodiments, the data transmission unit 348 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the controller 110 and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. For this purpose, in various embodiments, the data transmission unit 348 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0064] Although the data acquisition unit 342, CGR presentation unit 344, CGR map generation unit 346, and data transmission unit 348 are shown as residing on a single device (e.g., the display generation component 120 in Figure 1), it should be understood that in other embodiments, any combination of the data acquisition unit 342, CGR presentation unit 344, CGR map generation unit 346, and data transmission unit 348 may be located in separate computing devices.
[0065] Furthermore, Figure 3 is intended to illustrate the functionality of various features that may be present in a particular implementation, in contrast to the structural schematics of the embodiments described herein. As will be recognized by those skilled in the art, the separately shown items can be combined, and some items can be separated. For example, several functional modules shown separately in Figure 3 can be realized within a single module, and the various functions of a single functional block can be performed by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of certain functions and how functions are assigned between them, will vary depending on the implementation and, in some embodiments, will partially depend on a particular combination of hardware, software, and / or firmware selected for a particular implementation.
[0066] Figure 4 is a schematic diagram of an exemplary embodiment of the hand tracking device 140. In some embodiments, the hand tracking device 140 (Figure 1) is controlled by a hand tracking unit 244 (Figure 2) to track the position of one or more parts of the user's hand and / or the movement of one or more parts of the user's hand relative to a coordinate system defined for the scene 105 in Figure 1 (e.g., relative to parts of the physical environment surrounding the user, relative to the display generation component 120, or relative to parts of the user (e.g., the user's face, eyes, or head) and / or relative to the user's hand. In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).
[0067] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras) that captures three-dimensional scene information including at least the hand 406 of a human user. The image sensor 404 captures a hand image with sufficient resolution to allow for the distinction of fingers and their respective positions. The image sensor 404 can typically capture images of other parts of the user's body, or images of the entire body, and may have either a zoom function or a dedicated sensor with high magnification to capture an image of the hand at a desired resolution. In some embodiments, the image sensor 404 also captures a 2D color video image of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with other image sensors that capture the physical environment of the scene 105, or functions as an image sensor that captures the physical environment of the scene 105. In some embodiments, the image sensor 404 is positioned relative to the user or the user's environment such that the field of view of the image sensor or a portion thereof is used to define an interaction space in which hand movements captured by the image sensor are processed as input to the controller 110.
[0068] In some embodiments, the image sensor 404 outputs a sequence of frames containing 3D map data (and possibly color image data) to the controller 110, thereby extracting high-level information from the map data. This high-level information is typically provided to an application running on the controller via an application programming interface (API), which in turn drives the display generation component 120. For example, a user can interact with the software running on the controller 110 by moving their hand 406 and changing the hand's orientation.
[0069] In some embodiments, the image sensor 404 projects a spot pattern onto a scene including the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral shift of the spot in the pattern. This approach is advantageous in that the user does not need to hold or wear any kind of beacon, sensor, or other marker. This gives the depth coordinates of points in the scene relative to a given reference plane at a specific distance from the image sensor 404. In this disclosure, it is assumed that the image sensor 404 defines a set of orthogonal x, y, and z axes such that the depth coordinates of points in the scene correspond to the z component measured by the image sensor. Alternatively, the image sensor 404 (e.g., a hand tracking device) may use other 3D mapping methods such as stereoscopic imaging or time-of-flight measurement based on one or more cameras or other types of sensors.
[0070] In some embodiments, the hand tracking device 140 captures and processes a time sequence of depth maps containing the user's hand while the user moves their hand (e.g., the entire hand or one or more fingers). Software running on the image sensor 404 and / or the processor in the controller 110 processes the 3D map data to extract patch descriptors of the hand within these depth maps. Based on a previous learning process, the software matches these descriptors against patch descriptors stored in the database 408 to estimate the hand pose in each frame. The pose typically includes the 3D positions of the user's wrist and fingertips.
[0071] The software can also analyze the trajectory of the hand and / or fingers across multiple frames in a sequence to identify gestures. The posture estimation function described herein may be interleaved with the motion tracking function, so that patch-based posture estimation is performed only once every two (or more) frames, while tracking is used to detect changes in posture that occur over the remaining frames. Posture, motion, and gesture information is provided to an application program running on the controller 110 via the API described above. This program can, for example, move and modify an image presented on the display generation component 120 in response to the posture and / or gesture information, or perform other functions.
[0072] In some embodiments, the software may be downloaded electronically to the controller 110, for example, over a network, or instead, it may be provided on a tangible non-temporary medium such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is similarly stored in memory associated with the controller 110. Alternatively or additionally, some or all of the computer's described functions may be implemented in dedicated hardware such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP). Although the controller 110 is shown in Figure 4, for example, as a separate unit from the image sensor 404, some or all of the controller's processing functions may be associated with the image sensor 404 by a suitable microprocessor and software, or by dedicated circuitry within the housing of the image sensor 404 (e.g., a hand-tracking device), or in other ways. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with the display generation component 120 (e.g., in a television set, handheld device, or head-mounted device), or by any other suitable computerized device such as a game console or media player. The sensing function of the image sensor 404 can also be integrated into a computer or other computerized device controlled by the sensor output.
[0073] Figure 4 further includes schematic diagrams of depth maps 410 captured by image sensor 404 according to several embodiments. The depth map includes a matrix of pixels, each having a depth value, as described above. Pixels 412 corresponding to the hand 406 are segmented in this map from the background and the wrist. The brightness of each pixel in the depth map 410 is inversely proportional to the depth value, i.e., the measured z-distance from image sensor 404, with the gradation becoming richer as the depth increases. Controller 110 processes these depth values to identify and segment image components (i.e., adjacent pixel groups) that have features of a human hand. These characteristics may include, for example, the overall size, shape, and frame-to-frame motion of the depth map sequence.
[0074] Figure 4 also schematically shows the hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406, according to several embodiments. In Figure 4, the hand skeleton 414 is superimposed on the hand background 416, which has been segmented from the original depth map. In some embodiments, the hand (e.g., finger joints, fingertips, center of the palm, end of the hand connected to the wrist), and optionally major feature points on the wrist or arm connected to the hand, are identified and positioned on the hand skeleton 414. In some embodiments, the position and movement of these major feature points across multiple image frames are used by the controller 110 to determine, according to several embodiments, a hand gesture performed by the hand or the current state of the hand.
[0075] Figure 5 shows an exemplary embodiment of the eye-tracking device 130 (Figure 1). In some embodiments, the eye-tracking device 130 is controlled by an eye-tracking unit 243 (Figure 2) to track the position and movement of the user's gaze relative to the scene 105 or to the CGR content displayed via the display generation component 120. In some embodiments, the eye-tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, if the display generation component 120 is a head-mounted device such as a headset, helmet, goggles, or glasses, or a handheld device positioned in a wearable frame, the head-mounted device includes both a component for generating CGR content for user viewing and a component for tracking the user's gaze relative to the CGR content. In some embodiments, the eye-tracking device 130 is separate from the display generation component 120. For example, if the display generation component is a handheld device or a CGR chamber, the eye-tracking device 130 is optionally a separate device from the handheld device or CGR chamber. In some embodiments, the eye-tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, the head-mounted eye-tracking device 130 is optionally used with a display generation component that is mounted on the head or a display generation component that is not mounted on the head. In some embodiments, the eye-tracking device 130 is optionally used in combination with a head-mounted display generation component, rather than being a head-mounted device. In some embodiments, the eye-tracking device 130 is optionally part of a non-head-mounted display generation component, rather than being a head-mounted device.
[0076] In some embodiments, the display generation component 120 uses a display mechanism (e.g., left and right near-eye display panels) that displays frames containing left and right images in front of the user's eyes to provide the user with a 3D virtual view. For example, the head-mounted display generation component may include left and right optical lenses (referred to herein as eyepieces) positioned between the display and the user's eyes. In some embodiments, the display generation component may include, or be coupled to, one or more external video cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display on which the user can directly view the physical environment and display virtual objects on a transparent or translucent display. In some embodiments, the display generation component projects virtual objects onto the physical environment. The virtual objects are projected, for example, onto a physical surface or as holograms, so that the individual can use the system to observe the virtual objects superimposed on the physical environment. In such cases, separate display panels and image frames for the left and right eyes may not be required.
[0077] As shown in Figure 5, in some embodiments, the eye-tracking device 130 (e.g., gaze tracking device) includes at least one eye-tracking camera (e.g., an infrared (IR) camera or a near-IR (NIR) camera) and an illumination source (e.g., an IR or NIR light source such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eye. The eye-tracking camera may be directed toward the user's eye to receive reflected IR or NIR light from the light source directly from the eye, or alternatively, it may be directed toward a "hot" mirror positioned between the user's eye and a display panel that reflects IR or NIR light from the eye to the eye-tracking camera while allowing visual light to pass through. The eye-tracking device 130 optionally captures images of the user's eyes (e.g., as a video stream captured at 60-120 frames per second (fps)), analyzes the images to generate gaze tracking information, and communicates the gaze tracking information to the controller 110. In some embodiments, both of the user's eyes are tracked separately by their respective eye-tracking cameras and illumination sources. In some embodiments, only one of the user's eyes is tracked by a corresponding eye-tracking camera and light source.
[0078] In some embodiments, the eye-tracking device 130 is calibrated using a device-specific calibration process to determine the parameters of the eye-tracking device for a specific operating environment 100, e.g., the 3D geometric relationships and parameters of the LEDs, camera, hot mirror (if present), eyepiece, and display screen. The device-specific calibration process may be performed at the factory or another facility before delivery of the AR / VR device to the end user. The device-specific calibration process may be an automated calibration process or a manual calibration process. The user-specific calibration process may include estimating the eye parameters of a particular user, e.g., pupil position, central visual position, optical axis, visual axis, interpupillary distance. According to some embodiments, once the device-specific and user-specific parameters for the eye-tracking device 130 are determined, the images captured by the eye-tracking camera can be processed using a glint-assisted method to determine the user's current visual axis and viewpoint relative to the display.
[0079] As shown in Figure 5, the eye-tracking device 130 (e.g., 130A or 130B) includes an eyepiece (one or more) 520 and an eye-tracking system which includes at least one eye-tracking camera 540 (e.g., an infrared (IR) or near-IR (NIR) camera) positioned on the side of the user's face where eye tracking is performed, and an illumination source 530 (e.g., an IR or NIR light source such as an array or ring of NIR light-emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eyes (one or more) 592. The eye-tracking camera 540 is positioned between the user's eye(s) 592 and the display 510 (e.g., the left or right display panel of a head-mounted display, or the display or projector of a handheld device) and may be directed towards a mirror 550 that transmits visible light while reflecting IR or NIR light from the eye(s) 592 (e.g., as shown at the top of Figure 5), or may be directed towards the user's eye(s) 592 to receive reflected IR or NIR light from the user's eye(s) 592 (e.g., as shown at the bottom of Figure 5).
[0080] In some embodiments, the controller 110 renders AR or VR frames 562 (e.g., left and right frames of left and right display panels) and provides the frames 562 to the display 510. For various purposes, for example, when processing the frames 562 for display, the controller 110 uses eye-tracking input 542 from the eye-tracking camera 540. The controller 110 optionally uses a glint-assisted method or other appropriate method to estimate the user's viewpoint on the display 510 based on the eye-tracking input 542 obtained from the eye-tracking camera 540. The viewpoint estimated from the eye-tracking input 542 is optionally used to determine the direction the user is currently looking.
[0081] The following describes, but is not intended to be limiting, several possible use cases of the user's current gaze direction. As an exemplary use case, the controller 110 may render virtual content differently based on the determined user gaze direction. For example, the controller 110 may generate virtual content at a higher resolution in the central visual region determined from the user's current gaze direction than in the peripheral region. As another example, the controller may position or move virtual content within the view based at least partially on the user's current gaze direction. As yet another example, the controller may display specific virtual content within the view based at least partially on the user's current gaze direction. As another exemplary use case in an AR application, the controller 110 may capture the physical environment of the CGR experience and orient an external camera to focus in the determined direction. The external camera's autofocus mechanism can then focus on an object or surface in the environment that the user is currently viewing on the display 510. In another exemplary use case, the eyepiece 520 may be a focusing lens, and the controller uses eye-tracking information to adjust the focus of the eyepiece 520 so that the virtual object currently being viewed by the user has appropriate binocular coordination to match the convergence of the user's eye 592. The controller 110 can utilize the eye-tracking information to orient and adjust the focus of the eyepiece 520 so that the nearby object being viewed by the user appears at the correct distance.
[0082] In some embodiments, the eye-tracking device is part of a head-mounted device mounted on a wearable housing, which includes a display (e.g., display 510), two eyepieces (e.g., one or more eyepieces 520), an eye-tracking camera (e.g., one or more eye-tracking cameras 540), and a light source (e.g., a light source 530 (e.g., an IR LED or NIR LED)). The light source emits light (e.g., IR light or NIR light) toward the user's eye(s) 592. In some embodiments, the light sources may be arranged in a ring or circle around each lens, as shown in Figure 5. In some embodiments, eight light sources 530 (e.g., LEDs) are arranged around each lens 520 as an example. However, more or fewer light sources 530 may be used, and other arrangements and positions of the light sources 530 may be used.
[0083] In some embodiments, the display 510 emits light within the visible light range and does not emit light within the IR or NIR range, thus not introducing noise into the eye-tracking system. Note that the position and angle of the eye-tracking camera(s) 540 are given as examples and are not intended to be limiting. In some embodiments, a single eye-tracking camera 540 may be positioned on each side of the user's face. In some embodiments, two or more NIR cameras 540 may be used on each side of the user's face. In some embodiments, a camera 540 with a wider field of view (FOV) and a camera 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, a camera 540 operating at one wavelength (e.g., 850 nm) and a camera 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.
[0084] Embodiments of eye-tracking systems, such as those shown in Figure 5, can be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide users with computer-generated reality, virtual reality, augmented reality, and / or augmented virtual experiences.
[0085] Figure 6 shows glint-assisted eye-tracking pipelines according to several embodiments. In some embodiments, the eye-tracking pipeline is implemented by a glint-assisted eye-tracking system (e.g., an eye-tracking device 130 as shown in Figures 1 and 5). The glint-assisted eye-tracking system can maintain a tracking state. Initially, the tracking state is off or "no". When in tracking state, the glint-assisted eye-tracking system tracks the pupil contour and glint in the current frame by using prior information from previous frames when analyzing the current frame. When not in tracking state, the glint-assisted eye-tracking system attempts to detect the pupil and glint in the current frame, and if successful, initializes the tracking state to "yes" and continues in tracking state for the next frame.
[0086] As shown in Figure 6, the eye-tracking camera can capture left and right images of the user's left and right eyes. The captured images are then fed into the eye-tracking pipeline for processing, which is initiated at 610. As indicated by the arrow returning to element 600, the eye-tracking system can continue to capture images of the user's eyes at a rate of, for example, 60 to 120 frames per second. In some embodiments, each set of captured images may be fed into the pipeline for processing. However, in some embodiments, or under some conditions, not all captured frames are processed by the pipeline.
[0087] At 610, if the tracking status is yes for the currently captured image, the method proceeds to element 640. At 610, if the tracking status is no, the image is analyzed to detect the user's pupil and glint in the image, as shown in 620. At 630, if the pupil and glint are successfully detected, the method proceeds to element 640. If they are not successfully detected, the method returns to element 610 and processes the next image of the user's eyes.
[0088] At 640, if the process proceeds from element 610, the current frame is analyzed and the pupil and glint are tracked, based in part on prior information from the previous frame. At 640, if the process proceeds from element 630, the tracking state is initialized based on the detected pupil and glint in the current frame. The results of the processing at element 640 are checked to ensure that the tracking or detection results are reliable. For example, the results may be checked to determine whether a sufficient number of glints for pupil and gaze estimation are successfully tracked or detected in the current frame. At 650, if the results are unreliable, the tracking state is set to no at element 660, and the method returns to element 610 to process the next image of the user's eyes. At 650, if the results are reliable, the method proceeds to element 670. At 670, the tracking state is set to yes (if not already yes), and the pupil and glint information is passed to element 680 to estimate the user's gaze.
[0089] Figure 6 is intended to serve as an example of an eye-tracking technology that may be used in a particular implementation. As will be recognized by those skilled in the art, other eye-tracking technologies that currently exist or may be developed in the future may be used in the computer system 101 to provide the user with a CGR experience in various embodiments, either in place of or in combination with the glint-assisted eye-tracking technology described herein.
[0090] This disclosure describes various input methods for interaction with computer systems. Where one example is provided using one input device or method, and another example is provided using a different input device or method, each example may be compatible with the input device or method described in the other example, and their use should be considered optional. Similarly, various output methods for interaction with computer systems are described. Where one example is provided using one output device or method, and another example is provided using a different output device or method, each example may be compatible with the output device or method described in the other example, and their use should be considered optional. Similarly, various methods for interaction with virtual or mixed reality environments via computer systems are described. Where one example is provided using interaction with a virtual environment, and another example is provided using a mixed reality environment, each example may be compatible with the method described in the other example, and their use should be considered optional. Therefore, this disclosure discloses embodiments that are combinations of features of multiple examples, without exhaustively listing all features of the embodiments in the description of each exemplary embodiment. User interface and related processes
[0091] Here, we focus on embodiments of the user interface ("UI") and related processes that may be performed in a computer system such as a portable multifunction device or head-mounted device that communicates with a display generation component and (optionally) one or more sensors (e.g., a camera).
[0092] This disclosure relates to exemplary processes for representing a user as an avatar character in a CGR environment. Figures 7A–7C, 8A–8C, and 9 show examples of a user being represented in a CGR environment as a virtual avatar character having one or more display characteristics whose appearance varies based on the certainty of the user's physical body posture in a real environment. Figures 10A–10B, 11A–11B, and 12 show examples of a user being represented in a CGR environment as a virtual avatar character having appearances based on different appearance templates. The processes disclosed herein are implemented using a computer system (e.g., computer system 101 in Figure 1), as described above.
[0093] Figure 7A shows user 701 standing in the real environment 700 with both arms raised and the user's left hand holding the cup 702. In some embodiments, the real environment 700 is a motion capture studio including cameras 705-1, 705-2, 705-3, and 705-4 for capturing data (e.g., image data and / or depth data) that can be used to determine the posture of one or more parts of user 701. This may be referred to herein as capturing the posture of a part of user 701. The posture of a part of user 701 is used to determine the posture of a corresponding avatar character in the CGR environment (see, for example, avatar 721 in Figure 7C), such as an animated video set.
[0094] As shown in Figure 7A, camera 705-1 has a field of view 707-1 positioned over a portion 701-1 of user 701. The portion 701-1 includes the user's neck, collar area, and a portion of the user's face and head, including the user's right eye, right ear, nose, and mouth, but excluding the top of the user's head, the user's left eye, and the user's left ear, and other physical features of user 701 including a portion of the user's face and head. Since the portion 701-1 is within the camera's field of view 707-1, camera 705-1 captures the posture of the physical features of user 701 including the portion 701-1.
[0095] Camera 705-2 has a field of view 707-2 positioned over a portion 701-2 of user 701. The portion 701-2 includes the physical features of user 701, including the user's right hand, right wrist, and the distal portion of the user's forearm adjacent to the user's right wrist. Since the portion 701-2 is within the camera's field of view 707-2, camera 705-2 captures the posture of the physical features of user 701, including the portion 701-2.
[0096] Camera 705-3 has a field of view 707-3 positioned over a portion 701-3 of user 701. The portion 701-3 includes the user's physical features, including the user's left and right feet, as well as the left and right lower limb regions. Since the portion 701-3 is within the camera's field of view 707-3, camera 705-3 captures the posture of the physical features of user 701, including the portion 701-3.
[0097] Camera 705-4 has a field of view 707-4 positioned over a portion 701-4 of user 701. The portion 701-4 includes the user's physical features, including the user's left hand, left wrist, and the distal portion of the user's left forearm adjacent to the wrist. Since portion 701-4 is within the camera's field of view 707-4, camera 705-4 generally captures the pose of the physical features of user 701, including portion 701-4. However, as will be discussed in more detail below, some areas of portion 701-4, such as the palm of the user's left hand, are positioned behind the cup 702 and are therefore hidden from camera 705-4 and thus not considered to be within the field of view 707-4. Therefore, the pose of these areas (e.g., the palm of the user's left hand) is not captured by camera 705-4.
[0098] Any part of user 701 that is not within the camera's field of view is considered not to be captured by the camera. For example, the top of the user's head, the user's left eye, the user's upper arm, elbow, and the proximal end of the user's forearm, the user's upper limb and knee, and the user's torso are all outside the field of view of cameras 705-1 to 705-4, and therefore the position or orientation of these parts of user 701 is considered not to be captured by the camera.
[0099] Cameras 705-1 to 705-4 are described as non-limiting examples of devices for capturing the orientation of a part of the user 701, i.e., devices for capturing data that can be used to determine the orientation of a part of the user 701. Therefore, other sensors and / or devices can be used in addition to or instead of any of the cameras 705-1 to 705-4 to capture the orientation of a part of the user. Examples of such sensors include proximity sensors, accelerometers, GPS sensors, position sensors, depth sensors, thermal sensors, image sensors, other types of sensors, or any combination thereof. In some embodiments, these various sensors may be standalone components, such as wearable position sensors placed at different locations on the user 701. In some embodiments, various sensors can be integrated into one or more devices associated with user 701, such as the user's smartphone, the user's tablet, the user's computer, a motion capture suit worn by user 701, a headset worn by user 701 (e.g., HMD), a smartwatch (e.g., watch 810 in Figure 8A), another device worn by user 701 or otherwise associated with user 701, or any combination thereof. In some embodiments, various sensors can be integrated into one or more different devices associated with other users (users other than user 701), such as another user's smartphone, another user's tablet, another user's computer, a headset device worn by another user, another device worn by another user or otherwise associated with another user, or any combination thereof. Cameras 705-1 to 705-4 are shown as standalone devices in Figure 7A. However, one or more of the cameras may be integrated with other components, such as any of the sensors and devices described above. For example, one of the cameras may be a camera integrated with a second user's headset device located in the real environment 700.In some embodiments, data may be provided from face scans (e.g., using a depth sensor), media items such as photos and videos of the user 701, or other relevant sources. For example, depth data associated with the user's face can be collected when the user uses face recognition to unlock a personal communication device (e.g., a smartphone, smartwatch, or HMD). In some embodiments, the device that generates a replica of the user's part is separate from the personal communication device, and data from face scans is provided to the device that generates the replica of the user's part for use in constructing the replica of the user's part (e.g., securely and confidentially with one or more options for the user to decide whether or not to share data between devices). In some embodiments, the device that generates the replica of the user's part is the same as the personal communication device, and data from face scans is provided to the device that generates the replica of the user's part for use in constructing the replica of the user's part (e.g., face scans are used to unlock an HMD that also generates a replica of the user's part). This data can be used, for example, to improve the understanding of the posture of parts of user 701, or to increase the visual fidelity (discussed below) of the reproduction of parts of user 701 that are not detected using sensors (e.g., cameras 705-1 to 705-4). In some embodiments, this data can be used to detect changes in parts of the user that are visible when the device is unlocked (e.g., a new hairstyle, new glasses, etc.), thereby updating the representation of the user's parts based on the changes in the user's appearance.
[0100] In the embodiments described herein, the sensors and devices described above for capturing data that can be used to determine the posture of a part of a user are generally referred to as sensors. In some embodiments, the data generated using sensors is referred to as sensor data. In some embodiments, the term “posture data” is used to refer to data that can be used (for example, by a computer system) to determine the posture of at least a part of a user. In some embodiments, posture data may include sensor data.
[0101] In embodiments disclosed herein, a computer system uses sensor data to determine the posture of a part of user 701, and then represents the user as an avatar in the CGR environment with the avatar having the same posture as user 701. However, in some cases, the sensor data may not be sufficient to determine the posture of some parts of the user's body. For example, the user's body part may be outside the sensor's field of view, the sensor data may be corrupted, uncertain, or incomplete, or the user may be moving too fast for the sensor to capture the posture. In any case, the computer system determines (e.g., estimates) the posture of these parts of the user based on various datasets, which will be discussed in more detail below. Since these postures are estimates, the computer system calculates an estimate of the certainty (the confidence level of the accuracy of the determined posture) of each determined posture for parts of the user's body, particularly those not adequately represented by the sensor data. In other words, the computer system calculates an estimate of the certainty that the estimated posture of the corresponding part of user 701 is an accurate representation of the actual posture of the user's body part in the real environment 700. The estimation of certainty may be referred to herein as certainty (or uncertainty) or the amount of certainty (or uncertainty) of the estimated posture of a part of the user's body. For example, in Figure 7A, the computer system calculates with 75% certainty that the user's left elbow is raised to the side and bent at a 90° angle. The certainty (reliability) of the determined posture is represented using the certainty map 710, which is discussed below with respect to Figure 7B. In some embodiments, the certainty map 710 also represents the posture of the user 701 as determined by the computer system. The determined posture is represented using the avatar 721, which is discussed below with respect to Figure 7C.
[0102] Referring here to Figure 7B, the certainty map 710 is a visual representation of the computer system's certainty of the determined posture for a part of the user 701. In other words, the certainty map 710 represents a calculated estimate of the certainty of the position of different parts of the user's body determined using the computer system (e.g., the probability that the estimated position of a part of the user's body is accurate). The certainty map 710 is a human body template representing basic human features such as the head, neck, shoulders, torso, arms, hands, legs, and feet. In some embodiments, the certainty map 710 represents various sub-features such as fingers, elbows, knees, eyes, nose, ears, and mouth. The human features in the certainty map 710 correspond to the physical features of the user 701. Hatching 715 is used to represent the degree or amount of uncertainty in the determined position or posture of the part of the user 701's body corresponding to each hatched portion of the certainty map 710.
[0103] In the embodiments provided herein, each part of user 701 and the certainty map 710 are described at a level of granularity sufficient to illustrate the various embodiments disclosed herein. However, these features can be described at a further (or lower) level of granularity that further illustrates the user's posture and parts, as well as the corresponding certainty of the posture, without departing from the spirit and scope of this disclosure. For example, the user features of parts 701-4 captured using camera 705-4 can be further described as including the tips of the user's fingers but not including the base of the fingers positioned behind the cup 702. Similarly, the back of the user's right hand is facing away from camera 705-2, and therefore, image data acquired using camera 705-2 does not directly capture the back of the user's right hand and can be considered outside the field of view 707-2. As another example, the certainty map 710 can represent one amount of certainty regarding the user's left eye and different amounts of certainty regarding the user's left ear or the upper part of the user's head. However, for the sake of brevity, the details of these granularity variations will not be explained in any case.
[0104] In the embodiment shown in Figure 7B, the certainty map 710 includes portions 710-1, 710-2, 710-3, and 710-4 corresponding to the respective portions 701-1, 701-2, 701-3, and 701-4 of user 701. Thus, portion 710-1 represents the certainty of the posture of the user's neck, collar region, and portion of the user's face and head, including the user's right eye, right ear, nose, and mouth, but not the upper part of the user's head, the user's left eye, and the user's left ear. Portion 710-2 represents the certainty of the posture of the user's right hand, right wrist, and the distal portion of the user's forearm adjacent to the user's right wrist. Portion 710-3 represents the certainty of the posture of the user's left and right feet, and the left and right lower limb regions. Section 710-4 represents the certainty of the posture of the user's left hand, left wrist, and the distal portion of the user's left forearm adjacent to the wrist. In the embodiment shown in Figure 7B, the computer system determines with high certainty the posture of the physical features of sections 701-1, 701-2, and 701-3 of user 701 (for example, the physical features are determined to be within the camera's field of view and not moving), so sections 710-1, 710-2, and 710-3 are shown without hatching. However, section 710-4 is shown with slight hatching 715-1 on the section of the certainty map 710 corresponding to the palm of the user's left hand. This is because the computer system has lower certainty at the position of the palm of the user's left hand, which is hidden by the cup 702 in Figure 7A. In this example, since the user's left hand is hidden by the cup, the computer system has lower certainty about the posture of the user's left hand. However, there are other reasons why a computer system may have a certain degree of uncertainty in the determined posture. For example, a part of the user may be moving too fast for the sensor to detect its posture, or the lighting in the real environment may be insufficient (e.g., the user is backlit). In some embodiments, the certainty with which the computer system can determine the posture is based on, for example, the sensor's frame rate, the sensor's resolution, and / or the ambient light level.
[0105] The certainty map 710 represents the certainty of the determined pose for the parts of user 701 captured within the camera's field of view in Figure 7A, as well as for the parts of user 701 not captured within the camera's field of view. Therefore, the certainty map 710 further represents the computer system's certainty in estimating the determined pose for parts of user 701 that are outside the field of view of cameras 705-1 to 705-4, such as the top of the user's head, the user's left eye, the user's left ear, the user's upper arm, elbow, and the proximal end of the user's forearm, the user's upper limb, the user's upper limb and knee, the user's torso, and, as mentioned above, the user's left palm. Since these parts of user 701 are outside the field of view 707-1 to 707-4, the certainty of the determined pose for these parts is low. Therefore, the corresponding areas of the certainty map 710 are indicated by hatching 715 and show the estimated amount of uncertainty associated with the determined pose for each of these parts of user 701.
[0106] In the embodiment shown in Figure 7B, the density of the hatching 715 is directly proportional to the uncertainty represented by the hatching (or inversely proportional to the certainty represented by the hatching). Therefore, a higher density of hatching 715 indicates a higher degree of uncertainty (or lower certainty) in the posture of the corresponding body part of the user 701, and a lower density of hatching 715 indicates a lower degree of uncertainty (or higher certainty) in the posture of the corresponding body part of the user 701. No hatching indicates a high degree of certainty (e.g., 90%, 95%, or 99%) in the posture of the corresponding body part of the user 701. For example, in the embodiment shown in Figure 7B, hatching 715 is located on portion 710-5 of the certainty map 710 corresponding to the top of the user's head, the user's left eye, and the user's left ear, while there is no hatching on portion 710-1. Therefore, the certainty map 710 shows that in Figure 7B, the certainty of the posture of part 701-1 of user 701 is high, but the certainty of the posture of the upper part of the user's head, the user's left eye, and the user's left ear is low. In situations where the user is wearing a head-mounted device (for example, where the representation of the user's part is generated by integration with the head-mounted device, or generated at least partially based on sensors integrated with or attached to the head-mounted device), the appearance of the user's head and face covered by the head-mounted device is uncertain, but can be estimated based on sensor measurements of the visible parts of the user's head and / or face.
[0107] As described above, although some parts of user 701 are outside the camera's field of view, the computer system can determine the approximate posture of these parts of user 701 with a degree of uncertainty. This degree of uncertainty is represented in the certainty map 710 by showing hatching 715 of different densities. For example, the computer system estimates high certainty for the posture of the user's upper head region (the top of the user's head, the user's left eye, and the user's left ear). Therefore, this region is shown in Figure 7B as a portion 710-5 with a low hatching density, as indicated by the large amount of spacing between hatch lines. As another example, the computer system estimates lower certainty for the posture of the user's right elbow than for the posture of the user's right shoulder. Therefore, the right elbow portion 710-6 is shown with a higher hatching density than the right shoulder portion 710-7, as indicated by the decreasing amount of spacing between hatch lines. Finally, the computer system estimates a minimum amount of certainty for the posture of the user's torso (e.g., waist). Therefore, fuselage section 710-8 is shown with the highest hatching density.
[0108] In some embodiments, the posture of each part of user 701 can be determined or inferred using data from various sources. For example, in the embodiment shown in Figure 7B, the computer system determines the posture of part 701-2 of the user with high certainty based on the detection of this part of the user within the field of view 707-2 of camera 705-2, but the posture of the user's right elbow cannot be determined based solely on sensor data from camera 705-2. However, sensor data from camera 705-2 can be supplemented to determine the posture of parts of user 701 within or outside the field of view 707-2, including the posture of the user's right elbow. For example, if the computer system has high certainty about the posture of the user's neck and collar region (see part 710-1), the computer system can predict the posture of the user's right shoulder (see part 710-7) with high certainty (e.g., lower than that of the neck and collar region), based on known kinematics of the human body, and in this example in particular, based on knowledge of the position of the person's right shoulder relative to the collar region. In some embodiments, algorithms can be used to extrapolate or interpolate to approximate the posture of parts of the user, such as the upper part of the user's right arm and the user's right elbow. For example, the potential posture of the right elbow portion 710-6 depends on the position of the right upper arm portion 710-9, which is predicted from the position of the right shoulder portion 710-7. Furthermore, the position of the right upper arm portion 710-9 depends on the joint movement of the user's shoulder joint, which is outside the field of view of any of the cameras 705-1 to 705-4. Thus, the uncertainty of the posture of the right upper arm portion 710-9 is higher than that of the posture of the right shoulder portion 710-7. Similarly, the posture of the right elbow portion 710-6 depends on the estimated posture of the right upper arm portion 710-9 and the estimated posture of the proximal end of the user's right forearm. However, although the proximal end of the user's right forearm is not within the camera's field of view, the distal region of the user's right forearm is within the field of view 707-2. Therefore, the posture of the distal region of the user's right forearm can be determined with high certainty, as shown in portion 710-2. Thus, the computer system uses this information to estimate the posture of the proximal end of the user's right forearm and the right elbow portion 710-6. This information can also be used to further inform the system of the estimated posture of the right upper arm portion 710-9.
[0109] In some embodiments, the computer system uses an interpolation function to estimate the posture of a part of the user that is located between two or more physical features that have known postures. For example, referring to part 710-4 of Figure 7B, the computer system determines with high confidence the posture of the fingers of the user's left hand and the posture of the distal end of the user's left forearm because these physical features of user 701 are located within the field of view 707-4. However, the posture of the user's left palm cannot be determined with high confidence because it is hidden by the cup 702. In some embodiments, the computer system uses an interpolation algorithm to determine the posture of the user's left palm based on the known postures of the user's fingers and left forearm. Since the postures of many of the user's physical features that are close to or adjacent to the left palm are known, the computer system determines the posture of the left palm with relatively high confidence, as indicated by the relatively sparse hatching 715-1.
[0110] As described above, the computer system determines the pose of user 701 (or, in some embodiments, a set of poses of parts of user 701) and displays an avatar in the CGR environment representing user 701 with the determined pose(s)(s)(s)(s)(s). Figure 7C shows an example in which the computer system displays avatar 721 in 720 within the CGR environment. In the embodiment shown in Figure 7C, avatar 721 is presented as a virtual block character displayed using a display generation component 730. The display generation component 730 is similar to the display generation component 120 described above.
[0111] Avatar 721 includes parts 721-2, 721-3, and 721-4. Part 721-2 corresponds to part 701-2 of the certainty map 710 in Figure 7B and part 710-2 of user 701 in Figure 7A. Part 721-3 corresponds to part 701-3 of the certainty map 710 in Figure 7B and part 710-3 of user 701 in Figure 7A. Part 721-4 corresponds to part 701-4 of the certainty map 710 in Figure 7B and part 710-4 of user 701 in Figure 7A. Thus, part 721-2 represents the posture of the user's right hand, right wrist, and the distal portion of the user's forearm adjacent to the user's right wrist. Part 721-3 represents the posture of the user's left and right feet, as well as the left and right lower limb regions. Part 721-4 represents the posture of the user's left hand, left wrist, and the distal portion of the user's left forearm adjacent to the wrist. Other parts of avatar 721 represent the posture of corresponding parts of user 701, which will be discussed in more detail below. For example, the neck and collar region 727 of avatar 721 corresponds to part 701-1 of user 701, as will be discussed below.
[0112] The computer system displays an avatar 721 having visual display characteristics (referred to herein as variable display characteristics) that vary depending on the estimated certainty of the posture determined by the computer system. Generally, the variable display characteristics inform the viewer of the avatar 721 about the visual fidelity of the displayed avatar with respect to the posture of user 701, or about the visual fidelity of the displayed part of the avatar with respect to the posture(s)(s)(s) of the corresponding body part(s) of user 701. In other words, the variable display characteristics inform the viewer about the estimated amount or degree to which the rendered posture or appearance of a part of the avatar 721 conforms to the actual posture or appearance of the corresponding body part of user 701 in the real environment 700, which in some embodiments is determined based on the estimated certainty of the posture of each part of user 701. In some embodiments, when a computer system renders avatar features without variable display characteristics, or with variable display characteristics having a value indicating a high degree of certainty regarding the pose of the corresponding part of the user 701 (e.g., 75%, 80%, 90%, 95% certainty), the avatar features are said to be rendered with high fidelity. In some embodiments, when a computer system renders avatar features with variable display characteristics having a value indicating a lower degree of certainty regarding the pose of the corresponding part of the user 701, the avatar features are said to be rendered with low fidelity.
[0113] In some embodiments, the computer system adjusts the values of the variable display characteristics of each avatar feature to communicate changes in the estimated certainty of the pose of the user's body part represented by each avatar feature. For example, as the user moves, the computer system (e.g., continuously, continuously, automatically) determines the user's pose and updates the estimated certainty of the pose of each body part of the user accordingly. The computer system also modifies the appearance of the avatar by correcting the avatar's pose to match the user's newly determined pose and modifies the values of the variable display characteristics of the avatar features based on the changes in the certainty of the corresponding poses of the user's body part.
[0114] The variable display characteristics may include one or more visual features or parameters used to improve, degrade, or modify the appearance of the avatar 721 (e.g., the default appearance), as discussed in the following embodiments.
[0115] In some embodiments, the variable display characteristics include the display color of the avatar features. For example, the avatar may have a default color of green, and as the certainty of the user's part's posture changes, cooler colors may be used to represent parts of the avatar where the computer system has less certainty of the user's part's posture represented by the avatar features, and warmer colors may be used to represent parts of the avatar where the computer system has more certainty of the user's part's posture represented by the avatar features. For example, if the user moves their hand out of the camera's field of view, the computer system modifies the appearance of the avatar's hand by moving it from a posture with a high degree of certainty (e.g., position, orientation) (a posture that matches the posture of the user's hand when it was in the field of view) to a posture with a lower degree of certainty, based on the determination of the updated posture of the user's hand. As the avatar's hand moves from a posture with a high degree of certainty to a posture with very little certainty, the avatar's hand transitions from green to blue. Similarly, as the avatar's hand moves from a posture with a high degree of certainty to a posture with slightly less certainty, the avatar's hand transitions from green to red. As another example, as a hand moves from a very low-certainty posture to a relatively high-certainty posture, the hand transitions from blue to red, with various intermediate colors shifting from cool to warm as the hand moves to a posture with a higher degree of certainty. As yet another example, as a hand moves from a relatively high-certainty posture to a very low-certainty posture, the hand transitions from red to blue, with various intermediate colors shifting from warm to cool as the hand moves to a posture with a decreasing degree of certainty. In some embodiments, the change in the value of the variable display characteristic occurs at a faster rate when the avatar feature moves from an unknown posture (low-certainty posture) to a higher-certainty posture (e.g., a known posture), and at a slower rate when the avatar feature moves from a higher-certainty posture (e.g., a known posture) to a lower-certainty posture. For example, to continue with the example of the variable display characteristic being color, the avatar's hand changes from blue to red at a faster rate than it changes from red (or the default green) to blue.In some embodiments, the color change occurs at a speed faster than the speed at which the user is moving the hand corresponding to the avatar's hand.
[0116] In some embodiments, the variable display characteristic includes the amount of blurring effect applied to the displayed portion of the avatar feature. Conversely, the variable display characteristic may also be the amount of sharpness applied to the displayed portion of the avatar feature. For example, an avatar may have a default sharpness, and as the certainty of the user's pose changes, an increased blur (or decreased sharpness) can be used to represent portions of the avatar where the computer system's certainty of the user's pose represented by the avatar feature is low, and a decreased blur (or increased sharpness) can be used to represent portions of the avatar where the computer system's certainty of the user's pose represented by the avatar feature is high.
[0117] In some embodiments, the variable display characteristic includes the opacity of the displayed portion of the avatar feature. Conversely, the variable display characteristic may also be the transparency of the displayed portion of the avatar feature. For example, an avatar may have a default opacity, and as the certainty of the user's pose changes, a reduced opacity (or increased transparency) can be used to represent portions of the avatar where the computer system has low certainty of the user's pose represented by the avatar feature, and an increased opacity (or decreased transparency) can be used to represent portions of the avatar where the computer system has high certainty of the user's pose represented by the avatar feature.
[0118] In some embodiments, the variable display characteristics include the density and / or size of particles forming the displayed portion of the avatar feature. For example, an avatar may have a default particle size and particle spacing, and as the certainty of the user's part's pose changes, the particle size and / or particle spacing of the corresponding avatar feature changes based on whether the certainty has increased or decreased. For example, in some embodiments, the particle spacing increases (particle density decreases) to represent the portion of the avatar where the computer system has low certainty of the user 701's pose represented by the avatar feature, and the particle spacing decreases (particle density increases) to represent the portion of the avatar where the computer system has high certainty of the user's part's pose represented by the avatar feature. As another example, in some embodiments, the particle size increases (producing a more pixelated appearance and / or a lower resolution appearance) to represent the portion of the avatar where the computer system has low certainty of the user 701's pose represented by the avatar feature, and the particle size decreases (producing a less pixelated appearance and / or a higher resolution appearance) to represent the portion of the avatar where the computer system has high certainty of the user's part's pose represented by the avatar feature.
[0119] In some embodiments, the variable display characteristic includes one or more visual effects such as scale (e.g., fish scale), pattern, shading, and smoke effect. Examples are discussed in more detail below. In some embodiments, an avatar that smoothly transitions between higher and lower certainty regions can be displayed by varying, for example, the amount of the variable display characteristic (e.g., opacity, particle size, color, diffusion, etc.) along the transition region of each part of the avatar. For example, if the variable display characteristic is particle density, and the certainty of the user's forearm transitions from high certainty to low certainty at the user's elbow, the transition from high certainty to low certainty can be represented by displaying an avatar with a high particle density at the elbow and smoothly (gradually) transitioning to a lower particle density along the forearm.
[0120] The variable display characteristics are represented in Figure 7C by hatching 725 and may include one or more of the variable display characteristics described above. The density of the hatching 725 is used as a variability indicating the fidelity to which parts of the avatar 721 are rendered, based on the estimated certainty of the actual posture of the corresponding physical parts of the user 701. Thus, a higher hatching density is used to indicate parts of the avatar 721 where the variable display characteristics have a value indicating that the computer system determined the posture with lower certainty, and a lower hatching density is used to indicate parts of the avatar 721 where the variable display characteristics have a value indicating that the computer system determined the posture with higher certainty. In some embodiments, hatching is not used when the certainty of the posture of the avatar parts is of a high degree of certainty (e.g., 90%, 95%, or 99%).
[0121] In the embodiment shown in Figure 7C, avatar 721 is rendered as a block character having a similar pose to user 701. As described above, the computer system determines the pose of avatar 721 based on sensor data representing the pose of user 701. For parts of user 701 whose pose is known, the computer system renders the corresponding parts of avatar 721 using default or baseline values for variable display characteristics (e.g., resolution, transparency, particle size, particle density, color), or without variable display characteristics (e.g., without patterns or visual effects), as shown without hatching. For parts of user 701 whose pose falls below a predetermined certainty threshold (e.g., pose certainty is less than 100%, 95%, 90%, or 80%), the computer system renders the corresponding parts of avatar 721 using variable display characteristic values that vary from the default or baseline, or using variable display characteristics if they are not displayed to indicate high certainty, as shown by hatching 725.
[0122] In the embodiment shown in Figure 7C, the density of hatching 725 generally corresponds to the density of hatching 715 in Figure 7B. Therefore, the variable display characteristic value of avatar 721 generally corresponds to the certainty (or uncertainty) represented by hatching 715 in Figure 7B. For example, portion 721-4 of avatar 721 corresponds to portion 710-4 of the certainty map 710 in Figure 7B and portion 701-4 of user 701 in Figure 7A. Thus, portion 721-4 is indicated by hatching 725-1 on the left palm of the avatar (similar to hatching 715-1 in Figure 7B), indicating that the variable display characteristic value corresponds to a relatively high certainty of the posture of the user's left palm, but not as high as the other portions, and there is no hatching on the fingers and the distal end of the avatar's left forearm (indicating a high certainty of the posture of the corresponding portion of user 701).
[0123] In some embodiments, hatching is not used when the estimated certainty of the pose of a part of the avatar is relatively high (e.g., 99%, 95%, or 90%). For example, parts 721-2 and 721-3 of avatar 721 are shown without hatching. Since the pose of part 701-2 of user 701 is determined with high certainty, part 721-2 of avatar 721 is rendered without hatching in Figure 7C. Similarly, since the pose of part 701-3 of user 701 is determined with high certainty, part 721-3 of avatar 721 is rendered without hatching in Figure 7C.
[0124] In some embodiments, a portion of avatar 721 is displayed that has one or more avatar features not derived from user 701. For example, avatar 721 is rendered with avatar hair 726 determined based on the visual attributes of the avatar character, rather than the hairstyle of user 701. As another example, the hands and fingers of avatar portion 721-2 are blocky hands and blocky fingers that are not human hands or human fingers, but have the same posture as the user's hands and fingers in Figure 7A. Similarly, avatar 721 is rendered with a nose 722 that is different from user 701's nose but has the same posture. In some embodiments, different avatar features may be different human features (e.g., different human noses), features from a non-human character (e.g., a dog's nose), or abstract shapes (e.g., a triangular nose). In some embodiments, different avatar features can be generated using machine learning algorithms. In some embodiments, the computer system renders avatar features using features not derived from the user in order to save computational resources by avoiding additional operations performed to render avatar features with high visual fidelity regarding the corresponding parts of the user.
[0125] In some embodiments, a portion of avatar 721 having one or more avatar features derived from user 701 is displayed. For example, in Figure 7C, avatar 721 is rendered having a mouth 724 which is the same as the mouth of user 701 in Figure 7A. In another example, avatar 721 is rendered having a right eye 723-1 which is the same as the right eye of user 701 in Figure 7A. In some embodiments, such avatar features are rendered using video feeds of the corresponding portions of user 701 mapped onto a three-dimensional model of the virtual avatar character. For example, the avatar's mouth 724 and right eye 723-1 are based on video feeds of the user's mouth and right eye captured within the field of view 707-1 of camera 705-1 and displayed on avatar 721 (e.g., via video passthrough).
[0126] In some embodiments, the computer system renders certain features of the avatar 721 as having high (or increased) certainty of their poses, even if the computer system estimates that the pose of the corresponding part of the user is lower than high certainty. For example, in Figure 7C, the left eye 723-2 is shown without hatching, indicating high certainty of the pose of the user's left eye. However, in Figure 7A, the user's left eye is outside the field of view 707-1, and in Figure 7B, hatching 715 indicates that the certainty of the part 710-5 containing the user's left eye is lower than high certainty of the pose of the user's left eye. Rendering some features with high fidelity, even when the estimated certainty is lower than high, is sometimes done to improve the quality of communication using the avatar 721. For example, some features, such as eyes, hands, and mouth, may be considered important for communication purposes, and rendering such features using variable display characteristics may distract the user browsing and communicating with the avatar 721. Similarly, if user 701's mouth is moving too fast to capture sufficient pose data about the user's lips, teeth, tongue, etc., rendering the mouth using variable display characteristics such as blurring, transparency, color, increased particle spacing, or increased particle size would be distracting, so this computer system can render the avatar's mouth 724 with high certainty of mouth pose.
[0127] In some embodiments, if the certainty of the pose or appearance of a part of the user is low (e.g., 99%, 95%, 90% certainty), the pose and / or appearance of the corresponding avatar feature may be enhanced using data from different sources. For example, if the pose of the user's left eye is unknown, the pose of the avatar's left eye 723-2 can be estimated using a machine learning algorithm that determines the pose of the left eye 723-2 based on, for example, the mirrored pose of the user's right eye, which is known. Furthermore, since the user's left eye is outside the field of view 707-1, the sensor data from camera 705-1 may not include data for determining the appearance of the user's left eye (e.g., eye color). In some embodiments, the computer system may use data from other sources to obtain the data necessary to determine the appearance of the avatar feature. For example, the computer system may access previously captured face scan data, photographs, and videos of the user 701 associated with a personal communication device (e.g., a smartphone, smartwatch, or HMD) to obtain data for determining the appearance of the user's eyes. In some embodiments, other avatar features may be updated based on additional data accessed by the computer system. For example, if a recent photograph shows user 701 with a different hairstyle, the avatar's hair may be changed to match the hairstyle in user 701's recent photograph. In some embodiments, the device that generates the user's part replica is separate from the personal communication device, and data from face scans, photographs, and / or videos is provided to the user's part replica device for use in constructing the user's part replica (securely and confidentially, with one or more options for the user to decide whether or not to share data between devices). In some embodiments, the device that generates the user's part replica is the same as the personal communication device, and data from face scans, photographs, and / or videos is provided to the user's part replica device for use in constructing the user's part replica (for example, a face scan is used to unlock the HMD that generates the user's part replica).
[0128] In some embodiments, the computer system displays portions of the avatar 721 as having low visual fidelity with respect to the corresponding parts of the user, even when the pose of the corresponding user feature is known (or determined with high certainty). For example, in Figure 7C, the neck and collar region 727 of the avatar 721 is indicated by hatching 725, which shows that each avatar feature is displayed using variable display characteristics. However, the neck and collar region of the avatar 721 corresponds to a portion 701-1 of the user 701 within the field of view 707-1 and to a portion 710-1 which is shown as having high certainty in the certainty map 710. In some embodiments, the computer system displays avatar features as having low visual fidelity even when the pose of the corresponding parts of the user 701 is known (or determined with high certainty), in order to save computational resources used to render high-fidelity representations of the corresponding avatar features. In some embodiments, the computer system performs this action when the user features are considered less important for communication purposes. In some embodiments, a low-fidelity version of the avatar features is generated using a machine learning algorithm.
[0129] In some embodiments, the change in the value of the variable display characteristic is based on the movement speed of the corresponding part of the user 701. For example, the faster the user moves their hand, the lower the certainty of the hand's pose, and the variable display characteristic is presented with a value corresponding to the lower certainty of the hand's pose. Conversely, the slower the user moves their hand, the higher the certainty of the hand's pose, and the variable display characteristic is presented with a value corresponding to the higher certainty of the hand's pose. If the variable display characteristic is blurring, for example, if the user moves their hand at a faster speed, the avatar's hand will be rendered with a higher degree of blurring, and if the user moves their hand at a slower speed, it will be rendered with a lower degree of blurring.
[0130] In some embodiments, the display generation component 730 enables the display of the CGR environment 720 and the avatar 725 of the computer system user. In some embodiments, the computer system further displays a preview 735 within the CGR environment 720 via the display generation component 730, which includes a representation of the computer system user's appearance. In other words, the preview 735 shows the computer system user how they appear within the CGR environment 720 to other users browsing the CGR environment 720. In the embodiment shown in Figure 7C, the preview 735 shows the computer system user (e.g., a user different from user 701) how they appear as a female avatar character with variable display characteristics.
[0131] In some embodiments, the computer system calculates certainty estimates using data collected from one or more sources other than cameras 705-1 to 705-4. For example, these other sources can be used to supplement, or possibly replace, sensor data collected from cameras 705-1 to 705-4. These other sources may include different sensors, such as any of the sensors described above. For example, user 701 may wear a smartwatch or other wearable device that provides data indicating the position, movement, or location of the user's arm and / or other body parts. As another example, user 701 may have a smartphone in their pocket that provides data indicating the user's position, movement, posture, or other such data. Similarly, a headset device worn by another person in the real environment 700 may include sensors such as a camera that provide data indicating posture, movement, position, or other relevant information associated with user 701. In some embodiments, data may be provided from face scans (e.g., using depth sensors), media items such as photographs and videos of user 701, or other relevant sources as described above.
[0132] Figures 8A to 8C show an embodiment similar to Figure 7A, in which sensor data from cameras 705-1 to 705-4 is supplemented by data from the user's smartwatch. In Figure 8A, user 701 is currently positioned with their right hand on their hip and their left hand, wearing a smartwatch 810 on their left wrist and pressed tightly against the wall 805. For example, user 701 has moved from the posture in Figure 7A to the posture in Figure 8A. The user's right hand, right wrist, and right forearm are not currently positioned in the field of view 707-2. Furthermore, the user's left hand remains within the field of view 707-4, but the user's left wrist and left forearm are outside the field of view 707-4.
[0133] In Figure 8B, the certainty map 710 has been updated based on the new posture of user 701 in Figure 8A. Therefore, the uncertainty of the posture of the user's right elbow, right forearm, and right hand has been updated based on their new positions, resulting in increased hatching density 715 for these areas. Specifically, since the entire right arm of the user is outside the camera's field of view, the certainty of the posture of these parts of user 701 decreases from the posture certainty in Figure 7B. Thus, the certainty map 710 shows increased hatching density for these parts of the certainty map in Figure 8B. The posture of each sub-feature of the user's right arm depends on the position of adjacent sub-features, and since all sub-features of the right arm are outside the field of view of the camera or other sensors, the hatching density increases with each successive sub-feature, starting from the right upper arm portion 811-1, up to the right elbow portion 811-2, up to the right forearm portion 811-3, and up to the right hand and fingers 811-4.
[0134] While the user's left forearm is outside the field of view 707-4, the sensor data from camera 705-4 is supplemented by the sensor data from smartwatch 810, which provides posture data for the user's left forearm, so that the computer system can still determine the posture of the user's left forearm with high confidence. Therefore, the confidence map 710 shows the left forearm portion 811-5 and the left hand portion 811-6 (which is within the field of view 707-4) with a high degree of confidence in their respective postures.
[0135] Figure 8C shows the updated pose of avatar 721 in the CGR environment 720, based on the updated pose of user 701 in Figure 8A. In Figure 8C, the computer system renders the avatar's left arm 821 with a high degree of certainty as the user's left arm, as indicated by the absence of hatching on the left arm 821.
[0136] In some embodiments, the computer system renders the avatar 721 along with a representation of the object that the user 701 is interacting with. For example, if the user is holding an object, leaning against a wall, sitting in a chair, or otherwise interacting with an object, the computer system may render at least a portion (or a representation thereof) of the object within the CGR environment 720. For example, in Figure 8C, the computer system includes a wall rendering 825 positioned and shown adjacent to the avatar's left hand. This provides context for the user's posture so that the viewer can understand that the user 701 is in a posture with their left hand on a surface within the real environment 700. In some embodiments, the object that the user is interacting with may be a virtual object. In some embodiments, the object that the user is interacting with may be a physical object, such as a wall 805. In some embodiments, the rendered version of the object (e.g., the wall rendering 825) may be rendered as a virtual object or displayed as a video feed of a real object.
[0137] As described above, the computer system updates the pose of the avatar 721 in response to detected changes in the user 701's pose, which in some embodiments involves updating variable display characteristics based on the pose change. In some embodiments, updating variable display characteristics includes increasing or decreasing the value of the variable display characteristic (represented by increasing or decreasing the amount or density of hatching 725). In some embodiments, updating variable display characteristics includes introducing or removing variable display characteristics (represented by introducing or removing hatching 725). In some embodiments, updating variable display characteristics includes introducing or removing visual effects. For example, in Figure 8C, the computer system renders the avatar 721 with a right arm having a smoke effect 830 to indicate a relatively low certainty or reduction of certainty in the pose of the user's right arm. In some embodiments, visual effects may include other effects such as fish scales displayed on each avatar feature. In some embodiments, displaying visual effects includes replacing the display of each avatar feature with the displayed visual effect, as shown in Figure 8C. In some embodiments, displaying visual effects includes displaying each avatar feature using visual effects. For example, the avatar's right arm may be displayed using an arm fish scale. In some embodiments, multiple variable display features may be combined with other variable display features, such as the displayed visual effects. For example, the density of the arm scale decreases as the pose certainty decreases along a portion of the arm.
[0138] In the embodiment shown in Figure 8C, the user's right arm posture is represented by a smoke effect 830, which roughly represents the shape of the avatar's right arm having a posture that is lowered toward the side of the avatar's body. Although the avatar's arm posture does not accurately represent the actual posture of the user's right arm in Figure 8A, the computer system accurately determined that the user's right arm was lowered and not above the user's shoulder or directly to the side. This is because the computer system determined the current posture of the user's right arm based on sensor data acquired from cameras 705-1 to 705-4, and based on the user's arm's preceding posture and movement. For example, when the user moved their arm from the posture in Figure 7A to the posture in Figure 8A, the user's right arm moved downward as it moved out of the field of view 707-2. The computer system used this data to determine that the user's right arm posture was not raised or to the side, and therefore had to be lower than before. However, the computer system does not have enough data to accurately determine the posture of the user's right arm. Therefore, the computer system shows a low degree of certainty (reliability) of the user's right arm posture, as shown in the certainty map 710 in Figure 8B. In some embodiments, if the certainty of the posture of the user 701 falls below a threshold, the computer system uses visual effects to represent the corresponding avatar features, as shown in Figure 8C.
[0139] In Figure 8C, Preview 735 is updated to show that the current posture of the computer system user is waving within the CGR environment 720.
[0140] In some embodiments, the computer system provides a control feature that allows the user of the computer system to select how much of the avatar 721 is displayed (for example, which parts or features of the avatar 721 are displayed). For example, the control provides a spectrum in which, at one end, the avatar 721 is displayed using only avatar features for which the computer system has high certainty regarding the pose of the corresponding part of the user, and at the other end, the avatar 721 is displayed using all avatar features, regardless of the certainty of the pose of the corresponding part of the user 701.
[0141] Further explanations regarding Figures 7A-7C and 8A-8C are provided below with reference to Method 900 described with respect to Figure 9.
[0142] Figure 9 is a flowchart of an exemplary method 900 for presenting a virtual avatar character using display characteristics whose appearance varies based on the certainty of the user's body posture, according to several embodiments. In some embodiments, method 900 is performed in a computer system (e.g., computer system 101 in Figure 1) (e.g., smartphone, tablet, head-mounted display generation component) that communicates with a display generation component (e.g., display generation component 120 in Figures 1, 3, and 4) (e.g., display generation component 730 in Figures 7C and 8C) (e.g., visual output device, 3D display, transparent display, projector, head-up display, display controller, touchscreen, etc.). In some embodiments, method 900 is controlled by instructions stored in a non-temporary computer-readable storage medium and executed by one or more processors of the computer system, such as one or more processors 202 of computer system 101 (e.g., control unit 110 in Figure 1). Some operations of method 900 are optionally combined, and / or the order of some operations is optionally changed.
[0143] In method 900, a computer system (e.g., 101) receives posture data (e.g., physical position, orientation, gestures, movement, etc.) representing the posture (e.g., physical position, orientation, gestures, movement, etc.) of at least a first part of the user (e.g., 701-1, 701-2, 701-3, 701-4) (e.g., one or more physical features of the user (e.g., macro features such as arms, legs, hands, head, mouth, etc., and / or micro features such as fingers, face, lips, teeth, or other parts of each physical feature) (e.g., depth data, image data, image data from a sensor data camera (e.g., 705-1, 705-2, 705-3, 705-4, 705-5)) (902). In some embodiments, the posture data is received by the user The data includes a measure of certainty (e.g., reliability) that the determined posture of a part is accurate (e.g., an accurate representation of the posture of the part of the user in a real environment). In some embodiments, the posture data includes sensor data (e.g., image data from a camera, motion data from an accelerometer, location data from a GPS sensor, data from a proximity sensor, data from a wearable device (e.g., watch 810)). In some embodiments, the sensors may be connected to or integrated with a computer system. In some embodiments, the sensors may be external sensors (e.g., sensors from a different computer system (e.g., another user's electronic device)).
[0144] A computer system (e.g., 101) causes an avatar (e.g., 721) (e.g., a virtual avatar, a part of an avatar) (e.g., a virtual representation of at least a part of the user) to be presented (e.g., displayed, presented, projected) (904) (in a computer-generated reality environment (e.g., 720)) via a display generation component (e.g., 730), the avatar corresponding (e.g., anatomically) to a first part of the user (e.g., 701-1, 701-2, 701-3, 701-4), and (e.g., the appearance of the part of the user) The presenting (e.g., displaying) avatar features (e.g., 727, 724, 723-1, 723-2, 722, 726, 721-2, 721-3, 721-4) include each avatar feature (e.g., 727, 724, 723-1, 723-2, 722, 726, 721-2, 721-3, 721-4) that have variable display characteristics (e.g., 725, 830) indicating the certainty of the pose of a first part of the user (e.g., 710, 715), which are displayed with an estimated / predicted visual fidelity of the pose of each avatar feature with respect to the pose, and variable display characteristics (e.g., 725, 830) indicating the certainty of the pose of a first part of the user. By having the avatar present each avatar feature that corresponds to a first part of the user and includes each presented avatar feature that has a variable display characteristic indicating the certainty of the pose of a first part of the user, feedback indicating the reliability of the pose of the user's part represented by each avatar feature is provided to the user. By providing improved feedback, the usability of the computer system is enhanced, the user system interface is made more efficient (for example, by assisting the user in providing appropriate input when operating and / or interacting with the computer system, thereby reducing user errors), and power consumption is reduced and the battery life of the computer system is improved by enabling the user to use the system more quickly and efficiently.
[0145] In some embodiments, each avatar feature is superimposed on (or displayed on) the corresponding part of the user. In some embodiments, the variable display characteristics vary based on the certainty (e.g., reliability) of the pose of the user part. In some embodiments, the certainty of the pose of the user part is expressed as a certainty value (e.g., a value representing the certainty (reliability) that the determined pose of the user part is an accurate representation of the actual pose of the user part (e.g., in a real environment (e.g., 700)). In some embodiments, the certainty value is expressed using a range of values, e.g., a percentage range of 0% to 100%, where 0% indicates no certainty (minimum certainty) that the pose of each user feature (or part thereof) is accurate, and 100% indicates that the certainty of the estimation of each user feature is above a predetermined threshold certainty (e.g., 80%, 90%, 95%, or 99%) (e.g., even if the actual pose cannot be determined due to sensor limitations). In some embodiments, the certainty may be 0% if the computer system (e.g., 101) (or another processing device) does not have enough useful data to estimate the potential position or orientation of each user feature. For example, each user feature may not be within the field of view (e.g., 707-1, 707-2, 707-3, 707-4) of the image sensor (e.g., 705-1, 705-2, 705-3, 705-4), and each user feature may be as likely as any one of several different positions or orientations. Alternatively, for example, data generated using a proximity sensor (or some other sensor) may be uncertain or otherwise insufficient to accurately infer the orientation of each user feature. In some embodiments, the certainty can be high (e.g., 99%, 95%, 90% certainty) if the computer system (or another processing device) can clearly identify each user feature using orientation data and can determine the correct position of each user feature using orientation data.
[0146] In some embodiments, the variable display characteristic directly correlates to the certainty of the user's part's pose. For example, each avatar feature may be rendered using a variable display characteristic having a first value (e.g., a low value) to convey a first certainty (e.g., low certainty) of the user's part's pose (e.g., to the viewer). Conversely, each avatar feature may be rendered using a variable display characteristic having a second value (e.g., a high value) greater than the first value to convey a higher certainty of the user's part's pose (e.g., a high certainty, such as 80%, 90%, 95%, or 99%, exceeding a predetermined certainty threshold). In some embodiments, the variable display characteristic does not directly correlate to the certainty of the user's part's pose (e.g., a certainty value). For example, if the appearance of the user's part is important for communication purposes, the corresponding avatar feature may be rendered using a variable display characteristic having a value corresponding to high certainty, even if the certainty of the user's part's pose is low. This may be done, for example, when displaying each avatar feature using a variable display characteristic with a value corresponding to low certainty would be distracting (for example, for the viewer). Consider, for example, an embodiment in which an avatar includes an avatar head displayed on (e.g., superimposed on) the user's head, and the avatar head includes avatar lips (e.g., 724) representing the pose of the user's lips. A computer system (or another processing device) determines low certainty in the lip pose when the user's lips are moving (e.g., the lips are moving too fast for accurate detection, or the user's lips are partially covered), but displaying an avatar with lips rendered using a variable display characteristic with a value corresponding to low certainty or lower than high (e.g., maximum) certainty would be distracting, so the corresponding avatar lips can be rendered using a variable display characteristic with a value corresponding to high (e.g., maximum) certainty.
[0147] In some embodiments, a variable display characteristic (e.g., 725, 830) indicates the estimated visual fidelity of each avatar feature (e.g., 727, 724, 723-1, 723-2, 722, 726, 721-2, 721-3, 721-4) with respect to the pose of a first part of the user. In some embodiments, visual fidelity represents the authenticity of the displayed / rendered avatar (or part thereof) with respect to the pose of the user's part. In other words, visual fidelity is a measure of how closely the displayed / rendered avatar (or part thereof) is considered to fit the actual pose of the corresponding part of the user. In some embodiments, whether an increase or decrease in the value of a variable display characteristic indicates increased or decreased visual fidelity depends on the type of variable display characteristic being used. For example, if the variable display characteristic is a blurring effect, a larger value of the variable display characteristic (higher blurring) conveys decreased visual fidelity, and vice versa. Conversely, if the variable display characteristic is particle density, a larger value of the variable display characteristic (higher particle density) conveys increased visual fidelity, and vice versa.
[0148] In method 900, presenting an avatar (e.g., 721) is determined (e.g., based on posture data) that the posture of a first part of the user (e.g., 701-1, 701-2, 701-3, 701-4) is associated with a first certainty value (e.g., the certainty value is the first certainty value), and the computer system (e.g., 101) presents the avatar (e.g., 721) using each avatar feature (e.g., 727, 724, 723-1, 723-2, 722, 726, 721-2, 721-3, 721-4) which has a first value of variable display characteristics. This includes presenting (906) (for example, 727 has low-density hatching 725 in Figure 7C; 721-3, 721-2, 723-1, 723-2, 722, and / or 724 do not have hatching 725 in Figure 7C; part 721-4 has low hatching density 725-1 on the left hand in Figure 7C) (for example, each avatar feature is displayed having a first variable display characteristic value (e.g., blur degree, opacity, color, attenuation / density, resolution, etc.) that indicates a first estimated visual fidelity of each avatar feature with respect to the pose of the user's part). The user is provided with feedback indicating the certainty that the pose of the user's first part represented by each avatar feature corresponds to the actual pose of the user's first part, by presenting the avatar with each avatar feature having a first value of the variable display characteristic, in accordance with the determination that the pose of the user's first part is associated with a first certainty value. By providing improved feedback, the usability of the computer system is enhanced, the user system interface is made more efficient (for example, by assisting the user in providing appropriate input when operating and / or interacting with the computer system, thereby reducing user errors), and power consumption is reduced and the battery life of the computer system is improved by enabling the user to use the system more quickly and efficiently. In some embodiments, presenting an avatar using each avatar feature having a first value of variable display characteristics includes presenting each avatar feature having the same posture as the posture of the user's part.
[0149] In method 900, presenting an avatar (e.g., 721) includes, according to the determination that the posture of a first part of the user (e.g., 701-1, 701-2, 701-3, 701-4) is associated with a second certainty value that is different from (e.g., greater than) a first certainty value, the computer system (e.g., 101) presents the avatar using each avatar feature (e.g., 727, 724, 723-1, 723-2, 722, 726, 721-2, 721-3, 721-4) which has a second value of variable display characteristics different from a first value of variable display characteristics (908) (e.g., 721-2 is shown in Figure 8). In Figure 8C, the avatar is displayed using variable display characteristic 830; the avatar's left hand has no hatching; the avatar's left elbow has almost no hatching 725 in Figure 8C) (for example, each avatar feature is displayed having a second variable display characteristic value (e.g., blur degree, opacity, color, attenuation / density, resolution, etc.) that indicates a second estimated visual fidelity of each avatar feature with respect to the posture of the user's body part, and the second estimated visual fidelity is different from the first estimated visual fidelity (e.g., the second estimated visual fidelity indicates a higher estimated visual fidelity than the first estimated visual fidelity)). By presenting the avatar using each avatar feature having a second variable display characteristic value different from the first value of the variable display characteristic, in accordance with the determination that the posture of the user's first body part is associated with a second certainty, the user is provided with feedback indicating that the posture of the user's first body part represented by each avatar feature corresponds to a different certainty of the actual posture of the user's first body part. By providing improved feedback, the usability of the computer system is enhanced, the user system interface is made more efficient (for example, by assisting the user in providing appropriate input when operating and / or interacting with the computer system, thereby reducing user errors), and power consumption is reduced and the battery life of the computer system is improved by enabling the user to use the system more quickly and efficiently.In some embodiments, presenting an avatar using each avatar feature having a second value of the variable display characteristics includes presenting each avatar feature having the same posture as the user's part.
[0150] In some embodiments, a computer system (e.g., 101) receives second pose data representing the pose of at least a second part (e.g., 701-2) of a user (e.g., a part of the user different from a first part of the user) and causes an avatar (e.g., 721) to be presented via a display generation component (e.g., 730) (e.g., updating the presented avatar). The avatar includes a second avatar feature (e.g., 721-2) that corresponds to the second part of the user (e.g., the second part of the user is the user's mouth, and the second avatar feature is a representation of the user's mouth) and presents a second variable display feature (e.g., the same variable display feature as the first variable display feature) (e.g., a different variable display feature from the first variable display feature) indicating certainty of the pose of the second part of the user (e.g., 721-2 is displayed without hatching 725 indicating high certainty of the pose of 701-2). By presenting an avatar that includes a second avatar feature that corresponds to a second part of the user and has a second variable display characteristic indicating the certainty of the user's second part's posture, the user is provided with feedback indicating the reliability of the user's second part's posture represented by the second avatar feature, and further, the user is provided with feedback on the fluctuating reliability level of the posture of different parts of the user. By providing improved feedback, the usability of the computer system is improved, the user system interface is made more efficient (for example, by helping the user to provide appropriate input when operating and / or interacting with the computer system and reducing user errors), and in addition, power consumption is reduced and the battery life of the computer system is improved by allowing the user to use the system more quickly and efficiently.
[0151] In some embodiments, presenting an avatar (e.g., 721) includes presenting the avatar using a second avatar feature having a first value of a second variable display characteristic (e.g., 721-2 does not have hatching 725 in Figure 7C) according to the determination that the posture of a second part of the user (e.g., 701-2) is associated with a third certainty value (e.g., 710-2 does not have hatching 715 in Figure 7B). By presenting the avatar using a second avatar feature having a first value of a second variable display characteristic according to the determination that the posture of a second part of the user is associated with a third certainty value, the user is provided with feedback indicating the certainty that the posture of the second part of the user represented by the second avatar feature corresponds to the actual posture of the second part of the user, and further, feedback on the fluctuating confidence levels in the postures of different parts of the user. By providing improved feedback, the usability of the computer system is enhanced, the user system interface is made more efficient (for example, by assisting the user in providing appropriate input when operating and / or interacting with the computer system, thereby reducing user errors), and power consumption is reduced and the battery life of the computer system is improved by enabling the user to use the system more quickly and efficiently.
[0152] In some embodiments, presenting an avatar (e.g., 721) includes presenting the avatar using a second avatar feature having a second value of the second variable display feature different from the first value of the second variable display feature, according to the determination that the posture of a second part of the user (e.g., 701-2) is associated with a fourth certainty value different from a third certainty value (e.g., 811-3 and 811-4 have high-density hatching 715 in Figure 8B). (For example, avatar 721 is presented using the right arm of the avatar (including part 721-2) having a smoke effect 830 which is a variable display feature (e.g., similar to that represented by hatching 725 in some embodiments). By presenting an avatar having a second avatar feature with a value, the user is provided with feedback indicating a different degree of certainty that the pose of the second part of the user represented by the second avatar feature corresponds to the actual pose of the second part of the user, and further, feedback is provided on the fluctuating confidence level in the poses of different parts of the user. By providing improved feedback, the usability of the computer system is improved, the user system interface is made more efficient (for example, by helping the user to provide appropriate input when operating and / or interacting with the computer system and reducing user errors), and in addition, power consumption is reduced and the battery life of the computer system is improved by allowing the user to use the system more quickly and efficiently.
[0153] In some embodiments, presenting an avatar (e.g., 721) is determined according to the determination that the third certainty value corresponds to the first certainty value (e.g., is equal to, the same as) (for example, the certainty of the posture of the first part of the user is 50%, 55%, or 60%, and the certainty of the posture of the second part of the user is also 50%, 55%, or 60%) (for example, in Figure 7B, the certainty of the user's right elbow has a moderate hatching density as shown in part 710-6 of the certainty map 710, and the certainty of the user's left elbow also has a moderate hatching density), and the second variable table The first value of the display characteristic corresponds to the first value of the variable display characteristic (e.g., equal to, the same as) (e.g., in Figure 7C, both the left and right elbows of the avatar have a moderate hatching density) (e.g., both the variable display characteristic and the second variable display characteristic have values that indicate 50%, 55%, or 60% certainty of the posture of each part of the user (e.g., the first value of the variable display characteristic indicates 50%, 55%, or 60% certainty of the posture of the first part of the user, and the first value of the second variable display characteristic indicates 50%, 55%, or 60% certainty of the posture of the second part of the user)).
[0154] In some embodiments, presenting an avatar (e.g., 721) is determined to correspond to the second certainty value (e.g., equal to, the same as) the second certainty value (e.g., the certainty of the posture of the first part of the user is 20%, 25%, or 30%, and the certainty of the posture of the second part of the user is 20%, 25%, or 30%) (e.g., in Figure 7B, the certainty of the upper right arm of the user has a low hatching density as shown in part 710-9 of the certainty map 710, and the certainty of the upper left arm of the user also has a low hatching density), and the second variable display feature The second value of the variability characteristic corresponds to the second value of the variable display characteristic (e.g., it is equal to, the same as) (for example, in Figure 7C, both the left upper arm and the right upper arm of the avatar have a low hatching density) (for example, both the variable display characteristic and the second variable display characteristic have values that indicate 20%, 25%, or 30% certainty of the posture of each part of the user (for example, the second value of the variable display characteristic indicates 20%, 25%, or 30% certainty of the posture of the first part of the user, and the second value of the second variable display characteristic indicates 20%, 25%, or 30% certainty of the posture of the second part of the user)).
[0155] In some embodiments, the avatar features and the first parts of the user have the same relationship between certainty and variable display characteristic values (e.g., visual fidelity) as the second avatar features and the second parts of the user. For example, the certainty of the pose of the first part of the user directly corresponds to the value of the variable display characteristic, and the certainty of the pose of the second part of the user directly corresponds to the value of the second variable display characteristic. This is illustrated in the following embodiments illustrating different changes in certainty values. If the certainty of the pose of the first part of the user decreases by 5%, 7%, or 10%, the value of the variable display characteristic is adjusted by an amount indicating a 5%, 7%, or 10% decrease in certainty (e.g., directly or proportionally mapped between certainty and the value of the variable display characteristic), and if the certainty of the pose of the second part of the user increases by 10%, 15%, or 20%, the value of the second display characteristic is adjusted by an amount indicating a 10%, 15%, or 20% increase in certainty (e.g., directly or proportionally mapped between certainty and the value of the second variable display characteristic).
[0156] In some embodiments, Method 900 determines that the first value of the second variable display characteristic does not correspond to the first value of the variable display characteristic, according to the determination that the third certainty value corresponds to the first certainty value (e.g., is equal to, the same as) (e.g., the certainty of the posture of the first part of the user is 50%, 55%, or 60%, and the certainty of the posture of the second part of the user is also 50%, 55%, or 60%) (e.g., in Figure 7B, both part 710-2 and part 710-1 have no hatching as shown in the certainty map 710). Including (e.g., not equal to, different from) (e.g., part 721-2 of the avatar has no hatching, and part 727 of the avatar has hatching 725) (e.g., the variable display characteristic and the second variable display characteristic have values indicating different amounts of certainty of the posture of each part of the user (e.g., the first value of the variable display characteristic indicates 50%, 55%, or 60% certainty of the posture of the first part of the user, and the first value of the second variable display characteristic indicates 20%, 25%, or 30% certainty of the posture of the second part of the user)).
[0157] In some embodiments, Method 900 determines that the second value of the second variable display characteristic corresponds to the second value of the variable display characteristic, according to the determination that the fourth certainty value corresponds to the second certainty value (e.g., is equal to, the same as) (e.g., the certainty of the posture of the first part of the user is 20%, 25%, or 30%, and the certainty of the posture of the second part of the user is 20%, 25%, or 30%) (e.g., in Figure 7B, both parts 710-7 and 710-5 have low-density hatching 715 as shown in the certainty map 710). This includes not having (for example, not being equal to, or being different from) (for example, the left eye 723-2 of the avatar has no hatching, and the collar portion 727 of the avatar has hatching 725) (for example, the variable display characteristic and the second variable display characteristic have values that indicate different amounts of certainty of the posture of each part of the user (for example, the second value of the variable display characteristic indicates 20%, 25%, or 30% certainty of the posture of the first part of the user, and the second value of the second variable display characteristic indicates 40%, 45%, or 50% certainty of the posture of the second part of the user)).
[0158] In some embodiments, the relationship between the certainty of a second avatar feature and a second part of the user and the value of a variable display characteristic (e.g., visual fidelity) differs from that of the respective avatar feature and the first part of the user. For example, in some embodiments, the value of the variable display characteristic is selected to indicate a certainty different from the actual certainty of the posture.
[0159] For example, in some embodiments, the values of the variable display characteristics represent a higher degree of certainty than is actually associated with the corresponding user features. This may be done, for example, when the features are considered important for communication. In this example, rendering the corresponding avatar features using variable display characteristics with values that represent an accurate degree of certainty would be distracting, so rendering the features using variable display characteristics with values that represent a higher degree of certainty (e.g., a high-fidelity representation of each avatar feature) improves communication.
[0160] As another example, in some embodiments, the values of variable representation characteristics represent a lower degree of certainty than they actually are associated with the corresponding user features. This may be done, for example, when features are considered not critical to communication. In this example, rendering features using variable representation characteristics with lower certainty values (e.g., a low-fidelity representation of each avatar feature) typically saves computational resources that would otherwise be extended when rendering each avatar feature using variable representation characteristics with higher certainty values (e.g., a high-fidelity representation of each avatar feature). Since user features are considered not critical to communication, computational resources can be saved without sacrificing communication effectiveness.
[0161] In some embodiments, a computer system (e.g., 101) receives updated pose data representing a change in the posture of a first part (e.g., 701-1, 701-2, 701-3, 701-4) of a user (e.g., 701). In response to receiving the updated pose data, the computer system updates the representation of the avatar (e.g., 721), including updating the pose of each avatar feature based on the change in the posture of the first part of the user (e.g., at least one of its magnitude or direction) (e.g., the pose of each avatar feature is updated by a magnitude and / or direction corresponding to the magnitude and / or direction of the change in the posture of the first part of the user) (e.g., if the left hand of user 701 moves from the upright posture in Figure 7A to the position on the wall 805 in Figure 8A, the left hand of avatar 721 moves from the upright posture in Figure 7C to the position on the wall in Figure 8C). By updating the posture of each avatar feature based on the posture changes of the user's first body part, the user is provided with feedback indicating that the computer system modifies each avatar feature accordingly based on the movement of the user's first body part. This provides a control scheme for manipulating and / or synthesizing a virtual avatar using a display generation component, and the computer system processes morphological input of changes (as well as the magnitude and / or direction of those changes) to the user's physical features, including the user's first body part, and provides morphological feedback of the appearance of the virtual avatar. This improves the visual feedback to the user regarding posture changes of the user's physical features. This improves the usability of the computer system, makes the user system interface more efficient (for example, by helping the user to provide appropriate input when operating and / or interacting with the computer system and reducing user errors), and reduces power consumption and improves the battery life of the computer system by allowing the user to use the computer system more quickly and efficiently. In some embodiments, the avatar's position is updated based on the user's movement. For example, the avatar is updated in real time to mirror the user's movement.For example, if a user puts their arm behind their back, an avatar with the arm moved behind their back will be displayed, mirroring the user's movement.
[0162] In some embodiments, updating the representation of an avatar (e.g., 721) includes changing the position of at least a portion of the avatar based on the user's first-part posture change, in response to a change in the certainty of the user's first-part posture during a change in the user's first-part posture, as well as changing the variable display characteristics of each displayed avatar feature (for example, as the user's right hand moves from the posture in Figure 7A to the posture in Figure 8A, the corresponding avatar feature (the avatar's right hand) changes from no hatching in Figure 7C to having the smoked feature 830 in Figure 8C). By changing the variable display characteristics of each displayed avatar feature in addition to changing the position of at least a portion of the avatar based on the user's first-part posture change, feedback is provided to the user that the certainty of the posture of each avatar feature is affected by the user's first-part posture change. By providing improved feedback, the usability of the computer system is enhanced, the user system interface is made more efficient (for example, by assisting the user in providing appropriate input when operating and / or interacting with the computer system, thereby reducing user errors), and in addition, power consumption is reduced and the battery life of the computer system is improved by enabling the user to use the computer system more quickly and efficiently.
[0163] In some embodiments, updating the avatar representation includes modifying the current value of a variable display characteristic (e.g., increasing or decreasing it depending on the type of variable display characteristic) (e.g., modifying the variable display characteristic) based on the increased certainty in the pose of the user's first part, according to the determination that the certainty in the pose of the user's first part has increased based on updated pose data representing a change in the pose of the user's first part (e.g., the user's first part moves to a position where the certainty of the user's first part increases (e.g., the user's hand moves from a position outside the sensor's field of view (where the camera or other sensor cannot clearly capture the hand's position) to a position within the sensor's field of view (where the camera can clearly capture the hand's position)) (e.g., the user's hand moves from behind an object such as a cup to in front of an object)).
[0164] In some embodiments, updating the avatar representation includes modifying the current value of a variable display characteristic (e.g., increasing or decreasing it depending on the type of variable display characteristic) (e.g., modifying the variable display characteristic) based on the decrease in certainty in the posture of the user's first part, according to the determination that the certainty of the posture of the user's first part has decreased based on updated posture data representing a change in the posture of the user's first part (e.g., the user's first part moves to a position where the certainty of the user's first part decreases (e.g., the user's hand moves from a position within the sensor's field of view (where the camera or other sensor can clearly capture the hand's position) to a position outside the sensor's field of view (where the camera cannot clearly capture the hand's position)) (e.g., the user's hand moves from in front of an object such as a cup to behind an object).
[0165] In some embodiments, the certainty of the pose of a first part of the user (e.g., the user's hand) changes (e.g., increases or decreases) as the first part of the user moves, and the value of the variable display characteristic is updated in real time in accordance with the change in certainty. In some embodiments, the change in the value of the variable display characteristic is represented as a smooth, gradual change in the variable display characteristic applied to each avatar feature. For example, referring to an embodiment in which the certainty value changes as the user moves their hand to a position outside the field of view of the sensor, if the variable display characteristic corresponds to particle density, the density of particles including the avatar's hand increases gradually in harmony with the increase in certainty at the user's hand position as the user moves their hand into the field of view of the sensor. Conversely, the density of particles including the avatar's hand decreases gradually in harmony with the decrease in certainty at the user's hand position as the user moves their hand out of the field of view of the sensor. Similarly, if the variable display characteristic is a blurring effect, the amount of blur applied to the avatar's hand decreases gradually in harmony with the increase in certainty at the user's hand position as the user moves their hand into the field of view of the sensor. Conversely, the amount of blur applied to the avatar's hands gradually increases in harmony with the decrease in certainty regarding the user's hand position as the user moves their hands out of the sensor's field of view.
[0166] In some embodiments, presenting an avatar means presenting an avatar (e.g., 721) using each avatar feature having a third value of variable display characteristics, according to the determination that the posture data satisfies a first set of criteria that are satisfied when a first part of the user (e.g., the user's hands and / or face) is detected by a first sensor (e.g., 705-1, 705-2, 705-4) (e.g., the user's hands and / or face are visible to, detected by, or identified by a camera or other sensor) (e.g., 721-2, 721-3, 721-4, 723-1, 724, 722 are shown without hatching in Figure 7C) (e.g., since the first part of the user is detected by the sensor, each 1) Presenting the avatar using each avatar feature having a fourth variable display characteristic value that shows a lower certainty value than the third variable display characteristic value, according to the determination that the pose data does not satisfy the first set of criteria (e.g., the user's hand is outside the field of view 707-2 in Figure 8A) (e.g., the user's hand and / or face are not visible to the camera or other sensors, are not detected, or are not identified by them) (e.g., the avatar's right hand is displayed in Figure 8C as having a variable display characteristic indicated by the smoke effect 830) (e.g., since the first part of the user is not detected by the sensor, each avatar feature is represented with lower fidelity). Presenting the avatar using each avatar feature having a third or fourth variable display characteristic value depending on whether the first part of the user is detected by the first sensor provides the user with feedback that the certainty of the pose of each avatar feature is affected by whether or not the first part of the user is detected by the sensor.By providing improved feedback, the usability of the computer system is enhanced, the user system interface is made more efficient (for example, by assisting the user in providing appropriate input when operating and / or interacting with the computer system, thereby reducing user errors), and in addition, power consumption is reduced and the battery life of the computer system is improved by enabling the user to use the computer system more quickly and efficiently.
[0167] In some embodiments, while each avatar feature having a current value of a variable display characteristic is presented (e.g., the avatar hand 721-2 is not hatched in Figure 7C) (e.g., a first value of the variable display characteristic) (e.g., a second value of the variable display characteristic) (e.g., the current value of the variable display characteristic corresponds to the current visual fidelity of each avatar feature with respect to a first part of the user), a computer system (e.g., 101) receives updated posture data representing a change in the posture of the first part of the user (e.g., the user moves their hand).
[0168] In response to receiving updated pose data, the computer system (e.g., 101) updates the avatar's display. In some embodiments, updating the avatar representation involves decreasing the current value of a variable display characteristic according to the determination that the updated posture data represents a change in the posture of a first part of the user from a first position within the sensor's field of view (e.g., visible to the sensor within the sensor's field of view) to a second position outside the sensor's field of view (e.g., the user's right hand moves from within the field of view 707-2 in Figure 7A to outside the field of view 707-2 in Figure 8A) (e.g., outside the sensor's field of view, not visible to the sensor) (e.g., the hand moves from a position within the sensor's (e.g., camera's) field of view to a position outside the sensor's field of view) (e.g., the avatar's hand 721-2 is represented by no hatching in Figure 7C and by a smoke effect 830 in Figure 8C) (e.g., the decreased value of the variable display characteristic corresponds to the decreased visual fidelity of each avatar feature with respect to the first part of the user) (e.g., each avatar feature transitions to a display characteristic value that indicates a decrease in certainty in the updated posture of the user's hand). The system provides feedback to the user that moving the user's first part from within the sensor's field of view to outside the sensor's field of view causes a change in the certainty of the pose of each avatar feature (e.g., a decrease in certainty), by decreasing the current value of the variable display characteristic, according to the determination that the updated pose data represents a change in the pose of the user's first part from a first position within the sensor's field of view to a second position outside the sensor's field of view. By providing improved feedback, the system enhances the usability of the computer system, making the user system interface more efficient (e.g., by assisting the user in providing appropriate input when operating and / or interacting with the computer system, thereby reducing user errors), and also reduces power consumption and improves the battery life of the computer system by allowing the user to use the computer system more quickly and efficiently.
[0169] In some embodiments, updating the avatar representation includes increasing the current value of a variable display characteristic according to a determination that the updated pose data represents a change in the pose of a first part of the user from a second position to a first position (for example, the hand moves from a position outside the field of view of a sensor (e.g., a camera) to a position within the field of view of the sensor) (for example, moving part 701-2 from the position in Figure 8A to the position in Figure 7A) (for example, part 721-2 of the avatar is displayed without hatching in Figure 7C, and in Figure 8C, part 721-2 is replaced with a smoke effect 830) (for example, the increased value of the variable display characteristic corresponds to the increased visual fidelity of each avatar feature with respect to the first part of the user) (for example, each avatar feature transitions to a display characteristic value that indicates a decrease in certainty in the updated pose of the user's hand) (for example, each avatar feature transitions to a display characteristic value that indicates an increase in certainty in the updated pose of the user's hand). The system provides feedback to the user that moving the user's first part from outside the sensor's field of view into the sensor's field of view, by increasing the current value of the variable display characteristic according to the determination that the updated pose data represents a change in the pose of the user's first part from a second position to a first position, causes a change in the pose certainty (e.g., increased certainty) of each avatar feature. By providing improved feedback, the system enhances the usability of the computer system, making the user system interface more efficient (e.g., by assisting the user in providing appropriate input when operating and / or interacting with the computer system, thereby reducing user errors), and also reduces power consumption and improves the battery life of the computer system by allowing the user to use the computer system more quickly and efficiently.
[0170] In some embodiments, the current value of the variable display characteristic decreases at a first speed (e.g., if the hand moves to a position outside the sensor's field of view, the decrease occurs at a slow speed (e.g., slower than the detected hand movement)), and the current value of the variable display characteristic increases at a second speed higher than the first speed (e.g., if the hand moves to a position within the sensor's field of view, the increase occurs at a fast speed (e.g., faster than the decrease speed)). By decreasing the current value of the variable display characteristic at a first speed and increasing it at a second speed higher than the first speed, the user is provided with feedback regarding the timing at which increased certainty in the pose of each avatar feature is confirmed. By providing improved feedback, the usability of the computer system is improved, the user system interface is made more efficient (e.g., by helping the user to provide appropriate input when operating and / or interacting with the computer system, thereby reducing user errors), and in addition, power consumption is reduced and the battery life of the computer system is improved by enabling the user to use the computer system more quickly and efficiently.
[0171] In some embodiments, reducing the current value of a variable display characteristic involves reducing the current value of the variable display characteristic at a first speed according to the determination that a second location corresponds to a known position of a first part of the user (for example, the computer system knows the position of the hand even if the user's hand is positioned outside the sensor's field of view) (for example, the reduction of the variable display characteristic (e.g., visual fidelity) occurs at a slow speed (e.g., slower than the hand's movement to be detected) when the hand moves to a known position outside the sensor's field of view). In some embodiments, the position of the first part of the user can be determined by inferring the position (or approximate position) using other data, for example. For example, the user's hand may be outside the sensor's field of view, but the user's forearm is within the field of view, and therefore the position of the hand can be determined based on the known position of the forearm. In another example, the user's hand may be positioned behind the user and therefore outside the sensor's field of view, but the position of the hand can be determined (at least roughly) based on the position of the user's arm. In some embodiments, reducing the current value of a variable display characteristic includes reducing the current value of the variable display characteristic at a second speed higher than the first speed, according to the determination that the second position corresponds to a known location of the first part of the user (for example, the reduction of the variable display characteristic (e.g., visual fidelity) occurs at a high speed (e.g., faster than when the hand position is known) when the hand moves to an unknown position outside the sensor's field of view).
[0172] In some embodiments, a first value of the variable display characteristic represents the visual fidelity of each avatar feature with respect to the posture of the user's first part, which is higher than a second value of the variable display characteristic. In some embodiments, presenting the avatar includes: 1) associating the posture of the user's first part with a second certainty value according to the determination that the user's first part corresponds to a subset of physical features (e.g., the user's arms and shoulders) (e.g., each avatar feature is represented with lower fidelity because the user's first part corresponds to the user's arms and / or shoulders) (e.g., the avatar's neck and collar region 727 is presented using hatching 725 in Figure 7C); and 2) associating the posture of the user's first part with a first certainty value according to the determination that the user's first part does not correspond to a subset of physical features (e.g., each avatar feature is represented with higher fidelity because the user's first part does not correspond to the user's arms and / or shoulders) (e.g., the mouth 724 is rendered without hatching in Figure 7C). Depending on whether the first part of the user corresponds to a subset of physical features, the pose of the first part of the user is associated with a first or second certainty value. This allows for the saving of computational resources (resulting in reduced power consumption and improved battery life) by ceasing the generation and display of features with high fidelity (e.g., rendering features with lower fidelity) when these features are not particularly important for communication purposes, even if the certainty of these features is high. In some embodiments, the user's arms and shoulders are represented (via an avatar) using variable display characteristics that exhibit lower fidelity when they are not within the sensor's field of view or when they are not considered important features for communication. In some embodiments, features considered important for communication include the eyes, mouth, and hands.
[0173] In some embodiments, while an avatar (e.g., 721) is presented using each avatar feature having a first value of the variable display characteristic, a computer system (e.g., 101) updates the representation of the avatar, including: 1) presenting the avatar using each avatar feature having a first modified value of the variable display characteristic (e.g., reduced relative to the current value of the variable display characteristic (e.g., the first value)) (e.g., based on a reduced confidence value) according to a determination that the movement speed of the first part of the user is a first movement speed of the first part of the user; and 2) presenting the avatar using each avatar feature having a second modified value of the variable display characteristic (e.g., increased relative to the current value of the variable display characteristic (e.g., the first value)) (e.g., based on an increased confidence value) according to a determination that the movement speed of the first part of the user is a second movement speed of the first part of the user that is different from (e.g., lower than) the first movement speed. When the user's first part is moving at a first speed, the avatar is presented using each avatar feature having a first modified value of the variable display characteristic. When the user's first part is moving at a second speed, the avatar is presented using each avatar feature having a second modified value of the variable display characteristic. This provides the user with feedback indicating that variations in the user's first part's movement speed affect the certainty of the pose of each avatar feature. By providing improved feedback, the usability of the computer system is enhanced, the user system interface is made more efficient (for example, by assisting the user in providing appropriate input when operating and / or interacting with the computer system, thereby reducing user errors), and power consumption is reduced and the battery life of the computer system is improved by enabling the user to use the computer system more quickly and efficiently.
[0174] In some embodiments, as the movement speed of the user's first part increases, the certainty of the user's first part's position decreases (for example, based on sensor limitations, e.g., camera frame rate), resulting in a corresponding change in the value of the variable display characteristic. For example, if the variable display characteristic corresponds to the density of particles containing each avatar feature, a decrease in the certainty value indicates a decrease in particle density and a lower certainty in the user's first part's posture. In another example, if the variable display characteristic is a blurring effect applied to each avatar feature, a decrease in the certainty value indicates an increase in the blurring effect to indicate a lower certainty in the user's first part's posture. In some embodiments, as the movement speed of the user's first part decreases, the certainty of the user's first part's position increases, resulting in a corresponding change in the value of the variable display characteristic. For example, if the variable display characteristic corresponds to the density of particles containing each avatar feature, an increase in the certainty value indicates an increase in particle density and a higher certainty in the user's first part's posture. In another example, if the variable display characteristic is a blurring effect applied to each avatar feature, then as the certainty value increases, the blurring effect is reduced to indicate greater certainty in the pose of the first part of the user.
[0175] In some embodiments, the computer system (e.g., 101) includes changing the values of variable display characteristics (e.g., 725, 830) and changing one or more visual parameters of each avatar feature (e.g., changing the value of variable display characteristic 725 and / or visual effect 830, as discussed with respect to avatar 721 in Figures 7C and 8C). By changing one or more visual parameters of each avatar feature when changing the values of the variable display characteristics, the user is provided with feedback indicating a change in the reliability of the pose of the first part of the user represented by each avatar feature. By providing improved feedback, the usability of the computer system is improved, the user system interface is made more efficient (e.g., by assisting the user in providing appropriate input when operating and / or interacting with the computer system and reducing user errors), and in addition, power consumption is reduced and the battery life of the computer system is improved by enabling the user to use the computer system more quickly and efficiently.
[0176] In some embodiments, changing the value of a variable display characteristic corresponds to changing the value of one or more visual parameters (for example, increasing the value of a variable display characteristic may correspond to increasing the blur, pixelation, and / or color of each avatar feature). Thus, the variable display characteristic can be modified, altered, and / or adjusted using the methods provided herein for changing one or more visual parameters described herein.
[0177] In some embodiments, one or more visual parameters include blur (e.g., sharpness). In some embodiments, the blur or sharpness of each avatar feature is modified to indicate an increase or decrease in the certainty of the user's first body posture. For example, increasing the blur (decreasing the sharpness) indicates a decrease in the certainty of the user's first body posture, and decreasing the blur (increasing the sharpness) indicates an increase in the certainty of the user's first body posture. For example, in Figure 7C, the increased density of hatching 725 represents a higher degree of blur, and the decreased hatching density represents a lower degree of blur. Thus, avatar 721 is displayed with a blurred waist and a less blurred (sharper) chest.
[0178] In some embodiments, one or more visual parameters include opacity (e.g., transparency). In some embodiments, the opacity or transparency of each avatar feature is modified to indicate an increase or decrease in the certainty of the user's first body posture. For example, increasing opacity (decreasing transparency) indicates an increase in the certainty of the user's first body posture, and decreasing opacity (increasing transparency) indicates a decrease in the certainty of the user's first body posture. For example, in Figure 7C, increased density hatching 725 represents higher transparency (lower opacity), and decreased hatching density represents lower transparency (higher opacity). Thus, avatar 721 is displayed with a more transparent waist and a less transparent (more opaque) chest.
[0179] In some embodiments, one or more visual parameters include color. In some embodiments, the color of each avatar feature is modified to indicate an increase or decrease in the certainty of the user's first body posture, as described above. For example, a skin tone color may be presented to indicate an increase in the certainty of the user's first body posture, and a non-skin tone color (e.g., green, blue) may be presented to indicate a decrease in the certainty of the user's first body posture. For example, in Figure 7C, areas of avatar 721 without hatching 725, such as part 721-2, are displayed using a skin tone color (e.g., brown, black, tan, etc.), while areas with hatching 725, such as the neck and collar area 727, are displayed using a non-skin tone color.
[0180] In some embodiments, one or more visual parameters include the density of particles containing each avatar feature, as described above with respect to various variable display characteristics of the avatar 721.
[0181] In some embodiments, the density of particles containing each avatar feature includes the spacing between particles containing each avatar feature (e.g., average distance). In some embodiments, increasing the particle density includes decreasing the spacing between particles containing each avatar feature, and decreasing the particle density includes increasing the spacing between particles containing each avatar feature. In some embodiments, the particle density is changed to indicate an increase or decrease in the certainty of the user's first body posture. For example, increasing the density (reducing the spacing between particles) indicates an increase in the certainty of the user's first body posture, and decreasing the density (increasing the spacing between particles) indicates an increase in the certainty of the user's first body posture. For example, in Figure 7C, the hatching 725 with increased density represents a larger particle spacing, and the hatching density with decreased density represents a smaller particle spacing. Thus, in this example, avatar 721 is displayed with a high particle spacing at the waist and a small particle spacing at the chest.
[0182] In some embodiments, the density of particles containing each avatar feature includes the size of the particles containing each avatar feature. In some embodiments, increasing the particle density includes decreasing the size of the particles containing each avatar feature (e.g., producing a less pixelated appearance), and decreasing the particle density includes increasing the size of the particles containing each avatar feature (e.g., producing a more pixelated appearance). In some embodiments, the particle density is changed to indicate an increase or decrease in the certainty of the user's first part pose. For example, increasing the density indicates an increase in the certainty of the user's first part pose by reducing the particle size to provide a higher level of detail and / or resolution for each avatar feature. Similarly, decreasing the density indicates a decrease in the certainty of the user's first part pose by increasing the particle size to provide a lower level of detail and / or resolution for each avatar feature. In some embodiments, the density may be a combination of particle size and spacing, and these factors can be adjusted to indicate a higher or lower certainty of the user's first part pose. For example, smaller, more spaced particles may exhibit reduced density (and reduced certainty) compared to larger, more closely spaced particles (which exhibit higher certainty). For instance, in Figure 7C, increased density hatching 725 represents larger particle sizes, while decreased hatching density represents smaller particle sizes. Therefore, in this example, avatar 721 is displayed with larger particle sizes at the waist (the waist appears more pixelated) and smaller particle sizes at the chest (the chest appears less pixelated than the waist).
[0183] In some embodiments, the computer system (e.g., 101) includes changing the value of a variable display characteristic and presenting a visual effect (e.g., 830) associated with each avatar feature (e.g., introducing a visual effect display). By presenting a visual effect associated with each avatar feature when changing the value of the variable display characteristic, the user is provided with feedback that the confidence of the pose of the first part of the user represented by each avatar feature is below a threshold confidence level. By providing improved feedback, the usability of the computer system is improved, the user system interface is made more efficient (e.g., by helping the user to provide appropriate input when operating and / or interacting with the computer system, thereby reducing user errors), and in addition, power consumption is reduced and the battery life of the computer system is improved by allowing the user to use the computer system more quickly and efficiently. In some embodiments, the visual effect is displayed when the corresponding part of the user is below a threshold confidence value and not displayed when it exceeds a threshold confidence value. For example, if the confidence of the pose of the corresponding part of the user is below a 10% confidence value, a visual effect for each avatar feature is displayed. For example, visual effects include a smoke effect or fishscale that appears on the avatar's elbow when the user's elbow is outside the camera's field of view.
[0184] In some embodiments, the first part of the user includes a first physical feature (e.g., the user's mouth) and a second physical feature (e.g., the user's ears or nose). In some embodiments, a first value of the variable display characteristic represents the visual fidelity of each avatar feature with respect to the posture of the user's first physical feature, and is higher than a second value of the variable display characteristic. In some embodiments, presenting an avatar using each avatar feature having a second value of variable display characteristics includes presenting an avatar (e.g., 721) using renderings of first physical features based on the user's corresponding physical features (e.g., the avatar's mouth in Figure 7C is the same as the user's mouth 724 in Figure 7A) and renderings of second physical features based on corresponding physical features that do not belong to the user (e.g., the avatar's nose 722 in Figure 7C is not the same as the user's nose in Figure 7A) (e.g., the avatar includes the user's mouth (e.g., presented using a video feed of the user's mouth) and the avatar includes ears (or nose) that do not belong to the user (e.g., the ears (or nose) are from another person or are mimicked based on a machine learning algorithm)). By presenting an avatar using rendering of first physical features based on the user's corresponding physical features, and rendering of second physical features based on corresponding physical features that do not belong to the user, computational resources can be saved (resulting in reduced power consumption and improved battery life) by discontinuing high fidelity generation and display (for example, rendering some features as if they were based on corresponding physical features that do not belong to the user) when these features are not very important for communication purposes, even if the certainty of these features is high.
[0185] In some embodiments, when an avatar is presented with variable display characteristics exhibiting low fidelity (e.g., low certainty), the user's physical features considered important for communication (e.g., eyes, mouth, hands) are rendered based on the user's actual appearance, while other physical features are rendered based on something other than the user's actual appearance. That is, physical features important for communication are preserved with high fidelity, for example, using a video feed of the features, while physical features not considered important for communication (e.g., skeleton, hairstyle, ears, etc.) are rendered, for example, using similar features selected from a database of avatar features, or using computer-generated features derived based on a machine learning algorithm. Computational resources are conserved while ensuring the ability to effectively communicate with the user represented by the avatar by rendering the avatar with variable display characteristics exhibiting low fidelity (e.g., low certainty), such that important communication features are based on the appearance of the user's corresponding physical features, and unimportant features are based on different appearances. For example, computational resources are conserved by rendering unimportant features with variable display characteristics exhibiting low fidelity. However, when important communication features are rendered in this manner, they may not match the user's corresponding physical features, leading to communication disruption, which can be distracting and even confuse the user's identity represented by the avatar. Therefore, to ensure that communication is not disrupted, important communication features are rendered with high fidelity, thereby ensuring the ability to communicate effectively while still conserving computational resources.
[0186] In some embodiments, posture data is generated from multiple sensors (e.g., 705-1, 705-2, 705-3, 705-4, 810).
[0187] In some embodiments, the multiple sensors include one or more camera sensors (e.g., 705-1, 705-2, 705-3, 705-4) associated with (e.g., the computer system) a computer system (e.g., 101). In some embodiments, the computer system includes one or more cameras configured to capture posture data. For example, the computer system is a headset device having cameras positioned to detect the posture of one or more users in the physical environment around the headset.
[0188] In some embodiments, the multiple sensors include one or more camera sensors separate from the computer system. In some embodiments, posture data is generated using one or more cameras separate from the computer system and configured to capture posture data representing the user's posture. For example, cameras separate from the computer system may include a smartphone camera, a desktop camera, and / or a camera from a headset device worn by the user. Each of these cameras may be configured to generate at least a portion of the posture data by capturing the user's posture. In some embodiments, posture data is generated from multiple camera sources that provide the position or posture of the user (or a part of the user) from different angles and viewpoints.
[0189] In some embodiments, the multiple sensors include one or more non-visual sensors (e.g., 810) (e.g., sensors that do not include a camera sensor). In some embodiments, posture data is generated using one or more non-visual sensors (e.g., accelerometer, gyroscope, etc.), such as proximity sensors or sensors in a smartwatch.
[0190] In some embodiments, the posture of at least a first part of the user is determined using an interpolation function (for example, the posture of the avatar's left hand in part 721-4 of Figure 7C is interpolated). In some embodiments, the first part of the user is hidden by an object (e.g., a cup 702). For example, the user holds a coffee mug so that the tips of the user's fingertips are visible to a sensor capturing posture data, but the proximal region of the hand and fingers is positioned behind the mug so as to be hidden from the sensor capturing posture data. In this example, the computer system can determine the rough posture of the back of the user's fingers and hand by running an interpolation function based on the detected fingertips, the user's wrist, or a combination thereof. This information can be used to increase the confidence in the posture of these features. For example, these features can be represented using variable display characteristics that indicate very high confidence in the posture of these features (e.g., 99%, 95%, 90% confidence).
[0191] In some embodiments, posture data includes data generated from prior scan data (e.g., data from a previous body or face scan (e.g., depth data)) that captures information about the user's appearance to a computer system (e.g., 101). In some embodiments, scan data includes data generated from a scan of the user's body or face (e.g., depth data). For example, this could be data derived from a face scan to unlock or access a device (e.g., a smartphone, smartwatch, or HMD), or from a media library containing photos and videos with depth data. In some embodiments, posture data can be supplemented with posture data generated from prior scan data to increase understanding of the current posture of a part of the user (e.g., based on the posture of a known part of the user). For example, if the user's forearm is of a known length (e.g., 28 inches), then possible hand positions can be determined based on the position and angles of the elbow and part of the forearm. In some embodiments, pose data generated from prior scan data (e.g., by providing eye color, jaw shape, ear shape, etc.) can be used to increase the visual fidelity of a replica of a user part that is not visible to the computer system's sensors. For example, if the user is wearing a hat that covers their ears, prior scan data can be used to increase the certainty of the user's ear posture (e.g., position and / or orientation) based on prior scans that can be used to derive the user's ear posture data. In some embodiments, the device that generates the replica of the user part is separate from the device (e.g., smartphone, smartwatch, HMD), and data from the face scan is provided to the device that generates the replica of the user part for use in constructing the replica of the user part (securely and confidentially, for example, with one or more options for the user to decide whether or not to share data between devices).In some embodiments, the device that generates the user's part replica is the same as the device (e.g., a smartphone, smartwatch, or HMD), and data from a facial scan is provided to the device that generates the user's part replica for use in constructing the user's part replica (for example, the facial scan is used to unlock the HMD that generates the user's part replica).
[0192] In some embodiments, posture data includes data generated from prior media data (e.g., photos and / or videos of the user). In some embodiments, the prior media data includes image data and optionally includes depth data associated with previously captured photos and / or videos of the user. In some embodiments, prior media data can be used to supplement posture data and increase certainty about the current posture of parts of the user. For example, if a user is wearing glasses that obscure their eyes from sensors capturing posture data, prior media data can be used to increase certainty about the user's eye posture (e.g., position and / or orientation) based on existing photos and / or videos that can be accessed (e.g., securely and confidentially with one or more options for the user to decide whether or not to share data between devices) to derive the user's eye posture data.
[0193] In some embodiments, the pose data includes video data (e.g., a video feed) that includes at least a first portion of the user (e.g., 701-1). In some embodiments, presenting an avatar includes presenting a modeled avatar (e.g., 721) (e.g., a three-dimensional computer-generated model of the avatar) that includes each avatar feature (e.g., the avatar's mouth 724, the avatar's right eye 723-1) rendered using video data that includes the first portion of the user. In some embodiments, the modeled avatar is a three-dimensional computer-generated avatar, and each avatar feature is rendered as a video feed of the first portion of the user mapped onto the modeled avatar. For example, the avatar may be a three-dimensional imitation avatar model (e.g., green) that includes an eye shown as a video feed of the user's eye.
[0194] In some embodiments, presenting an avatar includes 1) presenting the avatar using each avatar feature and a first quantity (e.g., percentage, quantity) of avatar features other than each avatar feature, in accordance with the determination that an input indicating a first rendering value (e.g., a low rendering value) of the avatar has been received (e.g., before, during, or after representing the avatar), and 2) presenting the avatar using each avatar feature and a second quantity (e.g., percentage, quantity) of avatar features other than each avatar feature that is different from (e.g., greater than) the first quantity, in accordance with the determination that an input indicating a second rendering value (e.g., a high rendering value) of the avatar, which is different from the first rendering value, has been received (e.g., before, during, or after representing the avatar). By displaying a second rendering value of the avatar, which is different from the first rendering value, and presenting the avatar using each avatar feature, as well as a second quantity of avatar features other than each avatar feature, even if the pose of those features is highly certain, if the display of those features is not desired, the generation and display of avatar features can be stopped, thereby saving computational resources (reducing power consumption and improving battery life).
[0195] In some embodiments, the rendering value indicates the user's preference for rendering an amount of the avatar that does not correspond to a first part of the user. In other words, the rendering value is used to select the amount of the avatar to display (other than each avatar feature). For example, the rendering value can be selected from a sliding scale where, at one end of the scale, parts of the avatar other than those corresponding to the user's tracked features (e.g., each avatar feature) are not rendered, and at the other end of the scale, the entire avatar is rendered. For example, if the first part of the user is the user's hands and the lowest rendering value is selected, the avatar will appear as floating hands (each avatar feature becoming a hand corresponding to the user's hands). By selecting a rendering value, the user can customize the appearance of the avatar to increase or decrease the contrast between the parts of the avatar that correspond to the tracked user features and those that do not. This makes it easier for the user to identify which parts of the avatar are more credible and trustworthy.
[0196] In some embodiments, a computer system (e.g., 101) causes a representation (e.g., 735) of a user associated with a display generation component (e.g., the aforementioned user, or a second user if the aforementioned user is not associated with the display generation component) via a display generation component (e.g., 730), the representation of the user associated with the display generation component corresponding to the appearance of the user associated with the display generation component as presented to one or more other users. By causing a representation of the user associated with the display generation component corresponding to the appearance of the user associated with the display generation component as presented to one or more other users, the user is provided with feedback associated with the display generation component regarding the user's appearance as viewed by other users. By providing improved feedback, the usability of the computer system is improved, the user system interface is made more efficient (e.g., by helping the user to provide appropriate input when operating and / or interacting with the computer system and reducing user errors), and in addition, power consumption is reduced and the battery life of the computer system is improved by enabling the user to use the system more quickly and efficiently.
[0197] In some embodiments, a user viewing a display generation component is also presented to another user (for example, on a different display generation component). For example, two users communicate with each other as two different avatars presented in a virtual environment, and each user can view the other user on their respective display generation component. In this example, the display generation component of one of the users (e.g., the second user) displays the appearance of the other user (e.g., the first user) as well as the appearance of the second user as it appears when presented to the other user (e.g., the first user).
[0198] In some embodiments, presenting an avatar includes presenting the avatar using each avatar feature having a first appearance based on a first appearance of the user's first part (for example, if the user's chin is neatly shaved, each avatar feature represents the user's chin having a neatly shaved appearance). In some embodiments, a computer system (e.g., 101) receives data indicating an updated appearance of the user's first part (e.g., receives a recent photograph or video showing that the user currently has a beard on their chin) (e.g., receives recent face scan data showing that the user currently has a beard on their chin), and causes a display generation component (e.g., 730) to present an avatar (e.g., 721) using each avatar feature having an updated appearance based on the updated appearance of the user's first part (e.g., the user currently has a beard). By causing the avatar representation to be presented using each avatar feature having an updated appearance based on the updated appearance of the user's first part, an accurate representation of the user's first part is provided through the current representation of each avatar feature without manually manipulating the avatar or performing a registration process that incorporates updates to the user's first part. This provides an improved control scheme for editing or presenting custom appearances of avatars, which can reduce the input required to create custom appearances of avatars compared to when different control schemes are used (e.g., control schemes that require the manipulation of individual control points to build or correct avatars). By reducing the number of inputs required to perform tasks, it improves the usability of the computer system, makes the user system interface more efficient (e.g., by helping users provide appropriate inputs when operating / interacting with the computer system and reducing user errors), and reduces power consumption and improves the battery life of the computer system by allowing users to use the system more quickly and efficiently.
[0199] In some embodiments, the appearance of the avatar is updated to match the user's updated appearance (e.g., a new haircut, beard) based on recently captured images or scanned data. In some embodiments, data indicating the updated appearance of the user's first part is captured in a separate action, such as when unlocking a device. In some embodiments, the data is collected on a device separate from the computer system. For example, the data is collected when the user unlocks a device such as a smartphone, tablet, or other computer, and the data is communicated from the user's device to the computer system for subsequent use (securely or secretly), with one or more options for the user to decide whether or not to share the data (between devices). Such data is transmitted and stored in a manner that conforms to industry standards for securing personally identifiable information.
[0200] In some embodiments, posture data further represents objects associated with a first part of the user (e.g., 702, 805) (e.g., the user is holding a cup, the user is leaning against a wall, the user is sitting in a chair). In some embodiments, presenting the avatar includes presenting the avatar along with representations of objects adjacent to each avatar feature (e.g., overlapping with it, in front of it, touching it, interacting with it) (e.g., 825 in Figure 8C). By presenting the avatar along with representations of objects adjacent to each avatar feature, the user is provided with contextual feedback on the posture of the first part of the user. By providing improved feedback, the usability of the computer system is improved, the user system interface is made more efficient (e.g., by helping the user to provide appropriate input when operating and / or interacting with the computer system, thereby reducing user errors), and in addition, power consumption is reduced and the battery life of the computer system is improved by allowing the user to use the system more quickly and efficiently.
[0201] In some embodiments, the avatar is modified to include a representation of the object the user is interacting with (for example, in Figure 8C, avatar 721 also includes a wall rendering 825). This improves communication by providing context for the avatar's posture (which may be based on the user's posture). For example, if the user is holding a cup, each avatar feature (e.g., the avatar's arm) is modified to include a representation of the cup in the avatar's hand. As another example, if the user is leaning against a wall, each avatar feature (e.g., the avatar's shoulder and / or arm) is modified to include a representation of at least a portion of the wall. As yet another example, if the user is sitting in a chair or positioned at a desk, each avatar feature (e.g., the avatar's leg and / or arm) is modified to include a representation of at least a portion of the chair and / or desk. In some embodiments, the object the user is interacting with is a virtual object. In some embodiments, the object the user is interacting with is a real object. In some embodiments, the representation of the object is represented as a virtual object. In some embodiments, the representation of the object is represented using a video feed showing the object.
[0202] In some embodiments, the posture of a first part of the user is associated with a fifth certainty value. In some embodiments, presenting an avatar (e.g., 721) involves presenting the avatar using each avatar feature having a variable display characteristic value that indicates a lower certainty value than the fifth certainty value, according to the determination that the first part of the user is a first feature type (e.g., 701-1 includes the user's neck and collar region) (e.g., a type of feature considered not important for communication) (e.g., in Figure 7C, the avatar's neck and collar region 727 is shown using hatching 725) (e.g., even if the certainty of the user's first part posture is high, each avatar feature is presented using a variable display characteristic that indicates a low certainty value of the user's first part posture). When the first part of the user is of the first feature type, presenting the avatar using avatar features with variable display characteristic values that show a certainty value lower than the fifth certainty value allows for the saving of computational resources (e.g., rendering features with lower fidelity) when these features are not very important for communication purposes, even if the pose certainty of these features is high. This saves computational resources (resulting in reduced power consumption and improved battery life).
[0203] In some embodiments, if the pose of a user feature is highly certain but the feature is not considered important for communication, the corresponding avatar feature is rendered using a variable display characteristic that exhibits lower certainty (e.g., the avatar's neck and collar region 727 representing the variable display characteristic is shown using hatching 725 in Figure 7C) (or at least lower certainty than is actually associated with the user feature). For example, a camera sensor can capture the position of a user's knees (and therefore has high certainty of their position), but knees are generally not considered important for communication purposes, so the avatar's knees are presented using a variable display characteristic that exhibits lower certainty (e.g., lower fidelity). However, if a user feature (e.g., the user's knees) is important for communication, the variable display characteristic (e.g., fidelity) value for each avatar feature (e.g., the avatar's knees) is adjusted (e.g., increased) to accurately reflect the pose certainty of the user feature (e.g., knees). In some embodiments, rendering each avatar feature using a variable display feature with a lower certainty value when the feature is considered not critical to communication generally ensures that when rendering each avatar feature to a variable display feature, the computational resources required to extend each avatar feature to a variable display feature with a higher certainty value (e.g., a high-fidelity representation of each avatar feature). Since user features are considered not critical to communication, computational resources can be secured without sacrificing communication effectiveness.
[0204] In some embodiments, the pose of the user's first part is associated with a sixth certainty value. In some embodiments, presenting an avatar (e.g., 721) further includes presenting the avatar using each avatar feature having a variable display characteristic value that indicates a higher certainty value than the sixth certainty value, according to the determination that the user's first part is of a second feature type (e.g., the user's left eye in Figure 7A) (e.g., a type of feature considered important for communication (e.g., hands, mouth, eyes)) (e.g., in Figure 7C, the avatar's left eye 723-2 is presented without hatching) (e.g., even if the certainty of the user's first part's pose is not high, each avatar feature is presented using a variable display characteristic that indicates a high certainty of the user's first part's pose). This improves communication using a computer system by presenting the avatar with each avatar feature having a variable display characteristic value that exhibits a higher certainty value than the sixth certainty value, when the first part of the user is of the second feature type, thereby reducing the distraction caused by rendering avatar features (especially those important for communication purposes) with relatively low fidelity. This improves the usability of the computer system, makes the user system interface more efficient (for example, by assisting the user in providing appropriate input when operating and / or interacting with the computer system and reducing user errors), and reduces power consumption and improves the battery life of the computer system by allowing the user to use the computer system more quickly and efficiently.
[0205] In some embodiments, if the pose of a user feature is not highly certain but the feature is considered important for communication, the corresponding avatar feature is rendered using a variable display characteristic with a value indicating very high certainty (e.g., the avatar's left eye 723-2 is rendered without hatching in Figure 7C) (or at least with a higher certainty than is actually associated with the user feature). For example, the user's mouth is considered important for communication. Therefore, even if the certainty of the user's mouth pose could be 75%, 80%, 85%, or 90%, each avatar feature is rendered using a variable display characteristic that indicates high certainty (e.g., 100%, 97%, or 95% certainty) (e.g., high fidelity). In some embodiments, rendering the corresponding avatar feature using a variable display characteristic with a value indicating an exact certainty level would hinder communication (e.g., by distracting the user), so features considered important for communication despite lower certainty are rendered using a variable display characteristic with a value indicating higher certainty (e.g., a high fidelity representation of each avatar feature).
[0206] In some embodiments, very high certainty (or very high certainty) may refer to a certainty amount that exceeds a predetermined (e.g., a first) certainty threshold, such as 90%, 95%, 97%, or 99% certainty. In some embodiments, high certainty (or high certainty) may refer to a certainty amount that exceeds a predetermined (e.g., a second) certainty threshold, such as 75%, 80%, or 85% certainty. In some embodiments, high certainty is optionally lower than very high certainty but higher than relatively high certainty. In some embodiments, relatively high certainty (or relatively high certainty) may refer to a certainty amount that exceeds a predetermined (e.g., a third) certainty threshold, such as 60%, 65%, or 70% certainty. In some embodiments, relatively high certainty is optionally lower than high certainty and / or very high certainty but higher than low certainty. In some embodiments, a relatively low degree of certainty (or relatively low certainty) may refer to a certainty below a predetermined (e.g., a fourth) certainty threshold, such as 40%, 45%, 50%, or 55% certainty. In some embodiments, a relatively low degree of certainty is lower than a relatively high degree of certainty, a high degree of certainty, and / or a very high degree of certainty, but optionally higher than a low degree of certainty. In some embodiments, a low degree of certainty (or low degree of certainty, or slight certainty) may refer to a certainty below a predetermined certainty threshold, such as 45%, 50%, 55%, 60%, 65%, 70%, 75%, or 80% certainty. In some embodiments, a low degree of certainty is lower than a relatively low degree of certainty, a relatively high degree of certainty, a high degree of certainty, and / or a very high degree of certainty, but optionally higher than a very low degree of certainty. In some embodiments, very low certainty (or very low certainty, or very slight certainty) may refer to a certainty that falls below a predetermined (e.g., fifth) certainty threshold, such as 5%, 15%, 20%, 30%, 40%, or 45% certainty. In some embodiments, very low certainty is lower than low certainty, relatively low certainty, relatively high certainty, high certainty, and / or very high certainty.
[0207] Figures 10A-10B, 11A-11B, and 12 illustrate examples of how a user is represented in a CGR environment as a virtual avatar character having appearances based on different appearance templates. In some embodiments, the avatar has an appearance based on a character template or an abstract template. In some embodiments, the avatar changes its appearance by changing its posture. In some embodiments, the avatar has an appearance based on a transition between a character template and an abstract template. The processes disclosed herein are implemented using a computer system (e.g., computer system 101 in Figure 1), as described above.
[0208] Figure 10A shows four different users located in real-world environments. User A is in real-world environment 1000-1, User B is in real-world environment 1000-2, User C is in real-world environment 1000-3, and User D is in real-world environment 1000-4. In some embodiments, real-world environments 1000-1 to 1000-4 are different real-world environments (e.g., different locations). In some embodiments, one or more of real-world environments 1000-1 to 1000-4 are the same environment. For example, real-world environment 1000-1 may be the same real-world environment as real-world environment 1000-4. In some embodiments, one or more of real-world environments 1000-1 to 1000-4 may be different locations within the same environment. For example, real-world environment 1000-1 may be a first location in a room, and real-world environment 1000-4 may be a second location in the same room.
[0209] In some embodiments, the real-world environment 1000-1 to 1000-4 is similar to the real-world environment 700 and is a motion capture studio including cameras 1005-1 to 1005-5 similar to cameras 705-1 to 705-4 for capturing data (e.g., image data and / or depth data) that can be used to determine the posture of one or more parts of the user. The computer system uses the determined posture of one or more parts of the user to display an avatar having a specific posture and / or appearance template. In some embodiments, the computer system presents the user in the CGR environment as an avatar having an appearance template determined based on different criteria. For example, in some embodiments, the computer system determines the type of activity each user is performing (e.g., based on the posture of one or more parts of the user), and the computer system presents the user in the CGR environment as an avatar having an appearance template determined based on the type of activity the user is performing. In some embodiments, the computer system presents the user in the CGR environment as an avatar having an appearance template determined based on the computer system's focal position (e.g., gaze position) of the user. For example, when the user is focusing on the first avatar rather than the second avatar (e.g., looking at the first avatar), the first avatar with the first appearance template (e.g., character template) is presented, and then the second avatar with the second appearance template (e.g., abstract template) is presented. When the user shifts their focus to the second avatar, the second avatar with the first appearance template is presented (the second avatar transitions from the second appearance template to the first appearance template), and then the first avatar with the second appearance template is presented (the first avatar transitions from the first appearance template to the second appearance template).
[0210] Figure 10A shows User A in real-world environment 1000-1 with their arms raised and speaking. Cameras 1005-1 and 1005-2 capture the posture of User A in their respective fields of view 1007-1 and 1007-2 while User A is speaking. User B in real-world environment 1000-2 is shown reaching for item 1006 on a table. Camera 1005-3 captures the posture of User B in their field of view 1007-3 while User B is reaching for item 1006. User C in real-world environment 1000-3 is shown reading book 1008. Camera 1005-4 captures the posture of User C in their field of view 1007-4 while User C is reading. User D in real-world environment 1000-4 is shown with their arms raised and speaking. Camera 1005-5 captures the posture of the portion of user D within the field of view 1007-5 while user D is conversing. In some embodiments, users A and D are conversing with each other. In some embodiments, users A and D are conversing with a third party, such as a user of a computer system.
[0211] Cameras 1005-1 to 1005-5 are the same as cameras 705-1 to 705-4 described above. Therefore, cameras 1005-1 to 1005-5 capture the poses of user A, B, C, and D, respectively, as described above with respect to Figures 7A to 7C, Figures 8A to 8C, and Figure 9. For brevity, the details will not be repeated below.
[0212] Cameras 1005-1 to 1005-5 are described as non-limiting examples of devices for capturing the poses of user parts A, B, C, and D. Therefore, other sensors and / or devices can be used in addition to, or instead of, any of cameras 1005-1 to 1005-5 to capture the poses of user parts. Examples of such other sensors and / or devices are described above with respect to Figures 7A to 7C, 8A to 8C, and 9. For brevity, the details of these embodiments will not be repeated below.
[0213] Referring here to Figure 10B, the computer system receives sensor data generated using cameras 1005-1 to 1005-5 and displays avatars 1010-1 to 1010-4 within the CGR environment 1020 via the display generation component 1030. In some embodiments, the display generation component 1030 is similar to the display generation component 120 and the display generation component 730.
[0214] Avatars 1010-1, 1010-2, 1010-3, and 1010-4 represent users A, B, C, and D, respectively, within the CGR environment 1020. Figure 10B shows avatars 1010-1 and 1010-4 having appearances based on character templates, and avatars 1010-2 and 1010-3 having appearances based on abstract templates. In some embodiments, character templates include facial features such as the avatar's face, arms, hands, or other avatar features corresponding to parts of the human body, while abstract templates do not include such features, or such features are indistinguishable. In some embodiments, avatars having appearances based on abstract templates have amorphous shapes.
[0215] In some embodiments, the computer system determines the appearance of avatars 1010-1 to 1010-4 based on the poses of each user A, B, C, and D. For example, in some embodiments, if the computer system determines that a user is performing a first type of activity (for example, based on the user's pose), the computer system renders a corresponding avatar having an appearance based on a character template. Conversely, if the computer system determines that a user is performing a second (different) type of activity, the computer system renders a corresponding avatar having an appearance based on an abstract template.
[0216] In some embodiments, the first type of activity is an interactive activity (e.g., an activity involving interaction with other users), and the second type of activity is a non-interactive activity (e.g., an activity that does not involve interaction with other users). In some embodiments, the first type of activity is an activity performed in a specific location (e.g., the same location as the computer system user), and the second type of activity is an activity performed in a different location (e.g., a location remote from the computer system user). In some embodiments, the first type of activity is a manual activity (e.g., an activity involving the user's hands, such as touching an object, holding an object, moving an object, or using an object), and the second type of activity is a non-manual activity (e.g., an activity that does not generally involve the user's hands).
[0217] In the embodiments shown in Figures 10A and 10B, the computer system represents user A in the CGR environment 1020 as avatar 1010-1 having an appearance based on a character template. In some embodiments, the computer system presents avatar 1010-1 having an appearance based on a character template in response to detecting the posture of user A in Figure 10A. For example, in some embodiments, the computer system determines that user A is conversing based on posture data captured using camera 1005-1 and optionally camera 1005-2. In some embodiments, the computer system considers the conversation to be an interactive activity type and therefore presents avatar 1010-1 having an appearance based on a character template. As another example, in some embodiments, the computer system determines that user A is in the same location as the user of the computer system based on posture data (e.g., the positional component of the posture data) and therefore presents avatar 1010-1 having an appearance based on a character template. In some embodiments, the computer system determines that the user of the computer system is focusing on avatar 1010-1 (for example, looking at avatar 1010-1), and therefore the computer system presents avatar 1010-1 having an appearance based on a character template.
[0218] In the embodiments shown in Figures 10A and 10B, the computer system represents user B in the CGR environment 1020 as avatar 1010-2 having an appearance based on an abstract template. In some embodiments, the computer system presents avatar 1010-2 having an appearance based on an abstract template in response to detecting user B's posture in Figure 10A. For example, in some embodiments, the computer system determines that user B is reaching for object 1006 based on posture data captured using camera 1005-3. In some embodiments, the computer system considers this to be a non-interactive activity (e.g., user B is doing something other than interacting with another user) and therefore presents avatar 1010-2 having an appearance based on an abstract template. In some embodiments, the computer system considers this to be a non-manual activity (e.g., user B currently has nothing in their hand) and therefore presents avatar 1010-2 having an appearance based on an abstract template. As another example, in some embodiments, the computer system determines, based on posture data (e.g., the positional component of posture data), that user B is in a remote location from the user of the computer system, and therefore presents avatar 1010-2 having an appearance based on an abstract template. In some embodiments, the computer system determines that the user of the computer system is not focusing on avatar 1010-2 (e.g., not looking at avatar 1010-2), and therefore presents avatar 1010-2 having an appearance based on an abstract template.
[0219] In the embodiments shown in Figures 10A and 10B, the computer system represents user C in the CGR environment 1020 as avatar 1010-3 having an appearance based on an abstract template. In some embodiments, the abstract template may have different abstract appearances. Thus, avatars 1010-2 and 1010-3 may both be based on an abstract template but have different appearances from each other. For example, the abstract template may include different appearances, such as amorphous masses or different abstract shapes. In some embodiments, there may be different abstract templates that provide different abstract appearances. In some embodiments, the computer system presents avatar 1010-3 having an appearance based on an abstract template in response to detecting the posture of user C in Figure 10A. For example, in some embodiments, the computer system determines that user C is reading a book 1008 based on posture data captured using camera 1005-4. In some embodiments, the computer system considers reading to be a non-interactive activity and therefore presents avatar 1010-3 having an appearance based on an abstract template. As another example, in some embodiments, the computer system determines, based on posture data (e.g., the positional component of posture data), that user C is in a remote location from the user of the computer system, and therefore presents avatar 1010-3 having an appearance based on an abstract template. In some embodiments, the computer system determines that the user of the computer system is not focusing on avatar 1010-3 (e.g., not looking at avatar 1010-3), and therefore presents avatar 1010-3 having an appearance based on an abstract template.
[0220] In the embodiments shown in Figures 10A and 10B, the computer system represents user D in the CGR environment 1020 as an avatar 1010-4 having an appearance based on a character template. In some embodiments, the computer system presents the avatar 1010-4 having an appearance based on a character template in response to detecting the posture of user D in Figure 10A. For example, in some embodiments, the computer system determines that user D is conversing based on posture data captured using camera 1005-5. In some embodiments, the computer system considers the conversation to be an interactive activity type and therefore presents the avatar 1010-4 having an appearance based on a character template. As another example, in some embodiments, the computer system determines that user D is in the same location as the user of the computer system based on posture data (e.g., the positional component of the posture data) and therefore presents the avatar 1010-4 having an appearance based on a character template. In some embodiments, the computer system determines that the user of the computer system is focusing on avatar 1010-4 (for example, looking at avatar 1010-4), and therefore the computer system presents avatar 1010-4 having an appearance based on a character template.
[0221] In some embodiments, avatars having appearances based on character templates are presented with poses determined based on the poses of the corresponding users. For example, in Figure 10B, avatar 1010-1 has a pose that matches the pose determined for user A, and avatar 1010-4 has a pose that matches the pose determined for user D. In some embodiments, the poses of avatars having appearances based on character templates are determined in the same way as described above with respect to Figures 7A-7C, 8A-8C, and 9. For example, avatar 1010-1 is displayed with a pose determined based on the poses of parts of user A detected in fields of view 1007-1 and 1007-2. Since the entire body of user A is within the camera's field of view, the computer system determines the poses of the corresponding parts of avatar 1010-1 with maximum confidence, as indicated by the absence of dashed lines on avatar 1010-1 in Figure 10B. As another example, avatar 1010-4 is displayed with a pose determined based on the poses of parts of user D detected in field of view 1007-5 of camera 1005-5. Since user D's legs are outside the field of view 1007-5, the computer system determines the pose of avatar 1010-4's legs with a certain degree of uncertainty (lower than maximum certainty) and presents avatar 1010-4 with the determined pose and variable display characteristics represented by dashed lines 1015 on the avatar's legs, as shown in Figure 10B. For brevity, further details for determining the avatar's pose and displaying the variable display characteristics will not be repeated for all embodiments discussed with respect to Figures 10A, 10B, 11A, and 11B.
[0222] In some embodiments, the computer system updates the appearance of the avatar based on changes in the user's posture. In some embodiments, the appearance update includes changes in the avatar's posture without transitions between different appearance templates. For example, in response to detection that the user has raised and then lowered an arm, the computer system presents the avatar as an avatar character that changes posture along with the user (e.g., in real time) by raising and then lowering the corresponding arms of the avatar. In some embodiments, the avatar's appearance update includes transitions between different appearance templates based on changes in the user's posture. For example, in response to detection that the user has changed posture from one corresponding to a first activity type to one corresponding to a second activity type, the computer system presents the avatar transitioning from a first appearance template (e.g., a character template) to a second appearance template (e.g., an abstract template). Examples of avatars with updated appearances are discussed in more detail below with respect to Figures 11A and 11B.
[0223] Figures 11A and 11B are similar to Figures 10A and 10B, except that users A, B, C, and D have updated poses, and their corresponding avatars 1010-1 to 1010-4 have updated appearances. For example, a user moves from the pose in Figure 10A to the pose in Figure 11A, and the computer system responds by changing the appearance of the avatar from the appearance shown in Figure 10B to the appearance shown in Figure 11B.
[0224] Referring to Figure 11A, User A is shown still conversing, with their right arm down and their head turned forward. User B is shown holding item 1006 with their head facing forward. User C is waving and talking, with book 1008 hanging down beside them. User D is looking to the side with their arms down.
[0225] In Figure 11B, the computer system updates the appearance of avatars 1010-1 to 1010-4. Specifically, avatar 1010-1, which has an appearance based on a character template, remains displayed, while avatars 1010-2 and 1010-3 transition from appearances based on an abstract template to appearances based on a character template, and avatar 1010-4 transitions from appearances based on a character template to appearances based on an abstract template.
[0226] In some embodiments, the computer system, in response to detecting the posture of user A in Figure 11A, displays an avatar 1010-1 having an appearance based on a character template. For example, in some embodiments, the computer system determines, based on posture data captured using camera 1005-1 and optionally camera 1005-2, that user A is still conversing. In some embodiments, the computer system considers the conversation to be an interactive activity type and therefore presents an avatar 1010-1 having an appearance based on a character template. As another example, in some embodiments, the computer system determines, based on posture data (e.g., the positional component of the posture data), that user A is in the same location as the user of the computer system and therefore presents an avatar 1010-1 having an appearance based on a character template. In some embodiments, the computer system determines that the user of the computer system is focusing on avatar 1010-1 (e.g., looking at avatar 1010-1), and therefore presents an avatar 1010-1 having an appearance based on a character template.
[0227] In some embodiments, the computer system, in response to detecting the posture of user B in Figure 11A, displays avatar 1010-2 having an appearance based on a character template. For example, in some embodiments, the computer system determines, based on posture data captured using camera 1005-3, that user B is interacting with another user by shaking item 1006, and therefore presents avatar 1010-2 having an appearance based on a character template. In some embodiments, the computer system determines, based on posture data captured using camera 1005-3, that user B is performing a manual activity by holding item 1006, and therefore presents avatar 1010-2 having an appearance based on a character template. In some embodiments, the computer system determines that the user of the computer system is currently focusing on avatar 1010-2 (e.g., looking at avatar 1010-2), and therefore presents avatar 1010-2 having an appearance based on a character template.
[0228] In some embodiments, the computer system displays an avatar 1010-3 having an appearance based on a character template in response to detecting the posture of user C in Figure 11A. For example, in some embodiments, the computer system determines, based on posture data captured using camera 1005-4, that user C is interacting with another user by waving and / or speaking, and therefore presents an avatar 1010-3 having an appearance based on a character template. In some embodiments, the computer system determines that the user of the computer system is currently focusing on avatar 1010-3 (e.g., looking at avatar 1010-3), and therefore presents an avatar 1010-3 having an appearance based on a character template.
[0229] In some embodiments, the computer system, in response to detecting the posture of user D in Figure 11A, displays an avatar 1010-4 having an appearance based on an abstract template. For example, in some embodiments, the computer system determines, based on posture data captured using camera 1005-5, that user D is no longer interacting with other users because user D has their hands down, is looking away, and / or is no longer speaking. Therefore, the computer system presents an avatar 1010-4 having an appearance based on an abstract template. In some embodiments, the computer system determines that the user of the computer system is no longer focusing on avatar 1010-4 (e.g., not looking at avatar 1010-4), so the computer system presents an avatar 1010-4 having an appearance based on an abstract template.
[0230] As shown in Figure 11B, the computer system updates the pose of each avatar based on the user pose detected in Figure 11A. For example, the computer system updates the pose of avatar 1010-1 to match the pose of user A in Figure 11A. In some embodiments, the user of the computer system is not focusing on avatar 1010-1, but avatar 1010-1 is in the same location as the user of the computer system, or the user of the computer system is conversing, and in some embodiments the computer system considers this to be an interactive activity, so avatar 1010-1 still has an appearance based on the character template.
[0231] As shown in Figure 11B, the computer system also updates the appearance of avatar 1010-2 to have an appearance with a posture determined based on a character template and the posture of user B in Figure 11A. Since some parts of user B are outside the field of view 1007-3, the computer system determines the posture of the corresponding parts of avatar 1010-2 with a certain degree of uncertainty, as indicated by the dashed lines 1015 of avatar 1010-2's right arm and both legs. Also, since user B's right arm is outside the field of view 1007-3, the right arm shown in Figure 11A has a different posture than the actual posture of user B's right arm, and the posture shown in Figure 11B is the posture of avatar 1010-2's right arm determined (estimated) by the computer system. In some embodiments, since item 1006 is within the field of view 1007-3, item 1006 is represented in the CGR environment 1020 as indicated by reference 1006' in Figure 11B. In some embodiments, item 1006 is represented as a virtual object within the CGR environment 1020. In some embodiments, item 1006 is represented as a physical object within the CGR environment 1020. In some embodiments, since the computer system detects that item 1006 is in the hand of user B, the left hand of avatar 1010-2, and optionally, the representation of item 1006 within the CGR environment 1020, are shown with higher fidelity (for example, with higher fidelity than avatar 1010-2 and / or other parts of other objects within the CGR environment 1020).
[0232] As shown in Figure 11B, the computer system also updates the appearance of avatar 1010-3 to have an appearance having a posture determined based on a character template and the posture of user C in Figure 11A. Since some parts of user C are outside the field of view 1007-4, the computer system determines the posture of the corresponding parts of avatar 1010-3 with a certain degree of uncertainty, as indicated by the dashed lines 1015 of avatar 1010-3's left hand and both legs. In some embodiments, the computer system detects book 1008 within the field of view 1007-3 in Figure 10A, and no longer detects book 1008 or user C's left hand within the field of view 1007-3 in Figure 11A, so the computer system determines that book 1008 may be user C's left hand and therefore represents book 1008 in the CGR environment 1020 as indicated by reference 1008' in Figure 11B. In some embodiments, the computer system optionally represents the book 1008 as a virtual object having variable display characteristics that indicate uncertainty regarding the existence and orientation of the book 1008 within the real environment 1000-3.
[0233] In the embodiments described herein, a user of the computer system is referred to. In some embodiments, the user of the computer system is a user browsing the CGR environment 720 using the display generation component 730, or a user browsing the CGR environment 1020 using the display generation component 1030. In some embodiments, the user of the computer system can be represented in the CGR environment 720 or CGR environment 1020 as an avatar feature, similar to those disclosed herein for each of the avatars 721, 1010-1, 1010-2, 1010-3, and 1010-4. For example, in the embodiments discussed with respect to Figures 7A-7C and 8A-8C, the user of the computer system is represented as a female avatar character, as shown by preview 735. As another example, in the embodiments discussed with respect to Figures 10A to 10B and Figures 11A to 11B, the user of the computer system may be a user in the real environment, similar to any of users A to D in the real environment 1000-1 to 1000-4, and the user of the computer system can be represented in the CGR environment 1020 as an additional avatar character that is presented to and can interact with users A to D. For example, in Figure 11B, users A to C are interacting with the user of the computer system, but user D is not, and therefore avatars 1010-1, 1010-2, and 1010-3 are presented as having an appearance based on a character template, while avatar 1010-4 is presented as having an appearance based on an abstract template.
[0234] In some embodiments, each of users A to D participates in the CGR environment using avatars 1010-1 to 1010-4, and each user displays other avatars using a computer system similar to the computer system described herein. Thus, the avatars in the CGR environment 1020 can have different appearances for each user. For example, in Figure 11A, user D may interact with user A but not with user B, and therefore, in the CGR environment viewed by user A, avatar 1010-4 may have an appearance based on a character template, whereas in the CGR environment viewed by user B, avatar 1010-4 may have an appearance based on an abstract template.
[0235] The above is written with reference to specific embodiments for illustrative purposes. However, the above exemplary discussion is not intended to be exhaustive or to limit the invention to any specific form disclosed. For example, in some embodiments, the appearance of the avatar discussed with respect to Figures 10B and 11B may have varying levels of fidelity, as discussed with respect to Figures 7C and 8C. For example, if the user associated with the avatar is in the same location as the user of the computer system, the avatar may be presented with higher fidelity than an avatar associated with a user located in a different (remote) location than the user of the computer system.
[0236] Further explanations regarding Figures 10A-10B and 11A-11B are provided below with reference to Method 1200 described with respect to Figure 12.
[0237] Figure 12 is a flowchart of exemplary method 1200 for presenting avatar characters having appearances based on different appearance templates, according to several embodiments. In some embodiments, method 1200 is performed in a computer system (e.g., computer system 101 in Figure 1) (e.g., smartphone, tablet, head-mounted display generation component) that communicates with a display generation component (e.g., display generation component 120 in Figures 1, 3, and 4) (e.g., display generation component 1030 in Figures 10B and 11B) (e.g., visual output devices, 3D displays, transparent displays, projectors, head-up displays, display controllers, touchscreens, etc.). In some embodiments, method 1200 is controlled by instructions stored in a non-temporary computer-readable storage medium and executed by one or more processors of the computer system, such as one or more processors 202 of computer system 101 (e.g., control unit 110 in Figure 1). Some operations of method 1200 are optionally combined, and / or the order of some operations is optionally changed.
[0238] In method 1200, a computer system (e.g., 101) receives first data indicating that the current activity of one or more users (e.g., users A, B, C, and D in Figure 10A) (e.g., participants or objects represented in a computer-generated reality environment) is of a first type of activity (e.g., interaction (e.g., talking, making gestures, moving); physically being present in a reality environment; performing a non-manual activity (e.g., talking without moving)) (1202).
[0239] In response to receiving first data indicating that the current activity is of a first type, the computer system (e.g., 101) updates, via a display generation component (e.g., 1030), a representation of a first user having a first appearance (e.g., an animated character (e.g., a human; a cartoon character; a template of anthropomorphic constructs of non-human characters such as dogs, robots, etc.) (e.g., a first virtual representation) (e.g., a virtual avatar representing one of one or more users; a portion of an avatar representing one of one or more users) (e.g., within a computer-generated reality environment (e.g., 1020)) (e.g., displaying; virtually presenting; projecting; modifying) (1204). By updating a representation of a first user having a first appearance based on a first appearance template, the user is provided with feedback that the first user is interacting with them in response to receiving first data indicating that the current activity is a first type of activity. By providing improved feedback, the usability of the computer system is improved, the user system interface is made more efficient (for example, by helping the user to provide appropriate input when operating and / or interacting with the computer system and reducing user errors), and in addition, power consumption is reduced and the battery life of the computer system is improved by enabling the user to use the system more quickly and efficiently. In some embodiments, updating a representation of a first user having a first appearance includes presenting a representation of the first user performing a first type of activity while having an appearance based on a first appearance template. For example, if the first type of activity is an interactive activity such as talking, and the first appearance template is a dog character, the representation of the first user is presented as a talking interactive dog character.
[0240] While a representation of a first user having a first appearance is being presented (for example, avatar 1010-1 is presented in Figure 10B), the computer system (e.g., 101) receives second data (1206) indicating the current activities (e.g., activities being actively performed) of one or more users (e.g., users A, B, C, and D in Figure 11A). In some embodiments, the second data indicates the current activities of the first user. In some embodiments, the second data indicates the current activities of one or more users other than the first user. In some embodiments, the second data indicates the current activities of one or more users, including the first user. In some embodiments, the one or more users whose current activities are indicated by the second data are the same as the one or more users whose current activities are indicated by the first data.
[0241] In response to receiving second data indicating the current activities of one or more users, the computer system (e.g., 101) updates the appearance of the representation of the first user based on the current activities of one or more users (1208) (e.g., in Figure 11B, avatar 1010-1 is updated), and, according to the determination that the current activity is a first type of activity (e.g., the current activity is an interactive activity), includes causing the display generation component (e.g., 1030) to present the representation of the first user having a second appearance based on a first appearance template (1210) (e.g., in Figure 11B, avatar 1010-1 has a second pose while still being displayed using a character template) (e.g., a representation of the first user having a different appearance from the first appearance is presented (e.g., a first user performing a different activity from the first appearance is presented), but the representation of the first user is still based on the first appearance template when the current activity is a first type of activity). By presenting a representation of the first user having a second appearance based on a first appearance template, the user is provided with feedback that the first user is still interacting with the user, even if the first user's appearance has changed. By providing improved feedback, the usability of the computer system is improved, the user system interface is made more efficient (for example, by helping the user to provide appropriate input when operating and / or interacting with the computer system and reducing user errors), and in addition, power consumption is reduced and the battery life of the computer system is improved by enabling the user to use the system more quickly and efficiently. In some embodiments, presenting a representation of the first user having a second appearance based on a first appearance template includes presenting a representation of the first user performing a first type of current activity while having an appearance based on the first appearance template.For example, if the current activity is an interactive activity such as hand gestures, and the first appearance template is a dog character, the first user's representation will be presented as an interactive dog character making hand gestures.
[0242] In response to receiving second data indicating the current activity of one or more users, the computer system (e.g., 101) updates the appearance of the representation of the first user (e.g., avatars 1010-2, 1010-3, 1010-4) based on the current activity of one or more users, and, according to the determination that the current activity is a second type of activity different from the first type (e.g., the second type of activity is not an interactive activity; the second type of activity is performed while one or more users (e.g., optionally, including the first user) are not physically present in the real environment), generates a second appearance template different from the first via the display generation component (e.g., 1030). This includes presenting a representation of the first user having a third appearance based on an appearance template (1212) (for example, user B's avatar 1010-2 transitions from using the abstract template in Figure 10B to using the character template in Figure 11B) (for example, user C's avatar 1010-3 transitions from using the abstract template in Figure 10B to using the character template in Figure 11B) (for example, user D's avatar 1010-4 transitions from using the character template in Figure 10B to using the abstract template in Figure 11B) (for example, the second appearance template is a template for an inanimate character (e.g., a plant)) (e.g., a second visual representation). When the current activity is a second type of activity, presenting a representation of the first user having a third appearance based on a second appearance template provides the user with feedback that the first user has transitioned to performing a different activity than before, such as not interacting with the user. By providing improved feedback, the usability of the computer system is enhanced, the user system interface is made more efficient (for example, by assisting the user in providing appropriate input when operating and / or interacting with the computer system, thereby reducing user errors), and power consumption is reduced and the battery life of the computer system is improved by enabling the user to use the system more quickly and efficiently.In some embodiments, presenting a representation of a first user having a third appearance includes presenting a representation of a first user performing a second type of current activity while having an appearance based on a second appearance template. For example, if the second type of activity is a non-interactive activity such as sitting without moving, and the second appearance template is an amorphous mass, the representation of the first user is presented, for example, as an amorphous mass that does not move or interact with the environment. In some embodiments, presenting a representation of a first user having a third appearance includes presenting a representation of a first user transitioning from a first appearance to a third appearance. For example, the representation of the first user transitions from an interactive dog character to a non-interactive amorphous mass.
[0243] In some embodiments, a first type of activity is an activity performed by a first user (e.g., user A) at a first location (e.g., the first user is in the same conference room as the second user), and a second type of activity is an activity performed by the first user at a second location different from the first location (e.g., the first user is in a different location from the second user). In some embodiments, the appearance template used for the first user is determined based on whether the first user is located in the same physical environment as the user associated with the display generation component (e.g., the second user). For example, if the first user and the second user are in the same physical location (e.g., the same room), the first user is presented with a more realistic appearance template using the second user's display generation component. Conversely, if the first user and the second user are in different physical locations, the first user is presented with a less realistic appearance template. This is done, for example, to communicate to the second user that the first user is physically present with the second user. Furthermore, this distinguishes between users who are not physically present with the user (e.g., users who exist virtually). This frees up computing resources and streamlines the display environment by eliminating the need to display indicators that identify the location of different users or tags used to mark a particular user as present or remote.
[0244] In some embodiments, the first type of activity includes interaction between a first user (e.g., user A) and one or more users (e.g., user B, user C, user D) (e.g., the first user engaging with other users through utterances, gestures, movement, etc.). In some embodiments, the second type of activity does not include interaction between the first user and one or more users (e.g., the first user is not engaging with other users; e.g., the first user's focus is shifted to something other than other users). In some embodiments, when the first user is engaging with other users through utterances, gestures, movement, etc., the first user's representation has an appearance based on a first appearance template. In some embodiments, when the first user is not engaging with other users, or when the first user's focus is on something other than other users, the first user's representation has an appearance based on a second appearance template.
[0245] In some embodiments, the first appearance template corresponds to a first user appearance (e.g., a character template, such as the one represented by avatar 1010-1 in Figure 10B) that is more realistic than the second appearance template (e.g., an abstract template, such as the one represented by avatar 1010-2 in Figure 10B). By presenting a representation of the first user having the first appearance, which is a first user appearance that is more realistic than the second appearance template, the user is provided with feedback that the first user is interacting with them. By providing improved feedback, the usability of the computer system is improved, the user system interface is made more efficient (e.g., by helping the user to provide appropriate input when operating and / or interacting with the computer system, thereby reducing user errors), and in addition, power consumption is reduced and the battery life of the computer system is improved by enabling the user to use the system more quickly and efficiently.
[0246] In some embodiments, when a user is interacting with another user, for example by performing interactive activities such as conversation, gestures, or movement, the first user is represented with a highly realistic appearance (e.g., a high-resolution or high-fidelity appearance; e.g., an anthropomorphic shape such as an interactive virtual avatar character) (e.g., user A is represented by avatar 1010-1 using a character template). In some embodiments, when a user is not interacting with another user, the user is represented with a less realistic appearance (e.g., a low-resolution or low-fidelity appearance; e.g., an abstract shape (e.g., human-specific features removed, edges of representation softened)). For example, the first user's focus shifts from other users to other content or objects, or the first user is no longer performing interactive activities (e.g., the user is still sitting, not talking, not gesture, not moving, etc.).
[0247] In some embodiments, the first type of activity is a non-manual activity (e.g., an activity that involves little or no user hand movement (e.g., conversation)) (e.g., User A, User B, User D in Figure 10A, and User A, User D in Figure 11A). In some embodiments, the second type of activity is a manual activity (e.g., an activity that typically involves user hand movement (e.g., stacking toy blocks)) (e.g., User C in Figure 10A, and User B in Figure 11A). In some embodiments, the representation of the first user includes each presented hand portion corresponding to the first user's hand and having variable display characteristics (e.g., avatar 1010-2 in Figures 10B and 11B) (e.g., a set of one or more variable visual parameters for the rendering of each hand portion (e.g., blur degree, opacity, color, attenuation / density, resolution, etc.)). In some embodiments, the variable display characteristics indicate the estimated / predicted visual fidelity of the pose of each hand portion with respect to the pose of the first user's hand.In some embodiments, presenting a representation of a first user having a third appearance based on a second appearance template involves presenting a representation of the first user using each hand portion having a first value of variable display characteristics, according to the determination that 1) the second type of activity involves the first user's hands interacting with an object (e.g., in a physical or virtual environment) to perform a manual activity (e.g., the first user is holding a toy block in their hand) (e.g., user B is holding item 1006 in Figure 11A, avatar 1010-2 is presented in high fidelity (e.g., using a character template), and item 11B 1) Holding 1006' (for example, the variable display property has a value that indicates high visual fidelity of the posture of each hand part with respect to the posture of the first user's hand), 2) Presenting a representation of the first user using each hand part having a second value of the variable display property different from the first value of the variable display property, according to the determination that the first user's hand does not interact with an object (for example, in a physical or virtual environment) to perform a manual activity (for example, when user B does not hold item 1006, Figure 10B presents avatar 1010-2 using an abstract template) (for example, the variable display property has a value that indicates low visual fidelity of the posture of each hand part with respect to the posture of the first user's hand). The second type of activity provides the user with feedback that the first user has transitioned to performing an activity involving the use of their hands, by presenting a representation of the first user having a hand portion with a first or second value of the variable display characteristic, depending on whether the user's hands interact with an object to perform a manual activity.By providing improved feedback, the usability of the computer system is enhanced, the user system interface is made more efficient (for example, by assisting the user in providing appropriate input when operating and / or interacting with the computer system, thereby reducing user errors), and power consumption is reduced and the battery life of the computer system is improved by enabling the user to use the system more quickly and efficiently. In some embodiments, the representation of the first user's hand is rendered such that variable display characteristics show higher visual fidelity when the user's hand is interacting with an object (e.g., a physical or virtual object) and lower visual fidelity when the user's hand is not interacting with an object.
[0248] In some embodiments, the first type of activity is an activity associated with interaction with a first participant (for example, in Figure 10A, user A and user D are communicating) (for example, generally an activity associated with a user who is an active participant (e.g., in a conversation) interacting with others (e.g., the participant is speaking, making gestures, moving, focusing on other users, etc.)). In some embodiments, the second type of activity is an activity not associated with interaction with a first participant (for example, in Figure 10A, user B and user C are interacting) (for example, generally an activity associated with a user who is a bystander and not interacting with others (e.g., not in a conversation) (e.g., the bystander is not speaking, not making gestures, not moving, not focusing on other users, etc.)).
[0249] In some embodiments, presenting a representation of a first user having a second appearance based on a first appearance template includes presenting a representation of a first user having a first value of a variable display characteristic indicating a first visual fidelity of the representation of the first user (for example, in Figure 10B, user A is presented as avatar 1010-1 using a character template; in Figure 10B, user D is presented as avatar 1010-4 using a character template) (for example, a set of one or more visual parameters for rendering the representation of the first user that can be varied (e.g., blur degree, opacity, color, attenuation / density, resolution, etc.)) (for example, visual fidelity of the representation of the first user with respect to the pose of the corresponding part of the first user). When the representation of the first user has a second appearance based on a first appearance template, the user is provided with feedback that the first user's relevance to the current activity is higher because the first user is participating in the activity, by presenting the representation of the first user having a first value of a variable display characteristic indicating a first visual fidelity of the representation of the first user. By providing improved feedback, the usability of the computer system is improved, the user system interface is made more efficient (for example, by helping the user to provide appropriate input when operating and / or interacting with the computer system and reducing user errors), and in addition, power consumption is reduced and the battery life of the computer system is improved by enabling the user to use the system more quickly and efficiently. In some embodiments, the representation of the first user includes a first represented portion corresponding to a first part of the first user and is presented having a variable display characteristic indicating an estimated visual fidelity of the first represented portion with respect to the posture of the first part of the first user.
[0250] In some embodiments, presenting a representation of a first user having a third appearance based on a second appearance template includes presenting a representation of a first user having a second value of a variable display characteristic that indicates a second visual fidelity of the representation of the first user that is lower than a first visual fidelity of the representation of the first user (for example, in Figure 10B, user B is represented as avatar 1010-2 using an abstract template; in Figure 10B, user C is presented as avatar 1010-3 using an abstract template) (for example, when the user's activity is a bystander activity, a representation of the first user having a variable display characteristic that indicates lower visual fidelity is presented). When the representation of the first user has a third appearance based on a second appearance template, the user is provided with feedback that the first user's relevance to the current activity is lower because the first user is not participating in the activity, by presenting a representation of the first user having a second value of a variable display characteristic that indicates a second visual fidelity of the first user's representation that is lower than the first visual fidelity of the first user's representation. By providing improved feedback, the usability of the computer system is improved, the user system interface is made more efficient (for example, by helping the user to provide appropriate input when operating and / or interacting with the computer system and reducing user errors), and in addition, power consumption is reduced and the battery life of the computer system is improved by enabling the user to use the system more quickly and efficiently.
[0251] In some embodiments, users are considered active participants or bystanders based on their activity, and users considered active participants are represented with higher visual fidelity than users considered bystanders. For example, a user who is actively communicating, making gestures, moving, or focusing on other participants and / or users is considered a participant (or active participant) (based on this activity), and a representation of that user with variable display characteristic values indicating high visual fidelity is presented. Conversely, a user who is not actively communicating, not making gestures, not moving, or not focusing on other participants and / or users is considered a bystander (based on this activity), and a representation of that user with variable display characteristic values indicating low visual fidelity is presented.
[0252] In some embodiments, a display generation component (e.g., 1030) is associated with a second user (e.g., the second user is a viewer and / or user of the display generation component). In some embodiments, while presenting a representation of a third user having a first value of variable display properties (e.g., an avatar 1010-2 representing user B) (e.g., a representation of the first user) (e.g., alternatively, a representation of an object) via the display generation component (e.g., a set of one or more visual parameters for rendering the variable representation of the third user (e.g., blur degree, opacity, color, attenuation / density, resolution, etc.)), a computer system (e.g., 101) receives third data indicating the second user's focal position (e.g., data indicating the location or region in the physical or computer-generated environment that the second user's eyes are focusing on). In some embodiments, the third data is generated using a sensor (e.g., a camera sensor) that tracks the position of the second user's eyes and a processor that determines the focus of the user's eyes. In some embodiments, the focal point includes the focal point of the user's eye calculated in three dimensions (for example, the focal point may include a depth component).
[0253] In some embodiments, in response to receiving third data indicating the second user's focus position, the computer system (e.g., 101) updates the representation of the third user (e.g., alternatively, updates the representation of an object) based on the second user's focus position via a display generation component, and determines that 1) the second user's focus position corresponds to the position of the third user's representation (e.g., they are located together) (e.g., in Figure 11A, the computer system's user is looking at user B, and therefore avatar 1010-2 is displayed using a character template) (e.g., the second user is looking at the representation of the third user) (e.g., alternatively, the second user is looking at the representation of an object). This includes, 1) increasing the value of the variable display property of the third user's representation (for example, alternatively increasing the value of the variable display property of the object's representation), and 2) decreasing the value of the variable display property of the third user's representation according to the determination that the second user's focus position does not correspond to the position of the third user's representation (for example, they are not located together) (for example, the second user is not looking at the third user's representation) (for example, alternatively the second user is not looking at the object's representation) (for example, in Figure 11A, the computer system user is looking far away from user D, and therefore avatar 1010-4 is displayed using an abstract template) (for example, alternatively decreasing the value of the variable display property of the object's representation). Depending on whether the second user's focus position corresponds to the position of the third user's representation, the value of the variable display characteristics of the third user's representation is increased or decreased, thereby providing the second user with improved feedback that the computer system is detecting the focus position, and reducing computational resources by eliminating or reducing the amount of data that needs to be processed to render areas that the user is not focusing on.By providing improved feedback and reducing computational workload, the usability of the computer system is enhanced; the user system interface is made more efficient (for example, by assisting the user in providing appropriate input when operating and / or interacting with the computer system, thereby reducing user errors); and power consumption is reduced and the battery life of the computer system is improved by enabling the user to use the system more quickly and efficiently.
[0254] In some embodiments, variable display characteristics are adjusted for various items or users presented using the display generation component, depending on whether the user is focusing on each of those items or users. For example, as the user shifts their focus away from the representation of a user or object, the visual fidelity of the representation of the user or object (e.g., resolution, sharpness, opacity, and / or density) decreases (e.g., based on the amount the user's focus is away from the representation of the user). Conversely, as the user focuses on the representation of a user or object, the visual fidelity of the representation of the user or object (e.g., resolution, sharpness, opacity, and / or density) increases. This visual effect helps conserve computational resources by reducing the amount of data that needs to be rendered to the user using the display generation component. Specifically, objects and / or users presented outside the user's focus can be rendered with lower fidelity, thus consuming fewer computational resources. This can be done without sacrificing visual integrity, as the adjustments to visual fidelity mimic the natural behavior of the human eye. For example, when the human eye shifts its focus, the appearance of the object the eye is focusing on becomes sharper (e.g., higher fidelity), while objects outside the eye's focus become blurred (e.g., lower fidelity). In other words, objects rendered within the user's surrounding view are intentionally rendered at lower fidelity to conserve computational resources.
[0255] In some embodiments, the first appearance template corresponds to an abstract shape (e.g., 1010-2 in Figure 10B, 1010-3 in Figure 10B, 1010-4 in Figure 11B) (e.g., a representation of a shape that has no anthropomorphic features, or a representation of a shape that has fewer anthropomorphic features than the anthropomorphic shape used to represent the user when the user is engaged in a first type of activity). In some embodiments, the second appearance template corresponds to an anthropomorphic shape (e.g., 1010-1 in Figure 10B, 1010-4 in Figure 10B, 1010-1 in Figure 11B, 1010-2 in Figure 11B, 1010-3 in Figure 11B) (e.g., a human, an animal, or a representation of other features having facial features such as a face, eyes, mouth, limbs, etc., which enable the anthropomorphic shape to communicate the user's movements). By presenting a representation of the first user having a first appearance template corresponding to an abstract shape or a second appearance template corresponding to an anthropomorphic shape, the user is provided with feedback indicating that when presented as an anthropomorphic shape, the first user's relevance to the current activity is higher, and when presented as an abstract shape, the relevance to the current activity is lower. By providing improved feedback, the usability of the computer system is enhanced, the user system interface is made more efficient (for example, by assisting the user in providing appropriate input when operating and / or interacting with the computer system, thereby reducing user errors), and power consumption is reduced and the battery life of the computer system is improved by enabling the user to use the system more quickly and efficiently.
[0256] In some embodiments, aspects and / or operations of methods 900 and 1200 may be interchangeable, substituted, and / or added between these methods. For brevity, their details will not be repeated here.
[0257] The above is written with reference to specific embodiments for illustrative purposes. However, the above exemplary discussion is not intended to be exhaustive or to limit the invention to the exact form disclosed. Many modifications and variations are possible in light of the above teachings. These embodiments have been selected and described in order to best illustrate the principles of the invention and its practical applications, thereby enabling other persons skilled in the art to best use the invention and the various described embodiments with various modifications suitable for specific applications that may be conceived.
Claims
1. It is a method, In a computer system that communicates with a display generation component, Receiving posture data representing the posture of at least a first part of the user, The presenting of the avatar includes, via the display generation component, presenting an avatar that includes each presented avatar feature having variable display characteristics that correspond to the first part of the user and indicate the certainty of the posture of the first part of the user, and presenting the avatar is, In accordance with the determination that the posture of the first part of the user is associated with a first certainty value, the avatar is presented using the respective avatar features having the first value of the variable display characteristics. In accordance with the determination that the posture of the first part of the user is associated with a second certainty value different from the first certainty value, the avatar is presented using each of the avatar features having a second value of the variable display characteristic different from the first value of the variable display characteristic. Methods that include...
2. Receiving second posture data representing the posture of at least a second part of the user, The presenting of the avatar via the display generation component includes presenting a second avatar feature that corresponds to the second part of the user and has a second variable display characteristic that indicates the certainty of the posture of the second part of the user, and presenting the avatar is In accordance with the determination that the posture of the second part of the user is associated with a third certainty value, the avatar is presented using the second avatar feature having a first value of the second variable display characteristic, The process includes presenting the avatar using a second avatar feature having a second value of the second variable display feature that is different from the first value of the second variable display feature, in accordance with the determination that the posture of the second part of the user is associated with a fourth certainty value different from the third certainty value. The method according to claim 1.
3. Presenting the aforementioned avatar means In accordance with the determination that the third certainty value corresponds to the first certainty value, the first value of the second variable display characteristic corresponds to the first value of the variable display characteristic. In accordance with the determination that the fourth certainty value corresponds to the second certainty value, the second value of the second variable display characteristic corresponds to the second value of the variable display characteristic. The method according to claim 2, which includes the following:
4. In accordance with the determination that the third certainty value corresponds to the first certainty value, the first value of the second variable display characteristic does not correspond to the first value of the variable display characteristic, In accordance with the determination that the fourth certainty value corresponds to the second certainty value, the second value of the second variable display characteristic does not correspond to the second value of the variable display characteristic, The method according to claim 2, further comprising:
5. Receiving updated posture data representing the change in posture of the first part of the user, The method further includes updating the representation of the avatar in response to receiving the updated posture data, This includes updating the posture of each avatar feature based on the posture change of the first part of the user, The method according to any one of claims 1 to 4.
6. Updating the aforementioned representation of the avatar means In response to a change in the certainty of the posture of the first part of the user during the posture change of the first part of the user, in addition to changing the position of at least a portion of the avatar based on the posture change of the first part of the user, the variable display characteristics of each of the displayed avatar features are changed. The method according to claim 5, including the method described in claim 5.
7. The method according to any one of claims 1 to 6, wherein the variable display characteristics indicate the estimated visual fidelity of each avatar feature with respect to the posture of the first part of the user.
8. Presenting the aforementioned avatar means The avatar is presented using the respective avatar features having a third value of the variable display characteristics, in accordance with the determination that the posture data satisfies a first set of criteria that are satisfied when the first part of the user is detected by the first sensor. In accordance with the determination that the posture data does not satisfy the first set of criteria, the avatar is presented using each of the avatar features having a fourth value of the variable display characteristic that shows a lower certainty value than the third value of the variable display characteristic. The method according to any one of claims 1 to 7, including the method described in any one of claims 1 to 7.
9. While each of the avatar features having the current value of the variable display characteristics is presented, updated posture data representing the change in the posture of the first part of the user is received. The method further includes updating the representation of the avatar in response to receiving the updated posture data, In accordance with the determination that the updated posture data represents a change in the posture of the first part of the user from a first position within the sensor's field of view to a second position outside the sensor's field of view, the current value of the variable display characteristic is reduced. The current value of the variable display characteristic is increased in accordance with the determination that the updated posture data represents a change in the posture of the first part of the user from the second position to the first position. The method according to any one of claims 1 to 8, including
10. The current value of the variable display characteristic decreases at a first speed. The current value of the variable display characteristic increases at a second speed higher than the first speed. The method according to claim 9.
11. The first value of the variable display characteristic is higher than the second value of the variable display characteristic, representing the visual fidelity of each avatar feature with respect to the posture of the first part of the user. Presenting the aforementioned avatar means In accordance with the determination that the first part of the user corresponds to a subset of physical characteristics, the posture of the first part of the user is associated with the second certainty value, In accordance with the determination that the first part of the user does not correspond to the subset of physical characteristics, the posture of the first part of the user is associated with the first certainty value, The method according to any one of claims 1 to 10, including the method described in any one of claims 1 to 10.
12. The further includes updating the representation of the avatar while the avatar is being presented using each of the avatar features having the first value of the variable display characteristic, In accordance with the determination that the movement speed of the first part of the user is the first movement speed of the first part of the user, the avatar is presented using the respective avatar features having the first modified value of the variable display characteristics, The avatar is presented using the respective avatar features having a second modified value of the variable display characteristic, based on the determination that the movement speed of the first part of the user is different from the first movement speed, which is a second movement speed of the first part of the user. The method according to any one of claims 1 to 11.
13. Changing the value of the variable display characteristic, which includes changing one or more visual parameters of each of the aforementioned avatar features, The method according to any one of claims 1 to 12, further comprising:
14. The method according to claim 13, wherein the one or more visual parameters include a degree of blur.
15. The method according to claim 13, wherein the one or more visual parameters include opacity.
16. The method according to claim 13, wherein the one or more visual parameters include color.
17. The method according to claim 13, wherein the one or more visual parameters include the density of particles containing each of the avatar features.
18. The method according to claim 17, wherein the density of the particles containing each of the aforementioned avatar features includes the spacing between the particles containing each of the aforementioned avatar features.
19. The method according to claim 17, wherein the density of the particles containing each of the aforementioned avatar features includes the size of the particles containing each of the aforementioned avatar features.
20. This includes changing the values of the variable display characteristics, including presenting the visual effects associated with each of the aforementioned avatar features. The method according to any one of claims 1 to 19, further comprising:
21. The first part of the user includes a first physical characteristic and a second physical characteristic, The first value of the variable display characteristic is higher than the second value of the variable display characteristic, representing the visual fidelity of each avatar feature with respect to the posture of the user's first physical characteristic. Presenting the avatar using each of the avatar features having the second value of the variable display characteristics is, The process includes presenting the avatar using rendering of the first physical characteristics based on the user's corresponding physical characteristics, and rendering of the second physical characteristics based on corresponding physical characteristics that do not belong to the user. The method according to any one of claims 1 to 20.
22. The method according to any one of claims 1 to 21, wherein the attitude data is generated from a plurality of sensors.
23. The method according to claim 22, wherein the plurality of sensors include one or more camera sensors associated with the computer system.
24. The method according to any one of claims 22 to 23, wherein the plurality of sensors include one or more camera sensors separate from the computer system.
25. The method according to any one of claims 22 to 24, wherein the plurality of sensors include one or more non-visual sensors.
26. The method according to any one of claims 1 to 25, wherein the orientation of at least the first portion of the user is determined using an interpolation function.
27. The method according to any one of claims 1 to 26, wherein the posture data includes data generated from prior scan data capturing information about the appearance of the user of the computer system.
28. The method according to any one of claims 1 to 27, wherein the posture data includes data generated from prior media data.
29. The posture data includes video data that includes at least the first portion of the user, Presenting the avatar includes presenting a modeled avatar that includes each of the avatar features rendered using the video data which includes the first portion of the user. The method according to any one of claims 1 to 28.
30. Presenting the aforementioned avatar means In accordance with the determination that an input indicating a first rendering value of the avatar has been received, the avatar is presented using each of the avatar features and a first amount of avatar features other than each of the avatar features. In accordance with the determination that an input indicating a second rendering value of the avatar that is different from the first rendering value has been received, the avatar is presented using the respective avatar features and a second quantity of avatar features other than the respective avatar features, which is different from the first quantity. The method according to any one of claims 1 to 29, including the method described in any one of claims 1 to 29.
31. The display generation component is used to display a representation of the user associated with the display generation component, wherein the representation of the user associated with the display generation component corresponds to the appearance of the user associated with the display generation component, which is displayed to one or more users other than the user associated with the display generation component. The method according to any one of claims 1 to 30, further comprising:
32. Presenting the avatar includes presenting the avatar using each of the avatar features having a first appearance based on the first appearance of the first part of the user, and the method is Receiving data indicating the updated appearance of the first part of the user, The avatar is presented via the display generation component, using the respective avatar features having an updated appearance based on the updated appearance of the first part of the user. The method according to any one of claims 1 to 31, further comprising:
33. The posture data further represents objects associated with the first part of the user, Presenting the aforementioned avatar means This includes presenting the avatar together with representations of objects adjacent to each of the aforementioned avatar features. The method according to any one of claims 1 to 32.
34. The posture of the first part of the user is associated with a fifth certainty value. Presenting the aforementioned avatar means The method includes presenting the avatar using the respective avatar features that have a variable display characteristic value lower than the fifth certainty value, in accordance with the determination that the first part of the user is of the first feature type. The method according to any one of claims 1 to 33.
35. The posture of the first part of the user is associated with a sixth certainty value. Presenting the aforementioned avatar means The further includes presenting the avatar using the respective avatar features having a variable display characteristic value that shows a higher certainty value than the sixth certainty value, in accordance with the determination that the first part of the user is of the second feature type. The method according to any one of claims 1 to 34.
36. A non-temporary computer-readable storage medium for storing one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component, wherein the one or more programs include instructions for performing the method according to any one of claims 1 to 35.
37. A computer system, One or more processors, A memory for storing one or more programs configured to be executed by one or more processors, wherein the one or more programs include instructions for performing the method according to any one of claims 1 to 35, Computer system.
38. A computer system, Means for carrying out the method described in any one of claims 1 to 35 A computer system equipped with the following features.
39. A non-temporary computer-readable storage medium that stores one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component, wherein the one or more programs are The system receives posture data representing the posture of at least one part of the user, The display generation component causes an avatar to be displayed, which includes each presented avatar feature having variable display characteristics that correspond to the first part of the user and indicate the certainty of the posture of the first part of the user. Including commands, presenting the avatar is, In accordance with the determination that the posture of the first part of the user is associated with a first certainty value, the avatar is presented using the respective avatar features having the first value of the variable display characteristics. In accordance with the determination that the posture of the first part of the user is associated with a second certainty value different from the first certainty value, the avatar is presented using each of the avatar features having a second value of the variable display characteristic different from the first value of the variable display characteristic. Non-temporary computer-readable storage media, including [specific type of storage medium].
40. A computer system, One or more processors, A memory that stores one or more programs configured to be executed by one or more processors, An electronic device comprising, the one or more programs, The system receives posture data representing the posture of at least one part of the user, The system causes an avatar to be presented via a display generation component, which includes each presented avatar feature having variable display characteristics that correspond to the first part of the user and indicate the certainty of the pose of the first part of the user. Including commands, presenting the avatar is, In accordance with the determination that the posture of the first part of the user is associated with a first certainty value, the avatar is presented using the respective avatar features having the first value of the variable display characteristics. In accordance with the determination that the posture of the first part of the user is associated with a second certainty value different from the first certainty value, the avatar is presented using each of the avatar features having a second value of the variable display characteristic different from the first value of the variable display characteristic. A computer system, including
41. A computer system, Means for receiving posture data representing the posture of at least a first part of the user, The system includes means for presenting an avatar via a display generation component, which includes each presented avatar feature having variable display characteristics that correspond to the first part of the user and indicate the certainty of the pose of the first part of the user, and presenting the avatar is In accordance with the determination that the posture of the first part of the user is associated with a first certainty value, the avatar is presented using the respective avatar features having the first value of the variable display characteristics. In accordance with the determination that the posture of the first part of the user is associated with a second certainty value different from the first certainty value, the avatar is presented using each of the avatar features having a second value of the variable display characteristic different from the first value of the variable display characteristic. A computer system, including a computer system.
42. It is a method, In a computer system that communicates with a display generation component, Receiving first data indicating that the current activity of one or more users is of a first type, In response to receiving the first data indicating that the current activity is of the first type, the display generation component updates the representation of the first user having a first appearance based on the first appearance template, While the representation of the first user having the first appearance is being displayed, second data indicating the current activities of one or more users is received, The process includes updating the appearance of the first user's representation based on the current activities of the one or more users, in response to receiving the second data indicating the current activities of the one or more users, In accordance with the determination that the current activity is of the first type, the display generation component is used to present the representation of the first user having a second appearance based on the first appearance template, The process includes, in accordance with the determination that the current activity is of a second type of activity different from the first type, causing the display generation component to present the representation of the first user having a third appearance based on a second appearance template different from the first appearance template, method.
43. The first type of activity is an activity of the first user performed at a first location, The second type of activity is an activity of the first user performed at a second location different from the first location. The method according to claim 42.
44. The first type of activity includes an interaction between the first user and the one or more users, The second type of activity does not include interaction between the first user and the one or more users. The method according to any one of claims 42 to 43.
45. The method according to claim 44, wherein the first appearance template corresponds to the appearance of the first user which is more realistic than the second appearance template.
46. The first type of activity described above is a non-manual activity, The second type of activity described above is a manual activity, The representation of the first user includes each presented hand portion that corresponds to the hand of the first user and has variable display characteristics, To present the representation of the first user having the third appearance based on the second appearance template, The second type of activity involves presenting the first user's representation using the respective hand portion having a first value of the variable display characteristic, in accordance with the determination that the first user's hands interact with an object in order to perform the manual activity. The second type of activity includes presenting the first user's representation using each hand portion having a second value of the variable display characteristic different from a first value of the variable display characteristic, in accordance with the determination that the first user's hands do not interact with an object in order to perform the manual activity. The method according to any one of claims 42 to 45.
47. The first type of activity is an activity associated with interaction with the first participant, The second type of activity is an activity that is not associated with interaction with the first participant, Presenting the representation of the first user having the second appearance based on the first appearance template includes presenting the representation of the first user having a first value of a variable display characteristic indicating a first visual fidelity of the representation of the first user, Presenting the representation of the first user having the third appearance based on the second appearance template includes presenting the representation of the first user having a second value of the variable display characteristic that indicates a second visual fidelity of the representation of the first user that is lower than the first visual fidelity of the representation of the first user. The method according to any one of claims 42 to 46.
48. The display generation component is associated with a second user, and the method is While presenting a representation of a third user having a first value of variable display characteristics via the display generation component, the system receives third data indicating the focal position of the second user, The present invention further includes, in response to receiving the third data indicating the second user's focal position, updating the representation of the third user based on the second user's focal position via the display generation component, In accordance with the determination that the focal position of the second user corresponds to the position of the representation of the third user, the value of the variable display characteristic of the representation of the third user is increased, The process includes, in accordance with the determination that the focal position of the second user does not correspond to the position of the representation of the third user, reducing the value of the variable display characteristic of the representation of the third user, The method according to any one of claims 42 to 47.
49. The first appearance template corresponds to an abstract shape, The second appearance template mentioned above corresponds to the anthropomorphic shape, The method according to any one of claims 42 to 48.
50. A non-temporary computer-readable storage medium for storing one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component, wherein the one or more programs include instructions for performing the method according to any one of claims 42 to 49.
51. A computer system, One or more processors, A computer system comprising: a memory for storing one or more programs configured to be executed by one or more processors, wherein the one or more programs include instructions for performing the method according to any one of claims 42 to 49.
52. A computer system, Means for carrying out the method described in any one of claims 42 to 49 A computer system equipped with the following features.
53. A non-temporary computer-readable storage medium that stores one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component, wherein the one or more programs are Receive first data indicating that the current activity of one or more users is of a first type, In response to receiving the first data indicating that the current activity is of the first type, the display generation component updates the representation of the first user having a first appearance based on the first appearance template. While the representation of the first user having the first appearance is being displayed, second data indicating the current activities of one or more users is received. The command includes updating the appearance of the first user's representation based on the current activities of the one or more users, in response to receiving the second data indicating the current activities of the one or more users, In accordance with the determination that the current activity is of the first type, the display generation component is used to present the representation of the first user having a second appearance based on the first appearance template, The process includes, in accordance with the determination that the current activity is of a second type of activity different from the first type, causing the display generation component to present the representation of the first user having a third appearance based on a second appearance template different from the first appearance template, Non-temporary computer-readable storage medium.
54. A computer system, One or more processors, A memory that stores one or more programs configured to be executed by one or more processors, A computer system comprising, wherein one or more programs Receive first data indicating that the current activity of one or more users is of a first type, In response to receiving the first data indicating that the current activity is of the first type, the display generation component updates the representation of the first user having a first appearance based on the first appearance template. While the representation of the first user having the first appearance is being displayed, second data indicating the current activities of one or more users is received. In response to receiving the second data indicating the current activities of the one or more users, the appearance of the representation of the first user is updated based on the current activities of the one or more users. Including commands, In accordance with the determination that the current activity is of the first type, the display generation component is used to present the representation of the first user having a second appearance based on the first appearance template, The process includes, in accordance with the determination that the current activity is of a second type of activity different from the first type, causing the display generation component to present the representation of the first user having a third appearance based on a second appearance template different from the first appearance template, Computer system.
55. A computer system, Means for receiving first data indicating that the current activity of one or more users is of a first type, In response to receiving the first data indicating that the current activity is of the first type, means for updating a representation of a first user having a first appearance based on a first appearance template via the display generation component, Means for receiving second data indicating the current activities of one or more users while the representation of the first user having the first appearance is being displayed, The system includes means for updating the appearance of the representation of the first user based on the current activities of the one or more users, in response to receiving the second data indicating the current activities of the one or more users, In accordance with the determination that the current activity is of the first type, the display generation component is used to present the representation of the first user having a second appearance based on the first appearance template, The process includes, in accordance with the determination that the current activity is of a second type of activity different from the first type, causing the display generation component to present the representation of the first user having a third appearance based on a second appearance template different from the first appearance template, Computer system.