Presenting an avatar in a three-dimensional environment
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-20
- Publication Date
- 2026-08-11
AI Technical Summary
[0008]需要具有改进的方法和界面的电子设备来与三维环境进行交互。此类方法和界面可以补充或替换用于与三维环境进行交互的常规方法。此类方法和界面减少了来自用户的输入的数量、程度和/或性质,并且产生更高效的人机界面。对于电池驱动的计算设备,此类方法和界面节省功率,并且增大电池充电之间的时间间隔。
Smart Images

Figure CN115917474B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims priority to U.S. Provisional Application No. 63 / 036,411, filed June 8, 2020, entitled “PRESENTING AVATARS IN THREE-DIMENSIONAL ENVIRONMENTS,” the contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure relates in its entirety to computer systems that communicate with display generating components and optionally with one or more input devices that provide computer-generated experiences, including but not limited to electronic devices that provide virtual reality and mixed reality experiences via a display. Background Technology
[0004] In recent years, the development of computer systems for augmented reality has increased significantly. Exemplary augmented reality environments include at least some virtual elements that replace or enhance the physical world. Input devices for computer systems and other electronic computing devices (such as cameras, controllers, joysticks, touch-sensitive surfaces, and touchscreen displays) are used to interact with the virtual / augmented reality environment. Exemplary virtual elements include virtual objects such as digital images, videos, text, icons, and control elements (such as buttons and other graphics). Summary of the Invention
[0005] Some methods and interfaces for interacting with environments that include at least some virtual elements (e.g., applications, augmented reality environments, mixed reality environments, and virtual reality environments) are cumbersome, inefficient, and limited. For example, systems that provide insufficient feedback for performing actions associated with virtual objects, systems that require a series of inputs to achieve a desired result in an augmented reality environment, and systems where the manipulation of virtual objects is complex, tedious, and error-prone, impose a significant cognitive burden on users and detract from the immersive experience of virtual / augmented reality environments. Furthermore, these methods take longer than necessary, thus wasting the energy of the computer system. This latter consideration is particularly important in battery-powered devices.
[0006] Therefore, computer systems with improved methods and interfaces are needed to provide users with computer-generated experiences, making user interaction with the computer system more efficient and intuitive. Such methods and interfaces optionally complement or replace conventional methods for providing users with computer-generated realistic experiences. By helping users understand the relationship between the inputs provided and the device's responses to those inputs, such methods and interfaces reduce the quantity, extent, and / or nature of user input, thus creating a more effective human-computer interface.
[0007] The disclosed system reduces or eliminates the aforementioned defects and other problems associated with the user interface of a computer system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a laptop, tablet, or handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device, such as a watch or head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also referred to as a "touchscreen" or "touchscreen display"). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, in addition to display generation components, the computer system also has one or more output devices, including one or more haptic output generators and / or one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, memory, and one or more modules, and a program or instruction set stored in memory for performing multiple functions. In some implementations, the user interacts with the GUI through touch and gestures of a stylus and / or fingers on a touch-sensitive surface, movement of the user's eyes and hands in space relative to the GUI (and / or computer system) or the user's body (as captured by a camera and other motion sensors), and / or voice input (as captured by one or more audio input devices). In some implementations, functions performed through interaction optionally include image editing, drawing, presentations, word processing, spreadsheet creation, playing games, making and receiving phone calls, video conferencing, sending and receiving emails, instant messaging, test support, digital photography, digital video recording, web browsing, digital music playback, note-taking, and / or digital video playback. Executable instructions for performing these functions are optionally included in a transient and / or non-transitory computer-readable storage medium or other computer program product configured for execution by one or more processors.
[0008] Electronic devices with improved methods and interfaces are needed to interact with 3D environments. Such methods and interfaces can complement or replace conventional methods for interacting with 3D environments. They reduce the amount, extent, and / or nature of user input, resulting in more efficient human-computer interfaces. For battery-powered computing devices, such methods and interfaces save power and increase the time interval between battery charging sessions.
[0009] There is a need for electronic devices with improved methods and interfaces for interacting with other users in a 3D environment using avatars. Such methods and interfaces can complement or replace conventional methods for interacting with other users in a 3D environment using avatars. These methods and interfaces reduce the amount, degree, and / or nature of input from users and result in more efficient human-computer interfaces.
[0010] It should be noted that the various embodiments described above can be combined with any other embodiments described herein. The features and advantages described in this specification are not exhaustive; in particular, many additional features and advantages will be apparent to those skilled in the art from the accompanying drawings, description, and claims. Furthermore, it should be pointed out that the language used in this specification has been chosen in principle for readability and instruction purposes, and such choice may not be necessary to depict or define the subject matter of the invention. Attached Figure Description
[0011] To better understand the various embodiments described, reference should be made to the following detailed description in conjunction with the accompanying drawings, wherein similar reference numerals indicate corresponding parts in all the drawings.
[0012] Figure 1 The operating environment of a computer system for providing a CGR experience is shown according to some implementation schemes.
[0013] Figure 2 This is a block diagram illustrating a controller of a computer system configured to manage and coordinate the user's CGR experience according to some implementation schemes.
[0014] Figure 3 This is a block diagram illustrating a display generation component of a computer system configured to provide a CGR experience to a user, according to some embodiments.
[0015] Figure 4 A hand tracking unit of a computer system configured to capture user gesture input is shown according to some embodiments.
[0016] Figure 5 An eye-tracking unit of a computing system configured to capture a user's gaze input is shown according to some embodiments.
[0017] Figure 6 This is a flowchart illustrating a flare-assisted gaze tracking pipeline according to some implementation schemes.
[0018] Figures 7A to 7C The illustration shows a virtual avatar with display characteristics that change in appearance based on the determinism of the user's pose, according to some implementation schemes.
[0019] Figures 8A to 8CThe illustration shows a virtual avatar with display characteristics that vary in appearance based on the determinism of the user's pose, according to some implementation schemes.
[0020] Figure 9 This is a flowchart illustrating an exemplary method for presenting a virtual avatar character with display characteristics that change in appearance based on the determinism of the user's pose, according to some embodiments.
[0021] Figure 10A and Figure 10B The illustrations show virtual avatars with appearances based on different appearance templates according to some implementation schemes.
[0022] Figure 11A and Figure 11B The illustrations show virtual avatars with appearances based on different appearance templates according to some implementation schemes.
[0023] Figure 12 This is a flowchart illustrating an exemplary method for presenting an avatar with an appearance based on different appearance templates, according to some implementation schemes. Detailed Implementation
[0024] According to some implementations, this disclosure relates to a user interface for providing a computer-generated reality (CGR) experience to a user.
[0025] The systems, methods, and GUIs described in this paper improve user interface interactions with virtual / augmented reality environments in a variety of ways.
[0026] In some implementations, the computer system presents to a user an avatar with display characteristics that vary in appearance based on the deterministic pose of a portion of the user (e.g., in a CGR environment). The computer system receives pose data representing the pose of a portion of the user (e.g., from a sensor) and (e.g., via a display generation component) causes the presentation of the avatar, including avatar features corresponding to that portion of the user and having variable display characteristics indicating the deterministic pose of that portion of the user. The computer system assigns different values to the variable display characteristics of the avatar features depending on the deterministic pose of that portion of the user, which provides the user of the computer system with an indication of the estimated visual fidelity of the avatar's pose relative to the pose of that portion of the user.
[0027] In some implementations, the computer system presents an avatar to a user with appearances based on different appearance templates, which in some implementations change based on the activities performed by the user. The computer system (e.g., from sensors) receives data indicating that the current activity of one or more users is a first type of activity (e.g., an interactive activity). In response, the computer system (e.g., via a display generation component) updates the representation (e.g., the avatar) of the first user with the first appearance based on the first appearance template (e.g., a character template). Simultaneously with the presentation of the representation of the first user with the first appearance, the computer system (e.g., from sensors) receives second data indicating the current activity of one or more users. In response, the computer system (e.g., via a display generation component) causes the presentation of the representation of the first user with either a second appearance based on the first appearance template or a third appearance based on a second appearance template (e.g., an abstract template), depending on whether the current activity is a first type of activity or a second type of activity, thus providing the computer system's user with an indication of whether the user is performing a first type of activity or a second type of activity.
[0028] Figures 1 to 6 A description of an exemplary computer system for providing a CGR experience to a user is provided. Figures 7A to 7C as well as Figures 8A to 8C The illustration shows a virtual avatar character with display characteristics that vary in appearance based on the user's pose, according to some implementation schemes. Figure 9 This is a flowchart illustrating an exemplary method for presenting a virtual avatar character with display characteristics that change in appearance based on the determinism of the user's pose, according to various embodiments. Figures 7A to 7C as well as Figures 8A to 8C For showing Figure 9 The process in. Figures 10A to 10B as well as Figures 11A to 11B The illustrations show virtual avatars with appearances based on different appearance templates according to some implementation schemes. Figure 12 This is a flowchart illustrating an exemplary method for presenting an avatar with an appearance based on different appearance templates, according to some implementation schemes. Figures 10A to 10B as well as Figures 11A to 11B For showing Figure 12 The process in.
[0029] Furthermore, in methods described herein where one or more steps depend on the satisfaction of one or more conditions, it should be understood that the method may be repeated in multiple repetitions such that, during the repetitions, all conditions determining the steps in the method are satisfied in different repetitions of the method. For example, if a method requires performing a first step (if a condition is satisfied) and a second step (if a condition is not satisfied), those skilled in the art will know that the stated steps are repeated until both conditions are satisfied and not satisfied (in no particular order). Thus, a method described as having one or more steps depending on the satisfaction of one or more conditions can be rewritten as a method that repeats until each condition described in the method is satisfied. However, this does not require the system or computer-readable medium to declare that the system or computer-readable medium contains instructions for performing discretionary operations based on the satisfaction of the corresponding one or more conditions, and thus to determine whether possible conditions have been satisfied without explicitly repeating the steps of the method until all conditions determining the steps in the method are satisfied. Those skilled in the art will also understand that, similar to methods having discretionary steps, a system or computer-readable storage medium may repeat the steps of the method multiple times as needed to ensure that all discretionary steps have been performed.
[0030] In some implementation schemes, such as Figure 1 As shown, a CGR experience is provided to a user via an operating environment 100 including a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted display (HMD), a monitor, a projector, a touchscreen, etc.), one or more input devices 125 (e.g., an eye-tracking device 130, a hand-tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a haptic output generator 170, and other output devices 180), one or more sensors 190 (e.g., image sensors, light sensors, depth sensors, haptic sensors, orientation sensors, proximity sensors, temperature sensors, position sensors, motion sensors, speed sensors, etc.), and optionally one or more peripheral devices 195 (e.g., home appliances, wearable devices, etc.). In some embodiments, one or more of the input devices 125, output devices 155, sensors 190, and peripheral devices 195 are integrated with the display generation component 120 (e.g., in a head-mounted or handheld device).
[0031] In describing the CGR experience, various terms are used to distinguish several related but different environments that a user can sense and / or interact with (e.g., interacting with inputs detected by the computer system 101 that generates the CGR experience, causing the computer system to generate audio, visual, and / or haptic feedback corresponding to various inputs provided to the computer system 101). The following is a subset of these terms:
[0032] Physical environment: The physical environment refers to the physical world that people can sense and / or interact with without the aid of electronic systems. Physical environments, such as physical parks, include physical objects such as physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment through senses such as sight, touch, hearing, taste, and smell.
[0033] Computer-Generated Reality: Conversely, a computer-generated reality (CGR) environment refers to a fully or partially simulated environment that people sense and / or interact with via electronic systems. In CGR, a subset of a person's physical motion, or a representation thereof, is tracked, and in response, one or more characteristics of one or more virtual objects simulated in the CGR environment are modulated in a manner consistent with at least one physical law. For example, a CGR system can detect a person's head rotation and, in response, modulate the graphical content and sound field presented to the person in a manner similar to how such views and sounds change in a physical environment. In some cases (e.g., for accessibility reasons), the modulation of characteristics of virtual objects in the CGR environment can be done in response to a representation of physical motion (e.g., a voice command). A person can use any of their senses to sense and / or interact with CGR objects, including vision, hearing, touch, taste, and smell. For example, a person can sense and / or interact with audio objects that create a 3D or spatial audio environment that provides the perception of a point audio source in 3D space. For example, audio objects can enable audio transparency, which selectively introduces ambient sounds from the physical environment, with or without computer-generated audio. In some CGR environments, people can sense and / or interact only with audio objects.
[0034] Examples of CGR include virtual reality and mixed reality.
[0035] Virtual Reality: A virtual reality (VR) environment is a simulated environment designed to be entirely based on computer-generated sensory input for one or more senses. A VR environment includes multiple virtual objects that a person can sense and / or interact with. For example, trees, buildings, and computer-generated images representing human avatars are examples of virtual objects. A person can sense and / or interact with virtual objects in a VR environment through the simulation of a person's presence within the computer-generated environment and / or through the simulation of a subgroup of physical movements of a person within the computer-generated environment.
[0036] Mixed Reality: Compared to VR environments, which are designed to be entirely based on computer-generated sensory input, mixed reality (MR) environments refer to simulated environments designed to incorporate sensory input from the physical environment, or its representations, in addition to computer-generated sensory input (e.g., virtual objects). On the virtual continuum, a mixed reality environment is any state between, but not limited to, a purely physical environment as one end and a virtual reality environment as the other. In some MR environments, computer-generated sensory input can respond to changes in sensory input from the physical environment. Additionally, some electronic systems used to present an MR environment can track position and / or orientation relative to the physical environment to enable virtual objects to interact with real objects (i.e., physical objects or their representations from the physical environment). For example, a system can cause motion so that virtual trees appear stationary relative to the physical ground.
[0037] Examples of mixed reality include augmented reality and augmented virtual reality.
[0038] Augmented Reality (AR): An augmented reality (AR) environment is a simulated environment in which one or more virtual objects are overlaid on a physical environment or a representation of the physical environment. For example, an electronic system for presenting an AR environment may have a transparent or semi-transparent display through which a person can directly view the physical environment. The system can be configured to present virtual objects on the transparent or semi-transparent display, allowing a person to perceive the virtual objects overlaid on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture images or videos of the physical environment, which are representations of the physical environment. The system combines the images or videos with virtual objects and presents the combination on the opaque display. A person uses the system to indirectly view the physical environment via the images or videos of the physical environment and perceives the virtual objects overlaid on the physical environment. As used herein, video of the physical environment displayed on an opaque display is referred to as “pass-through video,” meaning that the system uses one or more image sensors to capture images of the physical environment and uses those images when presenting the AR environment on the opaque display. Alternatively, the system may have a projection system that projects virtual objects onto the physical environment, such as as holograms or on physical surfaces, allowing a person to perceive the virtual objects superimposed on the physical environment. Augmented reality environments also refer to simulated environments in which the representation of the physical environment is transformed by computer-generated sensory information. For example, in providing pass-through video, the system can transform images from one or more sensors to apply a selected viewpoint (e.g., viewpoint) different from the viewpoint captured by the imaging sensor. Furthermore, the representation of the physical environment can be transformed by graphically modifying (e.g., magnifying) portions of it, such that the modified portions are representative but not realistic versions of the original captured image. Still further, the representation of the physical environment can be transformed by graphically removing or blurring portions of it.
[0039] Augmented Virtual: An augmented virtual (AV) environment is a simulated environment in which a virtual or computer-generated environment incorporates one or more sensory inputs from a physical environment. Sensory inputs can be representations of one or more characteristics of the physical environment. For example, an AV park could have virtual trees and virtual buildings, but a person's face could be realistically reproduced from an image taken of a physical person. Similarly, virtual objects could adopt the shape or color of a physical object imaged by one or more imaging sensors. Furthermore, virtual objects could use shadows that correspond to the sun's position within the physical environment.
[0040] Hardware: Many different types of electronic systems enable people to sense and / or interact with various CGR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays shaped as lenses designed to be placed on a person's eyes (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. Head-mounted systems may have one or more speakers and an integrated opaque display. Alternatively, head-mounted systems may be configured to receive an external opaque display (e.g., a smartphone). Head-mounted systems may incorporate one or more imaging sensors for capturing images or video of the physical environment and / or one or more microphones for capturing audio of the physical environment. Head-mounted systems may have transparent or semi-transparent displays instead of opaque displays. Transparent or semi-transparent displays may have a medium through which light representing the image is directed to the person's eyes. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium can be an optical waveguide, holographic medium, optical combiner, optical reflector, or any combination thereof. In one embodiment, a transparent or translucent display can be configured to selectively become opaque. Projection-based systems can employ retinal projection technology, which projects graphic images onto a person's retina. Projection systems can also be configured to project virtual objects into a physical environment, such as as holograms, or onto a physical surface. In some embodiments, controller 110 is configured to manage and coordinate the user's CGR experience. In some embodiments, controller 110 includes a suitable combination of software, firmware, and / or hardware. References below... Figure 2The controller 110 is described in more detail. In some embodiments, the controller 110 is a computing device located locally or remotely relative to scene 105 (e.g., the physical environment). For example, the controller 110 is a local server located within scene 105. Alternatively, the controller 110 is a remote server (e.g., a cloud server, a central server, etc.) located outside scene 105. In some embodiments, the controller 110 is communicatively coupled to a display generation component 120 (e.g., an HMD, a monitor, a projector, a touchscreen, etc.) via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). In another example, controller 110 is included within the housing (e.g., physical enclosure) of display generation component 120 (e.g., HMD or portable electronic device including display and one or more processors), one or more input devices in input device 125, one or more output devices in output device 155, one or more sensors in sensor 190, and / or one or more peripheral devices in peripheral device 195, or shares the same physical housing or support structure with one or more of the aforementioned devices.
[0041] In some embodiments, the display generation component 120 is configured to provide a CGR experience to a user (e.g., at least the visual components of the CGR experience). In some embodiments, the display generation component 120 includes a suitable combination of software, firmware, and / or hardware. The following is relative to... Figure 3 The display generation component 120 is described in more detail. In some embodiments, the functionality of the controller 110 is provided by and / or combined with the display generation component 120.
[0042] According to some implementation schemes, when a user is virtually and / or physically present within scene 105, the display generation component 120 provides the user with a CGR experience.
[0043] In some embodiments, the display generating component is worn on a part of the user's body (e.g., on his / her head, his / her hand, etc.). Thus, the display generating component 120 includes one or more CGR displays provided for displaying CGR content. For example, in various embodiments, the display generating component 120 surrounds the user's field of view. In some embodiments, the display generating component 120 is a handheld device (such as a smartphone or tablet) configured to present CGR content, and the user holds the device having a display facing the user's field of view and a camera facing scene 105. In some embodiments, the handheld device is optionally placed within a housing worn on the user's head. In some embodiments, the handheld device is optionally placed on a support (e.g., a tripod) in front of the user. In some embodiments, the display generating component 120 is a CGR chamber, housing, or room configured to present CGR content, wherein the user does not wear or hold the display generating component 120. Many user interfaces described with reference to one type of hardware (e.g., a handheld device or a tripod-mounted device) for displaying CGR content can be implemented on another type of hardware (e.g., an HMD or other wearable computing device) for displaying CGR content. For example, a user interface illustrating interaction with CGR content triggered by an interaction occurring in the space in front of a handheld device or a tripod-mounted device can be similarly implemented using an HMD, where the interaction occurs in the space in front of the HMD and the response to the CGR content is displayed via the HMD. Similarly, a user interface illustrating interaction with CGR content triggered by movement of a handheld device or a tripod-mounted device relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eyes, head, or hand)) can be similarly implemented using an HMD, where the movement is caused by movement of the HMD relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eyes, head, or hand)).
[0044] Despite Figure 1 The relevant features of the operating environment 100 are shown in this disclosure, but those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure further relevant aspects of the exemplary embodiments disclosed herein.
[0045] Figure 2This is a block diagram of an example controller 110 according to some implementation schemes. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and to avoid obscuring further relevant aspects of the implementation schemes disclosed herein. Therefore, as a non-limiting example, in some embodiments, controller 110 includes one or more processing units 202 (e.g., microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), graphics processing units (GPUs), central processing units (CPUs), processing kernels, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), Bluetooth, ZigBee, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these components and various other components.
[0046] In some embodiments, one or more communication buses 204 include circuitry for interconnecting and communicating between system components. In some embodiments, one or more I / O devices 206 include at least one of a keyboard, mouse, touchpad, joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.
[0047] Memory 220 includes high-speed random access memory, such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate random access memory (DDR RAM), or other random access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. Memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. Memory 220 includes a non-transitory computer-readable storage medium. In some embodiments, memory 220 or the non-transitory computer-readable storage medium of memory 220 stores programs, modules, and data structures or subsets thereof, including optional operating system 230 and CGR experience module 240.
[0048] Operating system 230 includes instructions for handling various basic system services and for performing hardware-related tasks. In some embodiments, CGR experience module 240 is configured to manage and coordinate single or multiple CGR experiences for one or more users (e.g., single CGR experiences for one or more users, or multiple CGR experiences for corresponding groups of one or more users). To this end, in various embodiments, CGR experience module 240 includes a data acquisition unit 241, a tracking unit 242, a coordination unit 246, and a data transmission unit 248.
[0049] In some implementations, the data acquisition unit 241 is configured to acquire data from... Figure 1 The data acquisition unit 241 includes at least a display generation component 120 and optionally acquires data (e.g., presentation data, interactive data, sensor data, location data, etc.) from one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the data acquisition unit 241 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0050] In some implementations, the tracking unit 242 is configured to map scene 105, and the tracking at least shows the generated component 120 relative to... Figure 1 The tracking unit 242 tracks the location / position of scenario 105, and optionally the location / position relative to one or more of the tracking input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the tracking unit 242 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics. In some embodiments, the tracking unit 242 includes a hand tracking unit 244 and / or an eye tracking unit 243. In some embodiments, the hand tracking unit 244 is configured to track the location / position of one or more portions of the user's hand, and / or the location / position of one or more portions of the user's hand relative to... Figure 1 The movement of scene 105 relative to the display generating component 120 and / or relative to a coordinate system (defined relative to the user's hand). The following refers to the movement relative to... Figure 4 The hand tracking unit 244 is described in more detail. In some embodiments, the eye tracking unit 243 is configured to track the user's gaze (or more broadly, the user's eyes, face, or head) relative to scene 105 (e.g., relative to the physical environment and / or relative to the user (e.g., the user's hand)) or relative to CGR content displayed via display generation component 120. The following description is relative to... Figure 5 The eye-tracking unit 243 is described in more detail.
[0051] In some implementations, coordination unit 246 is configured to manage and coordinate the CGR experience presented to the user by display generation component 120, and optionally by one or more of output device 155 and / or peripheral device 195. For this purpose, in various implementations, coordination unit 246 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.
[0052] In some embodiments, the data transmission unit 248 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the display generation component 120, and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. For this purpose, in various embodiments, the data transmission unit 248 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.
[0053] Although the data acquisition unit 241, tracking unit 242 (e.g., including eye tracking unit 243 and hand tracking unit 244), coordination unit 246, and data transmission unit 248 are shown residing on a single device (e.g., controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 241, tracking unit 242 (e.g., including eye tracking unit 243 and hand tracking unit 244), coordination unit 246, and data transmission unit 248 may reside in a separate computing device.
[0054] also, Figure 2 This is used more as a functional description of various features that can exist in a specific implementation, and differs from the structural diagrams of the implementations described herein. As those skilled in the art will recognize, items shown individually can be combined, and some items can be separated. For example, Figure 2 Some functional modules shown individually may be implemented in a single module, and the various functions of a single functional block may be implemented in various implementations through one or more functional blocks. The actual number of modules and the division of specific functions and how features are allocated therein will vary depending on the specific implementation, and in some implementations, it depends in part on the specific combination of hardware, software and / or firmware chosen for that particular implementation.
[0055] Figure 3This is a block diagram illustrating an example of generating component 120 according to some embodiments. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and to avoid obscuring further relevant aspects of the embodiments disclosed herein. For this purpose, as a non-limiting example, in some embodiments, the display generating component 120 (e.g., HMD) includes one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, Firewire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, BlueTooth, ZigBee, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more CGR displays 312, one or more optional internal and / or external image sensors 314, memory 320, and one or more communication buses 304 for interconnecting these components and various other components.
[0056] In some embodiments, one or more communication buses 304 include circuitry for interconnecting and controlling communication between system components. In some embodiments, one or more I / O devices and sensors 306 include inertial measurement units (IMUs), accelerometers, gyroscopes, thermometers, one or more physiological sensors (e.g., blood pressure monitors, heart rate monitors, blood oxygen sensors, blood glucose sensors, etc.), one or more microphones, one or more speakers, haptic engines, and / or one or more depth sensors (e.g., structured light, time-of-flight, etc.).
[0057] In some embodiments, one or more CGR displays 312 are configured to provide a CGR experience to a user. In some embodiments, the one or more CGR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conducting electron emission display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical system (MEMS), and / or similar display types. In some embodiments, the one or more CGR displays 312 correspond to waveguide displays such as diffraction, reflection, polarization, and holography. For example, display generation component 120 (e.g., HMD) includes a single CGR display. In another example, display generation component 120 includes CGR displays for each of the user's eyes. In some embodiments, the one or more CGR displays 312 are capable of presenting MR and VR content.
[0058] In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face, including the user's eyes (and may be referred to as an eye-tracking camera). In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hand and optionally the user's arm (and may be referred to as a hand-tracking camera). In some embodiments, one or more image sensors 314 are configured to face forward in order to acquire image data corresponding to the scene that the user would see in the absence of a display generation component 120 (e.g., an HMD) (and may be referred to as a scene camera). One or more optional image sensors 314 may include one or more RGB cameras (e.g., having a complementary metal-oxide-semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), one or more infrared (IR) cameras, and / or one or more event-based cameras, etc.
[0059] Memory 320 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. Memory 320 includes a non-transitory computer-readable storage medium. In some embodiments, memory 320 or the non-transitory computer-readable storage medium of memory 320 stores programs, modules, and data structures, or subsets thereof, including optional operating system 330 and CGR rendering module 340.
[0060] Operating system 330 includes processes for handling various basic system services and for performing hardware-related tasks. In some embodiments, CGR presentation module 340 is configured to present CGR content to a user via one or more CGR displays 312. For this purpose, in various embodiments, CGR presentation module 340 includes a data acquisition unit 342, a CGR presentation unit 344, a CGR mapping generation unit 346, and a data transmission unit 348.
[0061] In some implementations, the data acquisition unit 342 is configured to acquire data from at least... Figure 1 The controller 110 acquires data (e.g., presentation data, interaction data, sensor data, location data, etc.). For the purposes described, in various embodiments, the data acquisition unit 342 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0062] In some embodiments, the CGR rendering unit 344 is configured to render CGR content via one or more CGR displays 312. For the aforementioned purposes, in various embodiments, the CGR rendering unit 344 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0063] In some implementations, the CGR mapping generation unit 346 is configured to generate a CGR mapping based on media content data (e.g., a 3D mapping of a mixed reality scene or a physical environment in which computer-generated objects may be placed to generate a computer-generated reality mapping). For the aforementioned purposes, in various implementations, the CGR mapping generation unit 346 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0064] In some embodiments, the data transmission unit 348 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the controller 110, and optionally one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. For the stated purpose, in various embodiments, the data transmission unit 348 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0065] Although the data acquisition unit 342, CGR rendering unit 344, CGR mapping generation unit 346, and data transmission unit 348 are shown residing in a single device (e.g., Figure 1 The data acquisition unit 342, the CGR rendering unit 344, the CGR mapping generation unit 346, and the data transmission unit 348 are located on the display generation component 120, but it should be understood that in other embodiments, any combination of the data acquisition unit 342, the CGR rendering unit 344, the CGR mapping generation unit 346, and the data transmission unit 348 may be located in a separate computing device.
[0066] also, Figure 3 This is used more as a functional description of various features that may exist in a particular implementation, and differs from the structural diagrams of the implementations described herein. As those skilled in the art will recognize, items shown individually can be combined, and some items can be separate. For example, Figure 3 Some functional modules shown individually may be implemented in a single module, and the various functions of a single functional block may be implemented in various implementations through one or more functional blocks. The actual number of modules and the division of specific functions and how features are allocated therein will vary depending on the specific implementation, and in some implementations, it depends in part on the specific combination of hardware, software and / or firmware chosen for that particular implementation.
[0067] Figure 4 This is a schematic illustration of an exemplary embodiment of the hand tracking device 140. In some embodiments, the hand tracking device 140 ( Figure 1 Controlled by hand tracking unit 244 Figure 2 To track the location / position of one or more parts of a user's hand, and / or the location of one or more parts of the user's hand relative to... Figure 1The movement of the hand tracking device 140 is relative to a portion of the user's surrounding physical environment, relative to the display generation component 120, or relative to a portion of the user (e.g., the user's face, eyes, or head), and / or relative to a coordinate system (defined relative to the user's hand). In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).
[0068] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras, etc.) that captures at least three-dimensional scene information including the human user's hand 406. The image sensor 404 captures images of the hand at sufficient resolution to distinguish the fingers and their corresponding positions. The image sensor 404 typically captures images of other parts of the user's body, or possibly all parts of the body, and may have scaling capabilities or be a dedicated sensor with increased magnification to capture images of the hand at the desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with other image sensors to capture the physical environment of scene 105, or serves as the image sensor for capturing the physical environment of scene 105. In some embodiments, the image sensor is positioned relative to the user or the user's environment in a way that uses the field of view of the image sensor 404 or a portion thereof to define an interaction space in which hand movements captured by the image sensor are considered input to the controller 110.
[0069] In some implementations, image sensor 404 outputs a sequence of frames containing 3D mapping data (and, in addition, possibly color image data) to controller 110, which extracts high-level information from the mapping data. This high-level information is typically provided via an application programming interface (API) to an application running on the controller, which in turn drives display generation component 120. For example, a user can interact with software running on controller 110 by moving his hand 406 and changing his hand pose.
[0070] In some embodiments, image sensor 404 projects a speckle pattern onto a scene including hand 406 and captures an image of the projected pattern. In some embodiments, controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) via triangulation based on the lateral offset of the specks in the pattern. This approach is advantageous because it does not require the user to hold or wear any kind of beacon, sensor, or other marker. This method gives the depth coordinates of points in the scene relative to a predetermined reference plane at a specific distance from image sensor 404. In this disclosure, it is assumed that image sensor 404 defines an orthogonal set of x-axis, y-axis, and z-axis such that the depth coordinates of points in the scene correspond to the z-component measured by the image sensor. Alternatively, image sensor 404 (e.g., a hand-tracking device) may use other 3D mapping methods, such as stereo imaging or time-of-flight measurement, based on a single or multiple cameras or other types of sensors.
[0071] In some implementations, hand tracking device 140 captures and processes time-series depth maps containing the user's hand as the user moves his hand (e.g., the entire hand or one or more fingers). Software running on a processor in image sensor 404 and / or controller 110 processes the 3D mapping data to extract image block descriptors of the hand from these depth maps. The software may match these descriptors with image block descriptors stored in database 408 based on a previous learning process to estimate the pose of the hand in each frame. The pose typically includes the 3D position of the user's hand joints and fingertips.
[0072] The software can also analyze the trajectories of the hand and / or fingers across multiple frames in a sequence to identify gestures. The pose estimation function described herein can be alternated with motion tracking, such that patch-based pose estimation is performed only once every two (or more) frames, while tracking is used to find pose changes occurring in the remaining frames. Pose, motion, and gesture information is provided to an application running on controller 110 via the aforementioned API. This application can, for example, move and modify the image presented on display generation unit 120 in response to the pose and / or gesture information, or perform other functions.
[0073] In some embodiments, the software may be downloaded to controller 110 electronically, for example, via a network, or alternatively, may be provided on a tangible, non-transitory medium such as an optical, magnetic, or electronic memory medium. In some embodiments, database 408 is also stored in memory associated with controller 110. Alternatively or additionally, some or all of the described functions of the computer may be implemented in dedicated hardware, such as custom or semi-custom integrated circuits or programmable digital signal processors (DSPs). Although in Figure 4The controller 110 is shown, but for example, as a separate unit from the image sensor 404, some or all of the controller's processing functions may be performed by a suitable microprocessor and software, or by dedicated circuitry within the housing of the image sensor 404 (e.g., a hand-tracking device), or by other devices associated with the image sensor 404. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with the display generation component 120 (e.g., in a television receiver, handheld device, or head-mounted device) or with any other suitable computerized device (such as a game console or media player). The sensing function of the image sensor 404 may also be integrated into a computer or other computerized device controlled by the sensor output.
[0074] Figure 4 It also includes a schematic diagram of a depth map 410 captured by image sensor 404 according to some embodiments. As described above, the depth map comprises a matrix of pixels with corresponding depth values. Pixel 412 corresponding to hand 406 has been segmented from the background and wrist in this map. The brightness of each pixel within the depth map 410 is inversely proportional to its depth value (i.e., the measured z-distance from image sensor 404), where gray shadows become darker as depth increases. Controller 110 processes these depth values to identify and segment components of the image that have human hand features (i.e., a group of adjacent pixels). These features may include, for example, overall size, shape, and frame-to-frame motion from the depth map sequence.
[0075] Figure 4 The diagram also schematically illustrates the hand skeleton 414 ultimately extracted by the controller 110 from the depth map 410 of the hand 406 according to some embodiments. Figure 4 In this configuration, the hand skeleton 414 is superimposed on the hand background 416, which has already been segmented from the original depth map. In some embodiments, key feature points of the hand, and optionally those on the wrist or arm connected to the hand (e.g., points corresponding to knuckles, fingertips, the center of the palm, the end of the hand connecting to the wrist, etc.), are identified and located on the hand skeleton 414. In some embodiments, the controller 110 uses the position and movement of these key feature points across multiple image frames to determine, according to some embodiments, the gesture performed by the hand or the current state of the hand.
[0076] Figure 5 An eye-tracking device 130 is shown. Figure 1 Exemplary implementations of ). In some implementations, the eye-tracking device 130 comprises an eye-tracking unit 243 ( Figure 2The eye-tracking device 130 is used to track the positioning and movement of a user's gaze relative to scene 105 or relative to CGR content displayed via display generation component 120. In some embodiments, the eye-tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, when the display generation component 120 is a head-mounted device (such as a headset, helmet, goggles, or glasses) or a handheld device placed in a wearable frame, the head-mounted device includes both components for generating CGR content for the user to view and components for tracking the user's gaze relative to the CGR content. In some embodiments, the eye-tracking device 130 is separate from the display generation component 120. For example, when the display generation component is a handheld device or a CGR chamber, the eye-tracking device 130 is optionally a separate device from the handheld device or CGR chamber. In some embodiments, the eye-tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, the head-mounted eye-tracking device 130 is optionally used in conjunction with a display generation component that is also head-mounted or not head-mounted. In some embodiments, the eye-tracking device 130 is not a head-mounted device and is optionally used in conjunction with head-mounted display generation components. In some embodiments, the eye-tracking device 130 is not a head-mounted device and is optionally part of non-head-mounted display generation components.
[0077] In some embodiments, the display generation unit 120 uses display mechanisms (e.g., a left near-eye display panel and a right near-eye display panel) to display frames including left and right images in front of the user's eyes, thereby providing the user with a 3D virtual view. For example, the head-mounted display generation unit may include left and right optical lenses (referred to herein as eye lenses) located between the display and the user's eyes. In some embodiments, the display generation unit may include or be coupled to one or more external cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation unit may have a transparent or semi-transparent display on which virtual objects are displayed, allowing the user to view the physical environment directly through the transparent or semi-transparent display. In some embodiments, the display generation unit projects virtual objects onto the physical environment. The virtual objects may be projected, for example, onto a physical surface or as holograms, allowing an individual to observe virtual objects superimposed on the physical environment using the system. In this case, separate display panels and image frames for the left and right eyes may not be necessary.
[0078] like Figure 5As shown, in some embodiments, eye-tracking device 130 (e.g., gaze tracking device) includes at least one eye-tracking camera (e.g., an infrared (IR) or near-infrared (NIR) camera) and an illumination source (e.g., an array or ring of IR or NIR light sources, such as LEDs) that emits light (e.g., IR or NIR light) toward the user's eye. The eye-tracking camera may be pointed at the user's eye to receive IR or NIR light reflected directly from the eye, or alternatively, it may be pointed at "hot" mirrors located between the user's eye and the display panel, which reflect the IR or NIR light from the eye back to the eye-tracking camera while allowing visible light to pass through. Eye-tracking device 130 optionally captures images of the user's eyes (e.g., as a video stream captured at 60-120 frames per second (fps), analyzes these images to generate gaze tracking information, and transmits the gaze tracking information to controller 110. In some embodiments, the user's two eyes are tracked separately using corresponding eye-tracking cameras and illumination sources. In some embodiments, only one of the user's eyes is tracked using corresponding eye-tracking cameras and illumination sources.
[0079] In some implementations, a device-specific calibration procedure is used to calibrate the eye-tracking device 130 to determine parameters for the eye-tracking device in a specific operating environment 100, such as the 3D geometry and parameters of the LEDs, camera, thermal mirror (if present), eye lenses, and display. The device-specific calibration procedure can be performed at a factory or another facility before the AR / VR equipment is delivered to the end user. The device-specific calibration procedure can be automated or manual. According to some implementations, a user-specific calibration procedure may include estimations of eye parameters for a specific user, such as pupil position, foveal position, optical axis, visual axis, interocular distance, etc. According to some implementations, once the device-specific and user-specific parameters for the eye-tracking device 130 are determined, a flash-assisted method can be used to process the images captured by the eye-tracking camera to determine the current visual axis and the user's gaze point relative to the display.
[0080] like Figure 5As shown, the eye-tracking device 130 (e.g., 130A or 130B) includes an eye lens 520 and a gaze tracking system. The gaze tracking system includes at least one eye-tracking camera 540 (e.g., an infrared (IR) or near-infrared (NIR) camera) positioned on the side of the user's face where eye tracking is performed, and an illumination source 530 (e.g., an array or ring of IR or NIR light sources, such as NIR light-emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eye 592. The eye-tracking camera 540 may be pointed toward a mirror 550 located between the user's eye 592 and a display 510 (e.g., the left or right display panel of a head-mounted display, or the display of a handheld device, projector, etc.). These mirrors reflect the IR or NIR light from the eye 592 while allowing visible light to pass through. Figure 5 (as shown in the top portion), or alternatively, it can be pointed towards the user's eye 592 to receive reflected IR or NIR light from the eye 592 (e.g., as shown in the top portion), Figure 5 (As shown in the bottom part).
[0081] In some implementations, controller 110 renders AR or VR frames 562 (e.g., left and right frames for the left and right display panels) and provides frames 562 to display 510. Controller 110 uses gaze tracking input 542 from eye-tracking camera 540 for various purposes, such as processing frame 562 for display. Controller 110 optionally estimates the user's gaze point on display 510 based on the gaze tracking input 542 obtained from eye-tracking camera 540 using a flash-assisted method or other suitable method. The gaze point estimated based on gaze tracking input 542 is optionally used to determine the direction the user is currently looking.
[0082] The following describes several possible use cases for the user's current gaze direction and is not intended to be limiting. As an example use case, controller 110 may render virtual content differently based on the determined direction of the user's gaze. For example, controller 110 may generate virtual content at a higher resolution in the concave region determined according to the user's current gaze direction than in the peripheral region. As another example, the controller may position or move virtual content in the view based at least partially on the user's current gaze direction. As yet another example, the controller may display specific virtual content in the view based at least partially on the user's current gaze direction. As another exemplary use case in an AR application, controller 110 may guide an external camera used to capture the physical environment for CGR experiences to focus in the determined direction. The external camera's autofocus mechanism may then focus on an object or surface in the environment that the user is currently looking at on display 510. As another exemplary use case, eye lens 520 may be a focusable lens, and the controller uses gaze tracking information to adjust the focus of eye lens 520 so that the virtual object the user is currently looking at has appropriate convergence / divergence to match the convergence of the user's eyes 592. The controller 110 can use gaze tracking information to guide the eye lens 520 to adjust its focus so that the nearby object that the user is looking at appears at the correct distance.
[0083] In some embodiments, the eye-tracking device is part of a head-mounted device that includes a display (e.g., display 510), two eye lenses (e.g., eye lens 520), an eye-tracking camera (e.g., eye-tracking camera 540), and a light source (e.g., light source 530 (e.g., IR or NIR LED)) mounted in a wearable housing. The light source emits light (e.g., IR or NIR light) toward the user's eyes 592. In some embodiments, the light source may be arranged in a ring or circle around each lens in the head-mounted device, such as... Figure 5 As shown in the diagram. In some embodiments, for example, eight light sources 530 (e.g., LEDs) are arranged around each lens 520. However, more or fewer light sources 530 may be used, and other arrangements and positions of the light sources 530 may be used.
[0084] In some embodiments, the display 510 emits light in the visible light range and does not emit light in the IR or NIR range, and therefore does not introduce noise into the gaze tracking system. It should be noted that the position and angle of the eye-tracking camera 540 are given by way of example and are not intended to be limiting. In some embodiments, a single eye-tracking camera 540 may be located on each side of the user's face. In some embodiments, two or more NIR cameras 540 may be used on each side of the user's face. In some embodiments, cameras 540 with a wider field of view (FOV) and cameras 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, cameras 540 operating at one wavelength (e.g., 850 nm) and cameras 540 operating at different wavelengths (e.g., 940 nm) may be used on each side of the user's face.
[0085] like Figure 5 The implementation of the gaze tracking system shown can be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide users with computer-generated reality, virtual reality, augmented reality, and / or augmented virtual experiences.
[0086] Figure 6 A strobe-assisted gaze tracking pipeline according to some embodiments is illustrated. In some embodiments, the gaze tracking pipeline uses a strobe-assisted gaze tracking system (e.g., such as...) Figure 1 and Figure 5 The eye-tracking device 130 shown is used to implement this. The flash-assisted gaze tracking system can maintain a tracking state. Initially, the tracking state is off or "no". When in tracking state, the flash-assisted gaze tracking system uses previous information from previous frames when analyzing the current frame to track the pupil outline and flash in the current frame. When not in tracking state, the flash-assisted gaze tracking system attempts to detect the pupil and flash in the current frame, and if successful, initializes the tracking state to "yes" and continues to the next frame in tracking state.
[0087] like Figure 6 As shown, the gaze-tracking camera captures left and right images of the user's left and right eyes. The captured images are then fed into a gaze-tracking pipeline for processing to begin at 610. As indicated by the arrow returning to element 600, the gaze-tracking system can continue capturing images of the user's eyes, for example, at a rate of 60 to 120 frames per second. In some embodiments, each set of captured images can be fed into the pipeline for processing. However, in some embodiments or under certain conditions, not all captured frames are processed by the pipeline.
[0088] At 610, for the currently captured image, if the tracking state is yes, the method proceeds to element 640. At 610, if the tracking state is no, the image is analyzed to detect the user's pupil and flash, as indicated at 620. At 630, if the pupil and flash are successfully detected, the method proceeds to element 640. Otherwise, the method returns to element 610 to process the next image of the user's eye.
[0089] At 640, if proceeding from element 610, the current frame is analyzed to track the pupil and flashes in part based on previous information from the previous frame. At 640, if proceeding from element 630, the tracking state is initialized based on the pupils and flashes detected in the current frame. The processing result at element 640 is checked to verify that the tracking or detection result is credible. For example, the result can be checked to determine whether a sufficient number of pupils and flashes used for gaze estimation were successfully tracked or detected in the current frame. At 650, if the result is not credible, the tracking state is set to no at element 660, and the method returns to element 610 to process the next image of the user's eye. At 650, if the result is credible, the method proceeds to element 670. At 670, the tracking state is set to YES (if not already yes), and the pupil and flash information is passed to element 680 to estimate the user's gaze point.
[0090] Figure 6 This is intended as an example of an eye-tracking technology that can be used in a particular specific implementation. As those skilled in the art will recognize, in various implementations, other existing or future-developed eye-tracking technologies may be used in place of or in combination with the flash-assisted eye-tracking technology described herein in a computer system 101 for providing a CGR experience to a user.
[0091] In this disclosure, various input methods are described in relation to interaction with a computer system. When an example is provided using one input device or method, and another example is provided using another input device or method, it should be understood that each example is compatible with and optionally utilizes the input device or method described with respect to the other example. Similarly, various output methods are described in relation to interaction with a computer system. When an example is provided using one output device or method, and another example is provided using another output device or method, it should be understood that each example is compatible with and optionally utilizes the output device or method described with respect to the other example. Similarly, various methods are described in relation to interaction with a virtual or mixed reality environment via a computer system. When an example is provided using interaction with a virtual environment, and another example is provided using a mixed reality environment, it should be understood that each example is compatible with and optionally utilizes the methods described with respect to the other example. Therefore, this disclosure discloses embodiments that are combinations of features of a plurality of examples without exhaustively listing all features of the embodiments in the description of each exemplary embodiment.
[0092] User interface and related processes
[0093] Now turn attention to implementations of user interfaces (“UIs”) and associated processes that can be implemented on a computer system (such as a portable multifunction device or a head-mounted device) that communicates with display generation components and (optionally) one or more sensors (e.g., a camera).
[0094] This disclosure relates to an exemplary process for representing a user as an avatar in a CGR environment. Figures 7A to 7C , Figures 8A to 8C and Figure 9 An example is depicted in which a user is represented in a CGR environment as a virtual avatar with one or more display characteristics that vary in appearance based on the deterministic pose of the user's physical body in the real environment. Figures 10A to 10B , Figures 11A to 11B and Figure 12 An example is depicted where a user is represented in a CGR environment as a virtual avatar character with an appearance based on different appearance templates. As described above, using a computer system (e.g., Figure 1 The computer system 101 in the document implements the process disclosed herein.
[0095] Figure 7AThe image depicts a user 701 standing in a real-world environment 700 with arms raised, and the user's left hand holding a cup 702. In some embodiments, the real-world environment 700 is a motion capture studio including cameras 705-1, 705-2, 705-3, and 705-4 for capturing data (e.g., image data and / or depth data), which can be used to determine the pose of one or more parts of the user 701. This is sometimes referred to herein as capturing the pose of a part of the user 701. The pose of the part of the user 701 is used to determine the corresponding avatar in a CGR environment (e.g., an animated film set) (see, for example, [link to relevant documentation]). Figure 7C The pose of the incarnation 721.
[0096] like Figure 7A As depicted, camera 705-1 has a field of view 707-1 located at a portion 701-1 of user 701. Portion 701-1 includes the physical features of user 701, including the user's neck, collar area, and a portion of the user's face and head, including the user's right eye, right ear, nose, and mouth, but excluding the top of the user's head, left eye, and left ear. Camera 705-1 captures the pose of the physical features of user 701's portion 701-1 because portion 701-1 is within the camera's field of view 707-1.
[0097] Camera 705-2 has a field of view 707-2 located at a portion 701-2 of user 701. The portion 701-2 includes the physical features of user 701, including the user's right hand, right wrist, and the distal portion of the user's forearm adjacent to the right wrist. Camera 705-2 captures the pose of the portion 701-2 of user 701 because the portion 701-2 is within the camera's field of view 707-2.
[0098] Camera 705-3 has a field of view 707-3 located at a portion 701-3 of user 701. The portion 701-3 includes the physical features of user 701, including the user's left and right feet, and the left and right lower leg regions. Camera 705-3 captures the pose of the portion 701-3 of user 701 because the portion 701-3 is within the camera's field of view 707-3.
[0099] Camera 705-4 has a field of view 707-4 located at a portion 701-4 of user 701. The portion 701-4 includes the physical features of user 701, including the user's left hand, left wrist, and the distal portion of the user's left forearm adjacent to the wrist. Camera 705-4 typically captures the pose of the physical features of the portion 701-4 of user 701 because the portion 701-4 is within the camera's field of view 707-4. However, as discussed in more detail below, some areas of the portion 701-4 (such as the palm of the user's left hand) are occluded by camera 705-4 because they are located behind cup 702 and are therefore considered not to be within the field of view 707-4. Therefore, the pose of these areas (e.g., the palm of the user's left hand) is not captured by camera 705-4.
[0100] Parts of user 701 that are not within the camera's field of view are considered not to be captured by the camera. For example, the top of the user's head, the user's left eye, the user's upper arm, elbow and proximal end of the user's forearm, the user's thigh and knee, and the user's torso are all outside the field of view of cameras 705-1 to 705-4, and therefore, the position or pose of these parts of user 701 is considered not to be captured by the camera.
[0101] Cameras 705-1 to 705-4 are described as non-limiting examples of devices for capturing the pose of a portion of user 701, i.e., devices for capturing data that can be used to determine the pose of a portion of user 701. Therefore, in addition to or in place of any of cameras 705-1 to 705-4, other sensors and / or devices may be used to capture the pose of a portion of the user. For example, such sensors may include proximity sensors, accelerometers, GPS sensors, positioning sensors, depth sensors, thermal sensors, image sensors, other types of sensors, or any combination thereof. In some embodiments, these various sensors may be separate components, such as wearable positioning sensors placed at different locations on user 701. In some embodiments, the various sensors may be integrated into one or more devices associated with user 701, such as the user's smartphone, user's tablet, user's computer, motion capture suit worn by user 701, headset worn by user 701 (e.g., HMD), smartwatch (e.g., Figure 8AThe various sensors may be integrated into one or more different devices associated with other users (e.g., users other than user 701), such as another user's smartphone, another user's tablet, another user's computer, a headset worn by another user, another device worn by a different user or otherwise associated with a different user, or any combination thereof. Cameras 705-1 to 705-4 are also included. Figure 7A The camera is shown as a standalone device. However, one or more of the cameras may be integrated with other components, such as any of the sensors and devices described above. For example, one of the cameras may be a camera integrated with a headset device of a second user present in the real environment 700. In some embodiments, data may be provided from facial scans (e.g., using a depth sensor), media items (such as pictures and videos of user 701), or other relevant sources. For example, depth data associated with the user's face may be collected when the user uses facial recognition to unlock a personal communication device (e.g., a smartphone, smartwatch, or HMD). In some embodiments, the device that generates the representation of that part of the user is separate from the personal communication device, and data from the facial scan is provided (e.g., securely and privately, with one or more options for the user to decide whether to share data between devices) to the device that generates the representation of that part of the user for constructing the representation of that part of the user. In some implementations, the device generating the representation of the user's part is the same as the personal communication device, and data from a facial scan is provided to the device generating the representation of the user's part for constructing the representation (e.g., a facial scan is used to unlock an HMD, which also generates a representation of the user's part). This data can be used, for example, to enhance the understanding of the pose of a part of the user 701, or to improve the visual fidelity of the representation of a part of the user 701 that was not detected using sensors (e.g., cameras 705-1 to 705-4) (discussed below). In some implementations, this data is used to detect changes in a part of the user that are visible when the device is unlocked (e.g., a new hairstyle, a new pair of glasses, etc.), so that the representation of the user's part can be updated based on changes in the user's appearance.
[0102] In the embodiments described herein, the sensors and devices discussed above for capturing data that can be used to determine the pose of a portion of a user are generally referred to as sensors. In some embodiments, the data generated using sensors is referred to as sensor data. In some embodiments, the term "pose data" is used to refer to data that can be used (e.g., by a computer system) to determine the pose of at least a portion of a user. In some embodiments, pose data may include sensor data.
[0103] In the embodiments disclosed herein, the computer system uses sensor data to determine the pose of parts of user 701 and then represents the user as an avatar in the CGR environment, where the avatar has the same pose as user 701. However, in some cases, the sensor data may be insufficient to determine the pose of some parts of the user's body. For example, a part of the user may be outside the sensor's field of view, the sensor data may be corrupted, uncertain, or incomplete, or the user may be moving too fast for the sensor to capture the pose. In any case, the computer system determines (e.g., estimates) the poses of these parts of the user based on various datasets, discussed in more detail below. Because these poses are estimates, the computer system computes a deterministic estimate (a confidence level of the accuracy of the determined pose) of each corresponding pose determined for parts of the user's body (particularly those not adequately represented by the sensor data). In other words, the computer system computes a deterministic estimate of the estimated pose of the corresponding part of user 701 as an accurate representation of the actual pose of the user's part in the real environment 700. The deterministic estimate is sometimes referred to in this paper as determinism (or uncertainty) or as a quantity of determinism of the estimated pose of a part of the user's body (or a quantity of uncertainty of the estimated pose of a part of the user's body). For example, in Figure 7A In this calculation, the computer system determined with 75% certainty that the user's left elbow was raised to one side and bent at a 90° angle. The certainty (confidence) of the determined pose is represented using deterministic diagram 710, which is discussed below. Figure 7B Discussion. In some implementations, the deterministic diagram 710 also represents the pose of user 701 determined by the computer system. The determined pose is represented using an avatar 721, which will be discussed below. Figure 7C Discussion.
[0104] Now for reference Figure 7B Deterministic diagram 710 is a deterministic visual representation of the determined pose of a portion of user 701 by the computer system. In other words, deterministic diagram 710 represents a calculated estimate (e.g., the probability that the estimated location of a portion of the user's body is correct) using the deterministic determination of the localization of different parts of the user's body by the computer system. Deterministic diagram 710 is a template of the human body, representing basic human features such as head, neck, shoulders, torso, arms, hands, legs, and feet. In some embodiments, deterministic diagram 710 represents various sub-features such as fingers, elbows, knees, eyes, nose, ears, and mouth. The human features of deterministic diagram 710 correspond to the physical features of user 701. Shading lines 715 are used to indicate the degree or amount of uncertainty in the determined localization or pose of the physical portion of user 701 corresponding to the corresponding shaded portion of deterministic diagram 710.
[0105] In the embodiments provided herein, the user 701 and the corresponding portions of the deterministic diagram 710 are described at a granularity sufficient to describe the various embodiments disclosed herein. However, without departing from the spirit and scope of this disclosure, these features may be described at an additional (or less) granularity to further describe the user's pose and portions, and the corresponding determinism of the pose. For example, user features of portions 701-4 captured using camera 705-4 may be further described as including the tips of the user's fingers, but excluding the base of the fingers located behind cup 702. Similarly, the back of the user's right hand faces away from camera 705-2 and can therefore be considered outside the field of view 707-2, since the image data obtained using camera 705-2 does not directly capture the back of the user's right hand. As another example, the deterministic diagram 710 may represent the determinism of a quantity for the user's left eye, and the determinism of a different quantity for the user's left ear or the top of the user's head. However, for the sake of brevity, the details of these granular variations are not described in all cases.
[0106] exist Figure 7B In the illustrated embodiment, deterministic diagram 710 includes portions 710-1, 710-2, 710-3, and 710-4 corresponding to the respective portions 701-1, 701-2, 701-3, and 701-4 of user 701. Therefore, portion 710-1 represents the deterministic pose of the user's neck, collar area, and a portion of the user's face and head, including the user's right eye, right ear, nose, and mouth, but excluding the top of the user's head, left eye, and left ear. Portion 710-2 represents the deterministic pose of the user's right hand, right wrist, and the distal portion of the user's forearm adjacent to the right wrist. Portion 710-3 represents the deterministic pose of the user's left and right feet, and the left and right lower leg areas. Portion 710-4 represents the deterministic pose of the user's left hand, left wrist, and the distal portion of the user's left forearm adjacent to the wrist. Figure 7B In the illustrated embodiment, portions 710-1, 710-2, and 710-3 are shown without shading because the computer system determines the pose of the physical features of portions 701-1, 701-2, and 701-3 of user 701 with a high degree of certainty (e.g., the physical features are within the camera's field of view and determined to be stationary). However, portion 710-4 is shown with a slight shading 715-1 on the portion of the deterministic diagram 710 corresponding to the user's left hand palm. This is because the computer system has lower certainty regarding the positioning of the user's left hand palm, where the left hand is in... Figure 7AThe left hand is obscured by a cup 702. In this example, the computer system is not very certain about the pose of the user's left hand because it is obscured by the cup. However, there are other reasons why the computer system may have a certain amount of uncertainty in the determined pose. For example, the user's part may move too fast for the sensor to detect its pose, or the lighting in the real environment may be poor (e.g., the user is backlit), etc. In some implementations, the computer system is able to determine the determinism of the pose based on, for example, the sensor's frame rate, the sensor's resolution, and / or the ambient light level.
[0107] Deterministic diagram 710 represents the situation in... Figure 7A The determinism diagram 710 represents the determinism of the poses of user 701 captured within the field of view of the cameras 705-1 to 705-4, as well as the determinism of the poses of user 701 not captured within the field of view of the cameras 705-1 to 705-4. Therefore, the determinism diagram 710 also represents the determinism of the computer system in estimating the poses of user 701 determined for portions outside the field of view of cameras 705-1 to 705-4 (e.g., the top of the user's head, the user's left eye, the user's left ear, the user's upper arm, elbow, and the proximal end of the user's forearm, the user's thigh and knee, the user's torso, and the user's left palm as previously described). Because these portions of user 701 are outside the field of view 707-1 to 707-4, the determinism of the poses determined for these portions is less certain. Therefore, the corresponding areas of the determinism diagram 710 are shown with shaded lines 715 to indicate the uncertainty of the estimates associated with the poses determined for each of these portions of user 701.
[0108] exist Figure 7B In the illustrated embodiment, the density of the shading 715 is proportional to the uncertainty represented by the shading (or inversely proportional to the certainty represented by the shading). Therefore, a higher density shading 715 indicates a greater uncertainty (or less certainty) in the pose of the corresponding physical part of user 701, and a lower density shading 715 indicates a smaller uncertainty (or greater certainty) in the pose of the corresponding physical part of user 701. No shading indicates a high degree of certainty (e.g., 90%, 95%, or 99% certainty) in the pose of the corresponding physical part of user 701. For example, in… Figure 7B In the illustrated embodiment, the shaded line 715 is located on portion 710-5 of the deterministic diagram 710, which corresponds to the top of the user's head, the user's left eye, and the user's left ear, while there is no shaded line on portion 710-1. Therefore, in Figure 7BIn the deterministic diagram 710, the deterministic nature of the pose of part 701-1 of user 701 is highly deterministic, while the deterministic nature of the poses of the top of the user's head, the user's left eye, and the user's left ear is less deterministic. In the case where the user wears a head-mounted device (e.g., where the representation of the user's parts is generated by the head-mounted device or at least partially based on sensors integrated into or mounted on the head-mounted device), the appearance of the parts of the user's head and face covered by the head-mounted device will be uncertain, but can be estimated based on sensor measurements of the visible parts of the user's head and / or face.
[0109] As described above, although some parts of user 701 are outside the camera's field of view, the computer system can determine the approximate pose of these parts of user 701 with varying degrees of certainty. These varying degrees of certainty are represented in the deterministic diagram 710 by showing different densities of the shading lines 715. For example, the computer system estimates the pose of the upper head region of the user (top of the user's head, left eye, and left ear) with high certainty. Therefore, this region is depicted in... Figure 7B In the diagram, portion 710-5 has a low shading density, as indicated by the large spacing between the shading lines. As another example, the computer system has less certainty in estimating the pose of the user's right elbow than in estimating the pose of the user's right shoulder. Therefore, the right elbow portion 710-6 is depicted with a greater shading density than the right shoulder portion 710-7, as indicated by the reduced spacing between the shading lines. Finally, the computer system estimates the pose of the user's torso (e.g., waist) with minimal certainty. Therefore, the torso portion 710-8 is depicted with the maximum shading density.
[0110] In some implementations, data from various sources can be used to determine or infer the pose of a corresponding part of user 701. For example, in Figure 7BIn the illustrated embodiment, the computer system determines the pose of the user's portion 701-2 based on detecting this portion of the user within the field of view 707-2 of camera 705-2; however, the pose of the user's right elbow is not known solely based on sensor data from camera 705-2. However, the sensor data from camera 705-2 can be supplemented to determine the pose of portions of the user 701 located inside or outside the field of view 707-2, including the pose of the user's right elbow. For example, if the computer system has high certainty regarding the pose of the user's neck and collar area (see section 710-1), the computer system can infer the pose of the user's right shoulder with high certainty (e.g., but slightly less than the neck and collar area) based on known mechanics of the human body (specifically, in this example, based on knowledge of the positioning of the person's right shoulder relative to the collar area) (see section 710-7). In some embodiments, it can use algorithms to extrapolate or interpolate to approximate the pose of portions of the user (such as the upper part of the user's right arm and the user's right elbow). For example, the potential pose of the right elbow portion 710-6 depends on the location of the right upper arm portion 710-9, which is inferred from the location of the right shoulder portion 710-7. Furthermore, the location of the right upper arm portion 710-9 depends on the joint motion of the user's shoulder joint, which is outside the field of view of any of the cameras 705-1 to 705-4. Therefore, the pose of the right upper arm portion 710-9 has greater uncertainty than the pose of the right shoulder portion 710-7. Similarly, the pose of the right elbow portion 710-6 depends on the estimated pose of the right upper arm portion 710-9 and the estimated pose of the proximal end of the user's right forearm. However, although the proximal end of the user's right forearm is not within the camera's field of view, the distal region of the user's right forearm is within the field of view 707-2. Therefore, as shown in portion 710-2, the pose of the distal region of the user's right forearm is known with high certainty. Therefore, the computer system uses this information to estimate the pose of the proximal end of the user's right forearm and right elbow portion 710-6. This information can also be used to further inform the estimated pose of the right upper arm portion 710-9.
[0111] In some implementations, the computer system uses an interpolation function to estimate the pose of a portion of the user's body, such as a portion located between two or more physical features with known poses. For example, referencing... Figure 7BIn section 710-4, the computer system determines the pose of the fingers on the user's left hand and the pose of the distal end of the user's left forearm with high certainty, because these physical features of the user 701 are within the field of view 707-4. However, the pose of the user's left palm is not known with high certainty because it is occluded by the cup 702. In some embodiments, the computer system uses an interpolation algorithm to determine the pose of the user's left palm based on the known poses of the user's fingers and left forearm. Because the pose is known for many physical features of the user near or adjacent to the left palm, the computer system determines the pose of the left palm with relatively high certainty, as indicated by the relatively sparse shading 715-1.
[0112] As described above, the computer system determines the pose of user 701 (in some embodiments, a set of poses of a portion of user 701) and displays an avatar representing user 701 having the determined pose in the CGR environment. Figure 7C An example of a computer system displaying avatar 721 in CGR environment 720 is shown. Figure 7C In the embodiment described herein, avatar 721 is presented as a virtual block character displayed using display generation component 730. Display generation component 730 is similar to display generation component 120 described above.
[0113] Incarnation 721 comprises parts 721-2, 721-3, and 721-4. Part 721-2 corresponds to... Figure 7B Part 710-2 of the deterministic diagram 710 in the diagram and Figure 7A Part 701-2 corresponds to user 701. Part 721-3 corresponds to Figure 7B Part 710-3 and the deterministic diagram 710 in the figure Figure 7A Part 701-3 corresponds to user 701. Part 721-4 corresponds to... Figure 7B Part 710-4 and the deterministic diagram in 710 Figure 7A Part 701-4 of user 701. Therefore, part 721-2 represents the pose of the user's right hand, right wrist, and the distal portion of the user's forearm adjacent to the user's right wrist. Part 721-3 represents the pose of the user's left foot, right foot, and the left and right lower leg regions. Part 721-4 represents the pose of the user's left hand, left wrist, and the distal portion of the user's left forearm adjacent to the wrist. Other parts of avatar 721 represent the poses of corresponding parts of user 701, and are discussed in more detail below. For example, the neck and collar region 727 of avatar 721 corresponds to part 701-1 of user 701, as described below.
[0114] The computer system displays an avatar 721 with visual display characteristics that vary depending on the estimation certainty of the pose determined by the computer system (referred to herein as variable display characteristics). Typically, variable display characteristics inform the viewer of avatar 721 about the visual fidelity of the displayed avatar relative to the pose of user 701, or about the visual fidelity of a displayed portion of the avatar relative to the pose of a corresponding physical portion of user 701. In other words, variable display characteristics inform the viewer of the extent to which the rendered pose or appearance of a portion of avatar 721 conforms to an estimate of the actual pose or appearance of the corresponding physical portion of user 701 in the real environment 700, which, in some embodiments, is determined based on the estimation certainty of the corresponding pose of that portion of user 701. In some embodiments, when the computer system renders an avatar feature without variable display characteristics or with variable display characteristics, the variable display characteristics having a value indicating the pose of a corresponding portion of user 701 with high certainty (e.g., greater than 75%, 80%, 90%, 95% certainty) are referred to as rendering with high fidelity. In some embodiments, when the computer system renders an avatar feature with variable display characteristics, the variable display characteristics having a value indicating the pose of a corresponding portion of user 701 with lower certainty are referred to as rendering with low fidelity.
[0115] In some implementations, the computer system adjusts the values of variable display characteristics of the corresponding avatar features to convey changes in the estimated determinism of the pose of a portion of the user represented by the corresponding avatar feature. For example, as the user moves, the computer system (e.g., continuously, constantly, automatically) determines the user's pose and updates the estimated determinism of the pose of the corresponding portion of the user accordingly. The computer system also modifies the appearance of the avatar by modifying the avatar's pose to match the newly determined user pose, and modifies the values of the variable display characteristics of the avatar features based on changes in the determinism of the corresponding pose of the user's portion.
[0116] Variable display characteristics may include one or more visual features or parameters used to enhance, degrade, or modify the appearance of avatar 721 (e.g., the default appearance), as discussed in the following examples.
[0117] In some implementations, variable display characteristics include the display color of the avatar features. For example, the avatar may have a default color of green, and as the determinism of the user's partial pose changes, cooler colors may be used to represent portions of the avatar where the computer system is less certain about the pose of the user's partial pose represented by the avatar features, and warmer colors may be used to represent portions of the avatar where the computer system is more certain about the pose of the user's partial pose represented by the avatar features. For example, when a user moves their hand out of the camera's field of view, the computer system modifies the avatar's appearance by moving the avatar's hand from a pose with a high degree of determinism (e.g., positioning, orientation) (a pose that matches the user's hand pose when it is in the field of view) to a pose with less determinism based on the updated determination of the user's hand pose. When the avatar's hand moves from a pose with a high degree of determinism to a pose with very little determinism, the avatar's hand changes from green to blue. Similarly, when the avatar's hand moves from a pose with a high degree of determinism to a pose with slightly less determinism, the hand changes from green to red. As another example, when the hand moves from a pose with very little determinism to a pose with relatively high determinism, the hand changes from blue to red, wherein various intermediate colors change from cool to warm as the hand moves to a pose with increasing determinism. As yet another example, when the hand moves from a pose with relatively high determinism to a pose with very little determinism, the hand changes from red to blue, wherein various intermediate colors change from warm to cool as the hand moves to a pose with decreasing determinism. In some implementations, the change in the value of the variable display characteristic occurs at a faster rate when the avatar feature moves from an unknown pose (a pose with very little determinism) to a more deterministic pose (e.g., a known pose), and at a slower rate when the avatar feature moves from a more deterministic pose (e.g., a known pose) to a less deterministic pose. For example, to continue with the example where the variable display characteristic is color, the avatar hand would change from blue to red at a faster rate than from red (or the default green) to blue. In some implementations, the color change occurs at a rate faster than the user moves the hand corresponding to the avatar's hand.
[0118] In some implementations, the variable display characteristic includes the amount of blurring effect applied to the displayed portion of the avatar feature. Conversely, the variable display characteristic can be a display sharpness measure applied to the displayed portion of the avatar feature. For example, the avatar may have a default sharpness, and as the pose of the user's portion changes with certainty, increased blur (or decreased sharpness) can be used to represent the portion of the avatar where the computer system is less certain about the pose of the user's portion represented by the avatar feature, and decreased blur (or increased sharpness) can be used to represent the portion of the avatar where the computer system is more certain about the pose of the user's portion represented by the avatar feature.
[0119] In some implementations, the variable display characteristic includes the opacity of the displayed portion of the avatar feature. Conversely, the variable display characteristic may be the transparency of the displayed portion of the avatar feature. For example, the avatar may have a default opacity, and as the pose of the user's portion changes with certainty, decreasing opacity (or increasing transparency) can be used to represent the portion of the avatar where the computer system is less certain about the pose of the user's portion represented by the avatar feature, and increasing opacity (or decreasing transparency) can be used to represent the portion of the avatar where the computer system is more certain about the pose of the user's portion represented by the avatar feature.
[0120] In some implementations, variable display characteristics include the density and / or size of the particles forming the display portion of the avatar feature. For example, an avatar may have a default granularity and spacing, and the granularity size and / or granular spacing of the corresponding avatar feature changes based on whether the determinism increases or decreases as the pose of the user's portion changes. For example, in some implementations, the particle spacing is increased (particle density decreased) to represent an avatar portion of the avatar where the computer system is less certain about the pose of the user 701's portion represented by the avatar feature, and the particle spacing is decreased (particle density increased) to represent an avatar portion of the avatar where the computer system is more certain about the pose of the user's portion represented by the avatar feature. As another example, in some implementations, the granularity is increased (creating a more pixelated and / or lower resolution appearance) to represent a portion of the avatar for which the computer system is less certain about the pose of the user 701's portion represented by the avatar feature, and the granularity is decreased (creating a less pixelated and / or higher resolution appearance) to represent a portion of the avatar for which the computer system is more certain about the pose of the user's portion represented by the avatar feature.
[0121] In some implementations, variable display characteristics include one or more visual effects, such as scales (e.g., fish scales), patterns, shadows, and smoke effects. Examples are discussed in more detail below. In some implementations, for example, by varying the amount of variable display characteristics (e.g., opacity, graininess, color, diffusion, etc.) along a transition region of the corresponding portion of the avatar, the avatar can be displayed as a smooth transition between regions of higher and lower determinism. For example, when the variable display characteristic is particle density and the determinism of the user's forearm transitions from high to low at the user's elbow, the transition from high to low determinism can be represented by displaying an avatar with high particle density at the elbow and smoothly (gradually) transitioning to low particle density along the forearm.
[0122] Variable display characteristics in Figure 7CThe shading is represented by shading line 725, which may include one or more of the variable display characteristics discussed above. The density of shading line 725 varies by a certain amount to indicate the fidelity of the portion of avatar 721 rendered based on the estimated certainty of the actual pose of the corresponding physical portion of user 701. Thus, a larger shading density is used to indicate a portion of the variable display characteristics of avatar 721 that indicates a value that instructs the computer system to determine the pose with lower certainty, and a lower shading density is used to indicate a portion of the variable display characteristics of avatar 721 that instructs the computer system to determine the pose with higher certainty. In some embodiments, shading is not used when the certainty of the pose of the portion of the avatar is high (e.g., 90%, 95%, or 99% certainty).
[0123] exist Figure 7C In the illustrated embodiment, avatar 721 is rendered as a block character with a pose similar to user 701. As described above, the computer system determines the pose of avatar 721 based on sensor data representing the pose of user 701. For portions of user 701 with known poses, the computer system renders the corresponding portion of avatar 721 with default or baseline values of variable display characteristics (e.g., resolution, transparency, particle size, particle density, color) or no variable display characteristics (e.g., no pattern or visual effects), as indicated by the unshaded lines. For portions of user 701 with poses below a predetermined determinism threshold (e.g., pose determinism less than 100%, 95%, 90%, 80%), the computer system renders the corresponding portion of avatar 721 with values of variable display characteristics different from the default or baseline, or renders the corresponding portion of the avatar with variable display characteristics if they are not displayed to indicate high determinism, as indicated by the shaded line 725.
[0124] exist Figure 7C In the illustrated implementation, the density of the shading line 725 typically corresponds to Figure 7B The density of the shaded line 715 in the image. Therefore, the value of the variable display characteristic of the incarnation 721 usually corresponds to the density of the shaded line 715 in the image. Figure 7B The shaded line 715 in the diagram represents certainty (or uncertainty). For example, part 721-4 of incarnation 721 corresponds to... Figure 7B Part 710-4 and the deterministic diagram in 710 Figure 7A Part 701-4 of user 701. Therefore, part 721-4 is depicted on the avatar's left palm with shading line 725-1 (similar to...). Figure 7B The shaded line 715-1 in the figure shows that the value indicating the variable display characteristic corresponds to a relatively high degree of certainty in the pose of the user's left palm, but not as high as in other parts, while there is no shaded line on the fingers and distal end of the avatar's left forearm (indicating a high degree of certainty in the pose of the corresponding part of the user 701).
[0125] In some implementations, shading is not used when the estimation certainty of the pose of a portion of the avatar is relatively high (e.g., 99%, 95%, 90%). For example, portions 721-2 and 721-3 of avatar 721 are shown without shading. Because the pose of portion 701-2 of user 701 was determined with high certainty, therefore... Figure 7C In the rendering, portion 721-2 of avatar 721 is rendered without a shadow. Similarly, because the pose of portion 701-3 of user 701 is determined using high determinism, therefore... Figure 7C In the middle, render part 721-3 of the avatar 721 without shadow lines.
[0126] In some implementations, a portion of avatar 721 is displayed, which has one or more avatar features not derived from user 701. For example, avatar 721 is rendered with avatar hair 726, which is determined based on the visual attributes of the avatar character rather than the hairstyle of user 701. As another example, the avatar hands and fingers in portion 721-2 are block hands and block fingers (which are not human hands or fingers), however, they have features similar to... Figure 7A The user's hand and fingers are in the same pose. Similarly, avatar 721 is rendered as having a nose 722 that is different from the user's nose but has the same pose. In some embodiments, different avatar features can be different human features (e.g., different human noses), features from non-human characters (e.g., a dog's nose), or abstract shapes (e.g., a triangular nose). In some embodiments, different avatar features can be generated using machine learning algorithms. In some embodiments, the computer system uses features not derived from the user to render avatar features in order to save computational resources by avoiding the additional operation of rendering avatar features with high visual fidelity relative to the user's corresponding part.
[0127] In some implementations, a portion of avatar 721 is displayed, which has one or more avatar characteristics derived from user 701. For example, in Figure 7C In China, the use of Figure 7A The mouth of user 701 is the same as the mouth of avatar 724 rendered in avatar 721. As another example, using... Figure 7A The avatar 721 is rendered with the same right eye 723-1 as the right eye of user 701. In some implementations, such avatar features are rendered using a video feed of the corresponding portion of user 701 mapped onto a 3D model of the virtual avatar character. For example, the avatar's mouth 724 and right eye 723-1 are based on a video feed of the user's mouth and right eye captured in the field of view 707-1 of camera 705-1 and are displayed (e.g., via video pass-through) on avatar 721.
[0128] In some implementations, the computer system renders some features of avatar 721 with high (or improved) certainty of its pose, even when the computer system estimates the pose of the corresponding part of the user with relatively low certainty. For example, in Figure 7C In the diagram, the left eye 723-2 is shown without a shaded line, indicating high certainty in the user's left eye pose. However, in... Figure 7A In the middle, the user's left eye is outside the field of view 707-1, and... Figure 7B In the diagram, the certainty of the portion 710-5 including the user's left eye is shown by a shaded line 715, which indicates a lower certainty of the pose of the user's left eye. Even when the estimated certainty is not very high, some features are rendered with high fidelity in some cases to improve the quality of communication using the avatar 721. For example, some features, such as the eyes, hands, and mouth, can be considered important for communication purposes, and rendering such features with variable display characteristics could be distracting for the user viewing the avatar 721 and communicating with it. Similarly, if the user 701's mouth moves too quickly to capture enough pose data for the user's lips, teeth, tongue, etc., the computer system can render the avatar's mouth 724 with high certainty of the mouth's pose because rendering the mouth with variable display characteristics, such as blurring, transparency, color, increased particle spacing, or increased particle size, could be distracting.
[0129] In some implementations, when the pose or appearance of a user's part is not very certain (e.g., less than 99%, 95%, or 90% certain), data from different sources can be used to enhance the pose and / or appearance of the corresponding avatar feature. For example, although the pose of the user's left eye is unknown, a machine learning algorithm can be used to estimate the pose of the avatar's left eye 723-2, for example, based on the known mirror pose of the user's right eye. Additionally, since the user's left eye is outside the field of view 707-1, the sensor data from camera 705-1 may not include data for determining the appearance (e.g., eye color) of the user's left eye. In some implementations, the computer system can use data from other sources to obtain the data needed to determine the appearance of the avatar feature. For example, the computer system can access previously captured facial scan data, images, and videos of user 701 associated with a personal communication device (e.g., smartphone, smartwatch, or HMD) to obtain data for determining the appearance of the user's eyes. In some implementations, other avatar features can be updated based on additional data accessed by the computer system. For example, if recent photos show a user 701 with a different hairstyle, the avatar's hair can be changed to match the hairstyle in user 701's recent photos. In some embodiments, the device generating the representation of the user's part is separate from the personal communication device, and data from facial scans, images, and / or videos is provided (e.g., securely and privately, with one or more options for the user to decide whether to share data between devices) to the device generating the representation of the user's part for the purpose of constructing the representation of the user's part. In some embodiments, the device generating the representation of the user's part is the same as the personal communication device, and data from facial scans, images, and / or videos is provided to the device generating the representation of the user's part for the purpose of constructing the representation of the user's part (e.g., a facial scan is used to unlock an HMD, which also generates a representation of the user's part).
[0130] In some implementations, the computer system displays a portion of avatar 721 with low visual fidelity relative to the corresponding portion of the user, even when the pose of the corresponding user feature is known (the pose is determined with high determinism). For example, in Figure 7CIn the diagram, the neck and collar areas 727 of avatar 721 are shown with a shading line 725, indicating that the corresponding avatar features are displayed with variable display characteristics. However, the neck and collar areas of avatar 721 correspond to portion 701-1 of user 701 (which is within the field of view 707-1) and portion 710-1 (which is shown as having high determinism in deterministic diagram 710). In some embodiments, even when the pose of the corresponding portion of user 701 is known (or determined with high determinism), the computer system displays the avatar features with low visual fidelity to conserve computational resources that would otherwise be used to render a high-fidelity representation of the corresponding avatar features. In some embodiments, the computer system does this when user features are considered less important for communication purposes. In some embodiments, a low-fidelity version of the avatar features is generated using a machine learning algorithm.
[0131] In some implementations, the change in the value of the variable display characteristic is based on the movement speed of the corresponding part of the user 701. For example, the faster the user moves their hand, the less certain the hand pose is, and the variable display characteristic is rendered with a value commensurate with the less certainty of the hand pose. Conversely, the slower the user moves their hand, the more certain the hand pose is, and the variable display characteristic is rendered with a value commensurate with the greater certainty of the hand pose. For example, if the variable display characteristic is blurriness, then when the user moves their hand at a faster speed, the avatar's hand is rendered with greater blurriness, and when the user moves their hand at a slower speed, the avatar's hand is rendered with less blurriness.
[0132] In some embodiments, the display generation component 730 enables the display of the CGR environment 720 and avatar 725 to a user of the computer system. In some embodiments, the computer system further displays a preview 735 via the display generation component 730, which includes a representation of how the user of the computer system appears within the CGR environment 720. In other words, the preview 735 shows the user of the computer system how they behave within the CGR environment 720 to other users viewing the CGR environment 720. Figure 7C In the illustrated implementation, Preview 735 shows users of the computer system (e.g., users other than user 701) that they are portrayed as female avatars with variable display characteristics.
[0133] In some implementations, the computer system uses data collected from one or more sources other than cameras 705-1 to 705-4 to calculate deterministic estimates. For example, these other sources may be used to supplement or, in some cases, replace the sensor data collected from cameras 705-1 to 705-4. These other sources may include different sensors, such as any of the sensors listed above. For example, user 701 may wear a smartwatch or other wearable device that provides data indicating the positioning, movement, or location of the user's arm and / or other body parts. As another example, user 701 may have a smartphone in their pocket that provides data indicating the user's location, movement, pose, or other such data. Similarly, a headset device worn by another person in the real-world environment 700 may include sensors, such as cameras, that provide data indicating pose, movement, location, or other relevant information associated with user 701. In some implementations, prior data may be provided from facial scans (e.g., using depth sensors), media items (such as pictures and videos of user 701), or other relevant sources discussed above.
[0134] Figures 8A to 8C Depicting something similar to Figure 7A The implementation scheme incorporates sensor data from cameras 705-1 to 705-4, supplemented by data from the user's smartwatch. Figure 8A In this setup, user 701 is now positioned with his right hand on his hip and his left hand flat against the wall 805, with the smartwatch 810 on his left wrist. For example, user 701 has already... Figure 7A pose movement in Figure 8A The user's pose is now positioned using field of view 707-2, instead of the right hand, right wrist, and right forearm. Additionally, the user's left hand remains within field of view 707-4, while the user's left wrist and left forearm are outside field of view 707-4.
[0135] exist Figure 8B In China, based on Figure 8A The new pose update deterministic diagram 710 for user 701 is shown. Therefore, the uncertainty of the pose of the user's right elbow, right forearm, and right hand is updated based on the new positioning, as indicated by the increased density of the shaded lines 715 in these areas. Specifically, because the user's entire right arm is outside any field of view in the camera's field of view, the deterministic nature of the pose of these parts of user 701 is increased from... Figure 7B The determinism of the pose is reduced. Therefore, for Figure 8BThese portions of the deterministic graph, in deterministic graph 710, show an increasing shadow density. Because the pose of each sub-feature of the user's right arm depends on the location of adjacent sub-features, and all sub-features of the right arm are outside the field of view of the camera or other sensors, the shadow density increases with each successive sub-feature, starting from the upper right arm portion 811-1, to the right elbow portion 811-2, to the right forearm portion 811-3, and to the right hand and fingers 811-4.
[0136] Even when the user's left forearm is outside the field of view 707-4, the computer system still determines the pose of the user's left forearm with high determinism because the sensor data from camera 705-4 is supplemented by sensor data from smartwatch 810, which provides the pose data of the user's left forearm. Therefore, deterministic diagram 710 shows the left forearm portion 811-5 and the left hand portion 811-6 (which is within the field of view 707-4), each with a high degree of determinism in its pose.
[0137] Figure 8C It shows the basis Figure 8A The updated pose of user 701 in CGR environment 720, and the updated pose of avatar 721. Figure 8C In the middle, the computer system renders the avatar's left arm 821 with a high degree of certainty in the pose of the user's left arm, as indicated by the unshaded lines on the left arm 821.
[0138] In some implementations, the computer system renders the avatar 721 as a representation of the object with which the user 701 interacts. For example, if the user holds the object, leans against a wall, sits in a chair, or otherwise interacts with the object, the computer system may render at least a portion of the object (or its representation) in the CGR environment 720. Figure 8C In this embodiment, the computer system includes a wall rendering 825, which is shown as the left-hand positioning of a proximity avatar. This provides context for the user's pose, allowing a viewer to understand that the user 701 is posing by resting his left hand against a surface in the real environment 700. In some embodiments, the object interacting with the user may be a virtual object. In some embodiments, the object interacting with the user may be a physical object such as wall 805. In some embodiments, a rendered version of the object (e.g., wall rendering 825) may be rendered as a virtual object or displayed as a video feed of a real object.
[0139] As described above, the computer system updates the pose of the avatar 721 in response to a detected change in the pose of the user 701. In some embodiments, this involves updating variable display characteristics based on the pose change. In some embodiments, updating variable display characteristics includes increasing or decreasing the value of the variable display characteristics (represented by increasing or decreasing the amount or density of the shading 725). In some embodiments, updating variable display characteristics includes introducing or removing variable display characteristics (represented by introducing or removing the shading 725). In some embodiments, updating variable display characteristics includes introducing or removing visual effects. For example, in... Figure 8C In this embodiment, the computer system renders the avatar 721 with a right arm featuring a smoke effect 830 to indicate a relatively low or reduced certainty in the pose of the user's right arm. In some implementations, the visual effect may include other effects, such as fish scales displayed on the corresponding avatar feature. In some implementations, displaying the visual effect includes replacing the display of the corresponding avatar feature with the displayed visual effect, such as... Figure 8C As shown. In some embodiments, displaying visual effects includes visually displaying corresponding avatar features. For example, the avatar's right arm could display fish scales on the arm. In some embodiments, multiple variable display characteristics can be combined, such as visual effects displayed with another variable display characteristic. For example, the scale density on the arm decreases as pose determinism decreases along a portion of the arm.
[0140] exist Figure 8C In the illustrated implementation, the user's right arm pose is represented by a smoke effect 830, which roughly resembles the shape of the avatar's right arm in a lowered pose facing the side of the avatar's body. Although in Figure 8A The pose of the avatar's arm does not accurately represent the actual pose of the user's right arm, but the computer system does accurately determine that the user's right arm is lowered, rather than above the user's shoulder or directly to the side. This is because the computer system determines the current pose of the user's right arm based on sensor data obtained from cameras 705-1 to 705-4 and based on the user's previous arm pose and movement. For example, when the user... Figure 7A In the middle position, move his arm to Figure 8A When the user's right arm is in a pose, it moves downwards as it moves out of the field of view 707-2. The computer system uses this data to determine that the user's right arm pose is not raised or moved to the side, and therefore must be lower than its previous position. However, the computer system does not have enough data to accurately determine the user's right arm pose. Therefore, the computer system indicates a low degree of certainty (confidence) in the user's right arm pose, such as... Figure 8B The determinism is indicated by figure 710. In some implementations, when the determinism of a portion of the pose of user 701 is below a threshold, the computer system represents the corresponding avatar features with visual effects, such as... Figure 8C As shown.
[0141] exist Figure 8C In the preview 735, the current pose of the user waving is shown on the computer system in the CGR environment 720.
[0142] In some implementations, the computer system provides a control feature where the user of the computer system can select how many avatars 721 to display (e.g., which parts or features of avatar 721 to display). For example, the control provides a spectrum where, at one end, avatar 721 includes only avatar features for which the computer system has high determinism regarding the pose of the corresponding part to the user, while at the other end of the spectrum, avatar 721 is displayed along with all avatar features regardless of the determinism of the pose of the corresponding part to the user 701.
[0143] The following text refers to the information below. Figure 9 Method 900 describes how to provide information about Figures 7A to 7C and Figures 8A to 8C Additional description.
[0144] Figure 9 This is a flowchart of an exemplary method 900 for presenting a virtual avatar character with display characteristics that change in appearance based on the deterministic pose of the user's body, according to some embodiments. In some embodiments, method 900 is in conjunction with a display generation component (e.g., Figure 1 , Figure 3 and Figure 4 The display generation component 120 in the middle (for example, Figure 7C and Figure 8C A computer system (e.g., display generation components 730) that communicates with a display generation component 730 (e.g., a visual output device, a 3D display, a transparent display, a projector, a head-up display, a display controller, a touch screen, etc.) Figure 1 The method 900 is executed at a computer system 101 (e.g., a smartphone, tablet, head-mounted display generating component). In some embodiments, the method 900 is performed by a processor stored in a non-transitory computer-readable storage medium and by one or more processors of the computer system (such as one or more processors 202 of computer system 101). Figure 1 The control unit 110 in the middle executes instructions to manage. Some operations in method 900 are optionally combined, and / or the order of some operations is optionally changed.
[0145] In method 900, a computer system (e.g., 101) receives (902) pose data (e.g., image data from a camera (e.g., 705-1; 705-2; 705-3; 705-4; 705-5)) representing the pose (e.g., physical orientation, gesture, movement, etc.) of at least a first part of a user (e.g., 701-1; 701-2; 701-3; 701-4) (e.g., corresponding user features) (e.g., one or more physical features of the user (e.g., macroscopic features such as arms, legs, hands, head, mouth, etc.; and / or microscopic features such as fingers, face, lips, teeth, or other parts of the corresponding physical features)). In some embodiments, the pose data includes data (e.g., deterministic values) indicating that the determined pose of that part of the user is accurate (e.g., confidence level) (e.g., an accurate representation of the pose of that part of the user in a real environment). In some embodiments, pose data includes sensor data (e.g., image data from a camera; motion data from an accelerometer; location data from a GPS sensor; data from a proximity sensor; data from a wearable device (e.g., a watch; 810)). In some embodiments, the sensor may be connected to or integrated with a computer system. In some embodiments, the sensor may be an external sensor (e.g., a sensor from a different computer system (e.g., another user's electronic device)).
[0146] A computer system (e.g., 101) causes (904) (e.g., in a computer-generated real environment (e.g., 720)) to present (e.g., display; visual presentation; projection) an avatar (e.g., 721) (e.g., a virtual avatar; a part of an avatar) (e.g., an avatar is a virtual representation of at least a part of a user) via a display generation component (e.g., 730), wherein the avatar includes a corresponding avatar feature (e.g., 727, 724, 723-1, 723-2, 722, 726, 721-2, 721-3, 721-4), the corresponding avatar feature (e.g., The anatomically corresponding first part of the user (e.g., 701-1; 701-2; 701-3; 701-4) is presented (e.g., displayed) as a set of one or more visual parameters (e.g., 725, 830) indicating the pose of the user's first part, which in turn indicates the determinism of the pose of the user's first part (e.g., 710; 715). The presentation of an avatar, including the corresponding avatar feature corresponding to the user's first part and presented as having a variable display characteristic indicating the determinism of the pose of the user's first part, provides the user with feedback indicating the confidence level of the pose of the user's first part represented by the corresponding avatar feature. Providing improved feedback enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating and / or interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling users to use the system more quickly and effectively.
[0147] In some implementations, the corresponding avatar feature is overlaid on (or displayed in place of) the corresponding portion of the user. In some implementations, the variable display characteristics vary based on the certainty (e.g., confidence) of the pose of the user's portion. In some implementations, the certainty of the pose of the user's portion is represented as a certainty value (e.g., a value indicating the certainty (confidence) that the determined pose of the user's portion is an accurate representation of the actual (e.g., in a real environment (e.g., 700)) pose of the user's portion). In some implementations, a range of values is used to represent the certainty value, such as a percentage range from 0% to 100%, where 0% indicates that there is no accurate (e.g., minimum) certainty in the pose of the corresponding user feature (or a portion thereof), while 100% indicates that the certainty of the estimated pose of the corresponding user feature is higher than a predetermined threshold certainty (e.g., 80%, 90%, 95%, or 99%) (e.g., even if the actual pose cannot be determined due to, for example, sensor limitations). In some implementations, the determinism can be 0% when the computer system (e.g., 101) (or another processing device) does not have sufficiently useful data from which to deduce the potential location or pose of the corresponding user feature. For example, the corresponding user feature may not be within the field of view (e.g., 707-1, 707-2, 707-3, 707-4) of the image sensors (e.g., 705-1, 705-2, 705-3, 705-4), and the corresponding user feature may also be in any of a plurality of different locations or poses. Or, for example, the data generated using a proximity sensor (or some other sensor) is indeterminate or insufficient to accurately deduce the pose of the corresponding user feature. In some implementations, the determinism can be high (e.g., 99%, 95%, 90% determinism) when the computer system (or another processing device) can explicitly identify the corresponding user feature using the pose data and can determine the correct location of the corresponding user feature using the pose data.
[0148] In some implementations, the variable display characteristics are directly related to the determinism of the pose of that part of the user. For example, a variable display characteristic with a first value (e.g., a low value) can be used to render a corresponding avatar feature to convey a first degree of determinism (e.g., low determinism) in the pose of that part of the user (e.g., to a viewer). Conversely, a variable display characteristic with a second value (e.g., a high value) greater than the first value can be used to render a corresponding avatar feature to convey greater determinism (e.g., high determinism such as determinism above a predetermined determinism threshold (e.g., 80%, 90%, 95%, or 99%) in the pose of that part of the user. In some implementations, the variable display characteristics are not directly related to the determinism (e.g., determinism value) of the pose of that part of the user. For example, if the appearance of that part of the user is important for communication purposes, a variable display characteristic with a value corresponding to high determinism can be used to render a corresponding avatar feature even if the pose of that part of the user has low determinism. For example, this can be done if displaying the corresponding avatar feature with a variable display characteristic having a value corresponding to low determinism would be distracting (e.g., for the viewer). Consider, for instance, an embodiment where the avatar includes an avatar head displayed on (e.g., superimposed on) the user's head, and the avatar head includes avatar lips (e.g., 724) representing the pose of the user's lips. While the computer system (or another processing device) determines the user's lip pose with low determinism as the user's lips move (e.g., the lips move too quickly to be accurately detected, the user's lips are partially covered, etc.), the corresponding avatar lips can be rendered with a variable display characteristic having a value corresponding to high (e.g., maximum) determinism, because otherwise, displaying the avatar with lips rendered with a variable display characteristic having a value corresponding to low determinism or less than high (e.g., maximum) determinism would be distracting.
[0149] In some embodiments, variable display characteristics (e.g., 725, 830) indicate the estimated visual fidelity of a corresponding avatar feature (e.g., 727, 724, 723-1, 723-2, 722, 726, 721-2, 721-3, 721-4) relative to the pose of a first part of the user. In some embodiments, visual fidelity represents the realism of the pose of the displayed / rendered avatar (or a portion thereof) relative to that part of the user. In other words, visual fidelity is a measure of how well the displayed / rendered avatar (or a portion thereof) is perceived to correspond to the actual pose of the corresponding part of the user. In some embodiments, whether an increase or decrease in the value of a variable display characteristic indicates improved or decreased visual fidelity depends on the type of variable display characteristic used. For example, if the variable display characteristic is a blurring effect, a larger value of the variable display characteristic (greater blur) indicates decreased visual fidelity, and vice versa. Conversely, if the variable display characteristic is particle density, then a larger value of the variable display characteristic (a larger particle density) conveys improved visual fidelity, and vice versa.
[0150] In method 900, presenting an avatar (e.g., 721) includes: determining, based on (e.g., pose data) a pose of a first portion of a user (e.g., 701-1; 701-2; 701-3; 701-4) associated with a first deterministic value (e.g., the deterministic value is a first deterministic value), and a computer system (e.g., 101) presents (906) an avatar (e.g., 721) having corresponding avatar features (e.g., 727, 724, 723-1, 723-2, 722, 726, 721-2, 721-3, 721-4), the corresponding avatar feature having a first value of a variable display characteristic (e.g., 727 in...). Figure 7C The low-density shading lines 725; 721-3, 721-2, 723-1, 723-2, 722 and / or 724 are in Figure 7C There is no shading line for 725; part 721-4 is in Figure 7CThe left hand in the image has a low shadow density 725-1 (e.g., the corresponding avatar feature is displayed as a first variable display characteristic value (e.g., amount of blur, opacity, color, attenuation / density, resolution, etc.) indicating the pose of the corresponding avatar feature relative to the user's part). Based on determining the pose of the user's first part and associating it with the first deterministic value, the avatar of the corresponding avatar feature presenting the first value of the variable display characteristics provides the user with deterministic feedback indicating that the pose of the user's first part, represented by the corresponding avatar feature, corresponds to the actual pose of the user's first part. Providing improved feedback enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating and / or interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively. In some embodiments, the avatar of the corresponding avatar feature presenting the first value of the variable display characteristics includes presenting the corresponding avatar feature having the same pose as the user's part.
[0151] In method 900, presenting an avatar (e.g., 721) includes: determining, based on (e.g., pose data) the pose of a first portion of the user (e.g., 701-1; 701-2; 701-3; 701-4) associated with a second deterministic value that is different from (e.g., greater than) the first deterministic value, and a computer system (e.g., 101) presents (908) an avatar (e.g., 721-2, 723-1, 723-2, 722, 726, 721-2, 721-3, 721-4) of the corresponding avatar features (e.g., 721-2 in...) having a second value of the variable display characteristics that is different from the first value of the variable display characteristics. Figure 8C The variable display feature 830 is used to represent this; the avatar's left hand is... Figure 8C There is no shading line in the middle; the incarnation's left elbow is... Figure 8C(e.g., the corresponding avatar feature is displayed with fewer shading lines 725) (e.g., the corresponding avatar feature is displayed with a second variable display characteristic value (e.g., amount of blur, opacity, color, attenuation / density, resolution, etc.), which indicates a second estimated visual fidelity of the corresponding avatar feature relative to the pose of that part of the user, wherein the second estimated visual fidelity is different from the first estimated visual fidelity (e.g., the second estimated visual fidelity indicates an estimate of visual fidelity higher than the first estimated visual fidelity)). Based on determining the pose of the first part of the user and associating it with a second deterministic value, the avatar of the corresponding avatar feature, which is presented with a second value of variable display characteristics that is different from the first value of the variable display characteristics, provides the user with feedback indicating that the pose of the first part of the user represented by the corresponding avatar feature corresponds to the actual pose of the first part of the user. Providing improved feedback enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating and / or interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively. In some implementations, an avatar that presents a second value of a corresponding avatar feature having variable display characteristics includes an avatar feature that presents a corresponding avatar feature having the same pose as the user's portion.
[0152] In some implementations, a computer system (e.g., 101) receives second pose data representing the pose of at least a second part (e.g., 701-2) of a user (e.g., a part of the user that differs from the first part of the user), and causes the presentation (e.g., updating of the presented avatar) of an avatar (e.g., 721) via a display generation component (e.g., 730). The avatar includes a second avatar feature (e.g., 721-2) corresponding to the second part of the user (e.g., the second part of the user is the user's mouth, and the second avatar feature is a representation of the user's mouth), and is presented with a second variable display characteristic (e.g., a variable display characteristic identical to the first variable display characteristic) indicating the determinism of the pose of the second part of the user (e.g., a variable display characteristic different from the first variable display characteristic) (e.g., 721-2 is displayed without a shading 725, thereby indicating a high degree of determinism in the pose of 701-2). The presentation of an avatar, which includes a second part corresponding to the user and is presented with variable display characteristics indicating the pose of the user's second part, provides the user with feedback indicating the confidence level of the pose of the user's second part represented by the second avatar feature. This further provides feedback on various levels of confidence in the poses of different parts of the user. Providing improved feedback enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating and / or interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively.
[0153] In some implementations, presenting the avatar (e.g., 721) includes associating the pose of a second part of the determined user (e.g., 701-2) with a third deterministic value (e.g., Figure 7B (710-2 without shaded line 715) presents an incarnation with a second incarnation characteristic, which has a first value of a second variable display characteristic (e.g., Figure 7C(See section 721-2, without shading 725). Based on the determination of the pose of the user's second part and its association with a third deterministic value, an avatar of a second avatar feature, presenting a first value with a second variable display characteristic, provides the user with deterministic feedback indicating that the pose of the user's second part, represented by the second avatar feature, corresponds to the actual pose of the user's second part. This further provides feedback on various levels of confidence in the poses of different parts of the user. Providing improved feedback enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating and / or interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively.
[0154] In some implementations, presenting an avatar (e.g., 721) includes: associating the pose of a second part of the determined user (e.g., 701-2) with a fourth deterministic value different from the third deterministic value (e.g., 811-3 and 811-4 in...). Figure 8B An avatar with a second avatar feature having a second value of a second variable display characteristic that is different from the first value of the second variable display characteristic (e.g., avatar 721 is presented as the avatar's right arm (which includes portions 721-2) having a smoke effect 830, which is a variable display characteristic (e.g., in some embodiments, similar to the variable display characteristic represented by the shading 725) is presented. Based on the determination of the pose of the user's second part and its association with a fourth certainty value, an avatar with a corresponding avatar feature having a second value of a second variable display characteristic that is different from the first value of the second variable display characteristic provides the user with feedback indicating that the pose of the user's second part, represented by the second avatar feature, corresponds to the actual pose of the user's second part. This further provides feedback on various levels of confidence in the poses of the different parts of the user. Providing improved feedback enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating and / or interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling users to use the system more quickly and effectively.
[0155] In some implementations, presenting an avatar (e.g., 721) includes: determining that a third certainty value corresponds to (e.g., equal to, identical to) a first certainty value (e.g., the certainty of the pose of the user's first part is 50%, 55%, or 60%, and the certainty of the pose of the user's second part is also 50%, 55%, or 60%) (e.g., in...). Figure 7BIn the figure, the deterministic representation of the user's right elbow has a moderate shading density as shown in portion 710-6 of the deterministic diagram 710, and the deterministic representation of the user's left elbow also has a moderate shading density. The first value of the second variable display characteristic corresponds to (e.g., equal to, the same as) the first value of the variable display characteristic (e.g., in the figure). Figure 7C In the image, both the left and right elbows of the avatar have moderate shadow density (e.g., the variable display characteristic and the second variable display characteristic both have values indicating 50%, 55%, or 60% certainty of the pose of the corresponding part of the user (e.g., the first value of the variable display characteristic indicates 50%, 55%, or 60% certainty of the pose of the first part of the user, and the first value of the second variable display characteristic indicates 50%, 55%, or 60% certainty of the pose of the second part of the user)).
[0156] In some implementations, presenting the avatar (e.g., 721) includes: determining that a fourth certainty value corresponds to (e.g., equal to, the same as) a second certainty value (e.g., the certainty of the pose of the user's first part is 20%, 25%, or 30%, and the certainty of the pose of the user's second part is also 20%, 25%, or 30%) (e.g., in...). Figure 7B In the figure, the user's right upper arm has a low shadow density as shown in portions 710-9 of the determination figure 710, and the user's left upper arm also has a low shadow density. The second value of the second variable display characteristic corresponds to (e.g., equal to, the same as) the second value of the variable display characteristic (e.g., in the figure). Figure 7C In the image, the upper left and upper right arms of the avatar both have low shadow density (e.g., the variable display characteristic and the second variable display characteristic both have values indicating 20%, 25%, or 30% certainty of the pose of the corresponding part of the user (e.g., the second value of the variable display characteristic indicates 20%, 25%, or 30% certainty of the pose of the first part of the user, and the second value of the second variable display characteristic indicates 20%, 25%, or 30% certainty of the pose of the second part of the user)).
[0157] In some implementations, the corresponding avatar feature and the first part of the user have the same relationship between determinism and the value of the variable display characteristic (e.g., visual fidelity) as the second avatar feature and the second part of the user. For example, the determinism of the pose of the first part of the user directly corresponds to the value of the variable display characteristic, and the determinism of the pose of the second part of the user also directly corresponds to the value of the second variable display characteristic. This is illustrated in the following example demonstrating different changes in the determinism value: when the determinism of the pose of the first part of the user decreases by 5%, 7%, or 10%, the value adjustment of the variable display characteristic indicates the amount by which the determinism decreases by 5%, 7%, or 10% (e.g., there is a direct or proportional mapping between the determinism and the value of the variable display characteristic), and when the determinism of the pose of the second part of the user increases by 10%, 15%, or 20%, the value adjustment of the second variable display characteristic indicates the amount by which the determinism increases by 10%, 15%, or 20% (e.g., there is a direct or proportional mapping between the determinism and the value of the second variable display characteristic).
[0158] In some implementations, method 900 includes: determining that a third certainty value corresponds to (e.g., equal to, identical to) a first certainty value (e.g., the certainty of the pose of a first part of the user is 50%, 55%, or 60%, and the certainty of the pose of a second part of the user is also 50%, 55%, or 60%) (e.g., in...). Figure 7B In the figure, neither part 710-2 nor part 710-1 has a shaded line as shown in the deterministic figure 710. The first value of the second variable display characteristic does not correspond to (e.g., is not equal to, is different from) the first value of the variable display characteristic (e.g., the avatar part 721-2 does not have a shaded line, while the avatar neck 727 has a shaded line 725). (e.g., the variable display characteristic and the second variable display characteristic have values that indicate different deterministic quantities of the pose of the corresponding part of the user (e.g., the first value of the variable display characteristic indicates 50%, 55%, or 60% determinism of the pose of the first part of the user, and the first value of the second variable display characteristic indicates 20%, 25%, or 30% determinism of the pose of the second part of the user)).
[0159] In some implementations, method 900 includes: determining that a fourth certainty value corresponds to (e.g., equal to, identical to) a second certainty value (e.g., the certainty of the pose of the first part of the user is 20%, 25%, or 30%, and the certainty of the pose of the second part of the user is also 20%, 25%, or 30%) (e.g., in... Figure 7BIn the figure, portions 710-7 and 710-5 both have low-density shading lines 715 as shown in the deterministic figure 710. The second value of the second variable display characteristic does not correspond to (e.g., is not equal to, is different from) the second value of the variable display characteristic (e.g., the left avatar eye 723-2 does not have a shading line, while the avatar collar portion 727 has a shading line 725). (e.g., the variable display characteristic and the second variable display characteristic have values that indicate different deterministic quantities of the pose of the corresponding portions of the user (e.g., the second value of the variable display characteristic indicates 20%, 25%, or 30% determinism of the pose of the first portion of the user, and the second value of the second variable display characteristic indicates 40%, 45%, or 50% determinism of the pose of the second portion of the user)).
[0160] In some implementations, the relationship between determinism and the values of variable display characteristics (e.g., visual fidelity) differs for the first part of the corresponding avatar feature and user compared to the second part of the second virtual feature and user. For example, in some implementations, the values of the variable display characteristics are selected to indicate determinism that differs from the actual determinism of the pose.
[0161] For example, in some implementations, the value of a variable display characteristic indicates a higher degree of certainty than it is actually associated with the corresponding user characteristic. This can be done, for example, when the characteristic is considered important for communication. In this example, rendering a characteristic with a variable display characteristic that indicates a higher degree of certainty (e.g., a high-fidelity representation of the corresponding avatar characteristic) enhances communication because rendering a corresponding avatar characteristic with a variable display characteristic that indicates an accurate level of certainty would be distracting.
[0162] As another example, in some implementations, the value of a variable display characteristic indicates a lower degree of certainty than it would actually be associated with the corresponding user characteristic. This can be done, for example, when the characteristic is not considered important for communication. In this example, rendering a characteristic with a variable display characteristic that indicates lower certainty (e.g., a low-fidelity representation of the corresponding avatar characteristic) saves the computational resources typically consumed when rendering the corresponding avatar characteristic with a variable display characteristic that indicates high certainty (e.g., a high-fidelity representation of the corresponding avatar characteristic). Because the user characteristic is not considered important for communication, computational resources can be saved without sacrificing communication efficiency.
[0163] In some implementations, the computer system (e.g., 101) receives updated pose data representing pose changes of a first portion (e.g., 701-1; 701-2; 701-3; 701-4) of a user (e.g., 701). In response to receiving the updated pose data, the computer system updates the presentation of the avatar (e.g., 721), including updating the pose of the corresponding avatar feature based on (e.g., at least one of the magnitude or direction of the change in pose of the user's first portion) (e.g., updating the pose of the corresponding avatar feature by an amplitude and / or direction corresponding to the magnitude and / or direction of the change in pose of the user's first portion) (e.g., the user 701's left hand from...). Figure 7A Moving from an upright position to Figure 8A The location on the wall 805, and the left hand that transforms into 721 from Figure 7C Moving from an upright position to Figure 8C (Positioning on the wall shown). The pose of the corresponding avatar feature is updated based on changes in the pose of the user's first part, providing feedback to the user indicating that movement of the user's first part causes the computer system to modify the corresponding avatar feature accordingly. This provides a control scheme for operating and / or constituting a virtual avatar using display generation components, where the computer system processes input including the user's first part in the form of changes in the user's physical characteristics (and the magnitude and / or direction of these changes) and provides feedback in the form of the virtual avatar's appearance. This provides the user with improved visual feedback regarding changes in the pose of the user's body characteristics. This enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating and / or interacting with the computer system), which in turn reduces power consumption and extends the computer system's battery life by enabling the user to use the computer system more quickly and efficiently. In some implementations, the avatar's positioning is updated based on the user's movement. For example, the avatar is updated in real time to mirror the user's movement. For instance, when a user places their arm behind their back, the display avatar moves its arm behind their back to reflect the user's movement.
[0164] In some implementations, updating the presentation of the avatar (e.g., 721) includes: in response to a deterministic change in the pose of the user's first part during a change in the pose of the user's first part, in addition to changing the positioning of at least a portion of the avatar based on the change in the pose of the user's first part, also changing the variable display characteristics of the corresponding avatar features displayed (e.g., when the user's right hand moves from...). Figure 7A pose movement in Figure 8A When in the pose, the corresponding incarnation feature (the incarnation's right hand) is from Figure 7C The shading line in the middle becomes Figure 8C(The image contains smoke features 830). In addition to altering the positioning of at least a portion of the avatar based on changes in the pose of the user's first part, variable display characteristics that alter the displayed avatar features provide feedback to the user, indicating that the determinism of the pose of the corresponding avatar features is affected by changes in the pose of the user's first part. Providing improved feedback enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by assisting the user in providing appropriate input and reducing user errors when operating and / or interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling the user to use the computer system more quickly and effectively.
[0165] In some implementations, updating the avatar's presentation includes: determining a certainty improvement in the pose of the user's first part based on updated pose data representing a change in the pose of the user's first part (e.g., the user's first part moves to a position of the user's first part with increased certainty (e.g., the user's hand moves from a position outside the sensor's field of view (where the camera or other sensor cannot clearly capture the position of the hand) to a position within the sensor's field of view (where the camera can clearly capture the position of the hand)) (e.g., the user's hand moves from behind an object such as a cup to in front of the object)); modifying the current value of a variable display characteristic based on the certainty improvement in the pose of the user's first part (e.g., increasing or decreasing depending on the type of variable display characteristic) (e.g., modifying the variable display characteristic) (e.g., since the position of the user's hand is more clearly captured by the camera in the updated position (within the cup's field of view or in front of the cup), the current value of the cup's variable display characteristic is clearly adjusted (e.g., increasing or decreasing depending on the type of variable display characteristic)).
[0166] In some implementations, updating the avatar's presentation includes: determining a decrease in the certainty of the user's first part's pose based on updated pose data representing a change in the pose of the user's first part (e.g., the user's first part moves to a position where the user's first part's pose is less certain (e.g., the user's hand moves from a position within the sensor's field of view (where the camera or other sensor can clearly capture the hand's position) to a position outside the sensor's field of view (where the camera cannot clearly capture the hand's position)) (e.g., the user's hand moves from in front of an object such as a cup to behind the object)); modifying (e.g., increasing or decreasing depending on the type of variable display characteristic) the current value of the variable display characteristic based on the decrease in the certainty of the user's first part's pose (e.g., modifying the variable display characteristic) (e.g., the user's hand's pose is less certain and the value of the variable display characteristic is adjusted (e.g., increasing or decreasing depending on the type of variable display characteristic) because the user's hand is outside the sensor's field of view or is occluded by a cup in the updated position).
[0167] In some implementations, as the user's first part moves, the determinism of the pose of the user's first part (e.g., the user's hand) changes (e.g., rises or falls), and the value of the variable display characteristic is updated in real time consistent with the change in determinism. In some implementations, the change in the value of the variable display characteristic is represented as a smooth, gradual change in the variable display characteristic applied to the corresponding avatar feature. For example, referring to an implementation where the determinism value changes when the user moves their hand to a location outside the sensor's field of view or moves their hand from a location outside the sensor's field of view, if the variable display characteristic corresponds to particle density, then as the user moves their hand into the sensor's field of view, the particle density including the avatar's hand gradually increases consistent with the increased determinism of the user's hand's location. Conversely, when the user moves from the sensor's field of view to a location outside the field of view, the particle density including the avatar's hand gradually decreases consistent with the decreased determinism of the user's hand's location. Similarly, if the variable display characteristic is a blurring effect, then when the user moves their hand into the sensor's field of view, the amount of blur applied to the avatar's hand gradually decreases consistent with the increased determinism of the user's hand's location. Conversely, as the user moves their hand out of the sensor's field of view, the amount of blur applied to the avatar's hand gradually increases in consistency with the reduced certainty of the user's hand positioning.
[0168] In some implementations, presenting an avatar includes: 1) based on determined pose data satisfying a first set of criteria (the first set of criteria being detected by a first sensor (e.g., 705-1; 705-2; 705-4) in a first part of the user (e.g., the user's hands and / or face) (e.g., the user's hands and / or face being visible to, detected by, or identified by a camera or other sensor)), presenting an avatar with corresponding avatar features having a third value having variable display characteristics (e.g., 721) (e.g., 721-2, 721-3, 721-4, 723-1, 724, 722 in...). Figure 7C (The following are examples of features not shown in shaded areas:) (e.g., corresponding avatar features are represented with higher fidelity because the user's first part is detected by sensors); and 2) the first set of criteria is not met based on the determined pose data (e.g., the user's hand is in...). Figure 8A Outside the field of view 707-2 (e.g., the user's hands and / or face are not visible to the camera or other sensors, and are not detected or recognized by the camera or other sensors), an avatar is presented with a corresponding avatar feature having a fourth value of a variable display feature that indicates a deterministic value lower than a third value of the variable display feature (e.g., the avatar's right hand is displayed as having...). Figure 8CThe smoke effect 830 indicates variable display characteristics (e.g., the corresponding avatar feature is represented with lower fidelity because the user's first part is not detected by the sensor). An avatar presents the corresponding avatar feature with a third or fourth value of variable display characteristics, depending on whether the first sensor detects the user's first part, providing feedback to the user; that is, the determinism of the pose of the corresponding avatar feature is affected by whether the sensor detects the user's first part. Providing improved feedback enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating and / or interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling the user to use the computer system more quickly and effectively.
[0169] In some implementations, when presenting a current value with variable display characteristics (e.g., in the case of Avatar 721-2), Figure 7C When the computer system (e.g., 101) receives the corresponding avatar feature representing a change in the pose of the user's first part (e.g., the user moves their hand), the computer system (e.g., 101) receives updated pose data representing a change in the pose of the user's first part (e.g., the user moves their hand).
[0170] In response to receiving updated pose data, the computer system (e.g., 101) updates the presentation of the avatar. In some embodiments, updating the avatar's presentation includes: determining, based on the updated pose data, the pose of a first portion of the user from a first position within the sensor's field of view (e.g., within the sensor's field of view; visible to the sensor) to a second position outside the sensor's field of view (e.g., the user's right hand from...). Figure 7A Move within the field of view 707-2 Figure 8A The change in the field of view 707-2 (e.g., outside the sensor's field of view; invisible to the sensor) (e.g., when the hand moves from a position within the sensor's (e.g., a camera's) field of view to a position outside the sensor's field of view) reduces the current value of the variable display characteristics (e.g., when the avatar hand 721-2 is outside the sensor's field of view). Figure 7C There is no shading line in it, and it is composed of... Figure 8CThe smoke effect 830 in the image represents (e.g., the decrease in the variable display characteristic corresponds to a decrease in the visual fidelity of the corresponding avatar feature relative to the user's first part) (e.g., the corresponding avatar feature becomes a reduced certainty in the updated pose indicating the user's hand). Based on the determined updated pose data, the current value of the variable display characteristic is reduced to provide feedback to the user, indicating a change in the pose of the user's first part from a first position within the sensor's field of view to a second position outside the sensor's field of view. That is, moving the user's first part from within the sensor's field of view to outside the sensor's field of view causes a change in the certainty of the pose of the corresponding avatar feature (e.g., a decrease in certainty). Providing improved feedback enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating and / or interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling the user to use the computer system more quickly and effectively.
[0171] In some implementations, updating the avatar's presentation includes: determining, based on updated pose data, representing a change in the pose of a first part of the user from a second position to a first position (e.g., moving a hand from a position outside the field of view of a sensor (e.g., a camera) to a position within the sensor's field of view) (e.g., moving part 701-2 from...). Figure 8A The location in the middle is moved to Figure 7A Positioning in the middle), increase the current value of the variable display characteristic (e.g., in the position). Figure 7C The avatar portion 721-2 is shown without a shading, while... Figure 8C The middle portion 721-2 is replaced with a smoke effect 830 (e.g., relative to the user's first part, the increased value of the variable display characteristic corresponds to an increased visual fidelity of the corresponding avatar feature) (e.g., the corresponding avatar feature transforms into a display characteristic value indicating a decreased certainty of the updated pose of the user's hand) (e.g., the corresponding avatar feature transforms into a display characteristic value indicating an increased certainty of the updated pose of the user's hand). Feedback is provided to the user based on the current value of the variable display characteristic, which indicates a change in the pose of the user's first part from a second position to a first position, according to the determined updated pose data. That is, moving the user's first part from outside the sensor's field of view into the sensor's field of view causes a change in the certainty of the pose of the corresponding avatar feature (e.g., increased certainty). Providing improved feedback enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating and / or interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling the user to use the computer system more quickly and effectively.
[0172] In some implementations, the current value of the variable display characteristic decreases at a first rate (e.g., if a hand moves to its position outside the sensor's field of view, the decrease occurs at a slow rate (e.g., slower than the detected hand movement)), and the current value of the variable display characteristic increases at a second rate greater than the first rate (e.g., if a hand moves to its position within the sensor's view, the increase occurs at a fast rate (e.g., faster than the rate of decrease)). Decreasing the current value of the variable display characteristic at the first rate and increasing it at a second rate greater than the first rate provides the user with timing feedback regarding the improved certainty in determining the pose of the corresponding avatar feature. Providing improved feedback enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating and / or interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling the user to use the computer system more quickly and effectively.
[0173] In some implementations, reducing the current value of the variable display characteristic includes: reducing the current value of the variable display characteristic at a first rate (e.g., the reduction of the variable display characteristic (e.g., visual fidelity) occurs at a slow rate (e.g., slower than the detected movement of the hand) based on a known location of the second position corresponding to the first part of the user's body (e.g., even if the user's hand is outside the sensor's field of view). In some implementations, the position of the first part of the user's body can be determined, for example, by deriving the position (or approximate position) using other data. For example, the user's hand may be outside the sensor's field of view, but the user's forearm is in the field of view, and therefore, the position of the hand can be determined based on the known location of the forearm. As another example, the user's hand may be located behind the user's back and therefore outside the sensor's field of view, but the position of the hand (at least approximately) is known based on the location of the user's arm. In some implementations, reducing the current value of the variable display characteristic includes: reducing the current value of the variable display characteristic at a second rate greater than the first rate based on determining that the second positioning corresponds to an unknown position of the first part of the user (e.g., when the hand moves to an unknown position outside the sensor's view, the reduction of the variable display characteristic (e.g., visual fidelity) occurs at a faster rate (e.g., a faster rate than when the hand's positioning is known).
[0174] In some embodiments, a first value of the variable display characteristic represents a higher visual fidelity of the corresponding avatar feature relative to the pose of the user's first part than a second value of the variable display characteristic. In some embodiments, presenting the avatar includes: 1) associating the pose of the user's first part with a second deterministic value based on determining that the user's first part corresponds to a subset of physical features (e.g., the user's arm and shoulder). (e.g., since the user's first part corresponds to the user's arm and / or shoulder, the corresponding avatar feature is represented with lower fidelity.) (e.g., the avatar neck and collar area 727 in...) Figure 7C (shown in shaded 725); and 2) associating the pose of the user's first part with a first deterministic value based on the subset of physical features determined to be non-corresponding to the user's first part (e.g., representing the corresponding avatar feature with higher fidelity because the user's first part does not correspond to the user's arm and / or shoulder) (e.g., in Figure 7C The mouth is rendered without shading (724). The pose of the user's first part is associated with a first or second deterministic value depending on whether the user's first part corresponds to a subset of physical features. This saves computational resources (thereby reducing power usage and improving battery life) by abandoning the generation and display of features that are less important for communication purposes, or even when the determinism of their poses is high (e.g., rendering the feature at a lower fidelity). In some embodiments, when the user's arms and shoulders are not within the sensor's field of view or when they are not considered important features for communication, the user's arms and shoulders (via avatar) are shown with variable display characteristics indicating lower fidelity. In some embodiments, features considered important for communication include the eyes, mouth, and hands.
[0175] In some implementations, while an avatar (e.g., 721) is presented as a corresponding avatar feature with a first value having variable display characteristics, a computer system (e.g., 101) updates the presentation of the avatar, including: 1) presenting an avatar with a corresponding avatar feature having a first change value having variable display characteristics (e.g., a decrease relative to the current value (e.g., the first value) of the variable display characteristics) based on determining that the movement speed of the user's first part is a first movement speed of the user's first part (e.g., the first movement speed); and 2) presenting an avatar with a corresponding avatar feature having a second change value having variable display characteristics (e.g., an increase relative to the current value (e.g., the first value) of the variable display characteristics) based on determining that the movement speed of the user's first part is a second movement speed of the user's first part that is different from the first movement speed (e.g., less than the first movement speed) based on determining that the movement speed of the user's first part is a second movement speed of the user's first part that is different from the first movement speed (e.g., less than the first movement speed) (e.g., the second ... (e.g., the second movement speed) (e.g., the second movement speed) based on increasing the second change value (e.g., the second change value) of the variable display characteristics (e.g., the second change value) based on increasing the second change value (e.g., the second change value) of the variable display characteristics. When the user's first part moves at a first speed, an avatar with a corresponding avatar feature having a first change value with variable display characteristics is presented; and when the user's first part moves at a second speed, an avatar with a corresponding avatar feature having a second change value with variable display characteristics is presented. This provides the user with deterministic feedback indicating how changes in the user's first part's movement speed affect the pose of the corresponding avatar feature. Providing improved feedback enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating and / or interacting with the computer system). This, in turn, reduces power consumption and extends the computer system's battery life by enabling users to use the computer system more quickly and effectively.
[0176] In some implementations, as the movement speed of the user's first part increases, the certainty of the user's first part's positioning decreases (e.g., due to sensor limitations—e.g., the camera's frame rate), causing a corresponding change in the value of the variable display characteristic. For example, if the variable display characteristic corresponds to the density of particles including the corresponding avatar feature, then as the certainty value decreases, the particle density decreases to indicate lower certainty of the user's first part's pose. As another example, if the variable display characteristic is a blurring effect applied to the corresponding avatar feature, then as the certainty value decreases, the blurring effect increases to indicate lower certainty of the user's first part's pose. In some implementations, as the movement speed of the user's first part decreases, the certainty of the user's first part's positioning increases, causing a corresponding change in the value of the variable display characteristic. For example, if the variable display characteristic corresponds to the density of particles including the corresponding avatar feature, then as the certainty value increases, the particle density increases to indicate higher certainty of the user's first part's pose. As another example, if the variable display characteristic is a blur effect applied to the corresponding avatar feature, then as the deterministic value increases, the blur effect decreases to indicate a higher degree of determinism in the pose of the user's first part.
[0177] In some implementations, the computer system (e.g., 101) changes the value of a variable display characteristic (e.g., 725; 830), including changing one or more visual parameters of the corresponding avatar characteristic (e.g., regarding...). Figure 7C and Figure 8C The values of the variable display characteristics 725 and / or visual effects 830 are described in the avatar 721. When the values of the variable display characteristics are changed, one or more visual parameters of the corresponding avatar feature are changed to provide feedback to the user indicating a change in the confidence level of the pose of the first part of the user represented by the corresponding avatar feature. Providing improved feedback enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating and / or interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling the user to use the computer system more quickly and effectively.
[0178] In some implementations, changing the value of a variable display characteristic corresponds to changing the value of one or more visual parameters (e.g., increasing the value of a variable display characteristic may correspond to increasing the amount of blur, pixelation, and / or color of the corresponding avatar feature). Therefore, the methods provided herein for changing one or more visual parameters described herein can be applied to modify, change, and / or adjust variable display characteristics.
[0179] In some implementations, one or more visual parameters include ambiguity (e.g., sharpness). In some implementations, the ambiguity or sharpness of a corresponding avatar feature is varied to indicate the certainty of an increase or decrease in the pose of the user's first part. For example, increasing ambiguity (decreasing sharpness) indicates the certainty of a decrease in the pose of the user's first part, and decreasing ambiguity (increasing sharpness) indicates the certainty of an increase in the pose of the user's first part. For example, in Figure 7C In the diagram, an increased density of the shading line 725 indicates greater blurriness, while a decreased shading density indicates less blurriness. Therefore, the avatar 721 is displayed with a blurred waist and a less blurred (sharper) chest.
[0180] In some implementations, one or more visual parameters include opacity (e.g., or transparency). In some implementations, the opacity or transparency of a corresponding avatar feature is varied to indicate the certainty of an increase or decrease in the pose of the user's first part. For example, increasing opacity (decreasing transparency) indicates the certainty of an increase in the pose of the user's first part, and decreasing opacity (increasing transparency) indicates the certainty of a decrease in the pose of the user's first part. For example, in Figure 7C In the diagram, an increased density of the shading line 725 indicates greater transparency (less opacity), and a decreased shading density indicates less transparency (greater opacity). Therefore, the avatar 721 is displayed with a more transparent waist and a less transparent (less opaque) chest.
[0181] In some implementations, one or more visual parameters include color. In some implementations, the color of a corresponding avatar feature is changed to indicate the certainty of an increase or decrease in the pose of the user's first part, as discussed above. For example, a skin-toned color may be presented to indicate the certainty of an increase in the pose of the user's first part, and a non-skin-toned color (e.g., green, blue) may be presented to indicate the certainty of a decrease in the pose of the user's first part. For example, in Figure 7C In the image, areas of the incarnation 721 that do not have a shading line 725 (such as parts 721-2) are displayed in skin-toned colors (e.g., brown, black, tan, etc.), while areas with a shading line 725 (such as the neck and collar areas 727) are displayed in non-skin-toned colors.
[0182] In some implementations, one or more visual parameters include the density of particles containing the corresponding avatar features, as discussed above regarding the various variable display characteristics of avatar 721.
[0183] In some embodiments, the density of particles containing the corresponding avatar feature includes the spacing between particles containing the corresponding avatar feature (e.g., average distance). In some embodiments, increasing the particle density includes decreasing the spacing between particles containing the corresponding avatar feature, and decreasing the particle density includes increasing the spacing between particles containing the corresponding avatar feature. In some embodiments, the particle density is changed to indicate the certainty of an increase or decrease in the pose of a user's first part. For example, increasing the density (decreasing the spacing between particles) indicates the certainty of an increase in the pose of the user's first part, and decreasing the density (increasing the spacing between particles) indicates the certainty of a decrease in the pose of the user's first part. For example, in Figure 7C In the diagram, an increased density of the shading line 725 indicates a larger particle spacing, while a decreased shading density indicates a smaller particle spacing. Therefore, in this example, the avatar 721 is shown as having a high particle spacing at the waist and a low particle spacing at the chest.
[0184] In some embodiments, the density of particles containing the corresponding avatar feature includes the size of the particles containing the corresponding avatar feature. In some embodiments, increasing the particle density includes decreasing the size of the particles containing the corresponding avatar feature (e.g., producing a less pixelated appearance), and decreasing the particle density includes increasing the size of the particles containing the corresponding avatar feature (e.g., producing a more pixelated appearance). In some embodiments, the particle density is varied to indicate an increase or decrease in the certainty of the pose of the user's first part. For example, increasing the density indicates an increase in the certainty of the pose of the user's first part by decreasing the particle size to provide greater detail and / or resolution of the corresponding avatar feature. Similarly, decreasing the density indicates a decrease in the certainty of the pose of the user's first part by increasing the particle size to provide less detail and / or resolution of the corresponding avatar feature. In some embodiments, the density can be a combination of particle size and spacing, and these factors can be adjusted to indicate more or less certainty in the pose of the user's first part. For example, smaller, spaced-out particles may indicate a decreased density (and decreased certainty) compared to larger, more closely spaced particles (indicating higher certainty). For example, in Figure 7C In the diagram, an increased density of the shading line 725 indicates a larger particle size, while a decreased shading density indicates a smaller particle size. Therefore, in this example, the avatar 721 is displayed as having a larger particle size at the waist (where the waist is highly pixelated) and a smaller particle size at the chest (where the chest is less pixelated than the waist).
[0185] In some implementations, the computer system (e.g., 101) changes the value of a variable display characteristic, including presenting a visual effect associated with a corresponding avatar feature (e.g., 830) (e.g., introducing a display of a visual effect). Presenting a visual effect associated with a corresponding avatar feature when the value of the variable display characteristic is changed provides feedback to the user that the confidence level of the pose of a first part of the user represented by the corresponding avatar feature is below a confidence threshold level. Providing improved feedback enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating and / or interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling the user to use the computer system more quickly and effectively. In some implementations, a visual effect is displayed when the user's corresponding part is at or below a threshold certainty value, and not displayed when it is above the threshold certainty value. For example, a visual effect is displayed for the corresponding avatar feature when the certainty of the pose of the user's corresponding part is below a 10% certainty value. For example, when the user's elbow is outside the camera's field of view, the visual effect is a smoke effect or fish scales displayed at the avatar's elbow.
[0186] In some embodiments, the user's first part includes a first physical feature (e.g., the user's mouth) and a second physical feature (e.g., the user's ears or nose). In some embodiments, a first value of the variable display characteristic represents a higher visual fidelity of the corresponding avatar feature relative to the pose of the user's first physical feature than a second value of the variable display characteristic. In some embodiments, presenting an avatar with a corresponding avatar feature having a second value of the variable display characteristic includes utilizing the user's corresponding physical feature (e.g., ...). Figure 7C The incarnation of the mouth 724 and Figure 7A The rendering of the first physical feature (which is the same as the user's mouth) and the rendering based on the corresponding physical feature that is not the user (e.g., the mouth of the user in the image ...). Figure 7C The incarnation of the nose 722 and Figure 7AThe avatar is rendered using a second physical feature (different from the user's nose) as shown in the image (e.g., 721). This includes, for example, the user's mouth (e.g., rendered using a video feed of the user's mouth) and ears (or nose) that are not the user's (e.g., ears (or nose) from another person or simulated based on a machine learning algorithm). Utilizing the rendering of a first physical feature based on the user's corresponding physical feature and a second physical feature based on a corresponding physical feature that is not the user's to save computational resources (thereby reducing power consumption and improving battery life) by forgoing the high-fidelity generation and display of these features (e.g., rendering some features based on physical features that are not the user's) when these features are less important for communication purposes, or even when the determinism of the pose of these features is high.
[0187] In some implementations, when an avatar is presented with variable display characteristics indicating low fidelity (e.g., low determinism), physical features of the user deemed important for communication (e.g., eyes, mouth, hands) are rendered based on the user's actual appearance, while other features are rendered based on something other than the user's actual appearance. That is, video feeds, for example, using features, preserve communication-important physical features with high fidelity, and physical features deemed unimportant for communication (e.g., skeletal structure, hairstyle, ears, etc.) are rendered, for example, using similar features selected from a database of avatar features or using computer-generated features derived from machine learning algorithms. By rendering the avatar with variable display characteristics indicating low fidelity (e.g., low determinism) in a manner that prioritizes the appearance of the user's corresponding physical features over important communication features, and renders the avatar with variable display characteristics indicating low fidelity based on different appearances, computational resources are conserved while maintaining the ability to communicate effectively with the user represented by the avatar. For example, computational resources are conserved by rendering unimportant features with variable display characteristics indicating low fidelity. However, if important communication features are rendered in this way, communication is suppressed because these features do not match the user's corresponding physical characteristics, which can be distracting and even confuse the identity of the user represented by the avatar. Therefore, to ensure that communication is not suppressed, important communication features are rendered with high fidelity, thus preserving the ability to communicate effectively while still conserving computational resources.
[0188] In some implementations, pose data is generated from multiple sensors (e.g., 705-1; 705-2; 705-3; 705-4; 810).
[0189] In some embodiments, the multiple sensors include one or more camera sensors (e.g., 705-1; 705-2; 705-3; 705-4) associated with (e.g., constituting) a computer system (e.g., 101). In some embodiments, the computer system includes one or more cameras configured to capture pose data. For example, the computer system is a headset device with a camera positioned to detect the poses of one or more users in the physical environment surrounding the headset.
[0190] In some embodiments, the multiple sensors include one or more camera sensors separate from the computer system. In some embodiments, pose data is generated using one or more cameras separate from the computer system and configured to capture pose data representing a user's pose. For example, the cameras separate from the computer system may include smartphone cameras, desktop computer cameras, and / or cameras from a headset device worn by the user. Each of these cameras may be configured to generate at least a portion of the pose data by capturing the user's pose. In some embodiments, pose data is generated from multiple camera sources providing the positioning or pose of the user (or parts of the user) from different angles and viewpoints.
[0191] In some implementations, the multiple sensors include one or more non-visual sensors (e.g., 810) (e.g., sensors excluding camera sensors). In some implementations, pose data is generated using one or more non-visual sensors, such as proximity sensors or sensors in a smartwatch (e.g., accelerometers, gyroscopes, etc.).
[0192] In some implementations, an interpolation function is used to determine at least a first part of the user's pose (e.g., interpolation). Figure 7C (The pose of the avatar of the left hand in parts 721-4). In some embodiments, the first part of the user is obscured by an object (e.g., a cup 702). For example, the user holds a coffee cup such that the tips of the user's fingers are visible to the sensor capturing pose data, but the proximal regions of the hand and fingers are behind the cup, making them hidden from the sensor capturing pose data. In this example, the computer system can perform an interpolation function based on the detected fingertips, the user's wrist, or a combination thereof to determine the approximate pose of the user's fingers and the back of the hand. This information can be used to improve the determinism of the pose of these features. For example, these features can be represented by variable display characteristics that indicate a very high degree of determinism (e.g., 99%, 95%, 90%) in the pose of these features.
[0193] In some embodiments, pose data includes data generated from previous scan data that captures information about the appearance of the user of the computer system (e.g., 101) (e.g., data from previous body or frontal scans (e.g., depth data)). In some embodiments, the scan data includes data generated from scans of the user's body or face (e.g., depth data). For example, this could be data derived from a facial scan used to unlock or access a device (e.g., a smartphone, smartwatch, or HMD), or data derived from a media library including photos and videos containing depth data. In some embodiments, pose data generated from previous scan data can be used to supplement pose data to increase the understanding of the current pose of a portion of the user (e.g., pose based on known portions of the user). For example, if the user's forearm is of known length (e.g., 28 inches), and therefore, the possible position of the hand can be determined based on the position and angle of a portion of the elbow and forearm. In some embodiments, pose data generated from previous scan data can be used to improve the visual fidelity of the reproduction of portions of the user that are not visible to the computer system's sensors (e.g., by providing eye color, chin shape, ear shape, etc.). For example, if a user wears a hat that covers their ears, previous scan data can be used to improve the certainty of the user's ear pose (e.g., positioning and / or orientation) based on previous scans that can be used to derive pose data of the user's ears. In some embodiments, the device generating the reproduction of the user's part is separate from the device (e.g., smartphone, smartwatch, HMD), and data from the facial scan is provided (e.g., securely and privately, with one or more options for the user to decide whether to share data between devices) to the device generating the reproduction of the user's part for the purpose of constructing the reproduction of the user's part. In some embodiments, the device generating the reproduction of the user's part is the same as the device (e.g., smartphone, smartwatch, HMD), and data from the facial scan is provided to the device generating the reproduction of the user's part for the purpose of constructing the reproduction of the user's part (e.g., the facial scan is used to unlock the HMD, which also generates the reproduction of the user's part).
[0194] In some embodiments, pose data includes data generated from prior media data (e.g., photos and / or videos of the user). In some embodiments, prior media data includes image data and optionally includes depth data associated with previously captured photos and / or videos of the user. In some embodiments, prior media data can be used to supplement pose data to improve the certainty of the current pose of a portion of the user. For example, if glasses worn by the user obstruct their eyes from the sensors capturing pose data, prior media data can be used to improve the certainty of the user's eye pose (e.g., positioning and / or orientation) based on pre-existing photos and / or videos that are accessible (e.g., securely and privately, utilizing one or more options for the user to decide whether to share the data between devices) to derive the user's eye pose data.
[0195] In some embodiments, the pose data includes video data (e.g., a video feed) that includes at least a first portion of the user (e.g., 701-1). In some embodiments, presenting the avatar includes presenting a modeled avatar (e.g., 721) (e.g., a three-dimensional computer-generated model of the avatar) that includes corresponding avatar features rendered using video data including the user's first portion (e.g., avatar mouth 724; right avatar eye 723-1). In some embodiments, the modeled avatar is a three-dimensional computer-generated avatar, and the corresponding avatar features are rendered as a video feed mapped onto the user's first portion of the modeled avatar. For example, the avatar could be a three-dimensional simulated avatar model (e.g., a green alien) including eyes shown in a video feed as the user's eyes.
[0196] In some implementations, presenting an avatar includes: 1) presenting an avatar having a corresponding avatar feature and a first amount (e.g., percentage; quantity) of avatar features other than the corresponding avatar feature, based on an input that indicates a first rendering value (e.g., a low rendering value) of the avatar, received (e.g., before, simultaneously with, or after the avatar's presentation); and 2) presenting an avatar having a second amount (e.g., percentage; quantity) of avatar features other than the corresponding avatar feature, based on an input that indicates a second rendering value (e.g., a high rendering value) of the avatar, different from the first rendering value, received (e.g., before, simultaneously with, or after the avatar's presentation), wherein the second amount is different from (e.g., greater than) the first amount. The second rendering value indicating the avatar is different from the first rendering value, and the presentation of an avatar having a corresponding avatar feature and a second amount different from the first amount of avatar features other than the corresponding avatar feature, saves computational resources (thereby reducing power usage and improving battery life) by abandoning the generation and display of these features when they are not needed, or even when the determinism of the pose of these features is high.
[0197] In some implementations, the render value indicates user preference and is used to render the amount of the avatar that does not correspond to the user's first part. In other words, the render value is used to select the amount of the avatar (in addition to the corresponding avatar feature) to display. For example, the render value can be selected from a sliding scale, where at one end of the scale, no part of the avatar is rendered except for the portion corresponding to the user's tracked feature (e.g., the corresponding avatar feature), and at the other end of the scale, the entire avatar is rendered. For example, if the user's first part is the user's hand, and the lowest render value is selected, the avatar will appear as a floating hand (the corresponding avatar feature would be the hand corresponding to the user's hand). Selecting the render value allows users to customize the appearance of the avatar to increase or decrease the contrast between the parts of the avatar that correspond to the tracked user feature and the parts that do not. This allows users to more easily identify which parts of the avatar can be trusted to be more realistic.
[0198] In some implementations, a computer system (e.g., 101) causes the presentation (e.g., 735) of a representation (e.g., 735) of a user associated with the display generation component (e.g., the user described above; a second user, wherein the aforementioned user is not associated with the display generation component) via a display generation component (e.g., 730), wherein the representation of the user associated with the display generation component corresponds to the appearance of the user associated with the display generation component, which is presented to one or more users other than the user associated with the display generation component. The presentation of the representation of the user associated with the display generation component, corresponding to the appearance of the user associated with the display generation component presented to one or more users, provides feedback to the user associated with the display generation component regarding the user appearance viewed by other users. Providing improved feedback enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating and / or interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling users to use the system more quickly and efficiently.
[0199] In some implementations, the user viewing the display-generated component is also presented to another user (e.g., on a different display-generated component). For example, two users communicate with each other as two different avatars presented in a virtual environment, and each user is able to view the other user on their own corresponding display-generated component. In this example, the display-generated component of one user (e.g., the second user) shows the appearance of the other user (e.g., the first user) as well as its own appearance (e.g., the second user) because it is presented to the other user (e.g., the first user) in the virtual environment.
[0200] In some embodiments, presenting an avatar includes presenting an avatar with corresponding avatar features having the first appearance based on a first appearance of a user's first part (e.g., the user's chin is cleanly shaved, and the corresponding avatar feature represents the user's chin having the appearance of a cleanly shaved avatar). In some embodiments, a computer system (e.g., 101) receives data indicating an updated appearance of the user's first part (e.g., receiving a recent photograph or video showing the user's chin now including a beard) (e.g., receiving recent facial scan data indicating the user's chin now including a beard), and causes an avatar with corresponding avatar features having the updated appearance based on the updated appearance of the user's first part (e.g., 721) via a display generation component (e.g., 730) (e.g., the avatar's chin now includes a beard). Presenting an avatar with corresponding avatar features having the updated appearance based on the updated appearance of the user's first part provides an accurate representation of the user's first part via the current representation of the corresponding avatar features, without having to manually manipulate the avatar or perform a registration process to incorporate updates to the user's first part. This provides an improved control scheme for editing or presenting avatars, which may require less input to generate a customized appearance for the avatar compared to using different control schemes (e.g., control schemes that require manipulating individual control points to build or modify the avatar). Reducing the amount of input required to perform the task enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating and / or interacting with the computer system), which in turn reduces power consumption and extends the system's battery life by enabling users to use the computer system more quickly and efficiently.
[0201] In some implementations, the avatar's appearance is updated based on recently captured image or scan data to match the user's updated appearance (e.g., new hairstyle, beard, etc.). In some implementations, data indicating the user's updated appearance (partially) is captured in a separate operation (e.g., when the device is unlocked). In some implementations, data is collected on a device separate from the computer system. For example, data is collected when a user unlocks a device such as a smartphone, tablet, or other computer, and the data is transferred from the user's device (e.g., securely and privately, utilizing one or more options for the user to decide whether to share data between devices) to the computer system for later use. Such data is transmitted and stored in a manner consistent with industry standards for fixing personally identifiable information.
[0202] In some implementations, the pose data further represents an object associated with the user's first part (e.g., 702; 805) (e.g., the user holding a cup; the user leaning against a wall; the user sitting in a chair). In some implementations, presenting an avatar includes presenting an object having characteristics associated with the corresponding avatar (e.g., Figure 8C An avatar representing a representation of an adjacent (e.g., overlapping, in front, touching, interacting) object (825 in the original text). Presenting an avatar representing an object adjacent to the corresponding avatar provides the user with contextual feedback about the user's pose in the first part. Providing improved feedback enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating and / or interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively.
[0203] In some implementations, the avatar is modified (e.g., in...). Figure 8C In this embodiment, avatar 721 also includes wall rendering 825 to include representations of objects interacting with the user. This improves communication by providing context of the avatar's pose, which can be based on the user's pose. For example, if the user is holding a cup, the corresponding avatar features (e.g., the avatar's arm) are modified to include a representation of the cup in the avatar's hand. As another example, if the user is leaning against a wall, the corresponding avatar features (e.g., the avatar's shoulder and / or arm) are modified to include a representation of at least a portion of the wall. As yet another example, if the user is sitting in a chair or at a desk, the corresponding avatar features (e.g., the avatar's legs and / or arm) are modified to include a representation of at least a portion of the chair and / or desk. In some embodiments, the objects interacting with the user are virtual objects. In some embodiments, the objects interacting with the user can be real objects. In some embodiments, the representation of the object is presented as a virtual object. In some embodiments, a video feed showing the object is used to present the object's representation.
[0204] In some implementations, the pose of the user's first part is associated with a fifth deterministic value. In some implementations, presenting the avatar (e.g., 721) includes, based on determining that the user's first part is a first feature type (e.g., 701-1 includes the user's neck and collar area) (e.g., a feature type considered unimportant to communication), presenting an avatar with a value having a variable display characteristic indicating a deterministic value less than the fifth deterministic value (e.g., avatar neck and collar area 727 in...). Figure 7C(Shown in shaded 725) (For example, even if the pose of the user's first part is highly deterministic, the corresponding avatar feature is presented as a variable display characteristic with a value indicating the pose of the user's first part). When the user's first part is a first feature type, the avatar of the corresponding avatar feature is presented with a variable display characteristic indicating a deterministic value less than a fifth deterministic value. This saves computational resources (thereby reducing power consumption and increasing battery life) by abandoning the generation and display of features with high fidelity (e.g., rendering the feature with lower fidelity) when these features are not important for communication purposes, or even when the pose of these features is highly deterministic.
[0205] In some implementations, when the pose of a user feature is highly deterministic, but the feature is not considered important for communication, variable display characteristics indicating low determinism (e.g., the avatar's neck and collar areas 727) are utilized. Figure 7C The corresponding avatar feature is rendered with a shaded line 725 indicating variable display characteristics (or at least a lower degree of certainty than actually associated with the user feature). For example, a camera sensor may capture the location of a user's knees (and therefore their location has high certainty), but because knees are generally not considered important for communication purposes, the avatar's knee is rendered with variable display characteristics indicating lower certainty (e.g., lower fidelity). However, if the user feature (e.g., the user's knee) becomes important for communication, the value of the variable display characteristics (e.g., fidelity) of the corresponding avatar feature (e.g., the avatar's knee) is adjusted (e.g., increased) to accurately reflect the certainty of the user feature's pose (e.g., the knee). In some embodiments, rendering the avatar with the variable display characteristics indicating lower certainty when the feature is not considered important for communication saves computational resources typically consumed when rendering the corresponding avatar feature with the variable display characteristics indicating high certainty (e.g., a high-fidelity representation of the corresponding avatar feature). Because user characteristics are not considered important for communication, computational resources can be saved without sacrificing communication efficiency.
[0206] In some implementations, the pose of the user's first part is associated with a sixth deterministic value. In some implementations, the avatar presentation (e.g., 721) also includes determining whether the user's first part is a second feature type (e.g., Figure 7A The user's left eye (e.g., feature types considered unimportant to communication (e.g., hand, mouth, eye)) presents an avatar with a value that has a variable display characteristic indicating a deterministic value greater than a sixth deterministic value (e.g., left avatar eye 723-2 in) that has a corresponding avatar feature. Figure 7C(The image is not shown with a shading) (For example, even if the pose of the user's first part is not highly deterministic, the corresponding avatar feature is presented as a variable display characteristic with a value indicating the pose of the user's first part). When the user's first part is of the second feature type, the avatar is presented with a variable display characteristic with a deterministic value indicating a value greater than the sixth deterministic value. This enhances communication using the computer system by reducing distractions caused by rendering avatar features (especially those important for communication purposes) at relatively low fidelity. This enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating and / or interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling users to use the computer system more quickly and effectively.
[0207] In some implementations, when the pose of a user feature does not have high determinism, but the feature is considered important for communication, the corresponding avatar feature is rendered using a variable display characteristic with a value indicating extremely high determinism (or at least higher determinism than actually associated with the user feature). Figure 7C In the example, the avatar's left eye (723-2) is rendered without a shading line. For example, the user's mouth is considered important for communication. Therefore, even if there might be 75%, 80%, 85%, or 90% certainty regarding the pose of the user's mouth, the corresponding avatar features are rendered using variable display characteristics that indicate high certainty (e.g., 100%, 97%, or 95% certainty) (e.g., high fidelity). In some embodiments, features considered important for communication are rendered using variable display characteristics with values indicating higher certainty (e.g., a high-fidelity representation of the corresponding avatar feature), even with lower certainty, because it will (e.g., by distracting the user) suppress communication. The corresponding avatar features are rendered using variable display characteristics with values indicating an accurate level of certainty.
[0208] In some embodiments, extremely high certainty (or very high certainty) may refer to a certainty amount above a predetermined (e.g., first) certainty threshold (such as 90%, 95%, 97%, or 99% certainty). In some embodiments, high certainty (or high certainty) may refer to a certainty amount above a predetermined (e.g., second) certainty threshold (such as 75%, 80%, or 85% certainty). In some embodiments, high certainty is optionally lower than extremely high certainty but higher than relatively high certainty. In some embodiments, relatively high certainty (or relatively high certainty) may refer to a certainty amount above a predetermined (e.g., third) certainty threshold (such as 60%, 65%, or 70% certainty). In some embodiments, relatively high certainty is optionally lower than high certainty and / or extremely high certainty but higher than low certainty. In some embodiments, relatively low certainty (or less certainty) may refer to a amount of certainty below a predetermined (e.g., fourth) certainty threshold (such as 40%, 45%, 50%, or 55% certainty). In some embodiments, relatively low certainty is lower than relatively high certainty, high certainty, and / or very high certainty, but optionally higher than low certainty. In some embodiments, low certainty (or less certainty or less certainty) may refer to a amount of certainty below a predetermined certainty threshold (such as 45%, 50%, 55%, 60%, 65%, 70%, 75%, or 80% certainty). In some embodiments, low certainty is lower than relatively low certainty, relatively high certainty, high certainty, and / or very high certainty, but optionally higher than very low certainty. In some embodiments, very low certainty (or extremely low or very little certainty) can refer to a amount of certainty below a predetermined (e.g., fifth) certainty threshold (such as 5%, 15%, 20%, 30%, 40%, or 45% certainty). In some embodiments, very low certainty is lower than low certainty, relatively low certainty, relatively high certainty, high certainty, and / or very high certainty.
[0209] Figures 10A to 10B , Figures 11A to 11B and Figure 12 Examples of users being represented as virtual avatars in a CGR environment based on appearances with different appearance templates are depicted. In some implementations, the avatar has an appearance based on a character template or an abstract template. In some implementations, the avatar changes its appearance by changing its pose. In some implementations, the avatar changes its appearance by transforming between a character template and an abstract template. As described above, using a computer system (e.g., Figure 1 The computer system 101 in the document implements the process disclosed herein.
[0210] Figure 10A Four different users are shown in a real-world environment. User A is in real-world environment 1000-1, user B is in real-world environment 1000-2, user C is in real-world environment 1000-3, and user D is in real-world environment 1000-4. In some embodiments, real-world environments 1000-1 to 1000-4 are different real-world environments (e.g., different locations). In some embodiments, one or more of real-world environments 1000-1 to 1000-4 are the same environment. For example, real-world environment 1000-1 may be the same real-world environment as real-world environment 1000-4. In some embodiments, one or more of real-world environments 1000-1 to 1000-4 may be different locations within the same environment. For example, real-world environment 1000-1 may be a first location in a room, and real-world environment 1000-4 may be a second location in the same room.
[0211] In some embodiments, the real-world environments 1000-1 to 1000-4 are motion capture studios similar to real-world environment 700 and include cameras 1005-1 to 1005-5, similar to cameras 705-1 to 705-4, for capturing data (e.g., image data and / or depth data) that can be used to determine the pose of one or more parts of a user. The computer system uses the determined pose of one or more parts of the user to display an avatar with a specific pose and / or appearance template. In some embodiments, the computer system presents the user in the CGR environment as an avatar of an appearance template determined based on different criteria. For example, in some embodiments, the computer system (e.g., based on the pose of one or more parts of the user) determines the type of activity performed by the respective user, and the computer system presents the user in the CGR environment as an avatar with an appearance template determined based on the type of activity performed by the user. In some embodiments, the computer system presents the user in the CGR environment as an avatar with an appearance template determined based on the user's focal position (e.g., eye gaze positioning) of the computer system. For example, when a user is focusing on (e.g., viewing) the first avatar rather than the second avatar, the first avatar is presented with a first appearance template (e.g., a character template), and the second avatar is presented with a second appearance template (e.g., an abstract template). When the user moves their focus to the second avatar, the second avatar is presented with the first appearance template (the second avatar changes from having the second appearance template to having the first appearance template), and the first avatar is presented with the second appearance template (the first avatar changes from having the first appearance template to having the second appearance template).
[0212] exist Figure 10AIn the real-world environment 1000-1, user A is depicted as having an arm raised and speaking. While user A is conversing, cameras 1005-1 and 1005-2 capture the pose of a portion of user A in their respective fields of view 1007-1 and 1007-2. In the real-world environment 1000-2, user B is depicted as reaching for object 1006 on a table. While user B is reaching for object 1006, camera 1005-3 captures the pose of a portion of user B in field of view 1007-3. In the real-world environment 1000-3, user C is depicted as reading a book 1008. While user C is reading, camera 1005-4 captures the pose of a portion of user C in field of view 1007-4. In the real-world environment 1000-4, user D is depicted as having an arm raised and speaking. While user D is conversing, camera 1005-5 captures a portion of user D's pose within field of view 1007-5. In some embodiments, users A and D converse with each other. In some embodiments, users A and D converse with a third party, such as a user of a computer system.
[0213] Cameras 1005-1 to 1005-5 are similar to cameras 705-1 to 705-4 as described above. Therefore, cameras 1005-1 to 1005-5 are similar to those mentioned above. Figures 7A to 7C , Figures 8A to 8C and Figure 9 The method described captures partial poses of the corresponding users A, B, C, and D. For the sake of brevity, details will not be repeated below.
[0214] Cameras 1005-1 to 1005-5 are described as non-limiting examples of devices for capturing the pose of portions A, B, C, and D of a user. Therefore, other sensors and / or devices may be used, in addition to or in place of any of cameras 1005-1 to 1005-5, to capture the pose of a portion of a user. (The above text is about...) Figures 7A to 7C , Figures 8A to 8C and Figure 9 Examples of other such sensors and / or devices are described below. For the sake of brevity, the details of these examples will not be repeated below.
[0215] Now for reference Figure 10B The computer system receives sensor data generated using cameras 1005-1 to 1005-5 and displays avatars 1010-1 to 1010-4 in the CGR environment 1020 via display generation unit 1030. In some embodiments, display generation unit 1030 is similar to display generation unit 120 and display generation unit 730.
[0216] Incarnations 1010-1, 1010-2, 1010-3, and 1010-4 represent users A, B, C, and D in CGR environment 1020, respectively. Figure 10B In this embodiment, avatars 1010-1 and 1010-4 are displayed with an appearance based on a character template, while avatars 1010-2 and 1010-3 are displayed with an appearance based on an abstract template. In some embodiments, the character template includes expressive features, such as an avatar's face, arms, hands, or other avatar features corresponding to human body parts, while the abstract template does not include such features or such features are indistinguishable. In some embodiments, avatars with an appearance based on an abstract template have irregular shapes.
[0217] In some implementations, the computer system determines the appearance of avatars 1010-1 to 1010-4 based on the poses of the corresponding users A, B, C, and D. For example, in some implementations, if the computer system determines (e.g., based on the user's pose) that the user is performing a first type of activity, the computer system renders a corresponding avatar with an appearance based on a role template. Conversely, if the computer system determines that the user is performing a second (different) type of activity, the computer system renders a corresponding avatar with an appearance based on an abstract template.
[0218] In some implementations, the first type of activity is an interactive activity (e.g., an activity involving interaction with other users), and the second type of activity is a non-interactive activity (e.g., an activity not involving interaction with other users). In some implementations, the first type of activity is an activity performed at a specific location (e.g., the same location as the user of the computer system), and the second type of activity is an activity performed at a different location (e.g., a location away from the user of the computer system). In some implementations, the first type of activity is a manual activity (e.g., an activity involving the user's hands, such as touching, holding, moving, or using an object), and the second type of activity is a non-manual activity (e.g., an activity that does not typically involve the user's hands).
[0219] exist Figure 10A and Figure 10B In the implementation described herein, the computer system represents user A in CGR environment 1020 as avatar 1010-1, which has an appearance based on a role template. In some implementations, the computer system responds to detection Figure 10AThe computer system presents an avatar 1010-1 with an appearance based on a persona template, based on the pose of user A. For example, in some embodiments, the computer system determines that user A is having a conversation based on pose data captured using camera 1005-1 and optionally camera 1005-2. In some embodiments, the computer system considers the conversation to be an interactive activity type and therefore presents an avatar 1010-1 with an appearance based on a persona template. As another example, in some embodiments, the computer system determines that user A is in the same location as the user of the computer system based on pose data (e.g., the positional components of the pose data) and therefore presents an avatar 1010-1 with an appearance based on a persona template. In some embodiments, the computer system presents an avatar 1010-1 with an appearance based on a persona template because the computer system determines that the user of the computer system is focusing on (e.g., viewing) avatar 1010-1.
[0220] exist Figure 10A and Figure 10B In the implementation described herein, the computer system represents user B in CGR environment 1020 as avatar 1010-2, which has an appearance based on an abstract template. In some implementations, the computer system responds to detection Figure 10A The computer system presents an avatar 1010-2 with an appearance based on an abstract template, based on the pose of user B. For example, in some embodiments, the computer system determines that user B is reaching for object 1006 based on pose data captured using cameras 1005-3. In some embodiments, the computer system considers this a non-interactive type of activity (e.g., user B is doing something other than interacting with another user), and therefore presents an avatar 1010-2 with an appearance based on an abstract template. In some embodiments, the computer system considers this a non-manual type of activity (e.g., because user B currently has nothing in their hand), and therefore presents an avatar 1010-2 with an appearance based on an abstract template. As another example, in some embodiments, the computer system determines that user B is in a user position away from the computer system based on pose data (e.g., the positional components of the pose data), and therefore presents an avatar 1010-2 with an appearance based on an abstract template. In some implementations, the computer system presents an avatar 1010-2 with an appearance based on an abstract template because the computer system determines that the user of the computer system is not focusing on (e.g., viewing) avatar 1010-2.
[0221] exist Figure 10A and Figure 10BIn the embodiments depicted, the computer system represents user C in CGR environment 1020 as avatar 1010-3, which has an appearance based on an abstract template. In some embodiments, the abstract template can have different abstract appearances. Therefore, although avatars 1010-2 and 1010-3 are both based on an abstract template, the avatars can look different from each other. For example, the abstract template can include different appearances, such as irregular spots or different abstract shapes. In some embodiments, different abstract templates providing different abstract appearances can exist. In some embodiments, the computer system responds to detection Figure 10A The computer system presents an avatar 1010-3 with an appearance based on an abstract template, according to the pose of user C. For example, in some embodiments, the computer system determines that user C is reading a book 1008 based on pose data captured using cameras 1005-4. In some embodiments, the computer system considers reading to be a non-interactive activity and therefore presents an avatar 1010-3 with an appearance based on an abstract template. As another example, in some embodiments, the computer system determines that user C is in a user position away from the computer system based on pose data (e.g., the positional components of the pose data) and therefore presents an avatar 1010-3 with an appearance based on an abstract template. In some embodiments, the computer system presents an avatar 1010-3 with an appearance based on an abstract template because the computer system determines that the user of the computer system is not focusing on (e.g., viewing) avatar 1010-3.
[0222] exist Figure 10A and Figure 10B In the implementation described herein, the computer system represents user D in CGR environment 1020 as avatar 1010-4, which has an appearance based on a role template. In some implementations, the computer system responds to detection Figure 10A The computer system presents an avatar 1010-4 with an appearance based on a persona template, according to the pose of user D. For example, in some embodiments, the computer system determines that user D is conversing based on pose data captured using cameras 1005-5. In some embodiments, the computer system considers the conversation to be an interactive activity type and therefore presents an avatar 1010-4 with an appearance based on a persona template. As another example, in some embodiments, the computer system determines that user D is in the same location as the user of the computer system based on pose data (e.g., the positional components of the pose data) and therefore presents an avatar 1010-4 with an appearance based on a persona template. In some embodiments, the computer system presents an avatar 1010-4 with an appearance based on a persona template because the computer system determines that the user of the computer system is focusing on (e.g., viewing) avatar 1010-4.
[0223] In some implementations, an avatar with an appearance based on a character template is presented as having a pose determined based on the pose of the corresponding user. For example, in Figure 10B In this context, avatar 1010-1 has a pose matching the pose determined for user A, and avatar 1010-4 has a pose matching the pose determined for user D. In some implementations, similar to the above description... Figures 7A to 7C , Figures 8A to 8C and Figure 9 The discussed method determines the pose of an avatar with an appearance based on a character template. For example, avatar 1010-1 is displayed with a pose determined based on the pose of a portion of user A detected in fields of view 1007-1 and 1007-2. Because user A's entire body is within the camera's field of view, the computer system determines the pose of the corresponding portion of avatar 1010-1 with maximum certainty, such as... Figure 10B In the example, avatar 1010-1 is not indicated by the dashed line. As another example, avatar 1010-4 is displayed as having a pose determined based on the pose of a portion of user D detected in the field of view 1007-5 of camera 1005-5. Because user D's legs are outside the field of view 1007-5, the computer system determines the pose of avatar 1010-4's legs with a degree of uncertainty (less than maximum certainty) and presents avatar 1010-4 with the determined pose and variable display characteristics indicated by the dashed line 1015 on the avatar's legs, as shown. Figure 10B As shown. For the sake of brevity, regarding... Figure 10A , Figure 10B , Figure 11A and Figure 11B All implementation schemes discussed will not repeat additional details used to determine the pose of the avatar and display variable display characteristics.
[0224] In some implementations, the computer system updates the avatar's appearance based on changes in the user's pose. In some implementations, the appearance update includes a change in the avatar's pose without transitioning between different appearance templates. For example, in response to detecting a user raising and then lowering their arm, the computer system presents the avatar as an avatar character that changes the user's pose (e.g., in real-time) by raising and then lowering the corresponding avatar arm. In some implementations, the avatar's appearance update includes transitions between different appearance templates based on changes in the user's pose. For example, in response to detecting a user changing from a pose corresponding to a first activity type to a pose corresponding to a second activity type, the computer system presents an avatar that transitions from a first appearance template (e.g., a character template) to a second appearance template (e.g., an abstract template). References below... Figure 11A and Figure 11B Examples of incarnations with updated appearances are discussed in more detail.
[0225] Figure 11A and Figure 11B Similar to Figure 10A and Figure 10B However, there exist users A, B, C, and D with updated poses, and corresponding avatars 1010-1 to 1010-4 with updated appearances. For example, users have already... Figure 10A pose movement in Figure 11A The pose in the image, and in response, the computer system has changed the avatar's appearance from... Figure 10B The appearance shown has been changed to Figure 11B The appearance shown.
[0226] refer to Figure 11A User A is now shown with their right arm lowered and head turned forward, while still conversing. User B is now picking up object 1006, their head facing forward. User C is now waving and speaking, with book 1008 lowered and to user C's side. User D is now looking to one side, their arm lowered.
[0227] exist Figure 11B In this update, the computer system has updated the appearance of avatars 1010-1 through 1010-4. Specifically, avatar 1010-1 retains its character-based template appearance, avatars 1010-2 and 1010-3 have been changed from an abstract template-based appearance to a character template-based appearance, and avatar 1010-4 has been changed from a character template-based appearance to an abstract template-based appearance.
[0228] In some implementations, the computer system responds to the detection Figure 11A The computer system depicts an avatar 1010-1 with an appearance based on a persona template, based on the pose of user A. For example, in some embodiments, the computer system determines that user A is still conversing based on pose data captured using camera 1005-1 and optionally camera 1005-2. In some embodiments, the computer system considers the conversation to be an interactive activity type and therefore presents an avatar 1010-1 with an appearance based on a persona template. As another example, in some embodiments, the computer system determines that user A is in the same location as the user of the computer system based on pose data (e.g., the positional components of the pose data) and therefore presents an avatar 1010-1 with an appearance based on a persona template. In some embodiments, the computer system presents an avatar 1010-1 with an appearance based on a persona template because the computer system determines that the user of the computer system is focusing on (e.g., viewing) avatar 1010-1.
[0229] In some implementations, the computer system responds to the detection Figure 11AThe computer system depicts an avatar 1010-2 with an appearance based on a character template, based on the pose of user B. For example, in some embodiments, the computer system determines that user B is interacting with other users by waving object 1006 based on pose data captured using camera 1005-3, and thus presents an avatar 1010-2 with an appearance based on a character template. In some embodiments, the computer system determines that user B is performing a manual activity by grasping object 1006 based on pose data captured using camera 1005-3, and thus presents an avatar 1010-2 with an appearance based on a character template. In some embodiments, the computer system presents an avatar 1010-2 with an appearance based on a character template because the computer system determines that the user of the computer system is currently focusing on (e.g., viewing) avatar 1010-2.
[0230] In some implementations, the computer system responds to the detection Figure 11A The computer system depicts an avatar 1010-3 with an appearance based on a persona template, based on the pose of user C. For example, in some embodiments, the computer system determines that user C is interacting with other users by waving and / or speaking based on pose data captured using camera 1005-4, and thus presents an avatar 1010-3 with an appearance based on a persona template. In some embodiments, the computer system presents an avatar 1010-3 with an appearance based on a persona template because the computer system determines that the user of the computer system is currently focusing on (e.g., viewing) avatar 1010-3.
[0231] In some implementations, the computer system responds to the detection Figure 11A The computer system depicts an avatar 1010-4 with an appearance based on an abstract template, based on the pose of user D. For example, in some embodiments, the computer system determines that user D is no longer interacting with other users based on pose data captured using cameras 1005-5 because the user's arm is lowered, because he is looking elsewhere, and / or because he is no longer speaking. Therefore, the computer system presents an avatar 1010-4 with an appearance based on an abstract template. In some embodiments, the computer system presents an avatar 1010-4 with an appearance based on an abstract template because the computer system determines that the user of the computer system is no longer focused on (e.g., viewing) avatar 1010-4.
[0232] like Figure 11B As shown, the computer system is based on Figure 11A The system updates the pose of the corresponding avatar based on the detected user pose. For example, the computer system updates the pose of avatar 1010-1 to match... Figure 11AThe pose of user A in the context of the computer system. In some implementations, the user of the computer system is not focusing on avatar 1010-1, but avatar 1010-1 still has an appearance based on the persona template because avatar 1010-1 is located in the same position as the user of the computer system, or because user A is still having a conversation. In some implementations, the computer system considers the conversation to be an interactive activity.
[0233] like Figure 11B As shown, the computer system also updates the appearance of avatar 1010-2 to have a character template-based appearance, and also has an appearance based on... Figure 11A The pose is determined by the pose of user B. Because some parts of user B are outside the field of view 1007-3, the computer system determines the pose of the corresponding parts of avatar 1010-2 with a certain degree of uncertainty, as indicated by the dashed lines 1015 on the right arm and two legs of avatar 1010-2. The right arm is also shown as having a pose similar to... Figure 11A The actual pose of User B's right arm is different from the actual pose of User B's right arm because User B's right arm is outside the field of view 1007-3, and Figure 11B The pose shown is the pose of the right arm of avatar 1010-2 determined (estimated) by the computer system. In some embodiments, because object 1006 is within the field of view 1007-3, object 1006 is represented in CGR environment 1020 as... Figure 11B Reference object 1006 is indicated in the diagram. In some embodiments, object 1006 is represented as a virtual object in the CGR environment 1020. In some embodiments, object 1006 is represented as a physical object in the CGR environment 1020. In some embodiments, the left hand of avatar 1010-2 and optionally the representation of object 1006 in the CGR environment 1020 are shown with higher fidelity (e.g., higher fidelity than other parts of avatar 1010-2 and / or other objects in the CGR environment 1020) because the computer system detects that object 1006 is in the hand of user B.
[0234] like Figure 11B As shown, the computer system also updates the appearance of avatar 1010-3 to have a character template-based appearance, and also has an appearance based on... Figure 11A The pose is determined by the pose of user C. Because some parts of user C are outside the field of view 1007-4, the computer system determines the pose of the corresponding parts of avatar 1010-3 with a certain degree of uncertainty, such as that indicated by the dashed lines 1015 on avatar 1010-3's left hand and two legs. In some embodiments, because the computer system detects in Figure 10A Book 1008 within the field of view 1007-3, and no longer detected. Figure 11AThe computer system determines that book 1008 is likely in user C's left hand within the field of view 1007-3, and therefore indicates that book 1008 in the CGR environment 1020 is, as shown by... Figure 11B Reference object 1008' is shown in the figure. In some embodiments, the computer system represents the book 1008 as a virtual object and optionally has variable display characteristics to indicate the uncertainty of the presence and pose of the book 1008 in the real environment 1000-3.
[0235] In the embodiments described herein, reference is made to a user of the computer system. In some embodiments, the user of the computer system is a user viewing CGR environment 720 using display generation component 730, or a user viewing CGR environment 1020 using display generation component 1030. In some embodiments, the user of the computer system can be represented as an avatar in CGR environment 720 or CGR environment 1020 in a manner similar to that disclosed herein with respect to the corresponding avatars 721, 1010-1, 1010-2, 1010-3, and 1010-4. For example, in relation to Figures 7A to 7C and Figures 8A to 8C In the proposed implementation scheme, the user of the computer system is represented as a female avatar, as depicted in Preview 735. As another example, in the discussion of... Figures 10A to 10B and Figures 11A to 11B In the described implementation scheme, the user of the computer system can be a user in a real environment, similar to any of users A to D in real environments 1000-1 to 1000-4, and the user of the computer system can be depicted as an additional avatar feature in CGR environment 1020, which is presented to users A to D and is able to interact with users A to D. For example, in Figure 11B In the game, users A through C interact with the computer system, and avatars 1010-1, 1010-2, and 1010-3 are thus presented with an appearance based on a role template, while avatar 1010-4 is presented with an appearance based on an abstract template.
[0236] In some implementations, each of users A through D participates in the CGR environment using avatars 1010-1 through 1010-4, and each of these users views the other avatars using a computer system similar to the computer system described herein. Therefore, the avatars in the CGR environment 1020 can have different appearances for each corresponding user. For example, in Figure 11A In this context, user D can interact with user A but not with user B. Therefore, in the CGR environment viewed by user A, avatar 1010-4 can have an appearance based on the role template, while in the CGR environment viewed by user B, avatar 1010-4 can have an appearance based on the abstract template.
[0237] For purposes of explanation, the foregoing description has been given by reference to specific embodiments. However, the illustrative statements above are not intended to be exhaustive or to limit the invention to the precise forms disclosed. For example, in some embodiments, regarding… Figure 10B and Figure 11B The appearance of the incarnation discussed can have different levels of fidelity, such as regarding Figure 7C and Figure 8C As discussed above, for example, when the user associated with an avatar is in the same location as the user of the computer system, a higher fidelity representation is possible than when an avatar is associated with a user in a (remote) location different from the user of the computer system.
[0238] about Figures 10A to 10B and Figures 11A to 11B Additional descriptions are provided below, with reference to the following: Figure 12 The method described is provided in 1200.
[0239] Figure 12 This is a flowchart of an exemplary method 1200 for presenting an avatar with an appearance based on different appearance templates, according to some embodiments. In some embodiments, method 1200 is used in conjunction with displaying generated components (e.g., Figure 1 The display generation component 120 in the middle (for example, Figure 10B and Figure 11B A computer system (e.g., display generation components 1030) that communicates with a display generation component 1030 (e.g., a visual output device, a 3D display, a transparent display, a projector, a head-up display, a display controller, a touch screen, etc.) Figure 1 The method 1200 is executed at a computer system 101 (e.g., a smartphone, tablet, head-mounted display generating component). In some embodiments, the method 1200 is performed by a processor stored in a non-transitory computer-readable storage medium and by one or more processors of the computer system (e.g., one or more processors 202 of computer system 101). Figure 1 The control unit 110 in the middle executes instructions to manage. Some operations in method 1200 are optionally combined, and / or the order of some operations is optionally changed.
[0240] In method 1200, the computer system (e.g., 101) receives (1202) first data, which (e.g., based on data (e.g., image data, depth data, orientation of user equipment, audio data, motion sensor data, etc.)) instructs one or more users (e.g., Figure 10AThe current activity of users A, B, C, D (e.g., participants or objects represented in a computer-generated real-world environment) is a first-type activity (e.g., interactive (e.g., speaking; gesturing; moving); physically present in a real-world environment; performing a non-manual activity (e.g., speaking without moving)).
[0241] In response to receiving first data indicating that the current activity is a first type of activity, the computer system (e.g., 101) updates (1204) via a display generation component (e.g., 1030) (e.g., display; visual presentation; projection; modification of the user's representation of the display) (e.g., in a computer-generated reality environment (e.g., 1020)) the first user (e.g., in a first visual representation) having a first appearance based on a first appearance template (e.g., a template of an animated character (e.g., a human; a cartoon character; an anthropomorphic construct of a non-human character such as a dog, robot, etc.)) (e.g., a first visual representation). Figure 10B The representation of 1010-1, 1010-2, 1010-3, and / or 1010-4 in the template (e.g., representing a virtual avatar of one or more users; representing a portion of an avatar of one or more users). In response to receiving first data indicating that the current activity is a first type of activity, updating the representation of the first user having a first appearance based on a first appearance template provides feedback to the user that the first user is interacting with the user. Providing improved feedback enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating and / or interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling the user to use the system more quickly and efficiently. In some embodiments, updating the representation of the first user having the first appearance includes presenting the representation of the first user performing the first type of activity while having an appearance based on the first appearance template. For example, if the first type of activity is an interactive activity such as speaking, and the first appearance template is a dog character, then the representation of the first user is presented as an interactive dog character that is speaking.
[0242] The presentation that evokes a first user having a first appearance (e.g., in) Figure 10B While the computer system (e.g., 101) is presenting the avatar 1010-1, the computer system (e.g., 101) receives (1206) instructions from one or more users (e.g., Figure 11AThe second data indicates the current activity (e.g., actively performed activity) of users A, B, C, and D. In some embodiments, the second data indicates the current activity of a first user. In some embodiments, the second data indicates the current activity of one or more users other than the first user. In some embodiments, the second data indicates the current activity of one or more users including the first user. In some embodiments, the current activity of one or more users indicated by the second data is the same as the current activity of one or more users indicated by the first data.
[0243] In response to receiving second data indicating the current activity of one or more users, the computer system (e.g., 101) updates (1208) the appearance of the representation of the first user (e.g., avatar 1010-1 in...) based on the current activity of the one or more users. Figure 11B (Updated in the middle), including, based on determining that the current activity is a first type of activity (e.g., the current activity is an interactive activity), causing (1210) via a display generation component (e.g., 1030) to present (1210) a representation of a first user with a second appearance based on a first appearance template (e.g., avatar 1010-1 in the still using Figure 11B The representation of a first user is presented with a second pose (e.g., the representation of the first user is presented with an appearance different from the first appearance (e.g., the representation of the first user is presented performing an activity different from the first appearance), but the representation of the first user is still based on the first appearance template). Presenting a representation of the first user with a second appearance based on the first appearance template when the current activity is a first type of activity provides feedback to the user; that is, the first user is still interacting with the user, even though the first user's appearance has changed. Providing improved feedback enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating and / or interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively. In some embodiments, presenting a representation of the first user with a second appearance based on the first appearance template includes presenting a representation of the first user performing a first type of current activity while having an appearance based on the first appearance template. For example, if the current activity is an interactive activity such as gesturing with hands, and the first appearance template is a dog character, then the representation of the first user is presented as an interactive dog character gesturing with its hands.
[0244] In response to receiving second data indicating the current activity of one or more users, a computer system (e.g., 101) updates the appearance of a first user's representation (e.g., avatars 1010-2, avatars 1010-3, avatars 1010-4) based on the current activity of one or more users. This includes, based on determining that the current activity is a second type of activity different from the first type (e.g., the second type of activity is not interactive; the second type of activity is performed when one or more users (e.g., optionally including the first user) are not physically present in a real environment), via a display generation component (e.g., 1030) causing (1212) the presentation of the first user's representation having a third appearance based on a second appearance template different from the first appearance template (e.g., user B's avatar 1010-2 from using...). Figure 10B Abstract templates in the middle are transformed into those using Figure 11B (e.g., User C's avatar 1010-3 from the user's...) Figure 10B Abstract templates in the middle are transformed into those using Figure 11B (e.g., User D's avatar 1010-4 from the user's...) Figure 10B The character template in the middle is changed to use Figure 11B The representation of a first user with a third appearance based on the second appearance template (e.g., an abstract template in the second appearance template) (e.g., a template for an inanimate character, such as a plant) (e.g., a second visual representation). When the current activity is of the second type, the presentation of a representation of a first user with a third appearance based on the second appearance template provides feedback to the user that the first user has switched to performing a different activity than before, such as not interacting with the user. Providing improved feedback enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating and / or interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling the user to use the system more quickly and efficiently. In some embodiments, presenting the representation of the first user with the third appearance includes presenting the representation of the first user performing the second type of current activity while having an appearance based on the first appearance template. For example, if the second type of activity is a non-interactive activity such as sitting still, and the second appearance template is an irregular blob, the representation of the first user is presented as an irregular blob, which, for example, does not move or interact with the environment. In some embodiments, the presentation of a representation of the first user with the third appearance includes presenting the representation of the first user transitioning from having the first appearance to having the third appearance. For example, the first user's representation changes from an interactive dog character to a non-interactive irregular spot.
[0245] In some implementations, the first type of activity is the activity of a first user (e.g., user A) performed at a first location (e.g., the first user and a second user are in the same meeting room), and the second type of activity is the activity of the first user performed at a second location different from the first location (e.g., the first user is in a different location than the second user). In some implementations, the appearance template used for the first user is determined based on whether the first user is in the same physical environment as a user associated with the display generation component (e.g., the second user). For example, if the first and second users are in the same physical location (e.g., the same room), the first user is presented with a more realistic appearance template using the display generation component for the second user. Conversely, if the first and second users are in different physical locations, the first user is presented with a less realistic appearance template. This is done, for example, to convey to the second user that the first user and the second user are physically together. Furthermore, this distinguishes between physically present people and non-present people (e.g., virtually present people). This saves computational resources and streamlines the displayed environment by eliminating the need for display indicators to identify the location of different users or labels that will be used to mark a particular user as present or remote.
[0246] In some implementations, the first type of activity includes interaction between a first user (e.g., user A) and one or more users (e.g., user B, user C, user D) (e.g., the first user interacting with other users by speaking, gesturing, moving, etc.). In some implementations, the second type of activity does not include interaction between the first user and one or more users (e.g., the first user is not interacting with other users; for example, the first user's focus shifts to something other than other users). In some implementations, when the first user interacts with other users by speaking, gesturing, moving, etc., the first user's representation has an appearance based on a first appearance template. In some implementations, when the first user is not interacting with other users or the user's focus is on something other than other users, the first user's representation has an appearance based on a second appearance template.
[0247] In some implementations, the first appearance template corresponds to a second appearance template for a first user (e.g., by...). Figure 10B The abstract template represented by Avatar 1010-2 in the image has a more realistic appearance (e.g., character templates, such as those created by...). Figure 10BThe avatar 1010-1 represents the character template. Presenting a representation of the first user with a first appearance provides feedback to the user that the first user is interacting with the user; this first appearance is a more realistic representation of the first user than the second appearance template. Providing improved feedback enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating and / or interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively.
[0248] In some implementations, when a user engages with other users, for example by performing interactive activities such as speaking, gesturing, moving, etc., the first user is represented with a more realistic appearance (e.g., user A is represented by an avatar 1010-1 using a persona template) (e.g., a high-resolution or high-fidelity appearance; e.g., an anthropomorphic shape such as an interactive virtual avatar). In some implementations, when a user is not engaging with other users, the user is represented with a less realistic appearance (e.g., a low-resolution or low-fidelity appearance; e.g., an abstract shape (e.g., human-specific features have been removed, and the edges of the representation have been softened)). For example, the first user's focus shifts from other users to other content or objects, or the first user ceases performing interactive activities (e.g., the user remains seated, not speaking, gesturing, moving, etc.).
[0249] In some implementations, the first type of activity is non-manual activity (e.g., activities that involve little or no movement of the user's hands (e.g., speaking)). Figure 10A Users A, B, and D are listed. Figure 11A User A and User D in the example). In some implementations, the second type of activity is a manual activity (e.g., an activity that typically involves the movement of the user's hands (e.g., stacking toy blocks)). Figure 10A User C in the middle; Figure 11A User B in the example. In some embodiments, the representation of the first user includes a hand corresponding to the first user and is presented as a corresponding hand portion with variable display characteristics (e.g., ...). Figure 10B and Figure 11B(e.g., a variable set of one or more visual parameters (e.g., amount of blur, opacity, color, attenuation / density, resolution, etc.) for rendering the corresponding hand portion). In some embodiments, the variable display characteristics indicate the estimated / predicted visual fidelity of the pose of the corresponding hand portion relative to the pose of the first user's hand. In some embodiments, causing a representation of the first user having a third appearance based on a second appearance template includes: 1) based on determining a second type of activity including the first user's hand interacting with an object (e.g., in a physical environment or a virtual environment) to perform a manual activity (e.g., the first user is holding a toy block in their hand), presenting a representation of the first user with a first value of the corresponding hand portion having variable display characteristics (e.g., when in...). Figure 11A When user B is holding object 1006, the avatar 1010-2 is presented in high fidelity (e.g., a character template) and... Figure 11B (e.g., the variable display characteristic has a high visual fidelity value indicating the pose of the corresponding hand portion relative to the pose of the first user's hand) and 2) based on determining that the second type of activity does not include the first user's hand interacting with the object (e.g., in a physical environment or a virtual environment) to perform a manual activity (e.g., the first user is not holding the toy block in their hand), presenting a representation of the first user with a second value of the variable display characteristic having a different value than the first value of the variable display characteristic (e.g., when user B is not holding object 1006, using...). Figure 10B The abstract template in the representation is embodied in 1010-2 (e.g., the variable display characteristic has a value indicating a low visual fidelity of the pose of the corresponding hand portion relative to the pose of the first user's hand). Depending on whether the second type of activity involves the user's hand interacting with an object to perform a manual activity, a representation of the first user with a corresponding hand portion having a first or second value of the variable display characteristic is presented, providing feedback to the user that the first user has transitioned to performing an activity involving the use of their hand. Providing improved feedback enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating and / or interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively. In some embodiments, the representation of the first user's hand is rendered such that the variable display characteristic indicates higher visual fidelity when the user's hand interacts with an object (e.g., a physical or virtual object), and lower visual fidelity when the user's hand is not interacting with an object.
[0250] In some implementations, the first type of activity is the activity associated with the interaction with the first participant (e.g., user A and user D in...). Figure 10A (e.g., activities typically associated with a user as an active participant in interaction with other users (e.g., during a conversation) – such as a participant speaking, gesturing, moving, focusing on other users, etc.). In some implementations, the second type of activity is not associated with activities related to interaction with the first participant (e.g., user B and user C in...). Figure 10A (e.g., activities that do not interact with other users) (e.g., activities that are typically associated with being an observer who does not engage with other users (e.g., talking) (e.g., the observer is not speaking, gesturing, moving, focusing on other users, etc.)).
[0251] In some implementations, the presentation of a representation of a first user having a second appearance based on a first appearance template includes presenting a representation of the first user with a first value indicating a first visual fidelity of the representation of the first user having variable display characteristics (e.g., using...). Figure 10B The character template in the file presents User A as Avatar 1010-1; using Figure 10B The role template in the image represents user D as an avatar 1010-4 (e.g., a variable set of one or more visual parameters of the rendering of the first user's representation (e.g., the amount of blur, opacity, color, attenuation / density, resolution, etc.)) (e.g., the visual fidelity of the first user's representation relative to the pose of the corresponding part of the first user). When the first user's representation has a second appearance based on the first appearance template, the first user's representation, which presents variable display characteristics indicating a first value of the first visual fidelity of the first user's representation, provides the user with feedback that is more relevant to the first user's current activity, as the first user is engaged in that activity. Providing improved feedback enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating and / or interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively. In some embodiments, the first user's representation includes a first representation portion corresponding to a first part of the first user and is presented with variable display characteristics indicating an estimated visual fidelity of the pose of the first representation portion relative to the first part of the first user.
[0252] In some implementations, the presentation of a representation of a first user having a third appearance based on a second appearance template includes presenting a representation of the first user with a second value of a second visual fidelity indicating the first user's representation, which has variable display characteristics and is less than the first visual fidelity of the first user's representation (e.g., using...). Figure 10B The abstract template in the file represents user B as avatar 1010-2; using Figure 10B The abstract template in the template represents user C as an avatar 1010-3 (e.g., when the user's activity is an observer activity, the representation of the first user is presented as a variable display characteristic value indicating lower visual fidelity). When the representation of the first user has a third appearance based on a second appearance template, the representation of the first user, which is presented as a second value of the second visual fidelity of the first user's representation, indicating a first visual fidelity less than the first visual fidelity of the first user's representation, provides feedback to the user that is less relevant to the current activity, since the first user is not actively participating in the activity. Providing improved feedback enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating and / or interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively.
[0253] In some implementations, users are considered either participants or observers based on their activities, and those considered participants are represented with higher visual fidelity than those considered observers. For example, users who actively communicate, gesture, move, focus on other participants and / or users (based on this activity) are considered participants (or active participants), and their representations are presented with variable display characteristic values indicating high visual fidelity. Conversely, users who do not actively communicate, gesture, move, focus on other participants and / or users (based on this activity) are considered observers, and their representations are presented with variable display characteristic values indicating low visual fidelity.
[0254] In some embodiments, the display generation component (e.g., 1030) is associated with a second user (e.g., the second user is the viewer and / or user of the display generation component). In some embodiments, when a representation of a third user (e.g., representing an avatar of user B, 1010-2) with variable display characteristics (e.g., a representation of the first user) (e.g., alternatively, a representation of an object) is presented via the display generation component (e.g., a variable set of one or more visual parameters of the rendering of the first user's representation (e.g., the amount of blur, opacity, color, attenuation / density, resolution, etc.)), the computer system (e.g., 101) receives third data indicating the focus position of the second user (e.g., data indicating the location or area in the physical or computer-generated environment where the second user's eyes are focused). In some embodiments, the third data is generated using a sensor (e.g., a camera sensor) to track the position of the second user's eyes and the focus of the user's eyes is determined using a processor. In some embodiments, the focus position includes the focus of the user's eyes calculated in three dimensions (e.g., the focus position may include a depth component).
[0255] In some implementations, in response to receiving third data indicating the focus position of a second user, a computer system (e.g., 101) updates the representation of a third user (e.g., alternatively, the representation of an updated object) based on the focus position of the second user via a display generation component, including: 1) determining the location of the second user's focus position corresponding to the location of the third user's representation (e.g., co-located with that location) (e.g., the user of the computer system in...). Figure 11A 1) Looking at user B, and therefore displaying avatar 1010-2 using a role template (e.g., the second user is looking at the representation of the third user) (e.g., alternatively, the second user is looking at the representation of an object), increasing the value of the variable display characteristics of the third user's representation (e.g., alternatively, increasing the value of the variable display characteristics of the representation of an object); and 2) determining that the focus position of the second user does not correspond to the position of the third user's representation (e.g., not co-located with that position) (e.g., the second user is not looking at the representation of the third user) (e.g., alternatively, the second user is not looking at the representation of an object), decreasing the value of the variable display characteristics of the third user's representation (e.g., the user of the computer system is looking at the representation of the third user). Figure 11AFrom the user D's perspective, the avatar 1010-4 is displayed using an abstract template (e.g., alternatively, the value of the variable display properties of the object's representation is reduced). Increasing or decreasing the value of the variable display properties of the third user's representation, depending on whether the second user's focus position corresponds to the position of the third user's representation, provides the second user with improved feedback that the computer system is detecting the focus position, and reduces computational resources consumed by eliminating or reducing the amount of data to be processed to render areas where the user is not focused. Providing improved feedback and reducing computational workload enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating and / or interacting with the computer system), which in turn reduces power consumption and extends system battery life by enabling users to use the computer system more quickly and effectively.
[0256] In some implementations, variable display characteristics are adjusted for various objects or users presented using display generation components, depending on whether the user is focusing on those corresponding objects or users. For example, when the user moves their focus away from the representation of the user or object, the visual fidelity (e.g., resolution, sharpness, opacity, and / or density, etc.) of the representation of the user or object decreases (e.g., based on the amount by which the user's focus deviates from the representation). Conversely, when the user moves their focus back to the representation of the user or object, the visual fidelity (e.g., resolution, sharpness, opacity, and / or density, etc.) of the representation of the user or object increases. This visual effect serves to conserve computational resources by reducing the amount of data to be rendered to the user using display generation components. Specifically, less computational resources are consumed because objects and / or users presented outside the user's focus can be rendered at lower fidelity. This can be done without sacrificing visual integrity, as the modification of visual fidelity mimics the natural behavior of the human eye. For example, when the human eye shifts focus, objects on which the eye is focused appear sharper (e.g., high fidelity), while objects outside the eye's focus become blurry (e.g., low fidelity). In short, objects rendered in the user's peripheral view are intentionally rendered at low fidelity to conserve computing resources.
[0257] In some implementations, the first appearance template corresponds to an abstract shape (e.g., Figure 10B 1010-2 in the middle; Figure 10B 1010-3 in; Figure 11B (e.g., 1010-4) (e.g., a representation of a shape that has no anthropomorphic characteristics or fewer anthropomorphic characteristics compared to the anthropomorphic shape used to represent a user when the user is engaged in a first type of activity). In some embodiments, the second appearance template corresponds to the anthropomorphic shape (e.g., Figure 10B 1010-1 in; Figure 10B1010-4 in the middle; Figure 11B 1010-1 in; Figure 11B 1010-2 in the middle; Figure 11B (e.g., a representation of a human, animal, or other character having expressive elements (such as a face, eyes, mouth, limbs, etc.) that allow the anthropomorphic shape to convey the user's movement). Presenting a first user representation with a first appearance template corresponding to an abstract shape or a second appearance template corresponding to an anthropomorphic shape provides feedback to the user; that is, when presented as an anthropomorphic shape, the first user representation is more relevant to the current activity, while when presented as an abstract shape, the first user representation is less relevant to the current activity. Providing improved feedback enhances the operability of the computer system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating and / or interacting with the computer system), which in turn reduces power consumption and extends the battery life of the computer system by enabling the user to use the system more quickly and effectively.
[0258] In some implementations, aspects and / or operations of methods 900 and 1200 may be interchanged, substituted, and / or added between these methods. For the sake of brevity, these details will not be repeated here.
[0259] For purposes of explanation, the foregoing description has been given by reference to specific embodiments. However, the illustrative discussion above is not intended to be exhaustive or to limit the invention to the precise forms disclosed. Many modifications and variations are possible based on the teachings above. The embodiments were chosen and described to best elucidate the principles of the invention and its practical application, thereby enabling others skilled in the art to best utilize the invention with various modifications suitable for the contemplated particular purpose, as well as the various described embodiments.
Claims
1. A method for presenting an avatar, comprising: At the computer system communicating with the display generation component: When a user of the computer system communicates with a user of a remote computer system in a computer-generated reality (CGR) experience, displaying the representation of the user of the remote computer system via the display generation component includes: Receive pose data representing at least a first portion of the pose of the user on the remote computer system; and The presentation of an avatar is caused by the display generation component, wherein the avatar includes a first portion corresponding to the user of the remote computer system and is presented as a corresponding avatar feature with variable display characteristics, the variable display characteristics indicating the determinism of the pose of the first portion of the user of the remote computer system, wherein the presentation of the avatar is caused by: Based on the determination that the pose of the first portion of the user of the remote computer system is associated with a first deterministic value, the presentation of the avatar with the corresponding avatar feature, the corresponding avatar feature having a first value of the variable display characteristic; and Based on the determination that the pose of the first part of the user of the remote computer system is associated with a second deterministic value different from the first deterministic value, the presentation of the avatar with the corresponding avatar feature is caused, the corresponding avatar feature having a second value of the variable display feature different from the first value of the variable display feature.
2. The method according to claim 1, further comprising: Receive second pose data representing at least a second portion of the pose of the user of the remote computer system; as well as The avatar is rendered via the display generation component, wherein the avatar includes a second portion corresponding to the user of the remote computer system and is rendered as a second avatar feature having a second variable display characteristic, the second variable display characteristic indicating the determinacy of the pose of the second portion of the user of the remote computer system, wherein rendering the avatar includes: Based on the association of the pose of the second portion of the user of the remote computer system with a third deterministic value, the avatar with the second avatar feature is rendered, the second avatar feature having a first value of the second variable display characteristic; and Associating the pose of the second part of the user of the remote computer system with a fourth deterministic value that is different from the pose of the second part of the user of the remote computer system, causes the presentation of the avatar with the second avatar feature, the second avatar feature having a second value of the second variable display feature that is different from the first value of the second variable display feature.
3. The method of claim 2, wherein causing the presentation of the avatar comprises: The third deterministic value of the pose of the second part of the user of the remote computer system corresponds to the first deterministic value, and the first value of the second variable display characteristic corresponds to the first value of the variable display characteristic; and The fourth deterministic value corresponds to the second deterministic value, and the second value of the second variable display characteristic corresponds to the second value of the variable display characteristic.
4. The method according to claim 2, further comprising: The third deterministic value of determining the pose of the second part of the user of the remote computer system corresponds to the first deterministic value, and the first value of the second variable display characteristic does not correspond to the first value of the variable display characteristic; and The determination is based on the fact that the fourth deterministic value corresponds to the second deterministic value, and the second value of the second variable display characteristic does not correspond to the second value of the variable display characteristic.
5. The method according to any one of claims 1 to 4, further comprising: Receive updated pose data representing pose changes of the first part of the user in the remote computer system; as well as In response to receiving the updated pose data, updating the presentation of the avatar includes: The pose of the corresponding avatar feature is updated based on the pose change of the first part of the user in the remote computer system.
6. The method of claim 5, wherein updating the presentation of the avatar comprises: In response to the deterministic change in the pose of the first part of the user in the remote computer system during the pose change of the first part of the user in the remote computer system, in addition to changing the positioning of at least a part of the avatar based on the pose change of the first part of the user in the remote computer system, the variable display characteristics of the corresponding avatar features displayed are also changed.
7. The method according to any one of claims 1 to 4, wherein the variable display characteristic indicates the estimated visual fidelity of the pose of the corresponding avatar feature relative to the first portion of the user of the remote computer system.
8. The method according to any one of claims 1 to 4, wherein The manifestation of the aforementioned incarnation includes: Based on determining that the pose data satisfies a first set of criteria when the first part of the user of the remote computer system is detected by a first sensor, the presentation of the avatar with the corresponding avatar feature is caused, the corresponding avatar feature having a third value of the variable display characteristic; as well as Based on the determination that the pose data fails to meet the first set of criteria, the avatar with the corresponding avatar feature is rendered, the corresponding avatar feature having a fourth value of the variable display feature indicating a lower deterministic value than the third value of the variable display feature.
9. The method according to any one of claims 1 to 4, further comprising: When the corresponding avatar feature is presented as the current value having the variable display characteristics, updated pose data representing the change in pose of the first part of the user of the remote computer system is received. as well as In response to receiving the updated pose data, updating the presentation of the avatar includes: Based on the determination that the updated pose data represents the change in the pose of the first part of the user of the remote computer system from a first position within the view of the sensor to a second position outside the view of the sensor, the current value of the variable display characteristic is reduced. as well as Based on the determination that the updated pose data represents the change in pose of the first part of the user of the remote computer system from the second position to the first position, the current value of the variable display characteristic is increased.
10. The method of claim 9, wherein: The current value of the variable display characteristic decreases at a first rate, and The current value of the variable display characteristic increases at a second rate greater than the first rate.
11. The method according to any one of claims 1 to 4, wherein: The first value of the variable display characteristic represents a higher visual fidelity of the corresponding avatar feature relative to the pose of the first part of the user in the remote computer system than the second value of the variable display characteristic. The manifestation of the aforementioned incarnation includes: Based on determining that the first portion of the user of the remote computer system corresponds to a subset of physical features, the pose of the first portion of the user of the remote computer system is associated with a second deterministic value; and Based on the determination that the first portion of the user of the remote computer system does not correspond to the subset of the physical features, the pose of the first portion of the user of the remote computer system is associated with the first deterministic value.
12. The method according to any one of claims 1 to 4, further comprising: When the avatar is presented as the corresponding avatar feature having the first value of the variable display characteristics, updating the presentation of the avatar includes: Based on determining that the movement speed of the first part of the user of the remote computer system is a first movement speed of the first part of the user of the remote computer system, the avatar with the corresponding avatar feature is displayed, the corresponding avatar feature having a first change value of the variable display characteristic; and Based on the determination that the first part of the user of the remote computer system has a movement speed that is different from the first movement speed of the first part of the user of the remote computer system, the presentation of the avatar with the corresponding avatar feature is caused, and the corresponding avatar feature has a second change value of the variable display characteristic.
13. The method according to any one of claims 1 to 4, further comprising: Changing the value of the variable display characteristic includes changing one or more visual parameters of the corresponding avatar feature.
14. The method of claim 13, wherein the one or more visual parameters include blurriness.
15. The method of claim 13, wherein the one or more visual parameters include opacity.
16. The method of claim 13, wherein the one or more visual parameters include color.
17. The method of claim 13, wherein the one or more visual parameters include the density of particles comprising the corresponding avatar feature.
18. The method of claim 17, wherein the density of the particles comprising the corresponding avatar feature includes the spacing between the particles comprising the corresponding avatar feature.
19. The method of claim 17, wherein the density of the particle comprising the corresponding avatar feature includes the size of the particle comprising the corresponding avatar feature.
20. The method according to any one of claims 1 to 4, further comprising: Changing the value of the variable display characteristic includes causing the presentation of visual effects associated with the corresponding avatar feature.
21. The method according to any one of claims 1 to 4, wherein: The first part of the user of the remote computer system includes a first physical characteristic and a second physical characteristic. The first value of the variable display characteristic represents a higher visual fidelity of the pose of the corresponding avatar feature relative to the first physical feature of the user in the remote computer system than the second value of the variable display characteristic. The presentation of the avatar, which causes the corresponding avatar feature to have the second value of the variable display characteristics, includes: The avatar is rendered by rendering a first physical feature based on the corresponding physical characteristics of the user of the remote computer system and a second physical feature based on the corresponding physical characteristics of the user who is not the user of the remote computer system.
22. The method according to any one of claims 1 to 4, wherein the pose data is generated by a plurality of sensors.
23. The method of claim 22, wherein the plurality of sensors includes one or more camera sensors associated with the computer system.
24. The method of claim 22, wherein the plurality of sensors includes one or more camera sensors separate from the computer system.
25. The method of claim 22, wherein the plurality of sensors comprises one or more non-visual sensors.
26. The method of any one of claims 1 to 4, wherein an interpolation function is used to determine the pose of at least the first portion of the user of the remote computer system.
27. The method of any one of claims 1 to 4, wherein the pose data comprises data generated from prior scan data that captures information about the appearance of the user on the remote computer system.
28. The method according to any one of claims 1 to 4, wherein the pose data includes data generated from prior media data.
29. The method according to any one of claims 1 to 4, wherein: The pose data includes video data, and the video data includes at least the first portion of the user's data from the remote computer system. Presenting the avatar includes presenting a modeled avatar, the modeled avatar including the corresponding avatar features rendered using the video data, the video data including the first portion of the user's data from the remote computer system.
30. The method according to any one of claims 1 to 4, wherein causing the presentation of the avatar comprises: Based on the determination that an input indicating a first rendering value for the avatar has been received, the avatar is rendered with the corresponding avatar features and a first amount of avatar features other than the corresponding avatar features; as well as Based on the determination that an input indicating a second rendering value different from the first rendering value has been received for the avatar, the avatar is rendered with the corresponding avatar feature and a second amount of avatar features other than the corresponding avatar feature, wherein the second amount is different from the first amount.
31. The method according to any one of claims 1 to 4, the method further comprising: The display generation component causes the presentation of a representation of a user associated with the display generation component, wherein the representation of the user associated with the display generation component corresponds to the appearance of the user associated with the display generation component, and the appearance is presented to one or more users other than the user associated with the display generation component.
32. The method according to any one of claims 1 to 4, wherein causing the presentation of the avatar comprises causing the presentation of the avatar with the corresponding avatar features having the first appearance based on the first appearance of the first portion of the user of the remote computer system, the method further comprising: Receive data indicating the updated appearance of the first part of the user's system on the remote computer system; as well as The display generation component, based on the updated appearance of the first part of the user of the remote computer system, causes the presentation of the avatar with the corresponding avatar features having the updated appearance.
33. The method according to any one of claims 1 to 4, wherein: The pose data further represents an object associated with the first portion of the user of the remote computer system; and The manifestation of the aforementioned incarnation includes: This causes the presentation of the avatar, which includes a representation of the object adjacent to the corresponding avatar feature.
34. The method according to any one of claims 1 to 4, wherein: The pose of the first part of the user in the remote computer system is associated with a fifth deterministic value, and The manifestation of the aforementioned incarnation includes: Based on the determination that the first part of the user of the remote computer system is a first feature type, the presentation of the avatar with the corresponding avatar feature having a value of less than the fifth deterministic value, which indicates the variable display characteristics, is caused.
35. The method according to any one of claims 1 to 4, wherein: The pose of the first part of the user in the remote computer system is associated with a sixth deterministic value, and The manifestation of the aforementioned incarnation also includes: Based on the determination that the first part of the user of the remote computer system is a second feature type, the presentation of the avatar with the corresponding avatar feature having a value greater than the sixth deterministic value, which indicates the variable display characteristics, is caused.
36. The method according to any one of claims 1 to 4, wherein: The presentation of the avatar having the corresponding avatar feature with the first value of the variable display characteristics includes: moving the corresponding avatar feature with the first value of the variable display characteristics as the pose of the first part of the user of the remote computer system changes over time; and The presentation of the avatar with the corresponding avatar feature having the second value of the variable display characteristics includes: moving the corresponding avatar feature having the second value of the variable display characteristics as the pose of the first part of the user of the remote computer system changes over time.
37. The method according to any one of claims 1 to 4, wherein causing the presentation of the avatar comprises: Based on a third deterministic value that associates the pose of the first part of the user of the remote computer system with the pose of the first part of the user of the remote computer system, the presentation of the avatar of the corresponding avatar feature with a third value of the variable display characteristic that is different from the first and second values of the variable display characteristic is caused.
38. The method of any one of claims 1 to 4, wherein the certainty of the pose of the first portion of the user of the remote computer system is based on sensor measurements from one or more sensors of the remote computer system.
39. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs including instructions for performing the method according to any one of claims 1 to 38.
40. A computer system comprising: One or more processors; as well as A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 1 to 38.
41. A computer system, comprising: Apparatus for performing the method according to any one of claims 1 to 38.
42. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs comprising instructions for performing the method according to any one of claims 1 to 38.
Citation Information
Patent Citations
Image generation system, image generation method, and information storage medium
US20110306420A1