User representation using depths relative to multiple surface points

By generating a set of values representing the 3D shape and appearance of a user's face using depth and appearance values relative to a surface, the system addresses the issue of inaccurate real-time user representation in electronic devices, achieving improved accuracy with reduced computational resources.

JP2025084733AActive Publication Date: 2025-06-03APPLE INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025004877
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-06-27
Filing Date
2025-01-14
Publication Date
2025-06-03
Estimated Expiration
2043-06-29

AI Technical Summary

Technical Problem

Existing electronic devices fail to accurately represent a user's real-time appearance, often using outdated images for avatar representations, leading to discrepancies between the displayed image and the user's current appearance.

Method used

The system generates a set of values representing the 3D shape and appearance of a user's face at a given point in time, using depth values and appearance values defined relative to a plurality of points on a surface, such as a cylindrical shape, to create a more accurate user representation.

Benefits of technology

This approach enables a more accurate user representation compared to RGBDA images, while requiring less computational effort and bandwidth than traditional 3D meshes or point clouds, and can be efficiently integrated with existing systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025084733000001_ABST
    Figure 2025084733000001_ABST
Patent Text Reader

Abstract

To provide a system, method and an electronic device for generating accurate or honest current representations of the appearances of users.SOLUTION: A method includes: obtaining sensor data (e.g., live data) of a user, the sensor data being associated with a certain point in time; generating a set of values representing the user based on the sensor data; and providing the set of values, a depiction of the user at the point in time being displayed based on the set of values. The set of values includes depth values that define three-dimensional (3D) positions of portions of the user relative to multiple 3D positions of points of a projected surface and appearance values (e.g., color, texture, opacity, etc.) that define appearances of the portions of the user.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to electronic devices, and more particularly, to systems, methods, and devices for representing a user in computer-generated content.

Background Art

[0002] In existing technologies, the current (e.g., real-time) representation of a user's appearance on an electronic device may not accurately or faithfully represent the user's appearance. For example, the device may provide an avatar representation of the user based on an image of the user's face that was acquired minutes, hours, days, or even years ago. Such a representation may not accurately represent the user's current (e.g., real-time) appearance. For example, the user's avatar may not be shown smiling when the user is smiling, or may not show the user's current beard. Thus, there may be a desire to provide a means for more accurately, faithfully, and / or currently representing the user efficiently.

Summary of the Invention

[0003] The various implementations disclosed herein include devices, systems, and methods that generate a set of values representing the three-dimensional (3D) shape and appearance of a user's face at a given point in time, which are used to generate user representations (e.g., avatars). In some implementations, a surface having a non-planar shape (e.g., a cylindrical shape) can be used to reduce distortion. The set of values includes depth values that define the depth of portions of the face relative to a plurality of points on the surface, e.g., points within a grid on a partially cylindrical surface. For example, the depth value of one point can define that the portion of the face is at a depth D1 behind the position of that point on the surface, e.g., at a depth D1 along a ray starting from that point. The techniques described herein use depth values that are different from the depth values in existing RGBDA images (e.g., red-green-blue-depth-alpha images). This is because RGBDA images define content depth for a single camera location, while the techniques described herein define depth as portions of the face relative to a plurality of points on a surface of a planar shape (e.g., a cylindrical shape).

[0004] Using a relatively simple set of values having depth values defined for a plurality of points on the surface can achieve several advantages. The set of values can enable a more accurate user representation than RGBDA images, while requiring less computational effort and bandwidth than using 3D meshes or 3D point clouds. Further, the set of values may be formatted / packaged in a similar manner to existing formats, e.g., RGBDA images, thereby enabling more efficient integration with systems based on such formats.

[0005] Generally, one inventive aspect of the subject matter described in this specification can be embodied in a method that includes, in a processor of a device, an action of obtaining sensor data of a user, where the sensor data is associated with a point in time. The action further includes generating a set of values representing the user based on the sensor data, the set of values including depth values defining 3D positions of portions of the user relative to 3D positions of points on a projection plane, and appearance values defining the appearance of portions of the user. The action further includes providing the set of values, and a depiction of the user at that point in time is displayed based on the set of values.

[0006] These and other embodiments can optionally include one or more of the following features.

[0007] In some aspects, the points are spaced apart at regular intervals along vertical and horizontal lines on the surface.

[0008] In some aspects, the surface is non-planar. In some aspects, the surface is at least partially cylindrical. In some aspects, the surface is planar.

[0009] In some aspects, the set of values is generated based on an alignment such that a subset of points on a central area of the surface corresponds to a central portion of the user's face. In some aspects, generating the set of values is further based on an image of the user's face captured while the user is expressing a plurality of different facial expressions.

[0010] In some aspects, the sensor data corresponds to only a first area of the user, and the set of image data corresponds to a second area including a third area different from the first area.

[0011] In some embodiments, the method further includes an action of obtaining additional sensor data of a user associated with a second period, an action of updating a set of values representing the user based on the additional sensor data of the second period, and an action of providing the updated set of values, wherein the depiction of the user is updated in the second period based on the updated set of values.

[0012] In some embodiments, providing the set of values includes transmitting a sequence of frames of 3D video data including a frame having the set of values during a communication session with a second device, and the second device renders an animated depiction of the user based on the sequence of frames of 3D video data.

[0013] In some embodiments, the electronic device includes a first sensor and a second sensor, and the sensor data is obtained from at least one partial image of the user's face from the first sensor from a first perspective and from at least one partial image of the user's face from the second sensor from a second perspective different from the first perspective.

[0014] In some embodiments, the depiction of the user is displayed in real time.

[0015] In some embodiments, generating the set of values representing the user is based on a machine learning model trained to generate the set of values.

[0016] In some embodiments, the depth value defines the distance between a part of the user and a corresponding point on a projection plane positioned along a ray perpendicular to the projection plane at the position of the corresponding point. In some embodiments, the appearance value includes a color value, a texture value, or an opacity value.

[0017] In some embodiments, the electronic device is a head-mounted device (HMD). In some embodiments, the HMD includes one or more inward-facing image sensors and one or more downward-facing image sensors, and the sensor data is captured by the one or more inward-facing sensors and the one or more downward-facing image sensors.

[0018] These and other embodiments can optionally include one or more of the following features.

[0019] According to some implementations, a non-transitory computer-readable storage medium stores computer-executable instructions that perform or cause the execution of any of the methods described herein. According to some implementations, the device includes one or more processors, a non-transitory memory, and one or more programs. The one or more programs are stored in the non-transitory memory and configured to be executed by the one or more processors, and the one or more programs include instructions to perform or cause the execution of any of the methods described herein.

Brief Description of the Drawings

[0020] As will be understood by those skilled in the art, the present disclosure may have a more detailed description by referring to aspects of some exemplary implementations, some of which are shown in the accompanying drawings.

[0021]

Figure 1

[0022]

Figure 2

[0023]

Figure 3A

[0024]

Figure 3B

[0025]

Figure 3C

[0026]

Figure 4

[0027]

Figure 5

[0028]

Figure 6

[0029]

Figure 7

[0030]

Figure 8

[0031] By convention, the various features shown in the drawings may not be drawn to scale. Accordingly, the dimensions of the various features may be arbitrarily enlarged or reduced for clarity. Additionally, some of the drawings may not depict all of the components of a given system, method, or device. Finally, throughout this specification and the figures, like reference numerals may be used to indicate like features. **DETAILED DESCRIPTION**

[0032] Numerous details are set forth in order to provide a thorough understanding of the exemplary embodiments shown in the drawings. However, the drawings merely illustrate some exemplary aspects of the present disclosure and should not be considered limiting. Those of ordinary skill in the art will understand that other effective aspects or variations may not include all of the specific details described herein. Furthermore, well-known systems, methods, components, devices, and circuits have not been described in exhaustive detail so as not to obscure more appropriate aspects of the exemplary embodiments described herein.

[0033] FIG. 1 shows an exemplary environment 100 of a real-world environment 105 (e.g., a room) that includes a device 10 having a display 15. In some embodiments, device 10 displays content 20 to user 25. For example, content 20 may be buttons, user interface icons, text boxes, graphics, avatars of the user or another user, etc. In some embodiments, content 20 may occupy the entire display area of display 15.

[0034] Device 10 obtains image data, motion data, and / or physiological data (e.g., pupil data, facial feature data, etc.) from user 25 via a plurality of sensors (e.g., sensors 35a, 35b, and 35c). For example, device 10 obtains line-of-sight characteristic data 40b via sensor 35b, upper facial feature characteristic data 40a via sensor 35a, and lower facial feature characteristic data 40c via sensor 35c.

[0035] This example and other examples considered in this specification show a single device 10 within the real-world environment 105, but the techniques disclosed herein are applicable to multiple devices as well as other real-world environments. For example, the functionality of device 10 may be performed by multiple devices where sensors 35a, 35b, and 35c are on each respective device, or may be divided among them in any combination.

[0036] In some implementations, the plurality of sensors (e.g., sensors 35a, 35b, and 35c) may include any number of sensors that acquire data related to the appearance of user 25. For example, when wearing a head-mounted device (HMD), one sensor (e.g., a camera inside the HMD) may acquire pupil data for eye tracking, and one sensor on a separate device (e.g., one camera such as a wide-view camera) may be able to capture all of the user's facial feature data. Alternatively, when device 10 is an HMD, a separate device may not be necessary. For example, when device 10 is an HMD, in one implementation, sensor 35b may be disposed inside the HMD to capture pupil data (e.g., gaze characteristic data 40b), and additional sensors (e.g., sensors 35a and 35c) may be disposed on the outer surface of the HMD that is on the HMD but faces towards the user's head / face to capture facial feature data (e.g., upper facial feature characteristic data 40a via sensor 35a and lower facial feature characteristic data 40c via sensor 35c).

[0037] In some implementations, as shown in FIG. 1, device 10 is a handheld electronic device (e.g., a smartphone or a tablet). In some implementations, device 10 is a laptop computer or a desktop computer. In some implementations, device 10 has a touchpad, and in some implementations, device 10 has a touch-sensitive display (also known as a "touch screen" or "touch screen display"). In some implementations, device 10 is a wearable device such as an HMD.

[0038] In some implementations, device 10 includes an eye-tracking system for detecting the position and movement of the eyes via eye characteristic data 40b. For example, the eye-tracking system may include one or more infrared (IR) light-emitting diodes (LEDs), an eye-tracking camera (e.g., a near-infrared (NIR) camera), and an illumination source (e.g., an NIR light source) that emits light (e.g., NIR light) towards the eyes of user 25. Further, the illumination source of device 10 can emit NIR light to illuminate the eyes of user 25, and the NIR camera can capture an image of the eyes of user 25. In some implementations, the image captured by the eye-tracking system may be analyzed to detect the position and movement of the eyes of user 25, or other information about the eyes such as color, shape, state (e.g., wide open, squinted, etc.), pupil dilation, or pupil diameter. Further, the viewpoint estimated from the eye-tracking image can enable viewpoint-based interaction with the content shown on the near-eye display of device 10.

[0039] In some implementations, device 10 includes a graphical user interface (GUI), one or more processors, memory, and one or more modules, programs, or sets of instructions stored in the memory to perform multiple functions. In some implementations, user 25 interacts with the GUI via finger contact and gestures on a touch-sensitive surface. In some implementations, the functions include image editing, drawing, presenting, word processing, website creation, disk authoring, spreadsheet creation, game play, making phone calls, video conferencing, sending emails, instant messaging, training support, digital photography, digital video recording, web browsing, playing digital music, and / or playing digital video. The executable instructions for performing these functions may be included in a computer-readable storage medium or other computer program product configured to be executed by one or more processors.

[0040] In some implementations, device 10 employs various physiological sensors, detection, or measurement systems. The physiological data detected may include, but is not limited to, electroencephalogram recording (EEG), electrocardiogram recording (ECG), electromyogram recording (EMG), functional near-infrared spectroscopy signal (fNIRS), blood pressure, skin conductance, or pupil reflex. Further, device 10 may simultaneously detect multiple forms of physiological data to benefit from the synchronous acquisition of physiological data. Further, in some implementations, the physiological data is involuntary data, e.g., representing responses not under conscious control. For example, pupil reflex may represent an involuntary movement.

[0041] In some implementations, one or both eyes 45 of user 25, including one or both pupils 50 of user 25, represent physiological data in the form of a pupil reflex (e.g., line-of-sight characteristic data 40b). The pupil reflex of user 25 causes a change in the size or diameter of pupil 50 via the optic nerve and oculomotor nerve. For example, the pupil reflex may include a constriction reflex (miosis), e.g., narrowing of the pupil, or a dilation reflex (mydriasis), e.g., enlargement of the pupil. In some implementations, device 10 may detect a pattern of physiological data representing a time-varying pupil diameter.

[0042] User data (e.g., upper face feature characteristic data 40a, lower face feature characteristic data 40c, and line-of-sight characteristic data 40b) may change over time, and device 10 can use the user data to generate and / or provide a representation of the user.

[0043] In some implementations, user data (e.g., upper face feature characteristic data 40a and lower face feature characteristic data 40c) includes texture data of facial features such as eyebrow movement, jaw movement, nose movement, cheek movement, etc. For example, when a person (e.g., user 25) smiles, the upper face features and lower face features (e.g., upper face feature characteristic data 40a and lower face feature characteristic data 40c) may include a number of muscle movements that can be reproduced by a representation of the user (e.g., an avatar) based on data captured by sensor 35.

[0044] According to some implementations, an electronic device (e.g., device 10) can generate and present an extended reality (XR) environment to one or more users during a communication session. In contrast to the physical environment that people can perceive and / or interact with without the assistance of an electronic device, the extended reality (XR) environment refers to an environment that is wholly or partially simulated and with which people can perceive and / or interact via the electronic device. For example, the XR environment may include augmented reality (AR) content, mixed reality (MR) content, virtual reality (VR) content, and the like. In an XR system, a subset of a person's body movements or their representations are tracked, and in response, one or more characteristics of one or more virtual objects simulated within the XR environment are adjusted to behave according to at least one physical law. As an example, the XR system can detect the rotation of a person's head and, in response, adjust the graphic content and sound field presented to the person in the same way as how such views and sounds would change in the physical environment. As another example, the XR system can detect the movement of an electronic device presenting the XR environment (e.g., a mobile phone, a tablet, a laptop, etc.) and, in response, adjust the graphic content and sound field presented to the person in the same way as how such views and sounds would change in the physical environment. In some situations (e.g., for accessibility reasons), the XR system can adjust one or more characteristics of the graphic content within the XR environment in response to a representation of a body movement (e.g., a voice command).

[0045] The existence of a wide variety of electronic systems enables a person to perceive and / or interact with various XR environments. Examples include head-mountable systems, projection-based systems, heads-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be placed over a person's eyes (similar to contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable controllers or handheld controllers with or without tactile feedback), smartphones, tablets, and desktop / laptop computers. A head-mountable system may have one or more speakers (singular or plural) and an integrated opaque display. Alternatively, a head-mountable system may be configured to accept an external opaque display (e.g., a smartphone). A head-mountable system may incorporate one or more imaging sensors for capturing an image or video of the physical environment and / or one or more microphones for capturing the sound of the physical environment. A head-mountable system may have a transparent or translucent display instead of an opaque display. The transparent or translucent display may have a medium through which light representing an image is directed towards a person's eyes. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scan light sources, or any combination of these technologies. The medium may be an optical waveguide, hologram medium, optical coupler, optical reflector, or any combination thereof. In some implementations, the transparent or translucent display may be configured to selectively become opaque. A projection-based system can employ retinal projection technology that projects a graphical image onto a person's retina. The projection system may also be configured to project virtual objects into the physical environment, for example, as holograms or onto a physical surface.

[0046] Figure 2 shows an exemplary environment 200 of the surface of a two-dimensional manifold for visualizing the parameterization of a user's facial expression according to some implementations. In particular, environment 200 shows a parameterized image 220 of a user's (e.g., user 25 of FIG. 1) facial expression. For example, the feature parameterization instruction set can obtain live image data of the user's face (e.g., image 210) and parameterize different points on the face based on a surface of a shape such as a cylindrical shape 215. In other words, the feature parameterization instruction set can generate a set of values representing the 3D shape and appearance of the user's face at a certain point in time, which is used to generate a user representation (e.g., an avatar). In some implementations, a surface having a non-planar shape (e.g., cylindrical shape 215) can be used to reduce distortion. The set of values includes depth values that define the depth of portions of the face for points within a grid on a partially cylindrical surface, such as an array of points 225 (e.g., vector arrows pointing towards the face of the user representation to represent depth values, similar to a height field or height map). The parameterization values may include fixed parameters such as light ray position, end points, direction, etc., and the parameterization values may also include changing parameters such as depth, color, texture, opacity, etc., which are updated with the live image data. For example, as shown within the extended portion 230 of the user's nose, the depth value of one point (e.g., point 232 at the tip of the user's nose) can be defined as being at a depth D1 behind the position of that point on the surface of the face portion, e.g., at a depth D1 along a light ray starting from that point and orthogonal to it.

[0047] The techniques described herein use depth values that are different from the depth values in an existing RGBDA image (e.g., a red-green-blue-depth-alpha image). This is because the RGBDA image defines the content depth for a single camera location, while the techniques described herein define depth as a face portion for multiple points on the surface of a planar shape (e.g., a cylindrical shape). A curved surface such as the cylindrical shape 215 implemented for the parameterized image 220 is used to reduce the distortion of a user representation (e.g., an avatar) in areas of the user representation that are not visible from a flat projection plane. In some implementations, the planar projection plane can be bent and shaped in any way to reduce distortion in a desired area based on the application of the parameterization. The use of different bending / curving shapes allows the user representation to be rendered clearly from more viewpoints. Additional examples of planar shapes and parameterized images are shown in FIGS. 3A-3E.

[0048] FIGS. 3A-3C show different examples of the surface of a 2D manifold for visualizing the parameterization of a user's face representation according to some implementations. In particular, FIG. 3A shows a plane such as a cylindrical surface similar to FIG. 2, but oriented around a different axis. For example, FIG. 3A shows a parameterized image 314 of a user's face representation (e.g., image 310) that includes an array of points (e.g., vector arrows pointing towards the face of the user's representation) on a cylindrical surface / shape (e.g., cylindrical shape 314) that is curved around the x-axis. FIG. 3B shows a parameterized image 326 of a user's face (e.g., image 320) that includes an array of points that includes a hemispherical surface / shape (e.g., hemispherical shape 324) as a 2D manifold.

[0049] Figures 2, 3A, and 3B show points (e.g., equidistant vector arrows pointing towards the face of a user's representation) on a surface (e.g., the surface of a 2D manifold) that are arranged at regular intervals along vertical and horizontal lines on the surface. In some implementations, the points may be unevenly distributed across the surface of the 2D manifold, such as not being regularly spaced along the vertical and horizontal grid lines around the surface, but may be focused on specific area(s) of the user's face. For example, some areas can have more points where there may be more detail / movement in the facial structure, and some points can have fewer points in areas with less detail / movement, such as the forehead (less detail) and nose (less movement). For example, as shown in FIG. 3C, a parameterized image 330 including an array of points 332 that are in a cylindrical surface / shape includes an area 334 showing a higher density of points around the eyes, and an area 336 showing a higher density of points around the mouth. For example, when generating a user's representation (e.g., generating an avatar) during a communication session, the techniques described herein can generate a more accurate representation of a person during the communication session because they are more selectively focused on the areas of the eyes and mouth that are more likely to move during the conversation. For example, the techniques described herein can render updates to the user representation around the mouth and eyes at a faster frame rate than other parts of the face that do not move as much during the conversation (e.g., the forehead, ears, etc.).

[0050] FIG. 4 shows an example of generating and displaying a portion of a user's facial expression according to some implementation forms. In particular, FIG. 4 shows an exemplary environment 400 for a process of generating user representation data 430 (e.g., avatar 435) by combining registered image data 410 and live image data 420. The registered image data 410 shows an image of a user (e.g., user 25 in FIG. 1) during the registration process. For example, the registered anthropomorphic representation may be generated when the system acquires image data (e.g., RGB image) of the user's face while the user provides different facial expressions. For example, the user may be told "raise your eyebrows", "smile", "frown", etc. to provide the system with the range of facial features for the registration process. A registered anthropomorphic preview may be shown to the user while the user provides a registered image to obtain a visualization of the state of the registration process. In this example, the registered image data 410 shows registered anthropomorphic representations with four different user expressions, but more or fewer different expressions may be utilized to obtain sufficient data for the registration process. The live image data 420 represents an example of an image of the user acquired while using the device, such as during an XR experience (e.g., live image data while using device 10 in FIG. 1 such as an HMD). For example, the live image data 420 represents an image acquired while the user wears device 10 in FIG. 1 as an HMD. For example, when device 10 is an HMD, in one implementation form, sensor 35b may be disposed inside the HMD to capture pupil data (e.g., gaze characteristic data 40b), and additional sensors (e.g., sensors 35a and 35c) may be disposed on the outer surface of the HMD that is on the HMD but facing the user's head / face to capture facial feature data (e.g., upper facial feature characteristic data 40a via sensor 35a and lower facial feature characteristic data 40c via sensor 35c).

[0051] User representation data 430 is an exemplary illustration of a user during an avatar display process. For example, (landscape) avatar 435A and front-facing avatar 435B are generated based on acquired registration data and are updated when the system acquires and analyzes real-time image data to update different values for the plane (e.g., the values of the vector points of array 225 are updated for each acquired live image data).

[0052] FIG. 5 is a system flow diagram of an exemplary environment 500 in which a system can generate a representation of a user's face based on parameterized data, according to some implementations. In some implementations, the system flow of exemplary environment 500 is executed on a device such as a mobile device, desktop, laptop, or server device (e.g., device 10 of FIG. 1). The image of exemplary environment 500 can be displayed on a device having a screen for displaying the image and / or a screen for viewing a stereoscopic image such as a head-mounted device (HMD) (e.g., device 10 of FIG. 1). In some implementations, the system flow of exemplary environment 500 is executed on processing logic including hardware, firmware, software, or a combination thereof. In some implementations, the system flow of exemplary environment 500 is executed on a processor that executes code stored in a non-transitory computer-readable medium (e.g., memory).

[0053] In some implementations, the system flow of exemplary environment 500 includes a registration process, a feature tracking / parameterization process, and an avatar display process. Alternatively, exemplary environment 500 may include only the feature tracking / parameterization process and the avatar display process and may obtain registration data from another source (e.g., previously stored registration data). In other words, there may be cases where the registration process has already been performed such that the user's registration data has already been provided because the registration process has already been completed.

[0054] The system flow of the registration process for the exemplary environment 500 acquires image data (e.g., RGB data) from sensors in the physical environment (e.g., the physical environment 105 of FIG. 1) and generates registration data. The registration data can include most of the texture, muscle activity, etc., even if not the entire face of the user. In some implementations, the registration data may be captured while different instructions for acquiring different poses of the user's face are provided to the user. For example, the user may be told "raise your eyebrows", "smile", "frown", etc. to provide the system with the range of facial features for the registration process.

[0055] The system flow of the avatar display process for the exemplary environment 500 acquires image data (e.g., RGB, depth, IR, etc.) from sensors in the physical environment (e.g., the physical environment 105 of FIG. 1), determines parameterized data of facial features, acquires and evaluates registration data, and generates and displays a portion of the representation of the user's face (e.g., a 3D avatar) based on the parameterized values. For example, generating and displaying a portion of the facial representation of the user technology described herein can be implemented on real-time sensor data streamed to the end user (e.g., a 3D avatar overlaid on an image of the physical environment within a CGR environment). In an exemplary implementation, the avatar display process occurs during real-time display (e.g., the avatar is updated in real-time when the user makes a facial gesture and changes their own facial features). Alternatively, the avatar display process may occur while analyzing the streamed image data (e.g., generating a 3D avatar for a person from a video).

[0056] In an exemplary implementation, environment 500 includes an image synthesis pipeline that acquires or obtains data of a physical environment (e.g., image data from image source(s) such as sensors 512A - 512N). An exemplary environment 500 is an example of acquiring image sensor data (e.g., light intensity data - RGB) for a registration process to generate registration data (e.g., image data 524 of different expressions), and acquiring image sensor data 515 (e.g., light intensity data, depth data, and position information) for a parameterization process for a plurality of image frames. For example, FIG. 506 (e.g., exemplary environment 100 of FIG. 1) represents a user acquiring image data when scanning their face and facial features in a physical environment (e.g., physical environment 105 of FIG. 1) during a registration process. Image(s) 516 represents a user acquiring image data when scanning their face and facial features in real - time (e.g., during a communication session). Image sensor(s) 512A, 512B - 512N (hereinafter referred to as sensor 512) may include a depth camera that acquires depth data, a light intensity camera (e.g., an RGB camera) that acquires light intensity image data (e.g., a sequence of RGB image frames), and a position sensor that acquires positioning information.

[0057] In the case of positioning information, some implementations include a Visual Inertial Odometry (VIO) system that uses continuous camera images (e.g., light intensity data) to estimate the movement distance and determine equivalent odometry information. Alternatively, some implementations of the present disclosure may include a SLAM system (e.g., a position sensor). The SLAM system may include a multi-dimensional (e.g., 3D) laser scan and distance measurement system that provides real-time simultaneous localization and mapping without relying on GPS. The SLAM system can generate and manage very accurate point cloud data resulting from the reflection of laser scans from objects in the environment. The movement of any of the points within the point cloud is accurately tracked over time, and as a result, the SLAM system can use the points within the point cloud as reference points for position and maintain an accurate understanding of its position and orientation while moving through the environment. The SLAM system may further be a Visual SLAM system that depends on light intensity image data to estimate the position and orientation of the camera and / or the device.

[0058] In an exemplary implementation, the environment 500 includes a registration instruction set 520 composed of instructions executable by a processor to generate registration data from sensor data. For example, the registration instruction set 520 acquires image data 506 such as light intensity image data (e.g., an RGB image from a light intensity camera) from the sensor and generates registration data 522 of the user (e.g., face feature data such as texture, muscle activity, etc.). For example, the registration instruction set generates registration image data 524 (e.g., the registration image data 410 of FIG. 4).

[0059] In an exemplary implementation, environment 500 includes a set of feature parameterization instructions 530 configured with instructions executable by a processor to generate, from live image data (e.g., sensor data 515), a set of values representing the 3D shape and appearance of a user's face at a given point in time (e.g., appearance values 534, depth values 536, etc.). For example, the set of feature parameterization instructions 530 obtains sensor data 515 from sensor 512, such as light intensity image data (e.g., a live camera feed such as RGB from a light intensity camera), depth image data (e.g., depth image data from a depth camera such as an infrared or time-of-flight sensor), and other sources of physical environment information (e.g., position and orientation data from a position sensor, such as camera positioning information such as pose data), for a user within a physical environment (e.g., user 25 within physical environment 105 of FIG. 1), and generates parameterization data 532 for face parameterization (e.g., muscle activity, geometric shape, latent space of facial expressions, etc.). For example, the parameterization data 532 can be represented by a parameterized image 538 (e.g., changing parameters such as appearance values such as texture data, color data, opacity, and depth values 536 at different points on the face based on sensor data 515). Face parameterization for the set of feature parameterization instructions 530 can include taking a partial view obtained from sensor data 515 and determining a small set of parameters (e.g., facial muscles) from a geometric model to update the user representation. For example, the geometric model may include a set of data for eyebrows, eyes, cheeks under the eyes, mouth area, jaw area, etc. The parameterization tracking of the set of parameterization instructions 530 can provide the geometric shape of the user's facial features.

[0060] In an exemplary implementation, environment 500 includes a set of representation instructions 540 configured with instructions executable by a processor to generate a representation of a user's face (e.g., a 3D avatar) based on parameterized data 532. Additionally, the set of representation instructions 540 is configured with instructions executable by a processor to display a portion of the representation based on corresponding values when the corresponding values are updated with live image data. For example, the set of representation instructions 540 obtains parameterized data 532 (e.g., appearance value 534 and depth value 536) from the set of feature parameterization instructions 530 and generates representation data 542 (e.g., a real-time representation of the user such as a 3D avatar). For example, the set of representation instructions 540 can generate a representation 544 (e.g., avatar 435 of FIG. 4).

[0061] In some implementations, the set of representation instructions 540 obtains texture data directly from sensor data (e.g., RGB, depth, etc.). For example, the set of representation instructions 540 first obtains image data 506 from sensor(s) 512 and / or sensor data 515 from sensor 512 to obtain texture data for generating a representation 544 (e.g., avatar 435 of FIG. 4), and then can update the user representation 544 based on updated values (parameterized data 532) obtained from the set of feature parameterization instructions 530.

[0062] In some implementations, the set of representation instructions 540 provides real-time inpainting. To process real-time inpainting, the set of representation instructions 540 utilizes registration data 522 to assist in filling in a representation (e.g., representation 544) when a device identifies a specific expression (e.g., via geometric matching) that matches the registration data. For example, part of the registration process may include registering the user's teeth when the user smiles. Thus, when the device identifies that the user is smiling in a real-time image (e.g., sensor data 515), the set of representation instructions 540 inpaints the user's teeth from the user's registration data.

[0063] In some implementations, the process for real-time inpainting of the expression instruction set 540 is provided by a machine learning model (e.g., a trained neural network) to identify patterns in the textures (or other features) within the registration data 522 and the parameterized data 532. Additionally, the machine learning model may be used to match the pattern to a learned pattern corresponding to the user 25, such as smiling, frowning, conversing, etc. For example, if it is determined that the smiling pattern shows teeth (e.g., geometric matching as described herein), there may also be a determination of other parts of the face that change in a similar manner to the user when smiling (e.g., cheek movement, eyebrows, etc.). In some implementations, the techniques described herein can learn patterns specific to a particular user 25.

[0064] In some implementations, the expression instruction set 540 may be repeated for each frame captured during each instance / frame of a live communication session or other experience. For example, for each iteration, while the user is using the device (e.g., wearing an HMD), the exemplary environment 500 may involve continuously acquiring parameterized data 532 (e.g., appearance values 534 and depth values 536) and updating the displayed portion of the representation 544 based on the updated values for each frame. For example, for each new frame of parameterized data, the system can update the display of the 3D avatar based on the new data.

[0065] FIG. 6 is a flowchart illustrating an exemplary method 600. In some implementations, a device (e.g., device 10 of FIG. 1) executes the techniques of method 200 to provide a set of values for depicting a user. In some implementations, the techniques of method 600 are executed on a mobile device, desktop, laptop, HMD, or server device. In some implementations, method 600 is executed on processing logic that includes hardware, firmware, software, or combinations thereof. In some implementations, method 600 is executed on a processor that executes code stored in a non-transitory computer-readable medium (e.g., memory).

[0066] At block 602, method 600 obtains sensor data of a user. For example, sensor data (e.g., live data such as video content including light intensity data (RGB) and depth data) is associated with a particular point in time, such as an image from an inward-facing / downward-facing sensor while the user is wearing an HMD (e.g., sensors 35a, 35b, 35c shown in FIG. 1) associated with a frame. In some implementations, the sensor data includes depth data (e.g., infrared, time-of-flight, etc.) and light intensity image data obtained during a scan process.

[0067] In some implementations, obtaining sensor data may include obtaining a first set of data (e.g., enrollment data) corresponding to a user's facial features (e.g., texture, muscle activity, shape, depth, etc.) in multiple configurations from a device (e.g., the enrolled image data 410 of FIG. 4). In some implementations, the first set of data includes unobstructed image data of the user's face. For example, an image of the face may be captured while the user is smiling, raising their eyebrows, puffing out their cheeks, etc. In some implementations, the enrollment data may be obtained by the user removing the device (e.g., HMD) and the device capturing an image without obstructing the face, or by using another device (e.g., a mobile device) that does not obstruct the face. In some implementations, the enrollment data (e.g., the first set of data) is obtained from light intensity images (e.g., RGB image(s)). The enrollment data may include most of the texture, muscle activity, etc., if not all, of the user's face. In some implementations, the enrollment data may be captured while different instructions for obtaining different poses of the user's face are provided to the user. For example, the user may be told by a user interface guide to "raise your eyebrows", "smile", "frown", etc. to provide the system with the range of facial features for the enrollment process.

[0068] In some implementations, obtaining sensor data may include obtaining a second set of data corresponding to one or more partial views of a face from one or more image sensors while the user is using (e.g., wearing) an electronic device (e.g., an HMD). For example, obtaining sensor data includes live image data 420 of FIG. 4. In some implementations, the second set of data may include a partial image of the user's face and thus may not represent all of the facial features represented in the enrollment data. For example, the second set of images may include some images of the forehead / inter-brow area (e.g., facial feature characteristic data 40a) from an upward-facing sensor (e.g., sensor 35a of FIG. 1). Additionally, or alternatively, the second set of images may include some images of the eyes (e.g., gaze characteristic data 40b) from an inward-facing sensor (e.g., sensor 35a of FIG. 1). Additionally, or alternatively, the second set of images may include some images of the cheeks, mouth, and jaw (e.g., facial feature characteristic data 40c) from a downward-facing sensor (e.g., sensor 35c of FIG. 1). In some implementations, the electronic device includes a first sensor (e.g., sensor 35a of FIG. 1) and a second sensor (e.g., sensor 35c of FIG. 1), and the second set of data is obtained from at least one partial image of the user's face from the first sensor from a first perspective (e.g., upper facial feature data 40a) and from at least one partial image of the user's face from the second sensor from a second perspective different from the first perspective (e.g., a plurality of IFC cameras that capture different perspectives of the user's face and body movements) (e.g., lower facial feature data 40c).

[0069] In block 604, method 600 generates a set of values representing a user based on sensor data, the set of values including i) depth values defining the 3D position of a portion of the user relative to the 3D positions of points on the projection plane, and ii) appearance values defining the appearance of the portion of the user. For example, generating a set of values representing a user based on sensor data (e.g., RGB values, alpha values, and depth values - RGBDA) may involve using both live sensor data from an inward-facing / downward-facing camera and registered data, such as images of faces with different expressions without an HMD being worn. In some implementations, generating the set of values may involve using a machine learning model trained to generate the set of values.

[0070] The set of values may include depth values defining the 3D position of a portion of the user relative to the 3D positions of points on the projection plane. For example, the depth value of one point can define that the facial portion is at a depth D1 behind the position of that point on the surface, e.g., at a depth D1 along a ray of light (e.g., ray 232 in FIG. 2) starting from that point. In some implementations, the depth value defines the distance between a portion of the user and the corresponding point on the projection plane positioned along a ray perpendicular to the projection plane at the position of the corresponding point. The techniques described herein use depth values different from the depth values in an existing RGBDA image that define the content depth for a single camera location. Appearance values may include values such as RGB data and alpha data that define the appearance of a portion of the user. For example, the appearance values may include color, texture, opacity, and the like.

[0071] In some implementations, the term "surface" refers to a 2D manifold that can be planar or non-planar. In some implementations, points on a surface (e.g., the surface of a 2D manifold) are spaced apart at regular intervals along vertical and horizontal lines on the surface. In some implementations, the points are regularly spaced apart along vertical and horizontal grid lines on a partially cylindrical surface, as shown in FIG. 2. Alternatively, other planar and non-planar surfaces may be utilized. For example, as shown in FIG. 3A, the plane of the cylindrical surface may be oriented / curved around a different axis (e.g., the y-axis in FIG. 2, the x-axis in FIG. 3A, etc.). Additionally, or alternatively, the plane may be a hemispherical manifold as shown by the parameterized image 326 in FIG. 3B.

[0072] In some implementations, the points may be unevenly distributed across the surface of the 2D manifold, such as not being regularly spaced apart along vertical and horizontal grid lines around the surface, but may be focused on a particular area(s) of the user's face. For example, some areas can have more points where there may be more detail / movement in the facial structure, and some points can have fewer points in areas with less detail / movement, such as the forehead (less detail) and nose (less movement). For example, as shown in FIG. 3C, area 334 shows a higher density of points around the eyes, and area 336 shows a higher density of points around the mouth.

[0073] In some implementations, a set of values is generated based on an alignment such that a subset of points on the central area of the surface corresponds to the central part of the user's face. For example, as shown in FIG. 2, the focused area of the user's nose in area 230, the feature point of ray 232 is the tip of the person's nose.

[0074] In some implementations, generating a set of values is further based on an image of the user's face captured while the user is expressing multiple different facial expressions. For example, the set of values is determined based on registered images of the face such as while the user is smiling, raising their eyebrows, puffing out their cheeks, etc. In some implementations, the sensor data corresponds only to a first area of the user (e.g., an area not blocked by a device such as an HMD), and the set of image data (e.g., registration data) corresponds to a second area including a third area different from the first area. For example, the second area may include some of the areas blocked by the HMD when the HMD is worn by the user.

[0075] In some implementations, determining user-specific parameterization (e.g., generating a set of values) may be adapted to each particular user. For example, the parameterization may be fixed based on registration identification information (e.g., to better cover the person's head size or nose shape), or the parameterization may be based on the current facial expression (e.g., where the parameterization may become longer when the mouth is open). In an exemplary implementation, method 600 may further include obtaining additional sensor data of the user associated with a second period, updating a set of values representing the user based on the additional sensor data of the second period, and providing the updated set of values, and the depiction of the user is updated in the second period based on the updated set of values (e.g., updating the set of values based on the current facial expression such that the parameterization also becomes longer when the mouth is open).

[0076] In some implementations, generating a set of values representing a user is based on a machine learning model trained to generate a set of values. For example, the process for generating parameterized data for the feature parameterization instruction set 530 is provided by a machine learning model (e.g., a trained neural network) to identify patterns in textures (or other features) within the registration data 522 and sensor data 515 (such as live image data like the image 516). Further, the machine learning model may be used to match the pattern to a learned pattern corresponding to the user 25, such as smiling, frowning, conversing, etc. For example, if it is determined that the smiling pattern shows teeth, there may also be a determination of other parts of the face that change in the same way as the user when smiling (e.g., cheek movement, eyebrows, etc.). In some implementations, the techniques described herein can learn patterns specific to the particular user 25 of FIG. 1.

[0077] In block 606, the method 600 provides a set of values, and a depiction of the user at that point in time is displayed based on the set of values. For example, the set of points is a frame of 3D video data transmitted during a communication session with another device, and the other device uses a set of values such as RGBDA information (along with information on how to interpret the depth values) to render a view of the user's face. In some implementations, consecutive frames of face data (a set of values representing the 3D shape and appearance of the user's face at different points in time) may be transmitted and used to display a live 3D video-like face depiction (e.g., a realistically moving avatar). In some implementations, the depiction of the user (e.g., an avatar shown to the second user on the display of the second user's device) is displayed in real time.

[0078] In some implementations, providing a set of values includes transmitting a sequence of frames of 3D video data including a frame having the set of values during a communication session with a second device, the second device rendering an animated depiction of a user based on the sequence of frames of 3D video data. For example, a set of points may be frames of 3D video data transmitted during a communication session with another device, and the other device uses the set of values (along with information on how to interpret depth values) to render a view of the user's face. Additionally, or alternatively, successive frames of face data (sets of values representing the 3D shape and appearance of the user's face at different points in time) may be transmitted and used to display a live 3D video-like depiction of the face.

[0079] In some implementations, a user's depiction may include sufficient data to enable a stereoscopic view of the user (e.g., left eye / right eye views) so that the face can be perceived with depth. In one implementation, a face depiction includes a 3D model of the face, as well as views of the representation from the left eye position and the right eye position, and is generated to provide a stereoscopic view of the face.

[0080] In some implementations, certain parts of the face that may be important for conveying a realistic or faithful appearance, such as the eyes and mouth, may be generated differently from other parts of the face. For example, parts of the face that may be important for conveying a realistic or faithful appearance may be based on current camera data, while other parts of the face may be based on previously acquired (e.g., registered) face data.

[0081] In some implementations, facial expressions are generated using textures, colors, and / or geometric shapes of various facial parts, and based on the depth values and appearance values of each frame of data, an estimated value of the confidence level of the generation technique that such textures, colors, and / or geometric shapes exactly correspond to the actual textures, colors, and / or geometric shapes of those facial parts is identified. In some implementations, the depiction is a 3D avatar. For example, the representation is a 3D model representing a user (e.g., user 25 in FIG. 1).

[0082] In some implementations, method 600 may be repeated for each frame captured during each instance / frame of a live communication session or other experience. For example, for each iteration, while the user is using the device (e.g., wearing an HMD), method 600 may involve continuously acquiring live sensor data (e.g., gaze characteristic data and facial feature data) and, for each frame, updating the displayed portion of the representation based on the updated parameterized values (e.g., RGBDA values). For example, for each new frame, the system can update the parameterized values and update the display of the 3D avatar based on the new data.

[0083] In some implementations, an estimator or statistical learning method is used to better understand or make predictions regarding physiological data (e.g., facial features and gaze characteristic data). For example, statistics about the characteristic data of gaze and facial features can be estimated by sampling the data set with replacement data (e.g., the bootstrap method).

[0084] FIG. 7 is a block diagram of an exemplary device 700. Device 700 shows an exemplary device configuration of device 10. Although certain features are shown, those skilled in the art will understand from this disclosure that various other features are not shown for the sake of brevity so as not to obscure more suitable aspects of the implementations disclosed herein. For that purpose, by way of non-limiting example, in some implementations, device 10 may include one or more processing units 702 (e.g., microprocessor, ASIC, FPGA, GPU, CPU, processing core, etc.), one or more input / output (I / O) devices and sensors 706, one or more communication interfaces 708 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, infrared, BLUETOOTH, ZIGBEE, SPI, I2C or similar types of interfaces), one or more programming (e.g., I / O) interfaces 710, one or more displays 712, one or more internal and / or external facing image sensor systems 714, memory 720, and one or more communication buses 704 for interconnecting these and various other components to each other.

[0085] In some implementations, one or more communication buses 704 include circuitry for interconnecting system components and controlling communication. In some implementations, one or more I / O devices and sensors 706 may include at least one of an inertial measurement unit (IMU), accelerometer, magnetometer, gyroscope, thermometer, one or more physiological sensors (e.g., blood pressure monitor, heart rate monitor, blood oxygen sensor, blood glucose sensor, etc.), one or more microphones, one or more speakers, a tactile engine, one or more depth sensors (e.g., structured light, time of flight, etc.).

[0086] In some implementations, one or more displays 712 are configured to present a view of the physical or graphic environment to the user. In some implementations, one or more displays 712 correspond to holographic, digital light processing (DLP), liquid-crystal display (LCD), liquid-crystal on silicon (LCoS), organic light-emitting field-effect transitory (OLET), organic light-emitting diode (OLED), surface-conduction electron-emitter display (SED), field-emission display (FED), quantum-dot light-emitting diode (QD-LED), micro-electro-mechanical system (MEMS), and / or similar display types. In some implementations, one or more displays 712 correspond to waveguide displays, such as diffractive, reflective, polarizing, holographic, etc. In one example, device 10 includes a single display. In another example, device 10 includes a display for each eye of the user.

[0087] In some implementations, one or more image sensor systems 714 are configured to obtain image data corresponding to at least a portion of the physical environment 105. Examples of the one or more image sensor systems 714 include one or more RGB cameras (e.g., those equipped with complementary metal-oxide semiconductor (CMOS) image sensors or charge-coupled device (CCD) image sensors), monochrome cameras, infrared cameras, depth cameras, event-based cameras, and the like. In various implementations, the one or more image sensor systems 714 further include an illumination light source that emits light such as a flash. In various implementations, the one or more image sensor systems 714 further include an on-camera image signal processor (ISP) configured to perform a plurality of processing operations on the image data.

[0088] Memory 720 includes high-speed random access memory such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some implementations, memory 720 includes non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile storage devices. Memory 720 optionally includes one or more storage devices located remotely from the one or more processing units 702. Memory 720 includes a non-transitory computer-readable storage medium.

[0089] In some implementations, memory 720 or the non-transitory computer-readable storage medium of memory 720 stores an optional operating system 730 and one or more instruction sets 740. The operating system 730 includes procedures for handling various basic system services and procedures for performing hardware-dependent tasks. In some implementations, the instruction set(s) 740 includes executable software defined by binary information stored in the form of electrical charges. In some implementations, the instruction set(s) 740 is software executable by the one or more processing units 702 to perform one or more of the techniques described herein.

[0090] The command set(s) 740 includes a registration command set 742, a feature parameterization command set 744, and a rendering command set 746. The command set(s) 740 may be embodied as a single software execution file or multiple software execution files.

[0091] In some implementations, the registration command set 742 is executable by the processing unit(s) 702 to generate registration data from the image data. The registration command set 742 (e.g., the registration command set 520 of FIG. 5) may be configured to obtain image information for generating a registration anthropomorphization (e.g., registration image data 524) and to provide instructions to the user to determine whether additional image information is needed to generate an accurate registration anthropomorphization for use by the avatar display process. For that purpose, in various implementations, the instructions include instructions and / or logic therefor, and heuristics and metadata therefor.

[0092] In some implementations, the feature parameterization command set 744 (e.g., the feature parameterization command set 530 of FIG. 5) is executable by the processing unit(s) 702 to parameterize the user's facial features and gaze characteristics (e.g., generate appearance values and depth values) by using one or more of the techniques described herein or otherwise as may be appropriate. For that purpose, in various implementations, the instructions include instructions and / or logic therefor, and heuristics and metadata therefor.

[0093] In some implementations, the feature representation instruction set 746 (e.g., the representation instruction set 540 of FIG. 5) is executable by a processing unit(s) 702 to generate and display a representation of a user's face (e.g., a 3D avatar) based on a first set of data (e.g., registration data) and a second set of data (e.g., parameterized data), and portions of the representation correspond to different parameterized values. For that purpose, in various implementations, the instructions include instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0094] The instruction set(s) 740 is shown as being present on a single device, but it should be understood that in other implementations, any combination of elements may be arranged within separate computing devices. Further, FIG. 7 is not a structural overview of the implementations described herein, but rather is more intended to illustrate the functions of the various features present in a particular implementation. As will be recognized by those of ordinary skill in the art, the separately shown items may be combined, and some items may be separated. The actual number of instruction sets, and how features are allocated among them, may vary from implementation to implementation and may depend in part on the particular combination of hardware, software, and / or firmware selected for a particular implementation.

[0095] FIG. 8 shows a block diagram of an exemplary head-mounted device 800 according to some implementations. The head-mounted device 800 includes a housing 801 (or enclosure) that houses various components of the head-mounted device 800. The housing 801 includes (or is coupled to) an eye pad (not shown) disposed at a proximal end of the housing 801 (with respect to user 25). In some implementations, the eye pad is a plastic or rubber piece that holds the head-mounted device 800 comfortably and snugly in place on the face of user 25 (e.g., in a position surrounding the eyes of user 25).

[0096] The housing 801 can accommodate a display 810 that displays an image and emits light towards or onto the eyes of the user 25. In various implementations, the display 810 emits light through an eyepiece having one or more optical elements 805 that refract the light emitted by the display 810, thereby making the display appear to the user 25 to be at a virtual distance greater than the actual distance from the eyes to the display 810. For example, the optical element(s) 805 can include one or more lenses, waveguides, other diffractive optical elements (DOEs), etc. In various implementations, the virtual distance is greater than at least the minimum focal distance of the eyes (e.g., 7 cm) so that the user 25 can focus on the display 810. Further, in various implementations, the virtual distance is greater than 1 meter to provide a better user experience.

[0097] The housing 801 also houses a tracking system that includes one or more light sources 822, a camera 824, cameras 832, 834, and a controller 880. The one or more light sources 822 emit light onto the eyes of the user 25 that reflects as a light pattern (e.g., a circle of glints) that can be detected by the camera 824. Based on the light pattern, the controller 880 can determine the eye tracking characteristics of the user 25. For example, the controller 880 can determine the line of sight direction and / or blink state (open or closed eyes) of the user 25. As another example, the controller 880 can determine the pupil center, pupil size, or fixation point. Thus, in various implementations, light is emitted by the one or more light sources 822, reflected by the eyes of the user 25, and detected by the camera 824. In various implementations, the light from the eyes of the user 25 is reflected by a hot mirror or passes through an eyepiece before reaching the camera 824.

[0098] The display 810 emits light in a first wavelength range, and one or more light sources 822 emit light in a second wavelength range. Similarly, the camera 824 detects light in the second wavelength range. In various implementations, the first wavelength range is the visible wavelength range (e.g., a wavelength range within the visible spectrum of approximately 400 - 700 nm), and the second wavelength range is the near-infrared wavelength range (e.g., a wavelength range within the near-infrared spectrum of approximately 700 - 1400 nm).

[0099] In various implementations, eye tracking (or specifically, a predetermined line of sight direction) is used to enable user interaction (e.g., the user 25 selects by looking at options on the display 810), provide foveated rendering (e.g., presented at a higher resolution within the area of the display 810 that the user 25 is looking at and at a lower resolution elsewhere on the display 810), or correct distortion (e.g., with respect to an image provided on the display 810).

[0100] In various implementations, one or more light sources 822 emit light towards the eyes of the user 25 that reflects in the form of a plurality of glints.

[0101] In various implementations, the camera 824 is a frame / shutter-based camera that generates an image of the eyes of the user 25 at a particular point in time or multiple points in time at a frame rate. Each image includes a matrix of pixel values corresponding to the pixels of the image, which correspond to the positions of the matrix of light sensors of the camera. In an implementation, each image is used to measure or track pupil dilation by measuring changes in pixel intensity associated with one or both of the user's pupils.

[0102] In various implementations, the camera 824 is an event camera that includes a plurality of light sensors (e.g., a matrix of light sensors) at a plurality of corresponding positions that generate an event message indicating a particular position of a particular light sensor in response to a change in the intensity of light detected by the particular light sensor.

[0103] In various implementations, cameras 832 and 834 are frame / shutter-based cameras that can generate images of the face of user 25 at a particular point in time or multiple points in time at a certain frame rate. For example, camera 832 captures an image of the user's face below the eyes, and camera 834 captures an image of the user's face above the eyes. The images captured by cameras 832 and 834 may include light intensity images (e.g., RGB) and / or depth image data (e.g., time-of-flight, infrared, etc.).

[0104] It should be understood that the above-described implementations are presented as examples, and the present invention is not limited to what has been specifically illustrated and described above. Rather, its scope includes both the combinations and sub-combinations of the various features described above, as well as those variations and modifications thereof that would occur to one of ordinary skill in the art upon reading the foregoing description and that are not disclosed in the prior art.

[0105] As described above, one aspect of the present technology is the collection and use of physiological data to improve the user experience of an electronic device with respect to interacting with electronic content. The present disclosure contemplates that, in some examples, the collected data may include personal information data that can be used to uniquely identify a particular person or to identify the interests, characteristics, or tendencies of a particular person. Such personal information data can include physiological data, demographic data, location-based data, phone numbers, email addresses, home addresses, device characteristics of personal devices, or any other personal information.

[0106] The present disclosure recognizes that the use of such personal information data in the present technology can be a use that benefits the user. For example, personal information data can be used to improve the interaction and control capabilities of an electronic device. Thus, the use of such personal information data enables calculated control of the electronic device. Further, other uses of personal information data that benefit the user are also contemplated by the present disclosure.

[0107] The present disclosure further contemplates that entities responsible for the collection, analysis, disclosure, transfer, storage, or other use of such personal information and / or physiological data will comply with well-established privacy policies and / or privacy practices. Specifically, such entities should implement and consistently use privacy policies and practices that are generally recognized as meeting or exceeding industry or government requirements for maintaining the confidentiality of personal information data. For example, personal information from users should be collected for legitimate and proper use by the entity and should not be shared or sold except for such legitimate use. Further, such collection should only be carried out after informing and obtaining consent from the user. In addition, such entities will take all necessary measures to protect and secure access to such personal information data and ensure that others with access rights to such personal information data comply with their privacy policies and procedures. Further, such entities may subject themselves to evaluation by third parties to demonstrate their compliance with widely accepted privacy policies and practices.

[0108] Notwithstanding the above, the present disclosure also contemplates implementations in which the user can selectively block the use of personal information data or access to personal information data. That is, the present disclosure contemplates that hardware elements or software elements may be provided to prevent or block access to such personal information data. For example, in the case of a user-adaptive content delivery service, the technology can be configured to allow the user to select an "opt-in" or "opt-out" to participate in the collection of personal information data during service registration. In another example, the user can choose not to provide personal information data to the content delivery service in question. In yet another example, the user can choose not to provide personal information but allow the transfer of anonymous information for the purpose of improving the functionality of the device.

[0109] Accordingly, while the present disclosure encompasses the use of personal information data for implementing one or more various disclosed embodiments, the present disclosure also contemplates that it is possible to implement those various embodiments without the need to access such personal information data. That is, the various embodiments of the present technology are not rendered inoperable by the absence of all or a portion of such personal information data. For example, content may be selected and delivered to a user by inferring preferences or settings based on non-personal information data, such as content requested by a device associated with the user, other non-personal information available in a content delivery service, or publicly available information, or only a minimal amount of personal information.

[0110] In some embodiments, data is stored using a public / private key system that enables only the owner of the data to decrypt the stored data. In some other implementations, data may be stored anonymously (e.g., without identification information and / or personal information about the user, such as legal name, username, time, and location data). In this way, other users, hackers, or third parties cannot determine the identification information of the user associated with the stored data. In some implementations, a user may access the user's stored data from a user device different from the user device used to upload the stored data. In these examples, the user may need to provide login authentication information to access the stored data.

[0111] Numerous specific details are set forth herein to provide a complete understanding of the claimed subject matter. However, one of ordinary skill in the art will understand that the claimed subject matter may be practiced without these specific details. In other instances, well-known methods, devices, or systems have not been described in detail so as not to obscure the claimed subject matter.

[0112] Unless otherwise specified, throughout the description of this specification, the use of terms such as "processing", "computing", "calculating", "judging", and "identifying" is understood to refer to actions or processes of a computing device. A computing device is one or more computers or similar electronic computing devices (singular or plural) that operate on or transform data represented as physical electronic or magnetic quantities within the memory, registers, or other information storage devices, transmission devices, or display devices of a computing platform.

[0113] The system(s) discussed in this specification is not limited to any particular hardware architecture or configuration. A computing device can include any suitable arrangement of components that provides a result conditioned on one or more inputs. Suitable computing devices include general-purpose computing devices to dedicated computing devices that implement one or more implementations of the subject matter, including general-purpose microprocessor-based computer systems that access software stored to program or configure a computing system. Any suitable programming, scripting, or other type of language or combination of languages can be used to implement the teachings contained herein within software for use in programming or configuring a computing device.

[0114] Implementations of the methods disclosed in this specification may be executed in the operation of such computing devices. The order of the blocks presented in the examples above can be changed; for example, the blocks can be rearranged, combined, or divided into sub-blocks. Certain blocks or processes can be executed in parallel.

[0115] The use of "adapted as" or "configured as" in this specification means non-limiting and inclusive language that does not exclude devices adapted or configured to perform additional tasks or steps. Further, the use of "based on" means non-limiting and inclusive in that a process, step, calculation, or other action "based on" one or more recited conditions or values may in fact be based on additional conditions or values beyond the recited conditions or values. Headings, lists, and numbering included in this specification are for ease of explanation and are not intended to be limiting.

[0116] In this specification, terms such as "first", "second", etc. may be used to describe various objects, but it should also be understood that these objects should not be limited by these terms. These terms are only used to distinguish one object from another. For example, as long as the name is consistently changed for all occurrences of "first node" and the name is consistently changed for all occurrences of "second node", the first node can be called the second node and, similarly, the second node can be called the first node without changing the meaning of the description. The first node and the second node are both nodes, but they are not the same node.

[0117] The terms used in this specification are for the purpose of describing particular implementations and are not intended to limit the scope of the claims. When used in the description of the implementations described and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. Also, as used herein, the term "or" shall be understood to refer to any and all possible combinations of one or more of the associated listed items and to include the same. When the terms "comprises" or "comprising" are used herein, they specify the presence of the stated features, integers, steps, operations, objects, or components, but it is further understood that they do not exclude the presence or addition of one or more other features, integers, steps, operations, objects, components, or groups thereof.

[0118] As used herein, the term "if" can be interpreted to mean "when" or "upon" or "in response to a determination" or "in accordance with a determination" or "in response to a detection" that the previously described condition is true, depending on the context. Similarly, the phrases "when it is determined that [the previously described condition is true]", "when [the previously described condition is true]", or "when [the previously described condition is true]" can be interpreted to mean "upon determination", "in response to a determination", "in accordance with a determination", "upon detection", or "in response to a detection" that the previously described condition is true.

[0119] The foregoing description and summary of the invention are to be understood as illustrative and exemplary in all respects and not as limiting, and it is to be understood that the scope of the invention disclosed herein is to be determined not only from the detailed description of the exemplary implementations, but also in accordance with the full breadth permitted by patent law. It is to be understood that the implementations shown and described herein are merely illustrative of the principles of the invention and that various modifications may be made by those skilled in the art without departing from the scope and spirit of the invention.

Claims

1. 1. A method comprising: In a processor of the device, acquiring sensor data of a user, the sensor data being associated with a point in time; a set of values ​​representing the user based on the sensor data, the set of values ​​comprising: a depth value defining a three-dimensional (3D) position of a portion of the user relative to a plurality of 3D positions of points of a projection plane; and generating a set of values ​​including an appearance value defining an appearance of the portion of the user; providing a set of values, wherein a representation of the user at the time point is displayed based on the set of values; A method comprising:

2. The method of claim 1 , wherein the points are spaced apart at regular intervals along vertical and horizontal lines on the projection plane.

3. The method of claim 1 , wherein the projection surface is non-planar.

4. The method of claim 1 , wherein the projection surface is at least partially cylindrical.

5. The method of claim 1 , wherein the projection surface is a plane.

6. The method of claim 1 , wherein the set of values ​​is generated based on an alignment such that the subset of points on a central area of ​​the projection surface corresponds to a central portion of the user's face.

7. The method of claim 1 , wherein generating the set of values ​​is further based on images of the user's face captured while the user is exhibiting a plurality of different facial expressions.

8. the sensor data corresponds to only a first area of ​​the user; the set of image data corresponds to a second area including a third area different from the first area; The method according to claim 7.

9. acquiring additional sensor data for the user associated with a second time period; updating the set of values ​​representative of the user based on the additional sensor data for the second time period; providing the updated set of values, wherein the depiction of the user is updated during the second time period based on the updated set of values. The method of claim 1.

10. 2. The method of claim 1 , wherein providing the set of values ​​includes transmitting a sequence of frames of 3D video data including frames having the set of values ​​during a communication session with a second device, the second device rendering an animated depiction of the user based on the sequence of frames of 3D video data.

11. 2. The method of claim 1, wherein the device includes a first sensor and a second sensor, and the sensor data is obtained from at least one partial image of the user's face from the first sensor from a first viewpoint and from at least one partial image of the user's face from the second sensor from a second viewpoint different from the first viewpoint.

12. The method of claim 1 , wherein the representation of the user is displayed in real time.

13. The method of claim 1 , wherein generating the set of values ​​representing the user is based on a machine learning model trained to generate the set of values.

14. The method of claim 1 , wherein the depth value defines a distance between a portion of the user and the corresponding point on the projection surface positioned along a ray perpendicular to the projection surface at the location of the corresponding point.

15. The method of claim 1 , wherein the appearance value comprises a color value, a texture value, or an opacity value.

16. The method of claim 1 , wherein the device is a head-mounted device (HMD).

17. The method of claim 16 , wherein the HMD includes one or more inward-facing image sensors and one or more downward-facing image sensors, and the sensor data is captured by the one or more inward-facing sensors and the one or more downward-facing image sensors.

18. A device, comprising: a non-transitory computer readable storage medium; and one or more processors coupled to the non-transitory computer-readable storage medium, the non-transitory computer-readable storage medium including program instructions that, when executed on the one or more processors, cause the device to: acquiring sensor data of a user, the sensor data being associated with a point in time; a set of values ​​representing the user based on the sensor data, the set of values ​​comprising: a depth value defining a three-dimensional (3D) position of a portion of the user relative to a plurality of 3D positions of points of a projection plane; and generating a set of values ​​including an appearance value defining an appearance of the portion of the user; providing a set of values, wherein a representation of the user at the time point is displayed based on the set of values; A device that causes an action to be performed, including

19. 20. The device of claim 18, wherein the points are spaced apart at regular intervals along vertical and horizontal lines on the projection surface.

20. A non-transitory computer readable storage medium storing program instructions executable on a device to perform operations, the operations comprising: acquiring sensor data of a user, the sensor data being associated with a point in time; a set of values ​​representing the user based on the sensor data, the set of values ​​comprising: a depth value defining a three-dimensional (3D) position of a portion of the user relative to a plurality of 3D positions of points of a projection plane; and generating a set of values ​​including an appearance value defining an appearance of the portion of the user; providing a set of values, wherein a representation of the user at the time point is displayed based on the set of values; A non-transitory computer readable storage medium comprising:

Citation Information

Patent Citations

  • Avatar generation that reflects the player's appearance

    JP2014522536A

  • Recording and sending emojis

    JP2020520030A

  • Representation of users based on current user appearance

    WO2022066450A1