Gaussian splat for user representation
3D Gaussian splats and UV mapping enable accurate, real-time user representation on electronic devices, addressing the challenge of outdated image-based representations by reducing computational demands and improving appearance fidelity.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2026-04-08
AI Technical Summary
Existing electronic devices fail to accurately represent a user's current appearance, often relying on outdated images, leading to inaccuracies in real-time representations.
Utilizing 3D Gaussian splats and UV mapping to generate user representations, which require less computational resources and bandwidth while enabling more accurate, real-time rendering of high-quality photorealistic scenes.
Provides efficient, accurate, and resource-friendly real-time user representation by combining sparse image data with live updates, enhancing user appearance fidelity in virtual communication.
Smart Images

Figure 2026060944000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to electronic devices, and more particularly, to systems, methods, and devices for representing a user in computer-generated content.
Background Art
[0002] In existing technologies, the current (e.g., real-time) representation of a user's appearance on an electronic device may not accurately or faithfully represent it. For example, the device may provide a representation of the user based on an image of the user's face acquired minutes, hours, days, or even years ago. Such a representation may not accurately represent the user's appearance. Therefore, there may be a desire to provide a means for efficiently providing a more accurate, faithful, and / or current representation of the user.
Summary of the Invention
[0003] The various implementations disclosed herein include devices, systems, and methods for generating views of user representations based on three-dimensional (3D) Gaussian splats. Gaussian splats enable real-time rendering of high-quality photorealistic scenes from a sparse set of images. In particular, a first set of captured user data (e.g., registration data) can be used in a first device (e.g., a transmitting device) to generate user representation data including splat parameter data (e.g., a 23-channel Gaussian UV map). The user representation data may be modified based on live user data. Views of user representations can be provided to a viewing device (e.g., rendering a live view of the sender's persona) by generating splats corresponding to the modified user representation data. A persona is a representation of the user, such as an avatar. Advantageously, splats avoid the need to use meshes to avoid the appearance of missing parts and offer other advantages. 3D representations of a user at multiple points in time can be generated on a viewing device by combining data and using the combined data to render views, for example, during a live communication (e.g., virtual communication or coexistence) session.
[0004] In some implementations, the data associated with each splat of a user representation can represent texture / color, position, splat shape, transparency level, covariance (e.g., how the splat is stretched / scaled), semantics (e.g., hair, mouth, skin, glasses, accessories, and / or other features), etc. The splats can be arranged in a two-dimensional (2D) grid structure corresponding to surface parameterization. Parameterization can correspond to human faces for high-quality face reconstruction. Parameterization of the splat distribution can provide higher quality data, be faster for training machine learning models, and provide faster (e.g., real-time) rasterization. In some implementations, 3D mapping information (e.g., identifying x, y, z positions corresponding to UV coordinates in a UV map) can be generated at registration (e.g., a Gaussian UV map).
[0005] Using 3D Gaussian splats and UV mapping offers several advantages. For example, 3D Gaussian splats require less computation, resources, and bandwidth than using 3D meshes, 3D point clouds, etc., while enabling more accurate user representation.
[0006] Generally, one innovative aspect of the subject matter described herein can be embodied in a device processor by an action to acquire user representation data of at least a portion of a user, wherein the user representation data is based on a first set of sensor data including images of the user acquired during a registration process, and the user representation data includes splat parameter data corresponding to a plurality of three-dimensional (3D) positions; an action to modify the user representation data based on a second set of sensor data acquired after the registration process; and an action to provide a view of the user representation based on the modified user representation data, wherein the action to provide the view includes generating a plurality of splats based on the splat parameter data of the modified user representation data.
[0007] These and other embodiments may optionally include one or more of the following features:
[0008] In some embodiments, at least a portion of the user includes the user's face and additional portions. In some embodiments, the user representation data is based on a UV map and 3D point cloud points associated with distribution data that defines the size and shape for rendering 3D point cloud points as splats corresponding to each point of the UV map.
[0009] In some embodiments, the splat parameter data includes 3D Gaussian parameters for each 3D position. In some embodiments, the 3D Gaussian parameters include at least one of the following: position information, color information, covariance information, transparency information, orientation, opacity information, range information on each axis, rotation data, scale, and semantics information.
[0010] In some embodiments, the user representation data includes 3D mapping information, which includes feature values and positional information for each map point. In some embodiments, the user representation data is modified based on body posture data acquired during the registration process, during a communication session with another device, or a combination thereof.
[0011] In some embodiments, the device is the viewer's device, and the user representation data is modified based on an additional set of sensor data acquired during a communication session with the sender's device associated with the user representation.
[0012] In some embodiments, user representation data is generated and updated during the registration process based on images of the user's face captured while the user is making multiple different facial expressions.
[0013] In some embodiments, the technology generates user representation data via a machine learning model trained using training data acquired through one or more sensors in one or more environments.
[0014] In some embodiments, providing a view of a user representation based on modified user representation data includes displaying the user representation in an extended reality (XR) environment.
[0015] In some embodiments, the action further includes modifying the view of the user representation by adjusting the user representation based on at least one color attribute among several color attributes of the environment, at least one light attribute among several light attributes of the environment, or a combination thereof.
[0016] In some embodiments, user representation data is acquired in a first physical environment, and the user representation is displayed in a view of a second physical environment different from the first physical environment. In some embodiments, the user representation is a 3D user representation.
[0017] According to some implementations, a non-temporary computer-readable storage medium stores computer-executable instructions that perform or cause execution of any of the methods described herein. According to some implementations, the device includes one or more processors, non-temporary memory, and one or more programs. The one or more programs are stored in the non-temporary memory and are configured to be executed by one or more processors, and the one or more programs include instructions that perform or cause execution of any of the methods described herein. [Brief explanation of the drawing]
[0018] This disclosure may have a more detailed description by reference to several exemplary implementations, some of which are shown in the accompanying drawings, as can be understood by those skilled in the art.
[0019] [Figure 1] This diagram shows a device that acquires sensor data from a user, in several different implementation configurations.
[0020] [Figure 2] This figure illustrates exemplary electronic devices operating in different physical environments during a communication session between a first user on a first device and a second user on a second device, in several implementation configurations, along with a combined 3D representation view of the second user for the first device.
[0021] [Figure 3A] This document presents examples of 3D Gaussian splats for use in generating 3D representations of views, using several different implementations. [Figure 3B] This document presents examples of 3D Gaussian splats for use in generating 3D representations of views, using several different implementations. [Figure 3C] This document presents examples of 3D Gaussian splats for use in generating 3D representations of views, using several different implementations. [Figure 3D] Examples of 3D Gaussian splats for use in generating views of 3D representations according to some implementations are shown.
[0022] [Figure 4] FIG. is a diagram showing an example of generating and displaying a splat stereo view on a device according to some implementations.
[0023] [Figure 5] An example of generating a user representation based on rendering splat parameter data according to some implementations is shown.
[0024] [Figure 6A] An example of generating a user representation based on rendering splat parameter data according to some implementations is shown. [Figure 6B] An example of generating a user representation based on rendering splat parameter data according to some implementations is shown. [Figure 6C] An example of generating a user representation based on rendering splat parameter data according to some implementations is shown.
[0025] [Figure 7] A flowchart representation of a method for providing a view of a user representation based on rendering splat parameter data according to some implementations.
[0026] [Figure 8] A block diagram showing device components of an exemplary device according to some implementations.
[0027] [Figure 9] A block diagram of one exemplary head-mounted device (HMD) according to some implementations.
[0028] By convention, various features shown in the drawings may not be depicted to scale. Therefore, the dimensions of various features may be arbitrarily enlarged or reduced for clarity. In addition, some drawings may not depict all components of a given system, method, or device. Finally, throughout this specification and the drawings, similar reference numerals may be used to indicate similar features. [Modes for carrying out the invention]
[0029] Numerous details are provided to give a full understanding of the exemplary implementations shown in the drawings. However, the drawings merely illustrate some exemplary embodiments of this disclosure and should not be considered limiting. Those skilled in the art will understand that other effective embodiments or variations do not include all of the specific details described herein. Furthermore, well-known systems, methods, components, devices, and circuits are not described in exhaustive detail so as not to obscure more suitable embodiments of the exemplary implementations described herein.
[0030] Figure 1 shows an exemplary environment 100 of an exemplary electronic device 105 operating in a physical environment 102. In some implementations, the electronic devices 105 can be enabled to share information with each other or with intermediate devices such as information systems. Furthermore, the physical environment 102 includes a user 110 wearing the device 105. In some implementations, the device 105 is configured to present a view of an extended reality (XR) environment based on the physical environment 102 and / or may include additional content such as virtual elements providing text narration.
[0031] In the example in Figure 1, the physical environment 102 is a room containing physical objects such as a wall hanging 120, a plant 125, and a desk 130. The electronic device 105 may include one or more cameras, microphones, depth sensors, motion sensors, or other sensors that can be used to capture and evaluate information about the physical environment 102 and the objects within it, as well as information about the user 110.
[0032] In the example in Figure 1, device 105 includes one or more sensors 116 (e.g., inward-facing sensors and outward-facing cameras) that capture light intensity images, depth sensor images, audio data, or other information about user 110. For example, one or more sensors 116 may capture images of the user's (e.g., user 110) forehead, eyebrows, eyes, eyelids, cheeks, nose, lips, chin, face, head, hands, wrists, arms, shoulders, torso, legs, or other body parts. Furthermore, one or more sensors 116 may capture images of elements / materials that are connected to or worn by user 110, such as glasses, earrings, and / or other accessories. For example, an inward-facing sensor may be able to see what is inside device 105 (e.g., the user's eyes and the area around the eyes), and other outward-facing cameras may be able to capture the user's face outside device 105 (e.g., a selfie camera facing user 110 outside device 105). For example, sensor data relating to the user's eyes 111 can reveal various user characteristics, such as the user's gaze direction 119 over time, the user's intermittent behavior over time, and the user's eye extension behavior over time. One or more sensors 116 can capture audio information, including the user's speech, sounds emitted by other users, and sounds within the physical environment 100.
[0033] In some implementations, device 105 includes an eye-tracking system for detecting eye position and movement via gaze characteristic data. For example, the eye-tracking system may include one or more infrared (IR) light-emitting diodes (LEDs), an eye-tracking camera (e.g., a near-infrared (NIR) camera), and an illumination source (e.g., an NIR light source) that emits light (e.g., NIR light) toward the user 110's eyes. Furthermore, the illumination source of device 105 can emit NIR light to illuminate the user 110's eyes, and the NIR camera can capture images of the user 110's eyes. In some implementations, the images captured by the eye-tracking system may be analyzed to detect the position and movement of the user 110's eyes, or to detect other information about the eyes, such as color, shape, state (e.g., wide open, strabismus), pupil dilation, or pupil diameter. Furthermore, the gaze estimated from the eye-tracking images can enable gaze-based interaction with content displayed on the device 105's near-eye display.
[0034] Furthermore, one or more sensors 116 can capture images of the physical environment 100 (e.g., outward-facing sensors). For example, one or more sensors 116 can capture images of the physical environment 100 including physical objects such as wall hangings 120, plants 125, and desks 130. Furthermore, one or more sensors 116 can capture images (e.g., light intensity images and / or depth data).
[0035] One or more sensors on device 105, such as one or more sensors 115, may identify user information based on proximity to or contact with a part of user 110. For example, one or more sensors 115 may capture sensor data that can provide biological information related to the user's circulatory system status (e.g., pulse), body temperature, respiratory rate, etc.
[0036] One or more sensors 116 or one or more sensors 115 may capture data that can determine the user's orientation 121 in the physical environment. In this example, the user's orientation 121 corresponds to the direction in which the user's torso is facing.
[0037] Several implementations disclosed herein determine user understanding based on sensor data acquired by a user-worn device, such as the first device 105. Such user understanding may indicate user states associated with providing user assistance. In some examples, an understanding of the user's appearance or behavior or environment can be used to recognize the need or desire for assistance and make such assistance available to the user. For example, based on determining such user states, enhancements to assist the user can be provided by enhancing or complementing the user's capabilities, such as providing guidance or other information about the environment to a person with a disability.
[0038] The content may be visible, for example, displayed on the display of device 105, or it may be audible, for example, produced as sound 118 by the speaker of device 105. In the case of sound content, the sound 118 may be produced in such a way that only user 110 can hear it, for example, through a speaker close to the user's ear 112, or at a volume below a threshold that makes it difficult for nearby people to hear. In some implementations, the sound mode (e.g., volume) is determined based on whether other people are within a threshold distance, or on how close other people are to user 110.
[0039] In some implementations, the content provided by device 105 and the sensor features of device 105 can be provided using components, sensors, or software modules that are small enough and efficient in terms of power consumption and use, so as to be compatible with or used in lightweight, battery-powered wearable products such as wireless earphones or other in-ear devices or head-mounted devices (HMDs) such as smart glasses / augmented reality (AR) glasses. Features can be facilitated by using a combination of multiple devices. For example, a smartphone (wirelessly connected to and interoperating with one or more wearable devices) can provide computing resources, connectivity to cloud or internet services, location services, and so on.
[0040] Figure 2 shows exemplary electronic devices operating in different physical environments during a communication session between a first user in a first device and a second user in a second device, in several implementation configurations, along with a view of the 3D representation of the second user for the first device. In particular, Figure 2 shows exemplary operating environments 200 of electronic devices 210, 265 operating in different physical environments 202, 250, respectively, while, for example, electronic devices 210, 265 share information with each other or with an intermediate device such as a communication session system / server during a communication session. In this example of Figure 2, physical environment 202 is a room including a wall hanging 212, a plant 214, and a desk 216 (e.g., physical environment 102 in Figure 1). Electronic device 210 includes one or more cameras, microphones, depth sensors, or other sensors that can be used to capture and evaluate information about the physical environment 202 and objects within it, as well as information about the user 225 of electronic device 210 (e.g., a handheld device). Information regarding the physical environment 202 and / or user 225 can be used during a communication session to provide visual content (e.g., for user representations) and audio content (e.g., for audible speech or text representations). For example, a communication session may provide one or more participants (e.g., users 225, 260) with a view of the 3D environment generated based on camera images and / or depth camera images of the physical environment 202, and may also provide a representation of user 225 generated based on camera images and / or depth camera images of user 225.
[0041] Furthermore, in this example in Figure 2, the physical environment 250 is a room including a wall hanging 252, a sofa 254, and a coffee table 256. The electronic device 265 includes one or more cameras, microphones, depth sensors, or other sensors which can be used to capture and evaluate information about the physical environment 250 and the objects within it, as well as information about the user 260 of the electronic device 265 (e.g., a user-worn device or HMD device such as device 105). Information about the physical environment 250 and / or the user 260 can be used to provide visual and audio content during a communication session. For example, a communication session can provide a view of the 3D environment generated based on camera images and / or depth camera images of the physical environment 250 (from the electronic device 265), as well as a representation of the user 260 based on camera images and / or depth camera images of the user 260 (from the electronic device 265). For example, the 3D environment can be transmitted by device 265 (e.g., via information system 290 via network connection 285) by communication session instruction set 280, which communicates with device 210 by communication session instruction set 282. Information system 290 can coordinate the encryption / decryption and pre-downloading of assets (e.g., 3D asset data such as data associated with user representations 240, 275) between two or more devices (e.g., electronic devices 210 and 265).
[0042] Figure 2 shows an example of a view 205 of a virtual environment (e.g., 3D environment 230) in device 210, where, if consent is given to view each user's representation during a particular communication session, the representation 232 on the wall 252 and the user representation 240 (e.g., persona of user 260) are provided. In particular, the user representation 240 of user 260 is generated based on one or more user representation techniques for a more realistic persona generated in real time. The generation of user representations will be further described herein.
[0043] Furthermore, the electronic device 265 within the physical environment 250 provides a view 266 that enables the user 260 to view the representation 272 on the wall 212 and the representation 275 of at least a portion of the user 225 in the 3D environment 270 (e.g., a persona) (e.g., from the middle of the torso upwards). In other words, the user representation 240 of the user 260 is generated in device 210 by generating a combined 3D representation of the user 260 for multiple points in time within a given period, based on data acquired from device 265 (e.g., a frame-specific 3D representation of the user 260). Alternatively, in some embodiments, the user representation 240 of the user 260 is generated in device 265 (e.g., a speaker's transmitting device) and transmitted to device 210 (e.g., a viewing device for viewing the speaker's persona). In some embodiments, each of the 3D representation 240 of user 260 and the 3D representation 275 of user 225 is generated by generating splats corresponding to the modified user representation data in accordance with the techniques described herein.
[0044] In the example in Figure 2, electronic device 210 is shown as a handheld device, and electronic device 265 is shown as a head-mounted device (HMD). However, either electronic device 210 or 265 may be a mobile phone, tablet, laptop, etc., or electronic device 265 may be worn by the user (e.g., a head-mounted device (glasses), headphones, ear-mounted device, etc.). In some implementations, the functions of devices 210 and 265 are achieved through two or more devices, for example, a mobile device and a base station, or a head-mounted device and an ear-mounted device. Various capabilities, including but not limited to power capacity, CPU capacity, GPU capacity, memory capacity, visual content display capacity, and audio content creation capacity, can be distributed among multiple devices. Multiple devices that can be used to achieve the functions of electronic devices 210 and 265 can communicate with each other via wired or wireless communication. In some implementations, each device communicates with a separate controller or server (e.g., a communication session server) to manage and coordinate the user experience. Such controllers or servers may be located within physical environment 202 and / or physical environment 250, or at a distance from them.
[0045] Furthermore, in the example in Figure 2, 3D environments 230 and 270 are XR environments based on a common coordinate system that can be shared with other users (e.g., virtual rooms for personas for a multi-person communication session). In other words, the common coordinate systems of 3D environments 230 and 270 are different from the coordinate systems of physical environments 202 and 250, respectively. For example, a common reference point can be used to align the coordinate systems. In some implementations, the common reference point can be a virtual object in the 3D environment that each user can visualize within their respective view. For example, a common centerpiece table around which user representations (e.g., user personas) are positioned within the 3D environment. Alternatively, the common reference point is not visible within each view. For example, the common coordinate system of the 3D environment can use a common reference point to position each individual user representation (e.g., around a table / desk). Thus, if the common reference point is visible, each view of the device can visualize the "center" of the 3D environment through perspective when viewing other user representations. Visualizing a common reference point can become more relevant to multi-user communication sessions, allowing each user's view to add a sense of perspective to the positions of other users during the communication session.
[0046] In some implementations, each user's representation can be realistic or unrealistic, and / or represent the user's current and / or previous appearance. For example, a realistic representation of user 225 or 260 can be generated based on a combination of live and previous images of the user. The previous images can be used to generate parts of the representation for which live image data is unavailable (e.g., parts of the user's face that are not in the field of view of the camera or sensor of electronic device 210 or 265, or that may be obscured by, for example, a headset or other means). In one example, electronic devices 210 and 265 are HMDs, and the live image data of the user's face includes a downward-facing camera that captures images of the user's cheeks and mouth, and an inward-facing camera image of the user's eyes, which can be combined with previous image data of the user's face, head, and other parts of the user's torso that are not currently observable from the device's sensors. Previous data on the user's appearance can be acquired earlier in the communication session, during previous use of the electronic device, during a registration process used to acquire sensor data of the user's appearance from multiple viewpoints and / or conditions, or by other means.
[0047] In some implementations, generating one or more user representations for a communication session (e.g., generating user representations 240 and 275), as shown in Figure 2, can be based on one or more rendering techniques, such as using a 3D mesh or a 3D point cloud. However, the technique described herein utilizes a rendering method of 3D Gaussian splats using UV mapping. Several advantages can be realized by using a simple set of values with defined depth values for multiple points, as represented by a 3D Gaussian splat using UV mapping. The set of values may require less computation and bandwidth than using a 3D mesh or a 3D point cloud, while enabling more accurate user representations than RGBDA images. Furthermore, the set of values may be formatted / packaged in a similar manner to existing formats, such as RGBDA images, thereby enabling more efficient integration with systems based on such formats.
[0048] Figures 3A to 3D show examples of 3D Gaussian splatters for use in generating 3D representations of views, in several implementation forms. For example, 3D Gaussian splatting (3DGS) can be used in 3D modeling to represent complex scenes as a combination of numerous color 3D Gaussians rendered to a camera view via splat-based rasterization. The position, size, rotation, color, and opacity of these Gaussian splatters can then be adjusted via differentiable rendering and gradient-based optimization techniques to represent the 3D scene given by the set of input images.
[0049] Figure 3A shows a 3D Gaussian splat 310, for example, an ellipsoidal shape formed by a 3D Gaussian distribution. The 3D Gaussian splat 310 can be used to represent position (μ), such as xyz coordinates. The 3D Gaussian splat 310 can further represent rotation and scale (e.g., Σ: covariance matrix), opacity (α), color (e.g., RGB values), anisotropic variance, spherical harmonic (SH) coefficients, and / or semantic classification (e.g., glasses, accessories). Figure 3B shows an environment 320 for rendering the splat 326 based on the visible direction of the camera 321. Figure 3C shows the ordering of the splats along the camera viewpoint direction along the ray 330 (e.g., splats 331, 332, 333, 334 are identified and ordered along the ray 330). For example, Gaussian splats is a technique in which individual 3D points are represented as a Gaussian distribution (like a "splat") having color values that change with the viewing angle, using spherical harmonics to model this view-dependent color variation, enabling real-time rendering of high-quality, photorealistic scenes from a sparse set of images. For example, each point has a color calculated based on its position relative to the camera, enabling realistic shading effects across different viewpoints. Figure 3D shows an environment 340 for blending splats 341, 342, 343, 344, 345 that can be seen along the direction of rays 330 from the camera viewpoint by composing splats 331, 332, 333, 334 on the image plane. Several implementations can use screen versus splat (similar to, for example, raycasting techniques), splat versus screen (similar to, for example, projection techniques), a combination thereof, or other techniques for composing the splats.
[0050] Figure 4 shows one exemplary environment 400 for generating a stereo view of a splat and displaying it on a device (e.g., an HMD) in several implementation forms. For example, device 410 is an HMD including a first display 420 for the left eye view and a second display 430 for the right eye view. The first display 420 and the second display 430 can then view the splat 450 rendered accordingly for each individual viewpoint. In some implementation forms, as shown in Figure 4, the images generated for each viewpoint can be rendered as a single mono image (e.g., for one eye), or the images can be presented as a stereo view. Additionally or alternatively, in some implementation forms, the images generated for each viewpoint can be rendered as a single grid mesh or a combination of stereo grid meshes.
[0051] Figure 5 shows an example of generating a user representation (e.g., a persona) based on rendering splat parameter data in several implementation forms. In particular, Figure 5 shows one exemplary user representation process 500 for obtaining registration data from a registration process 510 (e.g., registration images 512a, 512b, 512c) from a sending device for the sender and a viewpoint on a receiving device for the viewer, in order to generate a view 550 of a user representation 552 (e.g., a persona) using Gaussian splatting techniques.
[0052] The registration process 510 shows an image of the user (e.g., user 110 in Figure 1) during the registration process. The registration process 510 may include user enrollment registration (e.g., pre-registration of registration data) and acquisition of sensor data (e.g., live data registration). In some implementations, as shown in Figure 511, user enrollment registration may include the user (e.g., user 110) acquiring a full-view image of the user's face and part of the user's upper body using an external sensor on device 105, so that the user can remove device 105 (e.g., HMD) and point it at their face / body during the registration process. A registration personification may be generated when the system acquires image data (e.g., an RGB image) of the user's face while the user provides different facial expressions. For example, the user may be told to "raise your eyebrows," "smile," "frown," etc., to provide the system with a range of facial features for the registration process. A registration personification preview may be shown to the user while the user is providing a registration image to obtain a visualization of the state of the registration process. The registration image data 510 may have different user representations and may include registration personifications from different viewpoints (e.g., a front view in registration image 512a, a right side view in registration image 512b, and a left side view in registration image 512c). In some examples, more or fewer different representations and / or viewpoints may be used to obtain sufficient data for the registration process. In some implementations, at the final stage of registration, the user may be presented with selectable options used to generate the user representation, such as light induction / color correction, and the addition / removal of accessories (e.g., the user may select their best identifying information to represent their persona).
[0053] In some implementations, the conversion of registered image data to feature data 522 can be performed by converting it as part of a feature data process 520 for multiple representations (e.g., different sets of feature data for different representations) (e.g., via a converter). For example, feature data 522 may include trained feature information of user 110 obtained from registered images, such as per-pixel skin, color, and other semantic information. Feature data 522 may also include a list of positions for each feature value (e.g., multiple feature channels). Feature data 522 can then be decoded by a decoder to generate a 3D Gaussian UV map for each feature as part of a Gaussian UV map process 530. The 3D points of feature data 532 can be mapped to Gaussian parameters of the UV map 534 (e.g., 3D points + Gaussian parameters). For example, a UV map stores the x, y, and z positions of splat parameters (e.g., color (view-dependent information / harmonic information), covariance, alpha / transparency, orientation, opacity, range, rotation, scale on each axis, and semantic information (e.g., skin, hair, cheeks, nose, lips, eyebrows, accessories, etc.)). In other words, the Gaussian UV mapping process 530 can obtain 3D point information that contains sufficient information about what splat generation can be produced (e.g., 3D vector projection).
[0054] In some implementations, process 500 proceeds after generating 3D Gauss UV map data from the Gauss UV map process 530 (e.g., on the sender's device after registration), and (e.g., on the viewer's device) the system can obtain the 3D Gauss UV map data (e.g., feature data 532 mapped to Gauss parameters via UV map 534) and project the Gauss data using the current viewpoint (e.g., viewpoint data 536) to determine a 2D Gauss UV map 542 (e.g., 2D points + Gauss parameters) for the Gauss UV map process 540. Then, the Gauss splat can be used for rendering 545 to generate a view of the user representation 552 for the user representation generation process 550. For example, the 3D Gauss splat technique renders an image using the Gauss splat based on the viewer's current viewpoint (e.g., viewpoint data 536) with 2D points and associated Gauss parameters from the UV map.
[0055] Figures 6A-6C illustrate an example of generating a user representation (e.g., a persona) based on rendering splat parameter data, according to several implementations. In particular, Figure 6 shows one exemplary process 600 for obtaining Gaussian UV map data from the sender based on the viewpoint in the receiving device for the viewer, in order to generate a view 550 of a user representation 552 (e.g., a persona) associated with the sender using Gaussian splatting technology. As shown in Figure 6A, the exemplary rendering process 600 based on splat parameter data shows a registration phase 602 and a runtime phase 604, as divided by a timeline mark 603. In other words, the registration phase 602 can occur at some point before the communication session, and all other processes after the timeline mark 603 for the runtime phase 604 occur during the live communication session (e.g., generating a live persona of the user on the sending device for the user on the viewing device). Furthermore, as shown in Figure 6C, the exemplary rendering process 600 shows the data flow process between a transmitting device (e.g., sender phase 606) and a receiving device (e.g., receiver phase 608), as indicated by the transmit / receive timeline marks 607.
[0056] In an exemplary implementation, process 600 begins with registration phase 602. Registration phase 602 may include an offline registration process that can generate a user identification representation. The identification representation may be a set of latents extracted from a Gaussian UV map, several types of canonical representations generated for each user (e.g., a canonical (or base) Gaussian UV map), or a combination thereof.
[0057] In one exemplary implementation, registration phase 602 may begin with registration process 610, where registration process 210 shows an image of the user (e.g., user 110 in Figure 1) captured during registration. Registration process 610 may include user enrollment registration (e.g., pre-registration of registration data) and acquiring sensor data (e.g., live data registration) for capturing registration image 612. As shown in image 511 of Figure 5, user enrollment registration may include the user (e.g., user 110) acquiring a full-view image of the user's face using an external sensor on device 105, so that during the registration process, device 105 (e.g., HMD) can be removed and pointed at the user's face. Registration personification may be generated when the system acquires image data (e.g., RGB image) of the user's face while the user provides different facial expressions. For example, the user may be told to "raise your eyebrows," "smile," "frown," etc., to provide the system with a range of facial features for the registration process. While the user provides a registration image to obtain a visualization of the registration process status, a registration personification preview may be shown to the user. The registration image data may include registration personifications from different perspectives, each with a different user representation.
[0058] In some implementations, the optical normalization process 614 may be applied to the registration data. For example, the optical normalization process 614 may be provided to obtain a cropped image of the registration image 612 showing insufficient lighting conditions, such as those shown on the user's face (e.g., a segmented head or face of the user). The optical normalization process 614 can detect one or more attributes associated with the insufficient lighting conditions and adjust one or more registration images accordingly. The adjustments can then be applied to generate one or more post-processed registration images showing the removal of the attributes associated with the insufficient lighting conditions on the user's face. For example, because the registration image 612 may have been too dark, the optical normalization process 614 can brighten the user's face area, and the post-processing can be applied to the brightened face area across the entire post-processed registration image data.
[0059] In some implementations, after the registration image is normalized, the registration phase 602 proceeds to the Gaussian UV mapping process 620 to generate 3D Gaussian UV map data 622. For example, the conversion of the registration image data to feature data can be done as part of the feature data process (e.g., via a converter). For example, the feature data may include trained feature information of user 110 obtained from the registration image 612, such as skin, color, and other semantic information per pixel. The feature data may also include a list of positions for each feature value (e.g., 14 feature channels). The feature data can then be decoded by a decoder to generate 3D Gaussian UV map data 622 for each feature as part of the Gaussian UV mapping process 620. The 3D points of the feature data can be mapped to Gaussian parameters of the UV map (e.g., 3D points + Gaussian parameters). For example, a UV map can store the x, y, and z positions of splat parameters (e.g., color (view-dependent / harmonic information), covariance, alpha / transparency, orientation, opacity, range, rotation, scale on each axis, and semantic information (e.g., skin, hair, cheeks, nose, lips, eyebrows, etc.)).
[0060] In some implementations, the Gaussian UV mapping process 620 can generate a user's canonical representation, such as a canonical Gaussian UV map. The canonical Gaussian UV map can be synthesized from multiple registered representations and used as a stable, individualized reference for subsequent real-time updates. In some implementations, the canonical representation can be a neutral representation, average, or ideal composite and can be stored and used as a starting point for runtime animation. In runtime phase 604, the live feature latent (representation latent) can be used to calculate a delta or modification applied to the canonical Gaussian UV map to obtain the final frame-specific user representation.
[0061] In some implementations, the Gaussian UV mapping process 620 can also acquire additional feature data associated with the user's body (e.g., body tracking information 623). Body tracking information 623 may include one or more parts of the user's body other than the face (e.g., head, hands, upper / lower torso, etc.). Body tracking information 623 can be separated into different data streams based on tracking different parts of the body, such as the head / face as one data stream, the hands as another, and the upper and / or lower torso as yet another data stream. In some implementations, body tracking information 623 can be retouched via a retouching process 624, which provides the system with the ability to update (e.g., fine-tune) the skeletal model of body tracking information 623. For example, the retouching process 624 can be applied directly to the Gaussian UV map by modifying splat parameters to achieve specific goals, such as removing skin blemishes, changing hair color, slightly altering nose shape, and / or any other deformations including color, shape, and / or appearance. The registered body tracking information 623 can also be transmitted as skeletal blend data 627 (e.g., synthesized weight and asset data), which can be further analyzed for skeletal tracking during runtime phase 604, particularly at other stages of the process such as deformation and blurring of Gaussian UV maps, as well as in retouched evaluation body data.
[0062] The Identifier Latent Process 630 can obtain 3D Gaussian UV map data 622 extracted from the Gaussian UV map process 620. For example, the Identifier Latent Process 630 can generate a map 632 of the Identifier Latent that can be used as marker points, e.g., semantic points of the user that are likely to move between animation frames and can be determined at registration, and thus these points are updated at runtime (e.g., corners of the mouth, other features, etc.). In other words, the Gaussian UV map process 620 and the Identifier Latent Process 630 can determine 3D point information that contains sufficient information on what splat generation can be generated (e.g., 3D vector projection) and identify parts of the user (e.g., marker points 664) that need to be updated per frame identified by the Identifier Latent Process 630. In other words, “marker points” refer to semantic points (e.g., facial landmarks, corners of the mouth, tip of the nose, etc.) that are tracked and updated at runtime.
[0063] The registration phase 602 can then store these registration assets, such as the identification information latent map 632 and skeleton synthesis data 627, for future use during the runtime phase 604 (e.g., during a communication session with another device). During the runtime phase 604 (e.g., a communication session with another device has started), the exemplary process 600 proceeds to capture live data in the live data process 640, as shown in Figure 6B. In particular, the live data process 640 captures head / face tracking data 642 and body tracking data 650 (e.g., real-time sensor data for the user to update the user representation). The head / face tracking data 642 can be used to identify the representation latent 644 and head pose data 646 (e.g., HMD-versus-head pose information). The representation latent 644 can be used together with the identification information latent 632 for the animation and decoding process to update the Gaussian UV map data 622 with the user's current ("live") facial expression (e.g., an animated facial expression).
[0064] In some implementations, the skeletal tracking data 652 can be determined based on current ("live") sensor data (e.g., body tracking data 650) and skeletal synthesis data 627 obtained from the registration phase 602. Head pose data 646 can be used to update the skeletal tracking data 652 (e.g., to update data associated with the sender's head based on device pose data). The skeletal tracking data 652 can then be used by the body pose process 654 to identify skeletal joint data associated with multiple (digitized) skeletal joints of the sender. The skeletal joint data can be used in the animation and decoding process along with the identification information latent 632 to update the Gaussian UV map data 622 with the user's current ("live") skeletal movement (e.g., animated body pose). In some implementations, body tracking information can be sent to a decoder network to estimate complex body deformations.
[0065] The final stage of the transmitting device's runtime phase 604 (e.g., the sender phase 606) is the animation and decoding process of the Gaussian UV mapping process 660, as shown in Figure 6C. In the Gaussian UV mapping process 660, the Gaussian UV mapping data 622 and the identification information latent 632 are animated and decoded using the representation latent 644 to determine the Gaussian UV mapping data 662 containing the identified marker points 664. The "live" frame-by-frame Gaussian UV mapping data 662 is then sent to the receiving device as part of the receiver phase 608 during the communication session. In some implementations, the receiving device further acquires skeletal synthesis data 627 and skeletal joint data (e.g., animated body poses) and updates the Gaussian UV mapping data 662 during the deformation and blurring process of the Gaussian buffering process 670 by applying body pose data (e.g., neck, shoulders, hands, etc.) to generate the updated Gaussian UV mapping data 662 as Gaussian buffer data 672. The runtime phase 604 in the receiving device then acquires viewpoint data 674 (e.g., the viewpoint to be rendered) from the receiving device and can send the Gaussian buffer data 672 and viewpoint data 674 to a stereo proxy splat process 680 that generates stereo proxy geometry view data 682 (e.g., left and right viewpoints of RGBA splat data, as well as depth data). In some implementations, for more efficient rendering, the stereo proxy geometry view data 682 may be generated at 30 frames per second (FPS) while being rendered at 90 FPS. The stereo proxy geometry view data 682 can then be used in a 3D Gaussian splat rendering process (e.g., a user representation process 690) to generate a user representation 692.
[0066] In some implementations, the level of detail (LOD) can be determined when calling multiple people during a communication session. In this case, rendering a full-resolution persona for everyone in the communication session may be too resource-intensive. Therefore, the system may decide to fall back to a lower-resolution version (e.g., with fewer splats). For example, the system may decode multiple Gaussian UV maps with different resolutions as part of a feature data process 520. The recipient can then select the correct Gaussian UV map considering the current situation, based on one or more factors such as the distribution of people and their line of sight, among other things.
[0067] Figure 7 is a flowchart illustrating one exemplary method 700. In some implementations, a device (e.g., device 105 in Figure 1) performs the technique of method 700 to generate a view of a user representation based on rendering splat parameter data according to some implementations. In some implementations, the technique of method 700 is performed on a mobile device, desktop, laptop, HMD, or server device. In some implementations, method 700 is performed on processing logic, which includes hardware, firmware, software, or a combination thereof. In some implementations, method 700 is performed on a processor that executes code stored in a non-temporary computer-readable medium (e.g., memory). In some implementations, method 700 is performed on a processor in a device, such as a browsing device, that renders the user representation (e.g., device 210 in Figure 2 renders a 3D representation 240 of user 260 (persona) from data obtained from device 265).
[0068] In block 710, method 700 acquires user representation data of at least a portion of the user in the device's processor, the user representation data includes splat parameter data corresponding to multiple 3D positions, based on a first set of sensor data including images of the user acquired during the registration process. In some implementations, the device is a viewing device that renders the user representation (e.g., persona). For example, as shown in Figure 2, device 210 renders a user representation 240 of user 260 (e.g., sender). In some implementations, at least a portion of the user includes the user's face and additional parts (e.g., head, neck, clothing, hair, body, etc.). In some implementations, the user image acquired during the registration process is a two-dimensional (2D) image. Additionally or alternatively, in some implementations, the user image acquired during the registration process is a 2D image + depth data.
[0069] In some implementations, user representation data is based on a UV map and 3D point cloud points associated with distribution data that defines the size and shape for rendering 3D point cloud points as splats corresponding to each point in the UV map (e.g., a 3D Gaussian map). For example, as shown in Figure 6, a user representation 692 is generated for a particular viewpoint based on rendering points from a Gaussian UV map 662 that combine splat parameters obtained from registration data, and is updated for each frame based on one or more marker points 664 (e.g., a set of semantic points associated with facial features or other areas corresponding to a sender associated with the rendered user representation 692).
[0070] In some embodiments, the splat parameter data includes 3D Gaussian parameters for each 3D position. For example, the splat parameter data may include position information, color information, covariance information, transparency information, orientation, opacity information, range information on each axis, rotation data, scale, and semantic information (e.g., skin, hair, cheeks, nose, lips, eyebrows, etc.). For example, as shown in Figure 6, the Gaussian UV mapping process 620 generates 3D Gaussian UV map data 622 by converting registered image data into feature data that can include user-learned feature information obtained from the registered image, such as skin, color, and other semantic information per pixel. For example, the splat parameter position may identify where the splat is located based on xyz coordinates, the splat parameter covariance may identify how the splat is stretched / scaled (e.g., a 3x3 matrix), the splat parameter color may identify the RGB color, and the splat parameter alpha (α) may identify how transparent the splat is. In some implementations, user representation data includes 3D point cloud points associated with distribution data that defines the size and shape for rendering the 3D point cloud points as splats. For example, a splat model is generated using Gaussian splats, which include texture / color, position, and splat shape.
[0071] In some implementations, user representation data includes 3D mapping information, including feature values and position information for each map point. For example, the Gaussian UV map process 620 and the identification information latent process 630 can determine 3D point information that contains sufficient information about which splat generation can be generated (e.g., 3D vector projection), and the identification information latent process 630 identifies user parts (e.g., marker points 664) that need to be updated each frame (e.g., identify the x, y, z positions corresponding to the UV coordinates of the UV map).
[0072] In block 720, method 700 modifies user representation data based on a second set of sensor data acquired after the registration process. For example, splat parameter data is modified based on live sensor data. For example, splat parameter data can be acquired from a transmitting device, such as from the registration process, and the modification to the registered splat parameter data can be based on acquiring the sender's live sensor data to determine the sender's live representation (e.g., a live view of a realistic persona for a communication session). For example, as shown in Figure 6C, the receiving device acquires head / face tracking data as representation latent 644 in the Gaussian UV mapping process 660 to determine Gaussian UV mapping data 662 (for example, applying the user's live representation based on marker points 664, i.e., semantic points or feature points (e.g., facial landmarks, corners of the mouth, tip of the nose, etc.) that are tracked and updated during runtime). Furthermore, in some implementations, the receiving device acquires skeletal synthesis data 627 and skeletal joint data (e.g., animated body pose) in the Gaussian buffering process 670, and updates the Gaussian UV map data 662 during the deformation and blurring process of the Gaussian buffering process 670 by applying body pose data (e.g., neck, shoulders, hands, etc.), thereby generating the updated Gaussian UV map data 662 as Gaussian buffer data 672.
[0073] In various implementations, user representation data can be modified for the face but not the body, modified for the body but not the face, modified for both body and face, and / or modified on the sender's device or the receiver's device during registration. In some implementations, user representation data is modified based on body posture data acquired during the registration process, during a communication session with another device, or a combination thereof. In some implementations, the device is the viewer's device, and user representation data is modified based on an additional set of sensor data acquired during a communication session with the sender's device associated with the user representation. Alternatively, in some implementations, user representation data is generated and updated during the registration process based on images of the user's face captured while the user is making multiple different facial expressions (e.g., registration images of the face while the user is smiling, raising eyebrows, puffing out cheeks, etc.).
[0074] In some implementations, modifying user representation data generates a 3D Gaussian splat that includes texture, position, and splat shape based on at least some of the user's image data. For example, a 3D Gaussian distribution in a 2D space with color / density, e.g., a parameterization where the face is represented as a 2D grid and each element of the 2D grid contains a 3D Gaussian splat. In some implementations, the technique generates user representations via a machine learning model trained using training data acquired through one or more sensors in one or more environments. For example, a machine learning model that interprets image data and / or other sensor data captured during registration.
[0075] In block 730, method 700 provides a view of the user representation based on modified user representation data by generating multiple splats based on splat parameter data of the modified user representation data. For example, 3D Gaussian splats can be used to avoid or fill in missing parts, and body pose data can be applied to include additional areas of the user (e.g., neck area / shoulder area). For example, as shown in Figure 6C, the user representation 692 is generated for a particular viewpoint based on rendering points from a Gaussian UV map 662 that combines splat parameters obtained from registration data, and is updated for each frame based on one or more marker points 664 (e.g., a set of semantic points associated with facial features corresponding to the sender associated with the rendered user representation 692 or other areas). In some implementations, the system can generate the user representation by applying deltas to a canonical UV map, and in other implementations, the user representation can be generated directly from live data or other forms of identification information representation.
[0076] In some implementations, a second set of sensor data acquired after the registration process by a device (e.g., the viewer's device) includes a sequence of frames of a Gaussian UV map and corresponding marker points. These frames of the Gaussian UV map and corresponding marker points can be acquired from a second device (e.g., the sender's device) during a communication session. The device (e.g., the viewer's device) uses one or more splat techniques described herein to render an animated representation of the user (e.g., the sender) based on the sequence of frames of the Gaussian UV map and corresponding marker points. Marker points tracked and updated during runtime (e.g., semantic or feature points such as facial landmarks, corners of the mouth, or tip of the nose) can be used in conjunction with canonical representations and / or live latents to generate an animated persona.
[0077] In some implementations, a user identification information representation (e.g., a set of identification information latents, a canonical Gaussian UV map, or a combination thereof) and corresponding marker points are transmitted during a communication session with a second device and can be used to render a view of the user's (sender's) face (and upper body). Additionally or alternatively, face data (the appearance of the user's face at different points in time) and a sequence of frames of body tracking data can be transmitted and used to display a live 3D video-like depiction of the user (e.g., a "live" persona). For example, as shown in Figure 6C, the receiving device further acquires skeletal synthesis data 627 and skeletal joint data (e.g., animated body poses) and updates the Gaussian UV map data 662 during the deformation and blurring process of the Gaussian buffering process 670 by applying body pose data (e.g., neck, shoulders, hands, etc.) to generate the updated Gaussian UV map data 662 as Gaussian buffer data 672. The runtime phase 604 in the receiving device can then acquire viewpoint data 674 (e.g., the viewpoint to be rendered) from the receiving device and send the Gaussian buffer data 672 to a stereo proxy splat process 680 that generates stereo proxy geometry view data 682 (e.g., left and right viewpoints of RGBA splat data, as well as depth data). The stereo proxy geometry view data 682 can then be used in a 3D Gaussian splat rendering process (e.g., a user representation process 690) to generate a user representation 692.
[0078] In some embodiments, the second user representation is based on second image data acquired via a second set of sensors in a second physical environment having second lighting conditions (e.g., different lighting conditions from the first physical environment). For example, during the registration process, user representation data is acquired in a specific environment (also referred to herein as the “registration environment”), including some lighting condition information (e.g., luminance values and other lighting attributes), which may be different lighting data from live lighting data (e.g., two different physical environments between registration and persona generation based on “live” sensor data). In some implementations, Method 700 further includes providing a view of the user representation in a 3D environment. In some implementations, Method 700 further includes modifying the view of the user representation by adjusting the user representation based on at least one color attribute from a plurality of color attributes of the environment, at least one light attribute from a plurality of light attributes of the environment, or a combination thereof. For example, adjusting the color or lighting on the user representation, such as hair, face, or clothing, based on the color and / or light associated with the viewer’s environment and / or the sender’s environment. In other words, the lighting and / or color of a 3D representation (e.g., a persona) can be modified to match the lighting and / or color of the viewer's environment (for example, so that the reddish hue of the light shining in the viewer's room is reflected in the 3D representation). Alternatively, the lighting and / or color of a 3D representation (e.g., a persona) can be modified to match the lighting and / or color of the sender's environment (for example, so that the greenish hue of the light shining in the sender's room is reflected in the 3D representation for the viewer, even if the registered data does not reflect the greenish hue of the light).
[0079] In some implementations, user representation data of at least a portion of the user acquired during the registration process is based on images of the user's face captured in different poses and / or while the user is making multiple different facial expressions. For example, the images are registration images of the face while the user is facing the camera, while facing left of the camera, while facing right of the camera, and / or while the user is smiling, raising their eyebrows, puffing out their cheeks, etc. In some implementations, a first set of sensor data corresponds only to a first area of the user (e.g., the part not obscured by a device such as an HMD), and a second set of sensor data corresponds to a second area that includes a third area different from the first area. For example, the second area may include some of the parts obscured by the HMD when the HMD is worn by the user. For example, during the registration process, a larger portion of the user can be captured by image data (e.g., without the HMD) than during a live communication session with a user wearing an HMD.
[0080] In some implementations, as shown in Figure 2, rendering takes place during a communication session in which a second device (e.g., device 265) captures sensor data (e.g., partial image data of user 260 and environment 250) and, based on the sensor data, provides a series of frame-specific 3D representations corresponding to multiple moments in a period. For example, the second device 265 provides / transmits a series of frame-specific 3D representations to device 210, and device 210 generates a combined 3D representation to display a facial depiction (e.g., a realistic moving persona) such as a live 3D video of user 260 (e.g., representation 240 of user 260). Alternatively, in some implementations, the second device provides a 3D representation (e.g., a realistic moving persona) of the user (e.g., representation 140 of user 260) during the communication session. For example, the combined representation is determined in device 265 and transmitted to device 210. In some implementations, the view of the 3D representation is displayed on the device (e.g., device 210) in real time for multiple moments in a period. For example, the depiction of a user (e.g., an avatar shown to a second user on the display of a second user's second device) is displayed in real time and based on live lighting data.
[0081] In some implementations, the user representation view can include enough data to enable a stereo view of the user (e.g., left-eye / right-eye view) so that the face can be perceived with depth. In one implementation, the face representation is generated to provide a stereo view of the face, including a 3D model of the face, as well as views of the representation from the left-eye position and the right-eye position.
[0082] In some implementations, certain parts of the face that may be important for conveying a realistic or faithful appearance, such as the eyes and mouth, may be generated differently from other parts of the face (e.g., based on marker points). For example, parts of the face that may be important for conveying a realistic or faithful appearance may be based on current camera data, while other parts of the face may be based on previously acquired (e.g., registered) face data.
[0083] In some implementations, a facial representation is generated using various facial textures, colors, and / or geometric shapes, and an estimate of the confidence level of the generation technique is identified based on the depth and appearance values of each frame of data, indicating that such textures, colors, and / or geometric shapes accurately correspond to the actual textures, colors, and / or geometric shapes of those facial parts. In some implementations, the representation is a 3D persona. For example, the representation is a 3D model representing a user (e.g., user 110 in Figure 1).
[0084] In some implementations, a first set of sensor data and / or a second set of sensor data (e.g., live data such as video content including light intensity data (RGB) and depth data) are associated with a point in time, such as images from inward / downward sensors while the user is wearing a frame-associated HMD. In some implementations, the sensor data includes depth data (e.g., infrared, time of flight, etc.) and light intensity image data acquired during the scanning process.
[0085] In some implementations, acquiring a first set of sensor data during the registration process may include acquiring registration sensor data from a device (e.g., registration image data 610 in Figure 6) corresponding to the user's facial features in multiple configurations (e.g., texture, muscle activity, shape, depth, etc.). In some implementations, the first set of data may include unobstructed image data of the user's face. For example, images of the face may be captured while the user is smiling, raising their eyebrows, or puffing out their cheeks. In some implementations, registration data may be acquired by the user removing the device (e.g., HMD) and capturing images without the device obstructing the face, or by using another device (e.g., a mobile device) where the device (e.g., HMD) does not obstruct the face. In some implementations, registration data (e.g., the first set of data) is obtained from light intensity images (e.g., RGB images (one or more)). Registration data may include most, if not all, of the user's face, including texture, muscle activity, etc. In some implementations, registration data may be captured while the user is provided with different commands to acquire different poses of the user's face. For example, a user might be instructed by a user interface guide to "raise your eyebrows," "smile," or "frown" in order to provide the system with a range of facial features for the registration process.
[0086] In some implementations, Method 700 may be repeated for each frame captured during each instant / frame of a live communication session or other experience. For example, with each iteration, while the user is using the device (e.g., wearing an HMD), Method 700 may involve continuously acquiring live sensor data (e.g., face tracking data, body tracking data, etc.) and, frame by frame, updating the displayed portion of the user representation based on marker points using updated Gaussian UV maps and GPU Gaussian buffers. For example, with each new frame, the system may update the display of the 3D persona based on the new data.
[0087] Figure 8 is a block diagram of one exemplary device 800. Device 800 shows one exemplary device configuration for the devices described herein (e.g., devices 105, 210, 265, etc.). While certain features are shown, those skilled in the art will understand from this disclosure that various other features have been omitted for brevity so as not to obscure more appropriate embodiments of the implementations disclosed herein. For that purpose, in some non-limiting implementations, the device 800 may include one or more processing units 802 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 806, one or more communication interfaces 808 (e.g., USB, FireWire®, Thunderbolt®, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, Bluetooth®, ZIGBEE®, SPI, I2C, or similar types of interfaces), one or more programming (e.g., I / O) interfaces 810, one or more displays 812, one or more internal and / or external image sensor systems 814, memory 820, and one or more communication buses 804 for interconnecting these and various other components.
[0088] In some implementations, one or more communication buses 804 include circuits that interconnect system components and control communication. In some implementations, one or more I / O devices and sensors 806 include at least one of the following: an inertial measuring unit (IMU), an accelerometer, a magnetometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, one or more depth sensors (e.g., structured light, time of flight, etc.).
[0089] In some implementations, one or more displays 812 are configured to present the user with a view of a physical or graphic environment. In some implementations, one or more displays 812 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transitory (OLET), organic light-emitting diode (OLED), surface-conduction electron-emitter display (SED), field-emission display (FED), quantum dot light-emitting diode (QD-LED), micro-electro-mechanical system (MEMS), and / or similar display types. In some implementations, one or more displays 812 correspond to waveguide displays such as diffraction, reflection, polarization, and holographic displays. In one embodiment, device 10 includes a single display. In another embodiment, device 10 includes displays for each of the user's eyes.
[0090] In some implementations, one or more image sensor systems 814 are configured to acquire image data corresponding to at least a portion of the physical environment 102. Examples of one or more image sensor systems 814 include one or more RGB cameras (e.g., those with complementary metal-oxide-semiconductor (CMOS) image sensors or charge-coupled device (CCD) image sensors), monochrome cameras, IR cameras, depth cameras, and event-based cameras. In various implementations, one or more image sensor systems 814 further include an illumination source that emits light, such as a flash. In various implementations, one or more image sensor systems 814 further include an on-camera image signal processor (charge-coupled device, ISP) configured to perform a plurality of processing operations on the image data.
[0091] Memory 820 includes high-speed random-access memory such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory devices. In some implementations, memory 820 includes non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile storage devices. Memory 820 optionally includes one or more storage devices located remotely from one or more processing units 802. Memory 820 includes non-temporary computer-readable storage media.
[0092] In some implementations, memory 820 or a non-temporary computer-readable storage medium of memory 820 stores an optional operating system 830 and one or more instruction sets 840. The operating system 830 includes procedures for handling various basic system services and procedures for performing hardware-dependent tasks. In some implementations, the instruction sets 840 include executable software defined by binary information stored in the form of charges. In some implementations, the instruction sets 840 are executable software runnable by one or more processing units 802 to perform one or more of the techniques described herein.
[0093] The instruction set(s) 840 includes a registration instruction set(s) 842, an expression instruction set(s) 844, and a communication session instruction set(s) 846. The instruction set(s) 840 may be implemented as a single software executable file or as multiple software executable files.
[0094] In some implementations, the registration instruction set 842 is executable by a processing unit (one or more) 802 to generate registration data from image data. The registration instruction set 842 may be configured to provide instructions to the user to acquire image information for generating a registration personification (e.g., registration data 510) and to determine whether additional image information is needed to generate an accurate registration personification used by the persona display process. For this purpose, in various implementations, the instructions include instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0095] In some implementations, the representation instruction set 844 can be executed by a processing unit (one or more) 802 to generate a user representation (e.g., a Gaussian splat) based on registration data, either by using one or more of the techniques described herein or by other appropriate means. For this purpose, in various implementations, the instructions include instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0096] In some implementations, the communication session instruction set 846 can be executed by a processing unit (one or more) 802 to facilitate a communication session between two or more electronic devices (e.g., devices 210 and 265 as shown in Figure 2) using one or more of the techniques described herein, or otherwise appropriately done. For this purpose, in various implementations, the instructions include instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0097] While the instruction set(s) 840 is shown as existing on a single device, it should be understood that in other implementations, any combination of elements may be located within separate computing devices. Furthermore, Figure 8 is intended to illustrate the function of various features present in a particular implementation, rather than being a structural schematic of the implementations described herein. As will be recognized by those skilled in the art, the separately shown elements can be combined, and some elements can be separated. The actual number of instruction sets, and how features are assigned among them, may vary from implementation to implementation and may depend in part on a particular combination of hardware, software, and / or firmware selected for a particular implementation.
[0098] Figure 9 shows a block diagram of one exemplary head-mounted device 900 in several implementation configurations. The head-mounted device 900 includes a housing 901 (or enclosure) that houses the various components of the head-mounted device 900. The housing 901 includes (or is coupled to) an eyepad (not shown) located at the proximal end of the housing 901 (relative to the user 25). In some implementation configurations, the eyepad is a piece of plastic or rubber that comfortably and snugly holds the head-mounted device 900 in an appropriate position on the user 25's face (e.g., around the user 25's eyes 35).
[0099] The housing 901 can house a display 910 that displays an image and emits light toward or onto the eyes of the user 25. In various implementations, the display 910 emits light through an eyepiece having one or more optical elements 905 that refract the light emitted by the display 910, thereby making the display appear to the user 25 at a virtual distance greater than the actual distance from the eye to the display 910. For example, the optical elements (one or more) 905 may include one or more lenses, waveguides, other diffraction optical elements (DOEs), etc. In various implementations, the virtual distance is greater than at least the minimum focal distance of the eye (e.g., 7 cm) in order for the user 25 to be able to focus on the display 910. Furthermore, in various implementations, the virtual distance is greater than 1 meter in order to provide a better user experience.
[0100] The housing 901 also houses a tracking system including one or more light sources 922, cameras 924, 932, 934, and a controller 980. One or more light sources 922 emit light onto the user 25's eyes, which is reflected as a light pattern (e.g., a circle of glint) that can be detected by the camera 924. Based on the light pattern, the controller 980 can determine the user 25's eye tracking characteristics. For example, the controller 980 can determine the user 25's gaze direction and / or blinking state (eyes open or closed). As another example, the controller 980 can determine the pupil center, pupil size, or point of fixation. Thus, in various implementations, light is emitted by one or more light sources 922, reflected by the user 25's eyes, and detected by the camera 924. In various implementations, the light from the user 25's eyes is reflected by a hot mirror or passes through an eyepiece before reaching the camera 924.
[0101] The display 910 emits light in a first wavelength range, and one or more light sources 922 emit light in a second wavelength range. Similarly, the camera 924 detects light in the second wavelength range. In various implementations, the first wavelength range is the visible wavelength range (e.g., the wavelength range within the visible spectrum from approximately 400 to 700 nm), and the second wavelength range is the near-infrared wavelength range (e.g., the wavelength range within the near-infrared spectrum from approximately 700 to 1400 nm).
[0102] In various implementations, eye tracking (or in particular, a predetermined gaze direction) is used to enable user interaction (e.g., user 25 selects by looking at options on the display 910), provide foveal rendering (e.g., present at a higher resolution in the area of the display 910 that user 25 is looking at, and at a lower resolution elsewhere on the display 910), or correct distortion (e.g., for the image provided on the display 910). In various implementations, one or more light sources 922 emit light that is reflected in the form of multiple glints toward the user 25's eyes 35.
[0103] In various implementations, the camera 924 is a frame / shutter-based camera that generates images of the user's eyes 35 at specific or multiple points in time at a frame rate. Each image contains a matrix of pixel values corresponding to the pixels in the image, corresponding to the positions of the camera's light sensor matrix. In the implementation, each image is used to measure or track pupil dilation by measuring changes in pixel intensity associated with one or both of the user's pupils.
[0104] In various implementations, the camera 924 is an event camera that includes multiple light sensors (e.g., a matrix of light sensors) positioned at multiple locations, each of which generates an event message indicating the specific location of a particular light sensor in response to a specific light sensor detecting a change in light intensity.
[0105] In various implementations, cameras 932 and 934 are frame / shutter-based cameras capable of generating images of the user's face at specific or multiple points in time at a given frame rate. For example, camera 932 captures an image of the user's face below the eyes, and camera 934 captures an image of the user's face above the eyes. The images captured by cameras 932 and 934 may include light intensity images (e.g., RGB) and / or depth image data (e.g., time of flight, infrared, etc.).
[0106] The above-described implementations are examples only, and it should be understood that the present invention is not limited to those specifically illustrated and described above. Rather, its scope includes both combinations and partial combinations of the various features described above, as well as variations and modifications thereof not disclosed in the prior art, which a person skilled in the art would conceive of by reading the above description.
[0107] As described above, one aspect of the technology is the collection and use of physiological data to improve the user experience of electronic devices in relation to interacting with electronic content. In some examples, the disclosure intends that such collected data may include personal data that can be used to uniquely identify a particular person or to identify a particular person's interests, characteristics, or tendencies. Such personal data may include physiological data, demographic data, location-based data, telephone numbers, email addresses, home addresses, device characteristics of personal devices, or any other personal information.
[0108] This disclosure acknowledges that such use of personal data in the technology may be for the benefit of the user. For example, personal data can be used to improve the interaction and control capabilities of electronic devices. Thus, the use of such personal data enables computational control of electronic devices. Furthermore, other uses of personal data that may benefit the user are also conceivable in this disclosure.
[0109] This disclosure further intends that any entity responsible for collecting, analyzing, disclosing, transferring, storing, or otherwise using such personal and / or physiological data will adhere to well-established privacy policies and / or privacy practices. Specifically, such entities should implement and consistently use privacy policies and practices that are generally recognized as meeting or exceeding industry or government requirements for the strict confidentiality of personal data. For example, personal data from users should be collected for the entity's lawful and legitimate use and should not be shared or sold for any other purpose. Furthermore, such collection should only be carried out after informing and obtaining the user's consent. In addition, such entities will take all necessary steps to protect and secure access to such personal data and to ensure that others with access to such personal data comply with those privacy policies and procedures. Furthermore, such entities may undergo third-party evaluations to demonstrate their compliance with widely accepted privacy policies and practices.
[0110] Notwithstanding the foregoing, this disclosure also envisions implementations that allow users to selectively prevent the use of or access to personal data. That is, this disclosure envisions that hardware or software elements may be provided to prevent or prevent access to such personal data. For example, in the case of a user-conformable content delivery service, the technology may be configured to allow users to choose to "opt in" or "opt out" of participating in the collection of personal data during registration for the service. In another embodiment, the user may choose not to provide personal data to the content delivery service in question. In yet another embodiment, the user may choose not to provide personal data but to allow the transfer of anonymous information for the purpose of improving the functionality of the device.
[0111] Therefore, while this disclosure broadly covers the use of personal data to implement one or more of the disclosed embodiments, it is also conceivable that these embodiments could be implemented without requiring access to such personal data. In other words, the various embodiments of the technology would not be rendered inoperable by the absence of all or part of such personal data. For example, content could be selected and delivered to a user by inferring preferences or settings based on non-personal data, such as content requested by a device associated with the user, other non-personal data available in the content delivery service, or publicly available information, or by inferring preferences or settings based on only a minimal amount of personal information.
[0112] In some embodiments, data is stored using a public / private key system that allows only the data owner to decrypt the stored data. In some other implementations, data may be stored anonymously (without user identification and / or personal information, e.g., legal name, username, time and location data). In this way, other users, hackers, or third parties cannot determine the user identification information associated with the stored data. In some implementations, a user may be able to access their stored data from a user device different from the user device used to upload the stored data. In these examples, the user may be required to provide login credentials to access the stored data.
[0113] Numerous specific details are provided herein to provide a complete understanding of the subject matter of the claims. However, those skilled in the art will understand that the subject matter of the claims can be implemented without these specific details. In other examples, methods, apparatus, or systems known to those skilled in the art are not described in detail so as not to obscure the subject matter of the claims.
[0114] Unless otherwise specified, throughout this specification, terms such as “process,” “computing,” “calculating,” “determine,” and “identify” are understood to refer to actions or processes of a computing device. A computing device is one or more computers or similar electronic computing devices that manipulate or convert data expressed as physical electronic or magnetic quantities within the range of memory, registers or other information storage devices, transmitting devices, or display devices of a computing platform.
[0115] The systems (one or more) discussed herein are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement of components in which one or more inputs provide a conditioned result. Suitable computing devices include general-purpose microprocessor-based computer systems that access software stored to program or configure computing systems, ranging from general-purpose computing devices to dedicated computing devices that implement one or more implementations of the subject herein. The teachings contained herein may be implemented in software using any suitable programming, scripting, or other type of language or combination of languages for use in programming or configuring computing devices.
[0116] Implementations of the methods disclosed herein may be performed in the operation of such computing devices. The order of the blocks presented in the above examples is changeable; for example, blocks can be rearranged, combined, or divided into subblocks. Certain blocks or processes can be executed in parallel.
[0117] The use of “fitted to” or “configured to” in this specification means an unlimited and inclusive language that does not exclude devices that fit or are configured to perform additional tasks or steps. Furthermore, the use of “based on” means an unlimited and inclusive language that a process, step, calculation or other action “based on” one or more enumerated conditions or values may actually be based on additional conditions or values beyond those enumerated. The headings, lists and numbering included herein are for illustrative purposes only and are not intended to limit the scope of the explanation.
[0118] In this specification, terms such as “first,” “second,” etc., may be used to describe various objects, but it should be understood that these objects should not be limited by these terms. These terms are used solely to distinguish one object from another. For example, a “first node” can be called a “second node” without changing the meaning of the description, as long as the name is consistently changed for all occurrences of “first node” and for all occurrences of “second node.” Both the first node and the second node are nodes, but they are not the same node.
[0119] The terms used herein are for the purpose of describing specific implementations and are not intended to limit the scope of the claims. When used in the descriptions of the implementations described and in the appended claims, the singular forms "a," "an," and "the" are intended to include the plural form unless the context explicitly indicates otherwise. Furthermore, when used herein, the term "or" should be understood to mean and include any possible combination of one or more of the related enumerated items. When the terms "comprises" or "comprising" are used herein, they specify the presence of the described features, integers, steps, actions, objects, or components, but do not exclude the presence or addition of one or more other features, integers, steps, actions, objects, components, or groups thereof.
[0120] When used herein, the term "if" may be interpreted, depending on the context, as meaning "at the time" or "on the occasion of" or "in accordance with the determination" that the previously stated condition is true, or "in accordance with the determination" or "in accordance with the determination" or "in accordance with the determination." Similarly, the phrases "[when it is determined that the previously stated condition is true]," "[when the previously stated condition is true]," or "[when the previously stated condition is true]" may be interpreted as meaning "in the event of the determination that the previously stated condition is true," "in accordance with the determination," "in accordance with the determination," "when detected," or "in accordance with the determination."
[0121] The foregoing description and summary of the present invention should be understood to be illustrative and illustrative in all respects, but not restrictive, and the scope of the invention disclosed herein should be determined in accordance with the full scope permitted by patent law, not solely from the detailed description of exemplary implementations.
[0122] The implementations shown and described herein are merely illustrative of the principles of the present invention, and it should be understood that various modifications can be made by those skilled in the art without departing from the scope and spirit of the invention.
Claims
1. It is a method, In the device's processor, The acquisition of user representation data comprising user representation data of at least a portion of the user, wherein the user representation data is based on a first set of sensor data including an image of the user acquired during the registration process, and the user representation data includes splat parameter data corresponding to a plurality of three-dimensional (3D) positions. Modify the user representation data based on a second set of sensor data acquired after the registration process, A method comprising providing a view of a user representation based on the modified user representation data, wherein providing the view comprises generating a plurality of splats based on the splat parameter data of the modified user representation data.
2. The method according to claim 1, wherein the at least portion of the user includes the user's face portion and additional portions.
3. The method according to claim 1, wherein the user representation data is based on a UV map and the 3D point cloud points associated with distribution data that defines the size and shape for rendering the 3D point cloud points as splats corresponding to each point of the UV map.
4. The method according to claim 1, wherein the splat parameter data includes 3D Gaussian parameters for each 3D position.
5. The method according to claim 4, wherein the 3D Gaussian parameter includes at least one of position information, color information, covariance information, transparency information, orientation, opacity information, range information for each axis, rotation data, scale, and semantic information.
6. The method according to claim 1, wherein the user representation data includes 3D mapping information including feature values and position information for each map point.
7. The method according to claim 1, wherein the user representation data is modified based on body posture data acquired during the registration process, during a communication session with another device, or a combination thereof.
8. The method according to claim 1, wherein the device is a viewer's device, and the user representation data is modified based on an additional set of sensor data acquired during a communication session with the sender's device associated with the user representation.
9. The method according to claim 1, wherein the user expression data is generated and updated during the registration process based on images of the user's face captured while the user is making a number of different facial expressions.
10. The method according to claim 1, wherein the technique generates the user representation data via a machine learning model trained using training data acquired through one or more sensors in one or more environments.
11. The method according to claim 1, wherein providing the view of the user representation based on the modified user representation data includes displaying the user representation in an extended reality (XR) environment.
12. Modifying the view of the user representation by adjusting the user representation based on at least one color attribute among multiple color attributes of the environment, at least one light attribute among multiple light attributes of the environment, or a combination thereof. The method according to claim 1, further comprising:
13. The method according to claim 1, wherein the user representation data is acquired in a first physical environment, and the user representation is displayed in a view of a second physical environment different from the first physical environment.
14. The method according to claim 1, wherein the user representation is a 3D user representation.
15. It is a device, Non-temporary computer-readable storage medium and A device comprising one or more processors coupled to the non-temporary computer-readable storage medium, wherein the non-temporary computer-readable storage medium contains program instructions, and when the program instructions are executed on the one or more processors, the device causes the one or more processors to perform an operation, wherein the operation is The acquisition of user representation data comprising user representation data of at least a portion of the user, wherein the user representation data is based on a first set of sensor data including an image of the user acquired during the registration process, and the user representation data includes splat parameter data corresponding to a plurality of three-dimensional (3D) positions. Modify the user representation data based on a second set of sensor data acquired after the registration process, A device that provides a view of a user representation based on the modified user representation data, wherein providing the view includes generating a plurality of splats based on the splat parameter data of the modified user representation data.
16. The device according to claim 15, wherein the user representation data is based on a UV map and the 3D point cloud points associated with distribution data that defines the size and shape for rendering the 3D point cloud points as splats corresponding to each point of the UV map.
17. The device according to claim 15, wherein the splat parameter data includes 3D Gaussian parameters for each 3D position, and the 3D Gaussian parameters include at least one of position information, color information, covariance information, transparency information, orientation, opacity information, range information on each axis, rotation data, scale, and semantics information.
18. The device according to claim 15, wherein the user representation data is modified based on body posture data acquired during the registration process, during a communication session with another device, or a combination thereof.
19. The device according to claim 15, wherein the device is a viewer's device, and the user representation data is modified based on an additional set of sensor data acquired during a communication session with a sender's device associated with the user representation.
20. A non-temporary computer-readable storage medium that stores program instructions executable on a device to perform an action, wherein the action is: The acquisition of user representation data comprising user representation data of at least a portion of the user, wherein the user representation data is based on a first set of sensor data including an image of the user acquired during the registration process, and the user representation data includes splat parameter data corresponding to a plurality of three-dimensional (3D) positions. Modify the user representation data based on a second set of sensor data acquired after the registration process, A non-temporary computer-readable storage medium comprising providing a view of a user representation based on the modified user representation data, wherein providing the view includes generating a plurality of splats based on the splat parameter data of the modified user representation data.