Gaussian sputter for user representation
By using 3D Gaussian sputtering and UV mapping technology, high-quality user representations are generated, solving the problem of inaccurate user representations in existing technologies and achieving faster and more accurate user appearance rendering.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2026-03-31
AI Technical Summary
Existing technologies cannot accurately or truthfully represent a user's current appearance, leading to inaccurate user representations.
The method uses a three-dimensional Gaussian sputtering-based approach to generate user representations. It utilizes 3D Gaussian sputtering volumes and UV mapping technology to generate high-quality user representation data from sensor data on the device and render the user's view in real time.
It achieves more accurate and faster user representations, reduces computation and resource requirements, and improves the quality and real-time performance of user representations.
Smart Images

Figure CN121767524A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates generally to electronic devices, and more particularly to systems, methods and devices for representing a user in computer-generated content. Background Technology
[0002] Existing technologies may not be able to accurately or truthfully represent a current (e.g., real-time) representation of a user's appearance using an electronic device. For example, a device may provide a representation of a user based on images of the user's face obtained minutes, hours, days, or even years ago. Such representations may not accurately depict the user's appearance. Therefore, it may be desirable to provide a device that can effectively provide a more accurate, truthful, and / or current representation of the user. Summary of the Invention
[0003] The various specific embodiments disclosed herein include devices, systems, and methods for generating user representation views based on three-dimensional (3D) Gaussian sputtering. Gaussian sputtering enables the real-time rendering of high-quality, realistic scenes from a sparse set of images. Specifically, a first set of captured user data (e.g., registration data) can be used at a first device (e.g., a transmitting device) to generate user representation data including sputtering body parameter data (e.g., a 23-channel Gaussian UV map). The user representation data can be modified based on real-time user data. A view of the user representation can be provided to a viewing device (e.g., a real-time view rendering a sender's avatar) by generating sputtering bodies corresponding to the modified user representation data. A character is a representation of the user, such as an avatar. Advantageously, sputtering avoids the need to use meshes to avoid holes and offers other advantages. A 3D representation of the user at multiple moments can be generated on the viewing device, which combines the data and uses the combined data to render the view, for example, during a real-time communication (e.g., virtual communication or coexistence) session.
[0004] In some implementations, the data associated with each sputterant represented by the user can represent texture / color, location, sputterant shape, transparency level, covariance (e.g., how the sputterant is stretched / scaled), semantics (e.g., hair, mouth, skin, glasses, accessories, and / or other features), etc. Sputterants can be arranged in a parameterized two-dimensional (2D) raster structure corresponding to a surface. The parameterization can correspond to a human face used for high-quality face reconstruction. Parameterization of the sputterant distribution provides higher-quality data, allows for faster training of machine learning models, and can provide faster (e.g., real-time) rasterization. In some implementations, 3D mapping information (e.g., identifying the x, y, z positions corresponding to UV coordinates in a UV map) can be generated at registration (e.g., a Gaussian UV map).
[0005] Using 3D Gaussian sputtering volumes and UV mapping offers several advantages. For example, compared to using 3D meshes, 3D point clouds, etc., 3D Gaussian sputtering volumes may require less computation, resources, and bandwidth, while achieving a more accurate user representation.
[0006] Generally speaking, an innovative aspect of the subject matter described in this specification can be embodied in a method comprising the following actions: at a processor of a device, obtaining at least a portion of user representation data of a user, wherein the user representation data is based on a first set of sensor data including an image of the user obtained during a registration process, and the user representation data includes sputtering parameter data corresponding to multiple three-dimensional (3D) positions; modifying the user representation data based on a second set of sensor data obtained after the registration process; and providing a view of the user representation based on the modified user representation data, wherein providing the view includes generating multiple sputterings based on the sputtering parameter data of the modified user representation data.
[0007] These and other implementation schemes may optionally include one or more of the following features.
[0008] In some aspects, the user's at least part includes a facial portion and additional user portions. In some aspects, the user representation data is based on a UV map and 3D point cloud points associated with distribution data, which defines the size and shape of the sputtering body used to render the 3D point cloud points as corresponding to each point in the UV map.
[0009] In some aspects, the sputtering parameter data includes 3D Gaussian parameters for each 3D location. In other aspects, the 3D Gaussian parameters include at least one of the following: position information, color information, covariance information, transparency information, orientation, opacity information, range information in each axis, rotation data, scale, and semantic information.
[0010] In some aspects, user representation data includes 3D mapping information, which includes feature values and positional information for each mapping point. In other aspects, user representation data is modified based on body posture data obtained during the registration process, during a communication session with another device, or a combination thereof.
[0011] In some respects, the device is the viewer's device, where the user representation data is modified based on a set of additional sensor data acquired during a communication session with a device associated with the sender of the user representation.
[0012] In some respects, user representation data is generated and updated during the registration process based on images of the user's face captured while the user is making multiple different facial expressions.
[0013] In some respects, a technique generates user representation data via a machine learning model trained using training data obtained via one or more sensors in one or more environments.
[0014] In some respects, providing a view of the user representation based on the modified user representation data includes displaying the user representation in an extended reality (XR) environment.
[0015] In some aspects, the action also includes modifying the view of the user representation by adjusting the user representation based on at least one color attribute of a plurality of color attributes of the environment, at least one light attribute of a plurality of light attributes of the environment, or a combination thereof. In some aspects, In some respects, the user representation data is obtained in a first physical environment, and the user representation is displayed in a view of a second physical environment different from the first physical environment. In some respects, the user representation is a 3D user representation.
[0016] In some embodiments, a non-transitory computer-readable storage medium stores instructions that are computer-executable to perform or cause to perform any of the methods described herein. In some embodiments, an apparatus includes one or more processors, non-transitory memory, and one or more programs; the one or more programs are stored in the non-transitory memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing or causing to perform any of the methods described herein. Attached Figure Description
[0017] To enable those skilled in the art to understand this disclosure, more detailed descriptions can be made with reference to aspects of some exemplary embodiments, some of which are shown in the accompanying drawings.
[0018] Figure 1 Examples of devices for obtaining sensor data from users are shown, based on some specific implementations.
[0019] Figure 2 An exemplary electronic device is illustrated, which operates in different physical environments during a communication session between a first user at a first device and a second user at a second device, according to some specific implementation, and a combined 3D representation view of the second user at the first device.
[0020] Figures 3A to 3D Examples of 3D Gaussian sputtering bodies used in generating 3D representations of views are illustrated according to some specific implementations.
[0021] Figure 4 Examples are shown of generating and displaying a stereoscopic view of a sputtered body on a device according to some specific implementations.
[0022] Figure 5 Examples are shown of generating user representations based on rendering sputtering parameter data from some specific implementations.
[0023] Figures 6A to 6C Examples are shown of generating user representations based on rendering sputtering parameter data from some specific implementations.
[0024] Figure 7 It is a flowchart representation of a method for providing a user-represented view based on some specific implementations of rendering sputtering body parameter data.
[0025] Figure 8 This is a block diagram illustrating device components according to some specific implementations of exemplary devices.
[0026] Figure 9 It is a block diagram based on some specific implementation examples of head-mounted devices (HMDs).
[0027] As is customary practice, various features illustrated in the accompanying drawings may not be drawn to scale. Therefore, for clarity, the dimensions of various features may be arbitrarily expanded or reduced. Furthermore, some drawings may not depict all components of a given system, method, or apparatus. Finally, similar reference numerals may be used throughout the specification and drawings to denote similar features. Detailed Implementation
[0028] Numerous details have been described to provide a thorough understanding of the exemplary embodiments illustrated in the accompanying drawings. However, the drawings illustrate only some exemplary aspects of this disclosure and should not be considered limiting. Those skilled in the art will recognize that other effective aspects or variations do not include all the specific details set forth herein. Furthermore, well-known systems, methods, components, devices, and circuits have not been described exhaustively so as not to obscure further relevant aspects of the exemplary embodiments described herein.
[0029] Figure 1 An example environment 100 is illustrated, in which an exemplary electronic device 105 operates within a physical environment 102. In some embodiments, the electronic device 105 may be able to share information with each other or with intermediate devices, such as information systems. Additionally, the physical environment 102 includes a user 110 wearing the device 105. In some embodiments, the device 105 is configured to present views of extended reality (XR) environments, which may be based on the physical environment 102 and / or include added content, such as virtual elements providing textual narration.
[0030] exist Figure 1In the example, physical environment 102 is a room that includes physical objects such as wall hangings 120, plants 125, and tables 130. Electronic device 105 may include one or more cameras, microphones, depth sensors, motion sensors, or other sensors that can be used to capture information about physical environment 102 and the objects within it and to evaluate the physical environment and these objects, as well as to capture information about user 110.
[0031] exist Figure 1 In the example, device 105 includes one or more sensors 116 (e.g., inward-facing sensors and outward-facing cameras) that capture light intensity images, depth sensor images, audio data, or other information about user 110. For example, one or more sensors 116 may capture images of the user's (e.g., user 110's) forehead, eyebrows, eyes, eyelids, cheeks, nose, lips, chin, face, head, hands, wrists, arms, shoulders, torso, legs, or other body parts. Additionally, one or more sensors 116 may capture images of elements / materials attached to or worn by user 110 (e.g., glasses, earrings, and / or other accessories). For example, inward-facing sensors may see the interior of device 105 (e.g., the user's eyes and the area around the eyes), and other external cameras may capture the user's face outside device 105 (e.g., an egocentric camera pointing towards the exterior of device 105). As an example, sensor data about the user's eyes 111 can indicate various user characteristics, such as the user's gaze direction 119 over time, the user's saccade behavior over time, the user's eye expansion behavior over time, etc. One or more sensors 116 can capture audio information, including the user's voice and other sounds emitted by the user, as well as sounds within the physical environment 100.
[0032] In some embodiments, device 105 includes an eye-tracking system for detecting eye position and eye movement via eye gaze characteristic data. For example, the eye-tracking system may include one or more infrared (IR) light-emitting diodes (LEDs), an eye-tracking camera (e.g., a near-IR (NIR) camera), and an illumination source (e.g., an NIR light source) that emits light (e.g., NIR light) toward the user 110's eye. Furthermore, the illumination source of device 105 may emit NIR light to illuminate the user 110's eye, and the NIR camera may capture images of the user 110's eye. In some embodiments, the images captured by the eye-tracking system may be analyzed to detect the position and movement of the user 110's eye or to detect other information about the eye, such as color, shape, state (e.g., widening, strabismus, etc.), pupil dilation, or pupil diameter. Furthermore, the gaze point estimated from the eye-tracking images enables gaze-based interaction with content displayed on the near-eye display of device 105.
[0033] Additionally, one or more sensors 116 may capture images of the physical environment 100 (e.g., externally oriented sensors). For example, one or more sensors 116 may capture images of the physical environment 100 including physical objects such as wall hangings 120, plants 125, and tables 130. Furthermore, one or more sensors 116 may capture images (e.g., light intensity images and / or depth data).
[0034] One or more sensors (such as one or more sensors 115 on device 105) may identify user information based on proximity or contact with a portion of user 110. For example, one or more sensors 115 may capture sensor data that can provide biometric information related to the user’s cardiovascular status (e.g., pulse), body temperature, respiratory rate, etc.
[0035] One or more sensors 116 or one or more sensors 115 can capture data that can determine a user orientation 121 within the physical environment. In this example, user orientation 121 corresponds to the direction in which the user 110's torso is facing.
[0036] Some specific implementations disclosed herein determine user understanding based on sensor data obtained by a device worn by the user (such as the first device 105). Such user understanding can indicate the user's state of being associated with providing assistance. In some examples, the user's appearance or behavior, or their understanding of the environment, can be used to identify the need for or expectation of assistance, enabling the user to obtain such assistance. For example, based on determining such user state, augmentations can be provided to assist the user by enhancing or supplementing their capabilities (e.g., providing guidance or other information about the environment to disabled / impaired individuals).
[0037] The content can be visible (e.g., displayed on a display of device 105) or audible (e.g., generated as audio 118 by a speaker of device 105). In the case of audio content, audio 118 can be generated in a manner that makes it likely only user 110 will hear it (e.g., via a speaker close to user's ear 112 or at a volume below a threshold), making it unlikely that nearby people will hear it. In some specific implementations, the audio pattern (e.g., volume) is determined based on whether other people are within a threshold distance or based on how close other people are relative to user 110.
[0038] In some implementations, the content provided by device 105 and the sensor features of device 105 may be provided using components, sensors, or software modules that are small enough and efficient in terms of power consumption and use to be adapted to and otherwise used in lightweight, battery-powered wearable products, such as wireless earbuds or other ear-hook devices or head-mounted devices (HMDs), such as smart / augmented reality (AR) glasses. A combination of multiple devices may be used to facilitate the features. For example, a smartphone (which wirelessly connects to and interacts with wearable devices) may provide computing resources, connectivity to cloud or internet services, location services, etc.
[0039] Figure 2 Exemplary electronic devices operating in different physical environments during a communication session between a first user at a first device and a second user at a second device, according to some specific implementations, are illustrated, along with a 3D representation view of the second user at the first device. Specifically, Figure 2 Exemplary operating environments 200 are illustrated for electronic devices 210 and 265 operating in different physical environments 202 and 250 during a communication session (e.g., when electronic devices 210 and 265 share information with each other or with an intermediate device (such as a communication session system / server)). Figure 2 In this example, physical environment 202 is a room that includes wall mount 212, plant 214, and table 216 (e.g., Figure 1 The physical environment 202. Electronic device 210 includes one or more cameras, microphones, depth sensors, or other sensors that can be used to capture information about the physical environment 202 and objects within it, as well as information about the user 225 of electronic device 210 (e.g., a handheld device), and to evaluate the physical environment and these objects. Information about the physical environment 202 and / or the user 225 can be used to provide visual content (e.g., for user representation) and audio content (e.g., for audible speech or text transcription) during a communication session. For example, a communication session may provide one or more participants in a 3D environment (e.g., users 225, 260) with views generated based on camera images and / or depth camera images of the physical environment 202, and a representation of user 225 based on camera images and / or depth camera images of user 225.
[0040] Additionally, in Figure 2In this example, physical environment 250 is a room including wall mount 252, sofa 254, and coffee table 256. Electronic device 265 includes one or more cameras, microphones, depth sensors, or other sensors that can be used to capture information about physical environment 250 and objects within it, as well as information about user 260 (e.g., user wearable device or HMD device, such as device 105) and to evaluate the physical environment and these objects. Information about physical environment 250 and / or user 260 can be used to provide visual and audio content during a communication session. For example, the communication session can provide a view of a 3D environment generated based on camera images and / or depth camera images (from electronic device 265) of physical environment 250, and a representation of user 260 based on camera images and / or depth camera images (from electronic device 265). For example, in communication with device 265 via communication session command set 282 (e.g., through information system 290 via network connection 285), the 3D environment can be transmitted by device 210 via communication session command set 280. Information system 290 can coordinate the encryption / decryption and pre-download of assets (e.g., 3D asset data, such as data associated with user representations 240, 275) between two or more devices (e.g., electronic devices 210 and 265).
[0041] Figure 2 An example of a view 205 of a virtual environment (e.g., 3D environment 230) at device 210 is illustrated, wherein a representation 232 of a wall decoration 252 and a user representation 240 (e.g., a user 260's avatar) are provided, provided that the user representation of each user is viewed during a specific communication session. Specifically, user representation 240 of user 260 is generated based on one or more user representation techniques used for generating more realistic avatars in real time. The generation of user representations is further discussed herein.
[0042] Additionally, electronic device 265 within physical environment 250 provides view 266, which enables user 260 to view representation 272 of wall mount 212 and representation 275 (e.g., from mid-torso upwards) of at least a portion of user 225 (e.g., a portrait) within 3D environment 270. In other words, user representation 240 of user 260 is generated at device 210 by generating a combined 3D representation of user 260 for multiple moments within a time period based on data obtained from device 265 (e.g., a frame-specific 3D representation of user 260). Alternatively, in some embodiments, user representation 240 of user 260 is generated at device 265 (e.g., a speaker's transmitting device) and transmitted to device 210 (e.g., a viewing device viewing a speaker's avatar). In some embodiments, each of the 3D representation 240 of user 260 and the 3D representation 275 of user 225 is generated by generating a sputter corresponding to the modified user representation data according to the techniques described herein.
[0043] exist Figure 2 In the examples, electronic device 210 is exemplified as a handheld device, and electronic device 265 is exemplified as a head-mounted device (HMD). However, either electronic device 210 or 265 can be a mobile phone, tablet, laptop, etc., or, like electronic device 265, can be worn by a user (e.g., head-mounted devices (glasses), headphones, ear-mounted devices, etc.). In some implementations, the functionality of devices 210 and 265 is implemented via two or more devices (e.g., mobile device and base station or head-mounted device and ear-hook device). Various functions can be distributed across multiple devices, including but not limited to power functions, CPU functions, GPU functions, storage functions, memory functions, visual content display functions, audio content production functions, etc. The multiple devices used to implement the functionality of electronic devices 210 and 265 can communicate with each other via wired or wireless communication. In some implementations, each device communicates with a separate controller or server to manage and coordinate the user experience (e.g., a communication session server). Such a controller or server may be located in physical environment 202 and / or physical environment 250 or may be remote relative to that physical environment.
[0044] In addition, Figure 2In the examples, 3D environments 230 and 270 are XR environments based on a common coordinate system that can be shared with other users (e.g., a virtual room for avatars in a multi-user communication session). In other words, the common coordinate systems of 3D environments 230 and 270 are different from the coordinate systems of physical environments 202 and 250, respectively. For example, a common reference point can be used to align the coordinate systems. In some implementations, the common reference point can be a virtual object within the 3D environment that each user can visualize within their respective view. For example, a common center piece table around which a user representation (e.g., a portrait of a user) is positioned within the 3D environment. Alternatively, the common reference point may not be visible within each view. For example, the common coordinate system of the 3D environment can use the common reference point to position each respective user representation (e.g., around a table / desk). Therefore, if the common reference point is visible, each view of the device will be able to visualize the “center” of the 3D environment for perspective when viewing other user representations. The visualization of the common reference point can become more relevant to the multi-user communication session, allowing each user’s view to add perspective to the location of each other user during the communication session.
[0045] In some implementations, each user's representation can be real or unreal and / or represent the user's current and / or previous appearance. For example, a photorealistic representation of user 225 or user 260 can be generated based on a combination of real-time images and previous images. Previous images can be used to generate representations of portions of the user's face that are not available in real-time image data (e.g., portions of the user's face that are not in the field of view of the cameras or sensors of electronic devices 210 or 265 or that may be obscured, for example, by headphones or other means). In one example, electronic devices 210 and 265 are HMDs, and the real-time image data of the user's face includes images of the user's cheeks and mouth from a downward-facing camera and images of the user's eyes from an inward-facing camera. This real-time image data can be combined with previous image data of other parts of the user's face, head, and torso that are not currently visible from the device's sensors. Previous data about the user's appearance may be obtained at an earlier time during a communication session, during previous use of the electronic device, during a registration process for obtaining sensor data of the user's appearance from multiple viewpoints and / or conditions, or otherwise.
[0046] In some specific implementations, such as Figure 2The generation of one or more user representations for communication sessions illustrated herein (e.g., generating user representations 240, 275) can be based on one or more rendering techniques, such as using 3D meshes or 3D point clouds. However, the technique described herein utilizes a 3D Gaussian sputtering method using UV mapping. Several advantages can be achieved using a simple set of depth values defined relative to multiple points, as expressed by a 3D Gaussian sputtering using UV mapping. This set of values may require less computation and bandwidth than using 3D meshes or 3D point clouds, while achieving a more accurate user representation than RGBDA images. Furthermore, this set of values can be formatted / encapsulated in a manner similar to existing formats (e.g., RGBDA images), enabling more efficient integration with systems based on such formats.
[0047] Figures 3A to 3D Examples of 3D Gaussian sputtering bodies used in generating 3D representations of views are illustrated according to some specific implementations. For example, 3D Gaussian sputtering (3DGS) can be used for 3D modeling to represent a complex scene as a combination of numerous shaded 3D Gaussian bodies rendered into a camera view via sputtering-based rasterization. The position, size, rotation, color, and opacity of these Gaussian sputtering bodies can then be adjusted via differentiable rendering and gradient-based optimization such that they represent the 3D scene given by a set of input images.
[0048] Figure 3A An example of a 3D Gaussian sputtering body 310 is shown, for example, an elliptical shape formed by a 3D Gaussian distribution. The 3D Gaussian sputtering body 310 can be used to represent position (…). µ ), such as xyz coordinates. 3D Gaussian sputtering volumes 310 can also represent rotation and scaling (e.g., covariance matrix), opacity ( ⍺ Color (e.g., RGB values), anisotropic covariance, spherical harmonic (SH) coefficients, and / or semantic category (e.g., glasses, accessories, etc.). Figure 3B An example of an environment 320 for rendering a sputtering body 326 based on the visibility orientation of a camera 321 is shown. Figure 3C An example is shown where sputtering bodies are ordered along ray 330 in the camera viewpoint direction (e.g., sputtering bodies 331, 332, 333, and 334 are identified and ordered along ray 330). For example, Gaussian sputtering is a technique where individual 3D points are represented as a Gaussian distribution (e.g., "sputter") with color values that change according to the viewpoint. Spherical harmonics are used to model this view-dependent color variation, enabling the rendering of high-quality, realistic scenes in real-time from a sparse set of images. For example, each point has a color calculated based on its position relative to the camera, allowing for realistic shading across different viewpoints. Figure 3DAn environment 340 is illustrated for blending sputterants 341, 342, 343, 344, and 345, which can be viewed from a camera viewpoint along the direction of ray 330 by synthesizing sputterants 331, 332, 333, and 334 on an image plane. Some specific implementations may use screen-to-sputterant (e.g., similar to ray casting techniques), sputterant-to-screen (e.g., similar to projection techniques), combinations thereof, or other techniques for synthesizing sputterants.
[0049] Figure 4 An example environment 400 for generating and displaying a stereoscopic view of a sputtered object on a device (e.g., an HMD) is illustrated according to some specific embodiments. For example, device 410 is an HMD including a first display 420 for a left-eye view and a second display 430 for a right-eye view. The first display 420 and the second display 430 can then view the rendered sputtered object 450 for each corresponding viewpoint. In some specific embodiments, the image generated for each viewpoint can be rendered as a single monochrome image (e.g., rendered for one eye), or the image can be presented as a stereoscopic view, such as... Figure 4 As illustrated in the illustration. Additionally or alternatively, in some implementations, the generated image for each viewpoint may be rendered as a single raster grid or a combination of stereo raster grids.
[0050] Figure 5 Examples are illustrated for generating user (e.g., avatar) representations based on render sputterant parameter data in some specific implementations. Specifically, Figure 5 An example user representation process 500 is illustrated, which is used to obtain registration data (e.g., registration images 512a, 512b, 512c) from the viewpoints of the sending device and the receiving device for the sender, from the registration process 510, in order to generate a view 550 of user representation 552 (e.g., avatar) using Gaussian sputtering techniques.
[0051] The registration process 510 illustrates the user (e.g., Figure 1The registration process 510 may include user registration (e.g., pre-registration of registration data) and acquisition of sensor data (e.g., real-time data registration). In some implementations, as illustrated in Figure 511, user registration may include the user (e.g., user 110) using external sensors on device 105 to obtain a full-view image of his or her face and a portion of his or her upper body, and thus the user may remove and orient device 105 (e.g., HMD) toward his or her face / body during the registration process. A registration avatar may be generated when the system acquires image data (e.g., RGB image) of the user's face while the user is providing different facial expressions. For example, the user may be instructed to "raise your eyebrows," "smile," "frown," etc., to provide the system with a range of facial features for the registration process. A preview of the registration avatar may be shown to the user as he or she provides a registration image to visualize the state of the registration process. Registration image data 510 may include registration avatars with different user expressions and from different viewpoints (e.g., a front view in registration image 512a, a right view in registration image 512b, and a left view in registration image 512c). In some examples, more or fewer different expressions and / or viewpoints may be used to obtain sufficient data for the registration process. In some implementations, at the final stage of registration, the user may be presented with selection options for generating the user representation, such as light-induced correction / color correction, adding / removing attachments, etc. (e.g., the user can choose their best identity to represent as their avatar).
[0052] In some implementations, the transformation from registered image data to feature data 522 can occur as part of a feature data process 520 for multiple expressions (e.g., for different sets of feature data for different expressions) (e.g., via a transformer). For example, feature data 522 may include learned feature information of user 110 obtained from the registered image, such as skin, color, and other semantic information for each pixel. Feature data 522 may include a list of locations for each feature value (e.g., multiple feature channels). Then, as part of a Gaussian UV mapping process 530, feature data 522 can be decoded by a decoder to generate a 3D Gaussian UV map for each feature. The 3D points of feature data 532 can be mapped to Gaussian parameters of the UV map 534 (e.g., 3D points + Gaussian parameters). For example, the UV map stores the x, y, and z positions for sputterant parameters (e.g., color (view-related / harmonic information), covariance, α / transparency, orientation, opacity, range in each axis, rotation, scale, and semantic information (e.g., skin, hair, cheeks, nose, lips, eyebrows, attachments, etc.)). In other words, the Gaussian UV mapping process 530 can obtain 3D point information, which includes sufficient information (e.g., 3D vector projection) to determine which sputterant can be generated.
[0053] In some implementations, process 500 occurs after 3D Gaussian UV mapping data is generated from Gaussian UV mapping process 530 (e.g., at the sender's device after registration). The system (e.g., at the viewer's device) can obtain the 3D Gaussian UV mapping data (e.g., feature data 532 mapped to Gaussian parameters via UV mapping map 534) and project the Gaussian data using the current viewpoint (e.g., viewpoint data 536) to determine a 2D Gaussian UV mapping map 542 (e.g., 2D points + Gaussian parameters) for Gaussian UV mapping process 540. Gaussian sputtering can then be used for rendering 545 to generate a view for user representation 552 in user representation generation process 550. For example, 3D Gaussian sputtering techniques use 2D points from the UV mapping map and associated Gaussian parameters to render an image using Gaussian sputtering based on the viewer's current viewpoint (e.g., viewpoint data 536).
[0054] Figures 6A to 6C Examples of generating a user (e.g., avatar) representation based on rendering sputtering volume parameter data are illustrated. Specifically, Figure 6 illustrates an example process 600 for obtaining Gaussian UV mapping data from a transmitter based on the viewpoint of the receiver relative to the viewer, in order to generate a view 550 of a user representation 552 (e.g., avatar) associated with the transmitter using Gaussian sputtering techniques. The example rendering process 600 based on sputtering volume parameter data illustrates, as shown in the example... Figure 6AThe timeline marker 603 illustrates the division of the registration phase 602 and the runtime phase 604. In other words, the registration phase 602 can occur at some point before the communication session, and each other process after the timeline marker 603 for the runtime phase 604 occurs during the real-time communication session (e.g., generating a real-time avatar of the sending device's user for the viewing device). Furthermore, the example rendering process 600 illustrates the data flow process between the sending device (e.g., sender phase 606) and the receiving device (e.g., receiver phase 608), as illustrated by the send / receive timeline marker 607, as... Figure 6C exemplified in .
[0055] In an exemplary implementation, process 600 begins at registration phase 602. Registration phase 602 may include an offline registration process in which a user's identity representation may be generated. The identity representation may be a set of latencies extracted from a Gaussian UV map, some type of canonical representation (e.g., a canonical (or underlying) Gaussian UV map) that may be generated for each user, or a combination thereof.
[0056] In an exemplary implementation, registration phase 602 may begin with registration process 610, which exemplifies users captured during registration (e.g., ...). Figure 1 The registration process 610 may include user registration (e.g., pre-registration of registration data) and obtaining sensor data (e.g., real-time data registration) to capture the registration image 612. Figure 5 As illustrated in image 511, user registration may include a user (e.g., user 110) using an external sensor on device 105 to obtain a full-view image of his or her face, and thus removing and orienting device 105 (e.g., HMD) toward his or her face during the registration process. A registration avatar may be generated when the system acquires image data (e.g., RGB image) of the user's face while the user is providing different facial expressions. For example, the user may be instructed to "raise your eyebrows," "smile," "frown," etc., to provide the system with a range of facial features for the registration process. A preview of the registration avatar may be shown to the user as they provide a registration image to visualize the state of the registration process. The registration image data may include registration avatars with different user expressions and from different viewpoints.
[0057] In some implementations, the light normalization process 614 can be applied to the registration data. For example, a light normalization process 614 can be provided that obtains a cropped image of the registration image 612 (e.g., a segmented head or face of a user), illustrating poor lighting conditions as shown on the user's face. The light normalization process 614 can detect one or more attributes associated with the poor lighting conditions and adjust one or more registration images accordingly. The adjustments can then be applied to generate one or more post-processed registration images that illustrate the removal of attributes associated with the poor lighting conditions on the user's face. For example, the registration image 612 may be too dark, so the light normalization process 614 can brighten the user's facial area, and the post-processing can apply the brightened facial area to the entire post-processed registration image data.
[0058] In some implementations, after normalizing the registered image, registration phase 602 proceeds to Gaussian UV mapping process 620 to generate 3D Gaussian UV mapping data 622. For example, the transformation from registered image data to feature data can occur as part of a transformation (e.g., via a transformer). For instance, the feature data may include learned feature information of user 110 obtained from registered image 612, such as skin, color, and other semantic information for each pixel. The feature data may include a list of locations for each feature value (e.g., 14 feature channels). Then, as part of Gaussian UV mapping process 620, the feature data may be decoded by a decoder to generate 3D Gaussian UV mapping data 622 for each feature. The 3D points of the feature data can be mapped to Gaussian parameters of the UV mapping map (e.g., 3D points + Gaussian parameters). For example, a UV map can store the x, y, z positions for sputter body parameters such as color (view-related / harmonic information), covariance, α / transparency, orientation, opacity, range in each axis, rotation, scale, and semantic information (e.g., skin, hair, cheeks, nose, lips, eyebrows, etc.).
[0059] In some implementations, the Gaussian UV mapping process 620 can generate a canonical representation of the user, such as a canonical Gaussian UV map. The canonical Gaussian UV map can be synthesized from multiple registered expressions and can be used as a stable, personalized reference for subsequent real-time updates. In some implementations, the canonical representation can be a neutral expression, an average, or an idealized synthesis and can be stored and used as a starting point for runtime animation. In runtime phase 604, the real-time feature latency (expression latency) can be used to calculate the increments or modifications applied to the canonical Gaussian UV map, thereby producing the final frame-specific user representation.
[0060] In some implementations, the Gaussian UV mapping process 620 may also acquire additional feature data associated with the user's body (e.g., body tracking information 623). Body tracking information 623 may include one or more parts of the user's body other than the face (e.g., head, hands, upper / lower torso, etc.). Body tracking information 623 may be divided into different data streams based on the different parts of the body being tracked, such as the head / face as one data stream, the hands as another, and the upper and / or lower torso as yet another. In some implementations, body tracking information 623 may be refined via a refining process 624, which provides the system with the ability to update the skeletal model of body tracking information 623 (e.g., fine-tuning). For example, refining process 624 may be applied directly to the Gaussian UV map by changing sputtering parameters to achieve specific goals, such as removing skin blemishes, changing hair color, slightly altering the shape of the nose, and / or any other deformations involving color, shape, and / or appearance. The body tracking information 623 during registration can also send modified assessment body data as skeletal blending data 627 (e.g., blending weights and asset data), which can also be analyzed at another stage of the process, such as deformation and blurring of the Gaussian UV map and used for skeletal tracking during runtime stage 604.
[0061] The identity latency process 630 can obtain the extracted 3D Gaussian UV mapping data 622 from the Gaussian UV mapping process 620. For example, the identity latency process 630 can generate a mapping of the identity latency period 632, which can be used as markers, such as semantic points of the user that can be determined at registration, which may move during frame-by-frame animation and are therefore updated at runtime (e.g., corners of the mouth, other features, etc.). In other words, the Gaussian UV mapping process 620 and the identity latency process 630 can determine 3D point information that includes sufficient information to determine which sputtering body is generated (e.g., 3D vector projection) and identify portions of the user that may need to be identified by the identity latency process 630 for frame-by-frame updates (e.g., marker points 664). In other words, "marker points" refer to semantic or feature points (e.g., facial feature points, corners of the mouth, tip of the nose, etc.) that are tracked and updated during runtime.
[0062] Registration phase 602 can then store these registration assets, such as a mapping of identity latency 632 and skeletonized hybrid data 627, for future use during runtime phase 604 (e.g., during a communication session with another device). At runtime phase 604 (e.g., after a communication session with another device has been initiated), exemplary process 600 continues to capture real-time data at real-time data process 640, such as... Figure 6BAs shown. Specifically, the real-time data process 640 captures head / face tracking data 642 and body tracking data 650 (e.g., real-time sensor data for updating user representation). The head / face tracking data 642 can be used to identify expression latency 644 and head pose data 646 (e.g., HMD-to-head pose information). Expression latency 644 can be used in conjunction with identity latency 632 in animation and decoding processes to update Gaussian UV mapping data 622 using the user's current ("real-time") facial expression (e.g., animated facial expression).
[0063] In some implementations, skeletal tracking data 652 can be determined based on current (“real-time”) sensor data (e.g., body tracking data 650) and skeletal hybrid data 627 obtained from registration phase 602. Head pose data 646 can be used to update skeletal tracking data 652 (e.g., updating data associated with the sender’s head based on device pose data). Skeletal tracking data 652 can then be used by body pose process 654 to identify skeletal joint data associated with multiple (quantized) skeletal joints of the sender. Skeletal joint data can be used in conjunction with identity latency 632 in animation and decoding processes to update Gaussian UV mapping data 622 using the user’s current (“real-time”) skeletal movement (e.g., animated body pose). In some implementations, body tracking information can be sent to a decoder network to estimate complex body deformations.
[0064] The final stage of the transmitting device's runtime phase 604 (e.g., sender phase 606) is the animation and decoding process of the Gaussian UV mapping process 660, such as... Figure 6CAs shown. At Gaussian UV mapping process 660, Gaussian UV mapping data 622 and identity latency 632 are animated and decoded using expression latency 644 to determine Gaussian UV mapping data 662, which includes the identified marker points 664. Then, during the communication session, “real-time” frame-by-frame Gaussian UV mapping data 662 is sent to the receiving device as part of receiver phase 608. In some specific implementations, the receiving device also obtains skeletal blending data 627 and skeletal joint data (e.g., animated body poses) to update Gaussian UV mapping data 662 by applying body pose data (e.g., neck, shoulders, hands, etc.) during the deformation and blurring process of Gaussian buffering process 670, to generate updated Gaussian UV mapping data 662 as Gaussian buffer data 672. At runtime stage 604 at the receiving device, viewpoint data 674 (e.g., the viewpoint to be rendered) can then be obtained from the receiving device, and Gaussian buffer data 672 and viewpoint data 674 can be sent to a stereo proxy sputtering process 680, which generates stereo proxy geometry view data 682 (e.g., left and right viewpoints and depth data of RGBA sputtered data). In some implementations, for more efficient rendering, stereo proxy geometry view data 682 can be generated at 30 frames per second (FPS) but rendered at 90 FPS. Stereo proxy geometry view data 682 can then be used at a 3D Gaussian sputtering rendering process (e.g., user representation process 690) to generate user representation 692.
[0065] In some implementations, when multiple people are being called during a communication session, the level of detail (LOD) can be determined. In this case, rendering a full-resolution headshot for each person in the communication session might be too resource-intensive. Therefore, the system can determine to fall back to a lower resolution version (e.g., with fewer sputtering bodies). For example, as part of feature data process 520, the system can decode multiple Gaussian UV maps with different resolutions. The receiver can then be able to select the correct Gaussian UV map given the current situation based on one or more factors, such as, in particular, the distribution of people, viewing direction, etc.
[0066] Figure 7 This is a flowchart illustrating exemplary method 700. In some specific implementations, the device (e.g., Figure 1The technique of method 700 is performed on a device 105 to generate a view of a user representation based on rendering sputterant parameter data, according to some specific implementations. In some specific implementations, the technique of method 700 is performed on a mobile device, desktop computer, laptop computer, HMD, or server device. In some specific implementations, method 700 is performed on a processing logic component (including hardware, firmware, software, or a combination thereof). In some specific implementations, method 700 is performed on a processor executing code stored in a non-transitory computer-readable medium (e.g., memory). In some specific implementations, method 700 is implemented at a processor of a device (such as a viewing device) that presents a user representation (e.g., Figure 2 The device 210 presents a 3D representation 240 (avatar) of the user 260 based on data obtained from the device 265.
[0067] At box 710, method 700 obtains user representation data of at least a portion of the user at the device's processor. This user representation data is based on a first set of sensor data including an image of the user acquired during the registration process and includes sputtering parameter data corresponding to multiple 3D positions. In some specific implementations, the device is a viewing device that renders a user representation (e.g., an avatar). For example, as... Figure 2 As illustrated, device 210 renders a user representation 240 of user 260 (e.g., the sender). In some embodiments, this at least portion of the user includes a facial portion and additional portions of the user (e.g., head, neck, clothing, hair, body, etc.). In some embodiments, the image of the user acquired during the registration process is a two-dimensional (2D) image. Additionally or alternatively, in some embodiments, the image of the user acquired during the registration process is a 2D image plus depth data.
[0068] In some implementations, user representation data is based on a UV map and 3D point cloud points associated with distribution data that defines the size and shape of sputterants for rendering the 3D point cloud points to correspond to each point in the UV map (e.g., a 3D Gaussian map). For example, as illustrated in Figure 6, a user representation 692 is generated for a specific viewpoint based on points rendered from a Gaussian UV map 662, which combines sputterant parameters obtained from registration data and is updated for each frame based on one or more marker points 664 (e.g., a set of semantic points associated with facial features or other regions of the sender associated with the rendered user representation 692).
[0069] In some implementations, sputtering parameter data includes 3D Gaussian parameters for each 3D location. For example, sputtering parameter data may include location information, color information, covariance information, transparency information, orientation, opacity information, range information in each axis, rotation data, scale, and semantic information (e.g., skin, hair, cheeks, nose, lips, eyebrows, etc.). For instance, as illustrated in Figure 6, the Gaussian UV mapping process 620 generates 3D Gaussian UV mapping data 622 by transforming registered image data into feature data, which may include user-learned feature information obtained from the registered image, such as skin, color, and other semantic information for each pixel. For example, sputtering parameter location may identify where the sputtering is located based on xyz coordinates, sputtering parameter covariance may identify how the sputtering is stretched / scaled (e.g., a 3x3 matrix), sputtering parameter color may identify RGB colors, and sputtering parameter alpha (α) may identify the transparency of the sputtering. In some implementations, the user representation data includes 3D point cloud points associated with distributed data that defines the size and shape of the 3D point cloud points used to render the sputtering body. For example, a sputtering body model generated using a Gaussian sputtering body includes texture / color, position, sputtering body shape, etc.
[0070] In some implementations, the user representation data includes 3D mapping information, which includes feature values and position information for each mapping point. For example, the Gaussian UV mapping process 620 and the identity lurking process 630 can determine 3D point information that includes sufficient information (e.g., 3D vector projection) to identify which sputtering body was generated and to identify the user's portion that may need to be updated frame-by-frame by the identity lurking process 630 (e.g., marker point 664) (e.g., identifying the x, y, z position corresponding to the UV coordinates of the UV mapping map).
[0071] At box 720, method 700 modifies the user representation data based on a second set of sensor data obtained after the registration process. For example, the sputtering parameter data is modified based on real-time sensor data. For instance, the sputtering parameter data can be obtained from a transmitting device (such as from the registration process), and the modification of the registered sputtering parameter data can be based on obtaining real-time sensor data from the sender to determine the sender's real-time representation (e.g., a real-time view of the actual avatar of the communication session). For example, as... Figure 6CAs illustrated, at Gaussian UV mapping process 660, the receiving device acquires head / face tracking data as expression latency 644 to determine Gaussian UV mapping data 662 (e.g., applying the user's real-time expression based on markers 664, semantic or feature points (e.g., facial feature points, corners of the mouth, tip of the nose, etc.) tracked and updated during runtime). Furthermore, in some implementations, the receiving device acquires skeletal blending data 627 and skeletal joint data (e.g., animated body pose) at Gaussian buffering process 670 to update Gaussian UV mapping data 662 by applying body pose data (e.g., neck, shoulders, hands, etc.) during the deformation and blurring process of Gaussian buffering process 670, generating updated Gaussian UV mapping data 662 as Gaussian buffering data 672.
[0072] In various embodiments, user representation data may be modified for the face rather than the body, for the body rather than the face, or for both the body and the face, and / or may be modified during registration, on the sender-side device, or on the receiver-side device. In some embodiments, user representation data is modified based on body posture data obtained during the registration process, during a communication session with another device, or a combination thereof. In some embodiments, the device is the viewer's device, and the user representation data is modified based on a set of additional sensor data obtained during a communication session with a device associated with the sender of the user representation. Alternatively, in some embodiments, user representation data is generated and updated during the registration process based on images of the user's face captured while the user is expressing multiple different facial expressions (e.g., registration images of the face when the user is smiling, raising eyebrows, puffing out cheeks, etc.).
[0073] In some implementations, the modified user representation data generates a 3D Gaussian sputter based on image data for at least that portion of the user, where the Gaussian sputter includes texture, location, and sputter shape. For example, a 3D Gaussian distribution in a 2D space with color / density (e.g., parameterized), where the face is represented as a 2D grid, and each element of the 2D grid includes a 3D Gaussian sputter. In some implementations, a technique generates the user representation via a machine learning model trained using training data acquired via one or more sensors in one or more environments. For example, the machine learning model interprets image data and / or other sensor data captured during registration.
[0074] At box 730, method 700 generates multiple sputters based on the sputter parameter data of the modified user representation data, and provides a view of the user representation based on the modified user representation data. For example, 3D Gaussian sputtering can be used to avoid or fill cavities, and body pose data can be applied to additional areas including the user (e.g., neck / shoulder areas). For example, as... Figure 6C As illustrated, a user representation 692 is generated for a specific viewpoint based on points rendered from a Gaussian UV map 662, which combines sputtering parameters obtained from registration data and is updated for each frame based on one or more marker points 664 (e.g., a set of semantic points associated with facial features or other regions corresponding to the sender associated with the rendered user representation 692). In some implementations, the system can generate the user representation by applying increments to a canonical UV map, while in other implementations, the user representation can be generated directly from real-time data or other forms of identity representation.
[0075] In some implementations, a second set of sensor data, acquired after a registration process performed by a device (e.g., a viewer's device), includes a sequence of frames against a Gaussian UV map and corresponding marker points. The sequence of frames against the Gaussian UV map and corresponding marker points can be obtained from a second device (e.g., a sender's device) during a communication session. The device (e.g., the viewer's device) uses one or more sputtering techniques described herein to render an animated depiction of the user (e.g., the sender) based on the sequence of frames against the Gaussian UV map and corresponding marker points. Marker points (e.g., semantic or feature points, such as facial feature points, corners of the mouth, tip of the nose, etc.) tracked and updated during runtime can be used in conjunction with canonical representations and / or with real-time latency to generate an animated avatar.
[0076] In some implementations, a user's identity representation (e.g., a set of identity latency periods, a normalized Gaussian UV map, or a combination thereof) and corresponding marker points are transmitted during a communication session with a second device and can be used to render a view of the user's (sender's) face (and upper body). Additionally or alternatively, consecutive frames of facial data (the appearance of the user's face at different points in time) and body tracking data can be transmitted and used to display a real-time 3D video-like depiction of the user (e.g., a "live" avatar). For example, as... Figure 6CAs illustrated, the receiving device also acquires skeletal blending data 627 and skeletal joint data (e.g., animated body pose) to update the Gaussian UV mapping data 662 by applying the body pose data (e.g., neck, shoulders, hands, etc.) during the deformation and blurring process of the Gaussian buffer process 670, generating updated Gaussian UV mapping data 662 as Gaussian buffer data 672. The runtime phase 604 at the receiving device can then acquire the receiving device's viewpoint data 674 (e.g., the viewpoint to be rendered) and send the Gaussian buffer data 672 to a stereo proxy sputtering process 680, which generates stereo proxy geometry view data 682 (e.g., left and right viewpoints of RGBA sputtered data, and depth data). The stereo proxy geometry view data 682 can then be used at a 3D Gaussian sputtering rendering process (e.g., user representation process 690) to generate user representation 692.
[0077] In some embodiments, the second user representation is based on second image data acquired via a second set of sensors in a second physical environment with second lighting conditions (e.g., lighting conditions different from the first physical environment). For example, during the registration process, user representation data is acquired in a specific environment (also referred to herein as the "registration environment") that includes some lighting condition information (e.g., brightness values and other lighting attributes), which may be lighting data different from real-time lighting data (e.g., two different physical environments between registration and the generation of an avatar based on "real-time" sensor data). In some embodiments, method 700 also includes providing a view of the combined user representation in a 3D environment. In some embodiments, method 700 also includes modifying the view of the user representation by adjusting the user representation based on at least one color attribute of a plurality of color attributes of the environment, at least one light attribute of a plurality of light attributes of the environment, or a combination thereof. For example, adjusting the color or lighting on the user representation (such as hair, face, clothing, etc.) based on color and / or light associated with the viewer's environment and / or the sender's environment. In other words, the lighting and / or color of a 3D representation (e.g., an avatar) can be altered to match the lighting and / or color of the viewer's environment (e.g., a reddish hue emanating from the viewer's room will be reflected in the 3D representation). Alternatively, the lighting and / or color of a 3D representation (e.g., an avatar) can be altered to match the lighting and / or color of the sender's environment (e.g., even if the registration data does not reflect a green hue of light, a green hue emanating from the sender's room will be reflected in the 3D representation to the viewer).
[0078] In some implementations, at least a portion of the user representation data acquired during the registration process is based on images of the user's face captured in different poses and / or while the user is expressing multiple different facial expressions. For example, the images are registration images of the face when the user is facing the camera, to the left of the camera, and to the right of the camera, and / or when the user is smiling, raising eyebrows, puffing out cheeks, etc. In some implementations, a first set of sensor data corresponds only to a first region of the user (e.g., the portion not obscured by a device such as an HMD), and a second set of sensor data corresponds to a second region, which includes a third region different from the first region. For example, the second region may include portions of the portion obscured by the HMD when worn by the user. For example, during the registration process, a larger portion of the user may be captured by image data compared to a real-time communication session with a user wearing an HMD (e.g., without an HMD).
[0079] In some specific implementations, such as Figure 2 As illustrated, rendering occurs during a communication session in which a second device (e.g., device 265) captures sensor data (e.g., image data of a portion of user 260 and environment 250) and provides a sequence of frame-specific 3D representations corresponding to multiple moments within a time period based on the sensor data. For example, second device 265 provides / transmits the sequence of frame-specific 3D representations to device 210, and device 210 generates a combined 3D representation to display a live 3D video-like facial depiction of user 260 (e.g., a realistic moving avatar) (e.g., representation 240 of user 260). Alternatively, in some embodiments, the second device provides a 3D representation of the user (e.g., representation 140 of user 260) (e.g., a realistic moving avatar) during the communication session. For example, a combined representation is determined at device 265 and transmitted to device 210. In some embodiments, a view of the combined 3D representation is displayed in real-time on the device (e.g., device 210) relative to multiple moments within a time period. For example, a user's depiction is displayed in real time and based on real-time lighting data (e.g., an image shown to the second user on the display of the second user's second device).
[0080] In some implementations, the user-represented view may include sufficient data to achieve a stereoscopic view for the user (e.g., left / right eye views) that allows for a certain depth perception of the face. In one implementation, the depiction of the face includes a 3D model of the face, and generating representations from the left and right eye positions to provide a stereoscopic view of the face.
[0081] In some implementations, certain parts of the face (such as the eyes and mouth) that may be important for conveying a realistic or truthful appearance may be generated in a different way than other parts of the face (e.g., based on landmarks). For example, parts of the face that may be important for conveying a realistic or truthful appearance may be based on current camera data, while other parts of the face may be based on previously acquired (e.g., registered) facial data.
[0082] In some implementations, a facial representation is generated using the texture, color, and / or geometry of various facial features. This facial representation identifies the confidence level with which the generation technique accurately corresponds to the true texture, color, and / or geometry of those facial features for each data frame, based on depth and appearance values. In some implementations, the depiction is a 3D avatar. For example, the representation represents a user (e.g., Figure 1 3D model of user 110.
[0083] In some implementations, the first set of sensor data and / or the second set of sensor data (e.g., real-time data, such as video content including light intensity data (RGB) and depth data) are associated with points in time, such as images from inward-facing / downward-facing sensors being associated with frames when a user wears the HMD. In some implementations, the sensor data includes depth data (e.g., infrared, time-of-flight, etc.) and light intensity image data acquired during the scanning process.
[0084] In some embodiments, acquiring the first set of sensor data during the registration process may include acquiring registration sensor data (e.g., registration image data 610 of Figure 6) from the device corresponding to features of the user's face in various configurations (e.g., texture, muscle activation, shape, depth, etc.). In some embodiments, the first set of data includes unobstructed image data of the user's face. For example, images of the face may be captured when the user smiles, raises eyebrows, puffs out cheeks, etc. In some embodiments, registration data may be acquired by the user removing the device (e.g., HMD) and capturing images without the device obscuring the face, or by using another device (e.g., a mobile device) without the device (e.g., HMD) obscuring the face. In some embodiments, registration data (e.g., the first set of data) is acquired from light intensity images (e.g., RGB images). The registration data may include most (if not all) of the user's facial texture, muscle activation, etc. In some embodiments, registration data may be captured when the user is given different instructions for capturing different poses of the user's face. For example, user interface guidelines can instruct users to “raise your eyebrows,” “smile,” “frown,” etc., in order to provide the system with a range of facial features for the registration process.
[0085] In some implementations, method 700 may be repeated for each frame captured during each moment / frame of a live communication session or other experience. For example, for each iteration, as the user uses the device (e.g., a wearable HMD), method 700 may involve continuously acquiring real-time sensor data (e.g., face tracking data, body tracking, etc.), and for each frame, updating the portion of the display representing the user based on an updated Gaussian UV map and marker points using a GPU Gaussian buffer. For example, for each new frame, the system may update the display of the 3D avatar based on the new data.
[0086] Figure 8 This is a block diagram of example device 800. Device 800 illustrates an exemplary device configuration for devices described herein (e.g., devices 105, 210, 265, etc.). Although certain specific features are illustrated, those skilled in the art will recognize from this disclosure that various other features are not illustrated for the sake of brevity and so as not to obscure further relevant aspects of the specific implementations disclosed herein. Therefore, as a non-limiting example, in some specific implementations, device 800 includes one or more processing units 802 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 806, one or more communication interfaces 808 (e.g., USB, Firewire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, Bluetooth, ZigBee, SPI, I2C and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 810, one or more displays 812, one or more internal and / or external image sensor systems 814, memory 820, and one or more communication buses 804 for interconnecting these components and various other components.
[0087] In some embodiments, one or more communication buses 804 include circuitry that interconnects system components and controls communication between system components. In some embodiments, one or more I / O devices and sensors 806 include at least one of the following: an inertial measurement unit (IMU), an accelerometer, a magnetometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, or one or more depth sensors (e.g., structured light, time-of-flight, etc.).
[0088] In some embodiments, one or more displays 812 are configured to present a view of a physical or graphical environment to a user. In some embodiments, one or more displays 812 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conducting electron emission display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical systems (MEMS), and / or similar display types. In some embodiments, one or more displays 812 correspond to waveguide displays such as diffraction, reflection, polarization, and holography. In one example, device 10 includes a single display. In another example, device 10 includes displays for each of the user's eyes.
[0089] In some embodiments, one or more image sensor systems 814 are configured to acquire image data corresponding to at least a portion of the physical environment 102. For example, one or more image sensor systems 814 include one or more RGB cameras (e.g., having a complementary metal-oxide-semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), monochrome cameras, IR cameras, depth cameras, event-based cameras, etc. In various embodiments, one or more image sensor systems 814 also include an illumination source emitting light, such as a flash. In various embodiments, one or more image sensor systems 814 also include an on-camera image signal processor (ISP) configured to perform multiple processing operations on the image data.
[0090] Memory 820 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 820 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 820 optionally includes one or more storage devices remotely located to one or more processing units 802. Memory 820 includes a non-transitory computer-readable storage medium.
[0091] In some embodiments, memory 820 or a non-transitory computer-readable storage medium of memory 820 stores an optional operating system 830 and one or more instruction sets 840. Operating system 830 includes procedures for handling various basic system services and for performing hardware-related tasks. In some embodiments, instruction set 840 includes executable software defined by binary information stored in charge. In some embodiments, instruction set 840 is software executable by one or more processing units 802 to implement one or more of the techniques described herein.
[0092] Instruction set 840 includes registration instruction set 842, representation instruction set 844, and communication session instruction set 846. Instruction set 840 can be embodied in a single software executable file or multiple software executable files.
[0093] In some implementations, the registration instruction set 842 can be executed by the processing unit 802 to generate registration data from image data. The registration instruction set 842 can be configured to provide instructions to the user to acquire image information to generate a registration avatar (e.g., registration data 510) and to determine whether additional image information is needed to generate an accurate registration avatar to be used by the avatar display process. For these purposes, in various implementations, the instructions include instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0094] In some implementations, instruction set 844 is executed by processing unit 802 to generate a user representation based on registration data using one or more of the techniques discussed herein or other potentially suitable techniques (e.g., Gaussian sputtering). For these purposes, in various implementations, the instruction includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0095] In some specific implementations, the communication session instruction set 846 may be executed by the processing unit 802 to facilitate two or more electronic devices (e.g., such as) using one or more of the techniques discussed herein or otherwise suitable. Figure 2 The communication session between device 210 and device 265 shown. For these purposes, in various specific implementations, the instruction includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0096] Although instruction set 840 is shown as residing on a single device, it should be understood that in other specific implementations, any combination of elements may reside in separate computing devices. Furthermore, Figure 8This is intended more as a functional description of various features present in a particular implementation than as a structural diagram of the specific implementation described herein. As will be appreciated by those skilled in the art, the items shown individually can be combined, and some items can be separated. The actual number of instruction sets and how features are allocated therein will vary depending on the specific implementation and may depend in part on the specific combination of hardware, software, and / or firmware chosen for that particular implementation.
[0097] Figure 9 A block diagram illustrating an exemplary head-mounted device 900 according to some specific embodiments is shown. The head-mounted device 900 includes a housing 901 (or shell) housing various components of the head-mounted device 900. The housing 901 includes (or is coupled to) eye pads (not shown) disposed at a proximal end of the housing 901 (relative to the user 25). In various specific embodiments, the eye pads are plastic or rubber components that comfortably and snugly hold the head-mounted device 900 in an appropriate position on the face of the user 25 (e.g., around the eyes 35 of the user 25).
[0098] The housing 901 houses a display 910 that displays images, thereby emitting light toward or onto the eyes of the user 25. In various embodiments, the display 910 emits light through an eyepiece having one or more optical elements 905 that refract the light emitted by the display 910, so that the display appears to the user 25 at a virtual distance greater than the actual distance from the eye to the display 910. For example, the optical elements 905 may include one or more lenses, waveguides, other diffractive optical elements (DOEs), etc. In order for the user 25 to focus on the display 910, in various embodiments, the virtual distance is at least greater than the minimum focal length of the eye (e.g., 7 cm). Furthermore, to provide a better user experience, in various embodiments, the virtual distance is greater than 1 meter.
[0099] The housing 901 also houses a tracking system including one or more light sources 922, a camera 924, a camera 932, a camera 934, and a controller 980. One or more light sources 922 emit light onto the eyes of user 25, which is reflected as a light pattern (e.g., a flash) detectable by camera 924. Based on this light pattern, controller 980 can determine the eye-tracking characteristics of user 25. For example, controller 980 can determine the gaze direction and / or blinking state (open or closed eyes) of user 25. Also, controller 980 can determine the pupil center, pupil size, or point of focus. Thus, in various embodiments, light is emitted by one or more light sources 922, reflected from the eyes of user 25, and detected by camera 924. In various embodiments, light from the eyes of user 25 is reflected from a hot mirror or passes through an eyepiece before reaching camera 924.
[0100] Display 910 emits light within a first wavelength range, and one or more light sources 922 emit light within a second wavelength range. Similarly, camera 924 detects light within the second wavelength range. In various specific embodiments, the first wavelength range is the visible wavelength range (e.g., a wavelength range of approximately 400 nm to 700 nm within the visible spectrum), and the second wavelength range is the near-infrared wavelength range (e.g., a wavelength range of approximately 700 nm to 1400 nm within the near-infrared spectrum).
[0101] In various implementations, eye tracking (or specifically, a defined gaze direction) is used to enable user interaction (e.g., user 25 selects an option by looking at display 910), to provide foveated rendering (e.g., rendering a higher resolution in the area of display 910 that user 25 is viewing and a lower resolution elsewhere on display 910), or to correct distortion (e.g., for an image to be presented on display 910). In various implementations, one or more light sources 922 emit light toward user 25's eye 35, which is reflected in the form of multiple flashes.
[0102] In various implementations, camera 924 is a frame / shutter-based camera that generates images of user 25's eye 35 at specific time points or multiple time points at a certain frame rate. Each image includes a matrix of pixel values corresponding to the pixels in the image, which correspond to the positions of the camera's light sensor matrix. In specific implementations, each image is used to measure or track pupil dilation by measuring changes in pixel intensity associated with one or both of the user's pupils.
[0103] In various specific implementations, camera 924 is an event camera that includes multiple light sensors (e.g., a light sensor matrix) at multiple corresponding locations, which generates an event message indicating a specific location of a particular light sensor in response to a particular light sensor detecting a change in light intensity.
[0104] In various specific implementations, cameras 932 and 934 are frame / shutter-based cameras capable of generating images of the user 25's face at specific or multiple time points at a certain frame rate. For example, camera 932 captures an image of the user's face below the eyes, and camera 934 captures an image of the user's face above the eyes. The images captured by cameras 932 and 934 may include light intensity images (e.g., RGB) and / or depth image data (e.g., time-of-flight, infrared, etc.).
[0105] It should be understood that the specific embodiments described above are cited by way of example, and the invention is not limited to what has been specifically shown and described above. Rather, the scope includes both combinations and sub-combinations of the various features described above, as well as variations and modifications of the various features that would occur to those skilled in the art upon reading the foregoing description and that are not disclosed in the prior art.
[0106] As described above, one aspect of the present invention is the collection and use of physiological data to improve the user experience with electronic devices in interacting with electronic content. This disclosure envisions that, in some cases, the collected data may include personal information data that uniquely identifies a particular person or can be used to identify the interests, characteristics, or tendencies of a particular person. Such personal information data may include physiological data, demographic data, location-based data, telephone numbers, email addresses, home addresses, device characteristics of personal devices, or any other personal information.
[0107] This disclosure recognizes that the use of such personal information data in the present invention can benefit users. For example, personal information data can be used to improve the interactivity and controllability of electronic devices. Therefore, the use of such personal information data enables planned control of electronic devices. Furthermore, this disclosure also anticipates other uses of personal information data that benefit users.
[0108] This disclosure further envisions that entities responsible for the collection, analysis, disclosure, transmission, storage, or other use of such personal information and / or physiological data will comply with established privacy policies and / or privacy practices. Specifically, such entities should implement and adhere to privacy policies and measures recognized as meeting or exceeding industry or governmental requirements for maintaining the privacy and security of personal information data. For example, personal information from users should be collected for legitimate and reasonable purposes of the entity and not shared or sold outside of these legitimate purposes. Furthermore, such collection should only be conducted after receiving informed consent from users. Additionally, such entities should take any necessary steps to safeguard and protect access to such personal information data and ensure that others with access to such personal information data comply with their privacy policies and procedures. Furthermore, such entities may subject themselves to third-party assessments to demonstrate their compliance with widely accepted privacy policies and practices.
[0109] Regardless of the foregoing, this disclosure also contemplates specific implementations allowing users to selectively block the use or access to personal information data. That is, this disclosure contemplates providing hardware or software components to prevent or block access to such personal information data. For example, with regard to a content delivery service tailored to a user, the technology of this invention can be configured to allow a user to choose to "join" or "opt out" of the collection of personal information data during service registration. In another example, a user may choose not to provide personal information data for a target content delivery service. In yet another example, a user may choose not to provide personal information but allow the transmission of anonymous information for improving device functionality.
[0110] Therefore, while this disclosure broadly covers the use of personal information data to implement one or more of the various disclosed embodiments, it is also contemplated that various embodiments can be implemented without access to such personal information data. That is, various embodiments of the present invention will not become inoperable due to the absence of all or part of such personal information data. For example, preferences or settings can be inferred based on non-personal information data or an absolute minimum amount of personal information, such as content requested by a device associated with a user, other non-personal information available to the content delivery service, or publicly available information, thereby selecting content and delivering it to the user.
[0111] In some implementations, data is stored using a public / private key system that allows only the data owner to decrypt the stored data. In other implementations, data may be stored anonymously (e.g., without identification and / or without personal information about the user, such as legal name, username, time, and location data). This prevents other users, hackers, or third parties from identifying the user associated with the stored data. In some implementations, a user can access their stored data from a user device different from the device used to upload the stored data. In these cases, the user may need to provide login credentials to access their stored data.
[0112] This document sets forth numerous specific details to provide a comprehensive understanding of the claimed subject matter. However, those skilled in the art will understand that the claimed subject matter can be practiced without these specific details. In other instances, methods, apparatus, or systems known to a person of ordinary skill have not been described in detail so as not to obscure the claimed subject matter.
[0113] Unless otherwise specifically stated, it should be understood that throughout this specification, discussions using terms such as “processing,” “calculating,” “calculating,” “determining,” and “identifying” refer to the actions or processes of computing devices, such as one or more computers or similar electronic computing devices, which manipulate or convert data representing physical electronic or magnetic quantities within the memory, registers, or other information storage, transmission, or display devices of a computing platform.
[0114] The one or more systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement of components that provide results conditioned on one or more inputs. Suitable computing devices include computer systems based on multi-purpose microprocessors that access stored software that programs or configures the computing system from a general-purpose computing device to a special-purpose computing device that implements one or more specific embodiments of the subject matter of this invention. The teachings contained herein may be implemented in the software used for programming or configuring the computing device using any suitable programming, scripting, or other type of language or combination of languages.
[0115] Specific implementations of the methods disclosed herein can be performed in the operation of such computing devices. The order of the boxes presented in the above examples can be changed; for example, the boxes can be reordered, grouped, or divided into sub-boxes. Some boxes or procedures can be executed in parallel.
[0116] The use of "applies to" or "configured to" in this document implies open and inclusive language, which does not exclude applicability to or configuration to devices performing additional tasks or steps. Furthermore, the use of "based on" implies openness and inclusivity, as processes, steps, calculations, or other actions "based on" one or more of the stated conditions or values may in practice be based on additional conditions or values beyond those stated. The headings, lists, and numbering included herein are for illustrative purposes only and are not intended to be restrictive.
[0117] It will also be understood that while terms such as "first," "second," etc., may be used in this document to describe various objects, these objects should not be limited by these terms. These terms are merely used to distinguish one object from another. For example, a first node can be called a second node, and similarly, a second node can be called a first node, changing the meaning of the description, provided that all occurrences of "first node" are consistently renamed and all occurrences of "second node" are consistently renamed. First nodes and second nodes are both nodes, but they are not the same node.
[0118] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the claims. As used in the description of these embodiments and in the appended claims, the singular forms “a,” “an,” and “the” are intended to also cover the plural forms unless the context clearly indicates otherwise. It will also be understood that the term “or” as used herein refers to and covers any and all possible combinations of one or more of the associated listed items. It will also be understood that the terms “comprising” or “including” as used in this specification specify the presence of the stated features, integers, steps, operations, objects, or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, objects, components, or groups thereof.
[0119] As used herein, the term "if" can be interpreted as meaning "when the prerequisite is true" or "when the prerequisite is true" or "in response to determination" or "according to determination" or "in response to detection" that the prerequisite is true, depending on the context. Similarly, the phrase "if it is determined [the prerequisite is true]" or "if [the prerequisite is true]" or "when [the prerequisite is true]" can be interpreted as meaning "when it is determined that the prerequisite is true" or "in response to determination" or "according to determination" that the prerequisite is true or "when it is detected that the prerequisite is true" or "in response to detection" that the prerequisite is true, depending on the context.
[0120] The foregoing description and summary of the invention should be understood as exemplary and illustrative in every respect, and not restrictive, and the scope of the invention disclosed herein is determined not only by the detailed description of the exemplary specific implementations, but also by the full extent permitted by patent law.
[0121] It should be understood that the specific embodiments shown and described herein are merely illustrative of the principles of the invention, and those skilled in the art can make various modifications without departing from the scope and spirit of the invention.
Claims
1. A method, the method comprising: At the device's processor: At least a portion of user representation data of a user is obtained, wherein the user representation data is based on a first set of sensor data including an image of the user obtained during the registration process, and the user representation data includes sputtering parameter data corresponding to multiple three-dimensional (3D) locations; The user representation data is modified based on a second set of sensor data obtained after the registration process. as well as To provide a view of the user representation based on the modified user representation data, wherein providing the view includes generating a plurality of sputters based on the sputter parameter data of the modified user representation data.
2. The method of claim 1, wherein the at least part of the user comprises a facial portion and an additional portion of the user.
3. The method of claim 1, wherein the user representation data is based on a UV map and 3D point cloud points associated with distribution data, the distribution data defining the size and shape of the sputtering body corresponding to each point of the UV map.
4. The method of claim 1, wherein the sputtering parameter data includes 3D Gaussian parameters for each 3D location.
5. The method according to claim 4, wherein the 3D Gaussian parameters include at least one of position information, color information, covariance information, transparency information, orientation, opacity information, range information in each axis, rotation data, scale, and semantic information.
6. The method of claim 1, wherein the user representation data includes 3D mapping information, the 3D mapping information including feature values and position information for each mapping point.
7. The method of claim 1, wherein the user representation data is modified based on body posture data obtained during the registration process, during a communication session with another device, or a combination thereof.
8. The method of claim 1, wherein the device is a viewer's device, and wherein the user representation data is modified based on a set of additional sensor data obtained during a communication session with a device associated with the sender of the user representation.
9. The method of claim 1, wherein the user representation data is generated and updated during the registration process based on images of the user's face captured while the user is expressing multiple different facial expressions.
10. The method of claim 1, wherein the technique generates the user representation data via a machine learning model trained using training data obtained via one or more sensors in one or more environments.
11. The method of claim 1, wherein providing the view of the user representation based on the modified user representation data includes displaying the user representation in an extended reality (XR) environment.
12. The method according to claim 1, further comprising: The view of the user representation is modified by adjusting the user representation based on at least one color attribute of a plurality of color attributes of the environment, at least one light attribute of a plurality of light attributes of the environment, or a combination thereof.
13. The method of claim 1, wherein the user representation data is obtained in a first physical environment, and the user representation is displayed in a view of a second physical environment different from the first physical environment.
14. The method of claim 1, wherein the user representation is a 3D user representation.
15. An apparatus comprising: Non-transitory computer-readable storage medium; as well as One or more processors coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium includes program instructions that, when executed on the one or more processors, cause the one or more processors to perform operations, the operations including: At least a portion of user representation data of a user is obtained, wherein the user representation data is based on a first set of sensor data including an image of the user obtained during the registration process, and the user representation data includes sputtering parameter data corresponding to multiple three-dimensional (3D) locations; The user representation data is modified based on a second set of sensor data obtained after the registration process; and To provide a view of the user representation based on the modified user representation data, wherein providing the view includes generating a plurality of sputters based on the sputter parameter data of the modified user representation data.
16. The device of claim 15, wherein the user representation data is based on a UV map and 3D point cloud points associated with distribution data, the distribution data defining the size and shape of the sputtered body corresponding to each point of the UV map.
17. The apparatus of claim 15, wherein the sputtering body parameter data includes 3D Gaussian parameters for each 3D position, wherein the 3D Gaussian parameters include at least one of position information, color information, covariance information, transparency information, orientation, opacity information, range information in each axis, rotation data, scale, and semantic information.
18. The device of claim 15, wherein the user representation data is modified based on body posture data obtained during the registration process, during a communication session with another device, or a combination thereof.
19. The device of claim 15, wherein the device is a viewer's device, and wherein the user representation data is modified based on a set of additional sensor data obtained during a communication session with a device associated with the sender of the user representation.
20. A non-transitory computer-readable storage medium storing program instructions executable on a device to perform operations, the operations including: At least a portion of user representation data of a user is obtained, wherein the user representation data is based on a first set of sensor data including an image of the user obtained during the registration process, and the user representation data includes sputtering parameter data corresponding to multiple three-dimensional (3D) locations; The user representation data is modified based on a second set of sensor data obtained after the registration process. as well as To provide a view of the user representation based on the modified user representation data, wherein providing the view includes generating a plurality of sputters based on the sputter parameter data of the modified user representation data.