Context-based Gaussian scatter rendering for representations

By using 3D Gaussian sputtering technology, which uses Gaussian distributed scatter points to represent the user view, the problem of inaccurate user appearance representation in existing technologies is solved, achieving efficient and realistic user representation, especially accurate rendering of hair, thus improving rendering efficiency and visual effects.

CN121746558APending Publication Date: 2026-03-27APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies cannot accurately or truthfully represent a user's current appearance in electronic devices, especially the display of hair, resulting in inaccurate and unrealistic representations.

Method used

The method employs a 3D Gaussian sputtering approach, using Gaussian distributed scatter points to represent the user's view. By generating scatter point parameter data and selecting a rendering method based on the context of the viewing experience, the rendering efficiency and realism are optimized, mesh holes are avoided, and smooth transitions are achieved through Gaussian detail levels and dynamic region processing.

Benefits of technology

It achieves efficient and realistic user representation in electronic devices, especially accurate rendering of hair, improving rendering efficiency and visual effects, reducing visual artifacts, and adapting to different viewing distances and dynamic changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121746558A_ABST
    Figure CN121746558A_ABST
Patent Text Reader

Abstract

Various implementations disclosed herein include devices, systems, and methods for generating a user representation based on selecting a scatter rendering method. For example, a process may include obtaining representation data of at least a portion of an object. The representation data may include data for a plurality of sets of scatter points generated based on a scatter point generation technique. The process may also include determining a context of the viewing experience. The process may also include selecting a scatter rendering method for providing a view including a representation of the at least part of the object based on the determined context of the viewing experience. The process may also include providing a view of the representation of the at least a portion of the object based on a selected scatter rendering method, the selected scatter rendering method including rendering a subset of the scatters based on the scatter parameter data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates generally to electronic devices, and more particularly to systems, methods and apparatus for representing objects in computer-generated content. Background Technology

[0002] Existing technologies may not be able to accurately or truthfully represent the current (e.g., real-time) representation of the appearance of objects (such as users) on electronic devices. For example, a device may provide a representation of a user based on an image of the user's face obtained minutes, hours, days, or even years ago. Such a representation may not accurately depict the user's appearance; for example, it may not be possible to correctly display a person's hair for a realistic representation. Therefore, it may be desirable to provide means that efficiently provide a more accurate, truthful, and / or current representation of objects (such as users (e.g., characters)). Summary of the Invention

[0003] The various specific implementations disclosed herein include devices, systems, and methods for generating user-represented views based on three-dimensional (3D) Gaussian sputtering. Gaussian sputtering is a technique in which individual 3D points are represented as a Gaussian distribution (such as “scatter points”) with color values ​​that change according to the viewpoint. This technique uses spherical harmonics to model this view-dependent color variation and enables real-time rendering of high-quality, realistic scenes from sparse image sets. For example, each point has a color calculated based on its position relative to the camera, allowing for realistic shading effects across different viewpoints.

[0004] In an exemplary implementation, a first set of captured user data (e.g., registration data) can be used at a first device (e.g., a delivery device) to generate user representation data including scatter parameter data (e.g., multi-channel Gaussian UV mapping). A view of the user representation can be provided to a viewing device (e.g., rendering a live view of a delivery persona) by generating scatter points corresponding to the modified user representation data. A persona is a representation of the user, such as an avatar. Advantageously, sputtering avoids the need for a mesh to prevent holes and offers other advantages. A 3D representation of the user at multiple moments can be generated on the viewing device, which combines the data and uses the combined data to render a view, for example, during a live communication (e.g., virtual communication or coexistence) session.

[0005] In some implementations, to improve rendering efficiency, since not all scatter points are needed for every frame, multiple sets of scatter point parameter data can be generated, and the set of scatter point parameter data to be rendered can be selected based on the context of the viewing experience. In some implementations, the multiple sets of scatter point parameter data may correspond to different scatter point models with different numbers of scatter points, scatter point sizes, and / or scatter point qualities (e.g., scatter point parameter data sets with different resolutions / sizes ranging from 16k to 64k scatter points, etc.).

[0006] Additionally or alternatively, in some implementations, multiple sets of scatter parameter data may correspond to different Gaussian-based Levels of Detail (LODs). Gaussian-based LODs may be based on dividing scatter points into groups and merging specific groupings and / or hierarchical structures of scatter points based on 3DGS parameters. In some implementations, scatter points are arranged based on UV meshes or UV space topology to divide and combine / merge static (or less dynamic) scatter points. In other words, by keeping scatter points as static as possible during animation (e.g., there is typically little or no change / movement in shoulder scatter points during character rendering), the system can leverage the ability of temporal coherence and merge static scatter points. In some implementations, by identifying regions such as hair scatter points as static scatter points compared to skin scatter points, each group may have different rendering characteristics taken into account; for example, hair has a more refined volumetric structure, while skin has more opaque subsurface scattering properties. In some implementations, LODs may avoid merging specific dynamic regions, such as regions that require more detail and / or may move more during animation and rendering. For example, during a communication session, when a character is generated in real time, the mouth and / or eye areas may move during the session. Therefore, when merging scatter plots, those specific dynamic areas can be avoided.

[0007] Additionally or alternatively, in some implementations, the method may include providing dynamic, context-based selection and a process for transitions between LODs used for 3DGS rendering, with an emphasis on minimizing perceptible transition artifacts and simultaneously optimizing distance, gaze, and field of view. For example, systems and methods may provide dynamic, context-aware rendering of the 3D representation using Gaussian scatter points, where advanced LOD transition merging (e.g., merging scatter points during LOD transitions) is performed based on whether the character is inside or outside the gaze cone (e.g., within or around the fovea) (which can be determined based on the user's viewpoint, gaze, and / or the character's position relative to the field of view). For example, instead of abrupt LOD switching (e.g., from 64k scatter points to 16k scatter points), the system introduces a smooth, continuous reduction in scatter points. As the viewer moves away from the character, the scatter points gradually merge, creating a transition zone that avoids a noticeable "pop-up" effect. This merging can be tuned and incorporated into noise-based selections for further optimization. This method achieves a smooth, continuous LOD transition while merging scattered points, minimizes visual artifacts, and optimizes rendering efficiency and fidelity in regions of interest.

[0008] In some implementations, analysis of the context of the viewing experience (e.g., during a live communication session) used for selection can determine whether the view to be rendered is close to or far from the current viewpoint. For example, if the perceived distance of the representation to be rendered is greater than a perceived distance threshold (e.g., in the background of the current view), a lower-quality rendering of the representation (e.g., 32k scatter points or a LOD that has been grouped and merged or at least partially merged based on static scatter points of the identifier) ​​can be selected. On the other hand, if the perceived distance of the representation to be rendered is closer than a perceived distance threshold (e.g., a user representation (role) during a communication session), a higher-quality rendering of the representation (e.g., 64k scatter points or a LOD that has not yet been grouped / merged) can be selected. Additionally, the determined context of the viewing experience may indicate that the viewer is taking pictures of the viewing experience or recording video of the viewing experience, and therefore would expect a higher-quality representation to be rendered.

[0009] In some implementations, the data associated with each scatter point of the user representation can represent texture / color, 3D location, orientation and angular information (e.g., visibility cone), scatter point shape, transparency level, and covariance (e.g., how the scatter point is stretched / scaled), etc. The scatter points can be a 3D Gaussian distribution in a two-dimensional (2D) space with color / density (e.g., parameterized), where a person's face can be used as a raster, and multiple scatter points can be determined based on rays leaving the face / raster. Rasterization / parameterization of the scatter point distribution provides higher quality data, allows for faster training of machine learning models, and can provide faster (e.g., real-time) rasterization. In some implementations, 3D mapping information (e.g., identifying x, y, z positions corresponding to UV coordinates of a UV map) can be generated at registration (e.g., Gaussian UV mapping).

[0010] Typically, an innovative aspect of the subject matter described in this specification may be embodied in a method comprising the actions of: obtaining representation data representing at least a portion of an object at a processor of a device, wherein the representation data includes data for a plurality of sets of scatter points generated based on a scatter point generation technique, the data for each of the plurality of sets of scatter points including three-dimensional (3D) Gaussian data. The actions also include determining a context of a viewing experience. The actions further include selecting a scatter point rendering method for providing a view including a representation of the at least a portion of the object based on the determined context of the viewing experience. The actions further include providing a view of the representation of the at least a portion of the object based on the selected scatter point rendering method, wherein providing the view includes rendering a subset of the scatter points based on the scatter point parameter data.

[0011] These and other implementation schemes may optionally include one or more of the following features.

[0012] In some aspects, at least a portion of the object includes the user's facial features and additional features. In some aspects, the representation data is based on 3D point cloud points associated with distributed data that defines the size and shape for rendering these 3D point cloud points as scatter points. In some aspects, the 3D Gaussian data includes 3D Gaussian parameters. In some aspects, the 3D Gaussian parameters include at least one of the following: positional information, color information, covariance information, transparency information, orientation, opacity information, range information in each axis, rotation data, scaling factor, and semantic information.

[0013] In some aspects, the scatter point generation technique includes a scatter point criterion associated with the resolution of each scatter point associated with each of the plurality of sets of scatter points. In some aspects, the scatter point generation technique includes a scatter point criterion associated with the number of scatter points associated with each of the plurality of sets of scatter points. In some aspects, the scatter point generation technique includes a scatter point criterion associated with the size of each scatter point associated with each of the plurality of sets of scatter points.

[0014] In some respects, selecting the scatter rendering method includes choosing between a first method for rendering a first set of a plurality of sets of scatter points and a second method for rendering a second set of a plurality of sets of scatter points, wherein the first set of a plurality of sets of scatter points has more scatter points than the second set of a plurality of sets of scatter points.

[0015] In some aspects, each of the multiple sets of scatter points is based on a different level of detail (LOD). In some aspects, the scatter point generation technique includes scatter criteria associated with merging multiple subsets of scatter points based on the scatter point parameter data. In some aspects, the multiple subsets of scatter points are selected based on determining one or more static regions associated with representing at least a portion of the object. In some aspects, the multiple subsets of scatter points are determined during the registration process.

[0016] In some respects, selecting the scatter rendering method includes choosing a method to reduce the rendering level of detail by switching to rendering a lower level of detail (LOD) representation, which includes merging a subset of scatter points in a transition zone based on the location of at least a portion of the object to be depicted within the 3D environment.

[0017] In some aspects, during the merging of a subset of scatter points in the transition zone, the number of scatter points rendered decreases as the distance associated with the location of the at least part of the object to be depicted increases. In some aspects, the merging of the subset of scatter points is performed in groups, wherein each group corresponds to a range of distances associated with the location of the at least part of the object to be depicted.

[0018] In some aspects, merging scatter points is based on determining whether a viewpoint characteristic meets a criterion, wherein the viewpoint characteristic is determined based on the viewpoint and the location of the at least portion of the object to be depicted within the 3D environment. In some aspects, the criterion is based on a threshold angular distance from the gaze direction. In some aspects, the criterion is based on a threshold distance associated with the location of the at least portion of the object to be depicted within the 3D environment.

[0019] In some respects, merging scatter points in the peripheral region is performed using sparse UV mapping techniques. In some respects, the selection of scatter points for merging is based at least in part on a noise function. In some respects, the number of scatter points rendered to provide a representation of at least a portion of the object in the view is adjusted between a minimum and a maximum value based on at least one of the following: distance, gaze, or field of view associated with how at least a portion of the object should be depicted within the viewing experience.

[0020] In some aspects, the context of the viewing experience includes the viewpoint of the represented view, and the scatter rendering method for providing the view is selected based on a threshold distance associated with the viewpoint. In some aspects, the context of the viewing experience includes at least one frame recording content associated with the viewpoint of the represented view, and the scatter rendering method for providing the view is selected based on the at least one frame recording the content.

[0021] In some aspects, the view represented is provided for a first frame of a plurality of frames for a first viewpoint, and the method further includes: for a second viewpoint of a second frame of the plurality of frames: determining that the second viewpoint is equivalent to the first viewpoint; and reusing the view of the representation of at least a portion of the object from the first frame to the second frame.

[0022] In some aspects, the view of the representation is provided for a first frame of a plurality of frames for a first viewpoint, and the method further includes: for a second viewpoint of a second frame of the plurality of frames: determining that the second viewpoint is different from the first viewpoint; selecting an additional set of the plurality of scatter points based on the context of the viewing experience associated with the second frame; and updating the view of the representation of the at least part of the object based on the selected additional set of the plurality of scatter points.

[0023] In some aspects, the device is the viewer's device, wherein the scatter rendering method used to provide the view is selected based on data acquired during a communication session with the transmitter's device, which is associated with the representation of at least a portion of the object. In some aspects, the representation data is generated and updated during the registration process based on images of the user's face captured while the user is expressing multiple different facial expressions.

[0024] In some aspects, the scatter generation technique generates the representation data via a machine learning model trained using training data obtained via one or more sensors in one or more environments. In some aspects, the view of the representation of at least a portion of the object, based on a selected scatter rendering method, includes displaying the representation in an extended reality (XR) environment.

[0025] In some respects, representation data is obtained in a first physical environment and displayed in a view of a second physical environment different from the first physical environment. In some respects, the representation is a 3D representation.

[0026] Typically, another innovative aspect of the subject described in this specification can be embodied in the method, which includes the following actions: obtaining representation data representing at least a portion of an object at a processor of a device, wherein the representation data includes Gaussian data for rendering a plurality of scatter points for a first set of three-dimensional (3D) points to represent the 3D appearance of the object. The action also includes determining viewpoint characteristics based on a viewpoint and the location of the at least portion of the object to be depicted within a 3D environment. The action further includes generating modified representation data for a region by merging a subset of the representation data to represent a second set of 3D points, based on the viewpoint characteristics meeting a criterion, wherein the second set of 3D points is smaller in number than the first set of 3D points. The action also includes providing a view of the representation of the at least portion of the object by rendering scatter points based on the modified representation data.

[0027] These and other implementation schemes may optionally include one or more of the following features.

[0028] In some aspects, during the merging of the subset of representation data to represent a second set of 3D points, the number of rendered scatter points decreases as the distance associated with the location of the at least part of the object to be depicted increases. In some aspects, merging the subset of representation data to represent a second set of 3D points is performed in groups, where each group corresponds to a range of distances associated with the location of the at least part of the object to be depicted. In some aspects, merging the subset of representation data to represent a second set of 3D points is based on determining whether viewpoint characteristics meet a criterion, wherein the viewpoint characteristics are determined based on the viewpoint and the location of the at least part of the object to be depicted within the 3D environment.

[0029] In some respects, the standard is based on a threshold angular distance from the gaze direction. In other respects, the standard is based on a threshold distance associated with the location of at least a portion of the object to be depicted within the 3D environment.

[0030] In some aspects, merging the subset of representation data in the peripheral region to represent a second set of 3D points is performed using sparse UV mapping techniques. In some aspects, the selection of the subset of representation data for merging is based at least in part on a noise function. In some aspects, the number of scatter points rendered to provide a representation of at least a portion of the object is adjusted between a minimum and a maximum value based on at least one of the following: distance, gaze, or field of view associated with how at least a portion of the object should be depicted within the viewing experience. In some aspects, at least a portion of the object includes the user's facial portion and additional portions.

[0031] In some aspects, the representation data is based on 3D point cloud points associated with distributed data, which defines the size and shape used to render these 3D point cloud points as scatter points. In some aspects, the Gaussian data includes 3D Gaussian parameters. In some aspects, the Gaussian parameters include at least one of the following: positional information, color information, covariance information, transparency information, orientation, opacity information, range information in each axis, rotation data, scaling, and semantic information.

[0032] In some embodiments, a non-transitory computer-readable storage medium stores instructions that are computer-executable to perform or cause to perform any of the methods described herein. In some embodiments, an apparatus includes one or more processors, non-transitory memory, and one or more programs; the one or more programs are stored in the non-transitory memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing or causing to perform any of the methods described herein. Attached Figure Description

[0033] To enable those skilled in the art to understand this disclosure, more detailed descriptions can be made with reference to aspects of some exemplary embodiments, some of which are shown in the accompanying drawings.

[0034] Figure 1 Examples of devices for obtaining sensor data from users are shown, based on some specific implementations.

[0035] Figure 2 An exemplary electronic device is illustrated, which operates in different physical environments during a communication session between a first user at a first device and a second user at a second device, according to some specific implementation, and a combined three-dimensional (3D) representation of the second user at the first device.

[0036] Figures 3A to 3D Examples of 3D Gaussian scattering (3DGS) used in generating views with 3D representations are illustrated according to some specific implementations.

[0037] Figure 4 Examples of generating and displaying a 3D view of scatter points on a device are shown, based on some specific implementations.

[0038] Figure 5 Examples are given of generating multiple sets of scatter points based on different levels of detail (LOD) by merging groups of scatter points using 3DGS parameters, according to some specific implementations.

[0039] Figure 6 An example is shown of generating a user representation by rendering scatter parameter data using a scatter rendering method selected based on the viewing experience context, according to some specific implementation.

[0040] Figure 7 Exemplary electronic devices, according to some specific implementations, operate in different physical environments during communication sessions with multiple users, and include views of 3D representations of other users using a selected scatter rendering method.

[0041] Figure 8 Examples of gaze cones according to some specific implementations and including those composed of Figure 1 The device provides a view of the 3D environment for different roles.

[0042] Figure 9 Examples are given based on viewpoint characteristics according to some specific implementations. Figure 8 Examples of different regions of the gaze cone.

[0043] Figures 10A to 10C Examples of rendering characters in a 3D environment using different levels of detail based on viewpoint characteristics of some specific implementations are shown.

[0044] Figure 11 It is a flowchart representation of a method for providing a view representation by rendering scatter parameter data using a scatter rendering method selected based on the context of the viewing experience, according to some specific implementation.

[0045] Figure 12 It is a flowchart representation of a method for providing a view of representation by rendering scatter points based on a subset of merged representation data, according to some specific implementations.

[0046] Figure 13 This is a block diagram illustrating device components according to some specific implementations of exemplary devices.

[0047] Figure 14 It is a block diagram based on some specific implementation examples of head-mounted devices (HMDs).

[0048] As is customary practice, various features illustrated in the accompanying drawings may not be drawn to scale. Therefore, for clarity, the dimensions of various features may be arbitrarily expanded or reduced. Furthermore, some drawings may not depict all components of a given system, method, or apparatus. Finally, similar reference numerals may be used throughout the specification and drawings to denote similar features. Detailed Implementation

[0049] Numerous details have been described to provide a thorough understanding of the exemplary embodiments illustrated in the accompanying drawings. However, the drawings illustrate only some exemplary aspects of this disclosure and should not be considered limiting. Those skilled in the art will recognize that other effective aspects or variations do not include all the specific details described herein. Furthermore, well-known systems, methods, components, devices, and circuits have not been described exhaustively so as not to obscure further relevant aspects of the exemplary embodiments described herein.

[0050] Figure 1 An example environment 100 is illustrated, in which an exemplary electronic device 105 operates within a physical environment 102. In some embodiments, the electronic device 105 may be able to share information with another device or with an intermediary device, such as an information system. Additionally, the physical environment 102 includes a user 110 wearing the device 105. In some embodiments, the device 105 is configured to present views of extended reality (XR) environments, which may be based on the physical environment 102 and / or include added content, such as virtual elements.

[0051] exist Figure 1 In the example, physical environment 102 is a room that includes physical objects such as wall hangings 120, plants 125, and tables 130. Electronic device 105 may include one or more cameras, microphones, depth sensors, motion sensors, or other sensors that can be used to capture information about physical environment 102 and the objects within it and to evaluate the physical environment and these objects, as well as to capture information about user 110.

[0052] exist Figure 1In this example, device 105 includes one or more sensors 116 (e.g., inward-facing sensors and outward-facing cameras) that capture light intensity images, depth sensor images, audio data, or other information about user 110. For example, one or more sensors 116 may capture images of the user's (e.g., user 110's) forehead, eyebrows, eyes, eyelids, cheeks, nose, lips, chin, face, head, hands, wrists, arms, shoulders, torso, legs, or other body parts. For example, inward-facing sensors may see the interior of device 105 (e.g., the user's eyes and the area around them), and other external cameras may capture the user's face outside device 105 (e.g., an egocentric camera pointing outwards from device 105). As an example, sensor data about the user's eyes 111 may indicate various user characteristics, such as the user's gaze direction 119 over time, the user's saccade behavior over time, the user's eye expansion behavior over time, etc. One or more sensors 116 may capture audio information, including the user's voice and other sounds emitted by the user, as well as sounds within the physical environment 100.

[0053] In some embodiments, device 105 includes an eye-tracking system for detecting eye positioning and eye movement via eye gaze characteristic data. For example, the eye-tracking system may include one or more infrared (IR) light-emitting diodes (LEDs), an eye-tracking camera (e.g., a near-IR (NIR) camera), and an illumination source (e.g., an NIR light source) that emits light (e.g., NIR light) toward the user 110's eye. Furthermore, the illumination source of device 105 may emit NIR light to illuminate the user 110's eye, and the NIR camera may capture images of the user 110's eye. In some embodiments, the images captured by the eye-tracking system may be analyzed to detect the positioning and movement of the user 110's eye or to detect other information about the eye, such as color, shape, state (e.g., widening, strabismus, etc.), pupil dilation, or pupil diameter. Furthermore, the gaze point estimated from the eye-tracking images enables gaze-based interaction with content displayed on the near-eye display of device 105.

[0054] Additionally, one or more sensors 116 may capture images of the physical environment 100 (e.g., externally oriented sensors). For example, one or more sensors 116 may capture images of the physical environment 100 including physical objects such as wall hangings 120, plants 125, and tables 130. Furthermore, one or more sensors 116 may capture images (e.g., light intensity images and / or depth data).

[0055] One or more sensors (such as one or more sensors 115 on device 105) may identify user information based on proximity or contact with a portion of user 110. For example, one or more sensors 115 may capture sensor data that can provide biometric information related to the user’s cardiovascular status (e.g., pulse), body temperature, respiratory rate, etc.

[0056] One or more sensors 116 or one or more sensors 115 can capture data that can determine a user orientation 121 within the physical environment. In this example, user orientation 121 corresponds to the direction in which the user 110's torso is facing.

[0057] Some specific implementations disclosed herein determine user understanding or scene understanding based on sensor data obtained by a user wearable device (such as the first device 105). This user understanding can indicate the user's state associated with providing user assistance or facilitating a communication session. In some examples, the user's appearance or behavior, or their understanding of the environment (scene understanding), can be used to identify the need or expectation for assistance, making such assistance available to the user. For example, based on determining such scene understanding, enhancements can be provided to assist the user by enhancing or supplementing their capabilities (e.g., providing information about the environment to the person). Additionally, user understanding and / or scene understanding can be used to customize the viewing experience, such as selecting a rendering mode, which will be discussed further herein.

[0058] The content can be visible (e.g., displayed on a display of device 105) or audible (e.g., generated as audio 118 by a speaker of device 105). In the case of audio content, audio 118 can be generated in a manner that makes it likely only user 110 will hear it (e.g., via a speaker close to user's ear 112 or at a volume below a threshold), making it unlikely that nearby people will hear it. In some specific implementations, the audio pattern (e.g., volume) is determined based on whether other people are within a threshold distance or based on how close other people are relative to user 110.

[0059] In some implementations, the content provided by device 105 and the sensor features of device 105 can be provided using components, sensors, or software modules that are small enough and efficient in terms of power consumption and use to be adapted to and otherwise used in lightweight, battery-powered wearable products, such as wireless earbuds or other ear-hook devices or head-mounted devices (HMDs), such as smart / augmented reality (AR) glasses. A combination of multiple devices can be used to facilitate the features. For example, a smartphone (wirelessly connected and interacting with wearable devices) can provide computing resources, connectivity to cloud or internet services, location services, etc.

[0060] Figure 2An exemplary electronic device is illustrated, operating in different physical environments during a communication session between a first user at a first device and a second user at a second device. Specifically, Figure 2 Exemplary operating environments 200 are illustrated for electronic devices 210 and 265 operating in different physical environments 202 and 250 during a communication session (e.g., when electronic devices 210 and 265 share information with each other or with an intermediate device (such as a communication session system / server)). Figure 2 In this example, physical environment 202 is a room that includes wall mount 212, plant 214, and table 216 (e.g., Figure 1 The physical environment 202. Electronic device 210 includes one or more cameras, microphones, depth sensors, or other sensors that can be used to capture information about the physical environment 202 and objects within it, information about the user 225 of electronic device 210, and to assess the physical environment and objects within it. Information about the physical environment 202 and / or the user 225 can be used to provide visual content (e.g., for user representation) and audio content (e.g., for audible speech or text transcription) during a communication session.

[0061] Additionally, in Figure 2 In this example, physical environment 250 is a room including wall mount 252, sofa 254, and coffee table 256. Electronic device 265 includes one or more cameras, microphones, depth sensors, or other sensors that can be used to capture information about physical environment 250 and objects within it, as well as information about user 260 (such as wearable devices or HMD devices, like device 105) and to evaluate the physical environment and these objects. Information about physical environment 250 and / or user 260 can be used to provide visual and audio content during a communication session.

[0062] Information system 290 may use communication instruction set 280 and communication instruction set 282 to coordinate the sharing of assets (e.g., data associated with user representation) between two or more devices (e.g., electronic devices 210 and 265).

[0063] Figure 2 An example of a view 205 provided at device 210 is illustrated, wherein view 205 depicts a 3D environment 230 by providing pass-through video of a physical environment 202 enhanced with a user representation 240 (e.g., a character) of at least a portion of user 260. The provision / inclusion of the user representation 240 may be based on consent from user 260. Specifically, the user representation 240 of user 260 is generated based on one or more user representation techniques. The generation of the user representation is further discussed herein.

[0064] Figure 2An example of a view 266 provided at device 265 is illustrated, wherein the view depicts a 3D environment 270 by providing pass-through video of a physical environment 250 enhanced with a user representation 275 (e.g., a character) of at least a portion of user 225 (e.g., from above the mid-torso). The provision / inclusion of user representation 275 may be based on consent from user 225. User representations 240 and 275 may be generated at either device based on sensor data obtained at one of the respective devices 210 and 265. In some embodiments, each of the 3D representations 240 of user 260 and 275 of user 225 is generated by generating splats corresponding to the user representation data.

[0065] exist Figure 2 In the example, electronic devices 210 and 265 are exemplified as head-mounted devices (HMDs). However, either electronic device 210 or electronic device 265 can be a mobile phone, tablet, laptop, or any other form of wearable device (e.g., head-mounted devices (glasses), headphones, ear-worn devices, etc.). In some implementations, the functionality of each of devices 210 and 265 is implemented via two or more devices, such as a mobile device and a base station, or a head-mounted device and an ear-worn device. Various functions can be distributed across multiple devices, including but not limited to power functions, CPU functions, GPU functions, storage functions, memory functions, visual content display functions, audio content production functions, etc. The multiple devices used to implement the functions of electronic devices 210 and 265 can communicate with each other via wired or wireless communications. In some implementations, each device communicates with a separate controller or server to manage and coordinate the user experience (e.g., a communication session server). Such a controller or server may be located within physical environment 202 and / or physical environment 250 or may be remote relative to that physical environment.

[0066] Additionally, in Figure 2In the examples, 3D environments 230 and 270 may correspond to the physical environment of the respective viewing user (e.g., where the view shows pass-through video), or they may be a common environment corresponding to only one of the physical environments 202 and 250 or another real or virtual environment. The 3D environment may be based on a common coordinate system that can be shared among users (e.g., providing a virtual room for roles used in a multi-user communication session). In other words, a common coordinate system can be used for 3D environments 230 and 270. A common reference point can be used to align the coordinate system. In some implementations, the common reference point may be a virtual object within the 3D environment that each user can visualize within their respective view. For example, a common center piece table around which a user representation (e.g., a user's role) is positioned within the 3D environment. Alternatively, the common reference point may not be visible within each view. For example, the common coordinate system of the 3D environment may use the common reference point to position each respective user representation (e.g., around a table / desk). Therefore, if the common reference point is visible, each view of the device will be able to visualize the "center" of the 3D environment for perspective when viewing other user representations. The visualization of public reference points can become more relevant to multi-user communication sessions, allowing each user's view to add perspective to each other's points during the communication session.

[0067] In some implementations, the representation of each user can be realistic or unrealistic and / or represent the user's current and / or previous appearance. For example, a photorealistic representation of user 225 or user 260 can be generated based on a combination of a live image of the user and previous images. Previous images can be used to generate representations of portions of the live image data that are not available (e.g., portions of the user's face that are not in the view of the camera or sensor of electronic device 210 or 265, or that may be obscured by the respective device and / or, for example, by the user's hand). In one example, electronic devices 210 and 265 are HMDs, and the live image data of the user's face includes images of the user's cheeks and mouth from a downward-facing camera and images of the user's eyes from an inward-facing camera. This real-time image data can be combined with previous image data of other parts of the user's face, head, and torso that are not currently visible from the device's sensors. Previous data about the user's appearance may be obtained at an earlier time during a communication session, during previous use of the electronic device, during a registration process for obtaining sensor data of the user's appearance from multiple viewpoints and / or conditions, or otherwise.

[0068] In some specific implementations, such as Figure 2The views illustrating the generation of one or more user representations for a communication session (e.g., generating user representation 240, user representation 275) can be based on one or more rendering techniques, such as using 3D meshes, 3D point clouds, or 3D Gaussian scatter rendering methods. Such 3D Gaussian scatter rendering methods can use UV mapping and generate proxy mesh representations.

[0069] In some implementations, generating a view of one or more user representations during a communication session or otherwise involves relighting the one or more user representations. This may involve providing an appearance for the user representation that takes into account the lighting conditions of the environment in which it is to be viewed. For example, if the user representation is to be viewed within a virtual environment with virtual light sources, its appearance may be configured to take into account the lighting, such as the color of the lighting, the type of lighting, the location of the light source, etc. In another example, if the user representation is to be viewed within an extended reality (XR) environment in which some or all of the surrounding environment is based on the physical environment, its appearance may be configured to take into account real light sources in that environment. In yet another example, the environment may include both real and virtual content, such as both real and virtual light sources, and the user representation may be configured to take into account such combinations of real and virtual light sources.

[0070] Some specific implementations disclosed herein generate data (e.g., digital assets) representing the user appearance to be used in a user representation. Such data may have a format (e.g., defining 3D characteristics of a user's face at one or more points in time) that can be used to provide views of such a user representation from different viewpoint locations, such that a view of the face can be provided from a 3D viewpoint location relative to those 3D characteristics. Such data may also have a format that can be used to provide a user representation from different viewpoints while also taking lighting conditions into account (e.g., considering the light expected to be incident on the user representation in a given environment where the user representation is to be viewed). Taking lighting conditions into account can enhance the appearance and / or viewing experience of the user representation, for example, by providing a more realistic representation.

[0071] In some implementations, the data representing a user representation includes information about 3D points that define the appearance of the portion of the user representation corresponding to a 3D vertex of such a point. For example, data for a given 3D point may include spherical Gaussian representation data for displaying 3D Gaussian scatter plots (3DGS) from different viewpoints. This data may include parameters for each point, each identifying characteristics including, but not limited to, position, color (view-dependent / harmonic information), covariance, alpha / transparency, orientation, opacity, range in each axis, rotation, scaling, and / or semantic information (e.g., skin, hair, cheeks, nose, lips, eyebrows, etc.).

[0072] Figures 3A to 3DExamples of 3D Gaussian scatter points used in generating 3D representations of views are illustrated according to some specific implementations. For example, 3D Gaussian sputtering (3DGS) can be used for 3D modeling to represent a complex scene as a combination of a large number of shaded 3D Gaussian volumes rendered into a camera view via sputtering-based rasterization. The position, size, rotation, color, and opacity of these Gaussian scatter points can then be adjusted via differentiable rendering and gradient-based optimization such that they represent the 3D scene given by a set of input images.

[0073] Figure 3A An example of a 3D Gaussian scatter plot 310 is shown, for example, an elliptical shape formed by a 3D Gaussian distribution. The 3D Gaussian scatter plot 310 can be used to represent position (…). µ ), such as xyz coordinates. 3D Gaussian scatter plot 310 can be used to represent direction and angle information for the visibility cone. 3D Gaussian scatter plot 310 can also represent rotation and scaling (e.g., q: covariance matrix), opacity ( α Color (e.g., RGB values), anisotropic covariance, and spherical harmonic (SH) coefficients. Figure 3B An example of environment 320 is shown for rendering scatter points 326 based on the visibility direction of camera 321. Figure 3C An example is shown where scattered points are ordered along the camera's viewing direction along ray 330. For example, scattered points 331, 332, 333, and 334 are identified and ordered along ray 330. Gaussian sputtering is a technique in which individual 3D points are represented as a Gaussian distribution (such as "scattered points") with color values ​​that change according to the viewpoint. This technique uses spherical harmonics to model this view-dependent color variation and enables real-time rendering of high-quality, realistic scenes from sparse image sets. For example, each point has a color calculated based on its position relative to the camera, allowing for realistic shading effects across different viewpoints. Figure 3D An environment 340 is illustrated for blending scatter points 341, 342, 343, 344, and 345, which can be viewed from a camera view along the direction of ray 330 by synthesizing scatter points 331, 332, 333, and 334 on an image plane. Some specific implementations may use screen-to-scatter (e.g., similar to ray casting), scatter-to-screen (e.g., similar to projection), combinations thereof, or other techniques for synthesizing scatter points.

[0074] Figure 4An example environment 400 for generating and displaying a stereoscopic view of scatter points on a device (e.g., an HMD) is illustrated according to some specific implementations. For example, device 410 is an HMD including a first display 420 for a left-eye view and a second display 430 for a right-eye view. The first display 420 and the second display 430 can then display rendered scatter points 450 for each corresponding viewpoint. In some specific implementations, the image generated for each viewpoint can be rendered as a single monochrome image (e.g., rendered for one eye), or the image can be presented as a stereoscopic view, such as... Figure 4 As illustrated in the example. Additionally or alternatively, in some implementations, the image generated for each viewpoint may be rendered as a single raster grid or a combination of stereo raster grids.

[0075] In some implementations, the culling techniques used to select a subset of scatter data for rendering a stereoscopic view can be applied individually to each eye, similar to those described herein, or can be applied based on viewpoint for culling of the overall viewer's FoV, fixation, and / or occlusion. For example, a particular scatter point for the left-eye viewpoint might be occluded and therefore not selected for rendering, but the same scatter point for the right-eye viewpoint might need to be rendered to avoid any holes / gap from the rendering for the right-eye viewpoint. Alternatively, in some implementations, if only one right-eye or left-eye viewpoint requires rendering a particular scatter point, the system described herein will render that scatter point for both viewpoints to avoid any mismatch or blanking in the viewer's overall stereoscopic view.

[0076] In some implementations, modifying the rendering method for different resolutions based on the context of the viewing experience, scatter plot culling, and other techniques can be used to improve rendering efficiency. As further described herein, a scatter plot rendering method can be selected based on one or more criteria associated with determining the context of the viewing experience and continuously updating the rendering method as the viewing experience may change (e.g., using 3D Gaussian sputtering rendering techniques to change the variance, color, size, etc. of the scatter plots).

[0077] Figure 5 Examples are given of generating multiple sets of scatter points based on different levels of detail (LOD) by merging specific groups of scatter points using 3DGS parameters, according to some specific implementations. Specifically, Figure 5 An example 3DGS LOD process 500 is illustrated, which obtains 3DGS data 512 from 3DGS process 510 to generate multiple sets of 3DGS data based on different LODs by merging specific groupings / hierarchies of scatter points based on 3DGS parameters. In some specific implementations, Figure 5The example procedure of 3DGS LOD process 500 illustrates an example data stream procedure used in the registration phase for the transmitting device to generate multiple sets of scatter parameter data that can be used to render the transmitter's representation at the receiving device, which will be referenced in this document. Figure 6 Further description.

[0078] After acquiring 3DGS data 512, the 3DGS LOD process 500 proceeds to the scatter point identification process 520. The scatter point identification process 520 can be used to identify scatter points that can be grouped into static scatter points during runtime (e.g., when rendering a portrait) and dynamic scatter points that can move during runtime. For example, the scatter point identification process 520 can identify specific regions that can move on the character during a communication session, such as, in particular, region 522 (e.g., the eye region) and region 524 (e.g., the mouth region). In other words, by keeping scatter points as static as possible during animation (e.g., the shoulder scatter points in a rendered character do not change), the system can utilize the ability of temporal coherence and identify only static scatter points to be grouped, mapped, and merged, avoiding the merging of specific dynamic regions, such as those requiring more detail and / or potentially moving more during animation and rendering. For example, during a communication session, when the character is generated in real time, the mouth region (e.g., region 524) and / or the eye region (e.g., region 522) may move during the session. Therefore, when merging scatter points, those specific dynamic regions can be avoided.

[0079] In some implementations, dynamic regions (e.g., moving parts of a character, such as eyes and mouth) may be identified based on scatter data, semantics, and / or determined during the registration phase. For example, the system may identify scatter points with the largest scatter parameter changes within a threshold time period, indicating the presence of dynamic moving parts represented by those Gaussian bodies. The system may utilize semantic models to identify which regions are associated with eyes and mouth, etc., and / or the system may identify certain parts of the face that change most significantly from one expression to another during registration (e.g., the mouth and eye regions may move more frequently when making expressions, and these regions may indicate dynamic regions). In some implementations, using temporal coherence to identify static scatter points to be grouped, mapped, and merged may be determined based on a predetermined model and / or based on learning over time (e.g., some scatter points never change within a threshold time period or a threshold time period of a communication session).

[0080] After identifying static and dynamic scatter points in the scatter point identification process 520, the 3DGS LOD process 500 continues to the LOD UV grouping process 530 (e.g., identifying scatter points in the UV space). The LOD UV grouping process 530 divides and merges static scatter points into groups (or grids) and uses them for different LODs in the UV space. For example, LOD-0 UV grouping 532 exemplifies three exemplary groups G1, G2, and G3, such that each group includes multiple static scatter points (e.g., scatter point 533). However, group G3 (e.g., an L-shaped group) excludes dynamic scatter points 531 (e.g., identified dynamic scatter points from region 522 or region 524). LOD-1 UV grouping 534 also exemplifies three exemplary groups G1, G2, and G3, but at this LOD-1 level, one or more scatter points are combined based on one or more 3DGS parameters. For example, four scatter points from group G1 of LOD-0 UV group 532 can be combined into a single scatter point group 535 in G1 based on opacity, another 3DGS parameter, or a combination of 3DGS parameters. LOD-2 UV group 536 exemplifies two groups G1 and G2, but at this LOD-2 level, more scatter points can be combined based on one or more 3DGS parameters. For example, scatter points from groups G1 and G2 of LOD-1 UV group 534 can be combined into a single scatter point group 537 in G1 based on opacity, another 3DGS parameter, or a combination of 3DGS parameters. Furthermore, each level of grouping (e.g., LOD-1 UV group 534 and LOD-2 UV group 536, etc.) excludes identified dynamic scatter points 531.

[0081] After different LOD UV groups are identified in the LOD UV grouping process 530, the 3DGS LOD process 500 may optionally visualize the LOD UV groups in 3D space via the LOD 3D visualization process 540 (e.g., visualize identified scatter points in 3D space). For example, the LOD UV grouping process 530 identifies groups of scatter points in UV space, and then the LOD 3D visualization process 540 visualizes the scatter points in 3D space. The LOD 3D visualization process 540 maps the identified static scatter points to each level group in different level groupings in 3D space. For example, LOD-0 3D data 542 identifies three groups of scatter points G1, G2, and G3 at the LOD-0 level, LOD-1 3D data 544 identifies three groups of scatter points G1, G2, and G3 at the LOD-1 level, and LOD-2 3D data 546 identifies two groups of scatter points G1 and G2 at the LOD-2 level. LOD UV grouping process 530 can identify which scatter points can be merged into a group (e.g., merged via LOD merging process 550).

[0082] The 3DGS LOD process 500 then proceeds to the LOD merging process 550 for merging techniques (e.g., UV space and 3D space merging). For example, the 3DGS LOD process 500 groups nearby Gaussian volumes (e.g., static scatter points) together into single units in 2×2, 4×4, or similar clusters based on an identifier of which scatter points can be merged into a group through the previous process.

[0083] In some implementations, the merging process can increase the size of each scatter point, which can distort the overall appearance of the rendering (e.g., a character). Some merging techniques (i.e., especially 3D spatial voxel grouping) can introduce noticeable holes after merging. However, the grouping process described herein (e.g., LOD UV grouping process 530) and LOD merging process 550 minimize holes or gaps between Gaussian scatter points without requiring rescaling to fill these holes and maintain a consistent representation, as the 3DGSLOD process 500 follows the object's surface. Therefore, the LOD merging process 550 uses the obtained LOD mapping data and, for each level, generates a UV map for each LOD that resolves gaps in the rendering.

[0084] In some specific implementations, 3D points of feature data can be mapped to Gaussian parameters of a UV map (e.g., 3D points and Gaussian parameters are mapped to allow scattering based on a UV grid) to generate multiple LODs of the 3D Gaussian map (e.g., LOD-0 UVM 552, LOD-1 UVM 554, and LOD-2 UVM 556, etc.) for rendering using specific scattering generation techniques (e.g., models / parameters that specify the total number of scattering points, scattering size attributes, etc. for different resolutions). For example, the UV map stores x, y, z positions for scatter parameters (e.g., orientation and angle information (visibility cone information), color (view-dependent / harmonic information), covariance, α / transparency, orientation, opacity, range in each axis, rotation, scaling, and semantic information (e.g., skin, hair, cheeks, nose, lips, eyebrows, etc.). In an exemplary implementation, the 3DGS LOD data 560 generated by the LOD merging process 550 can be used by a viewer's device to receive a 3D Gaussian UV map associated with the transmitter in order to render a representation of the transmitter at the viewer's device (e.g., generating a transmitter persona during a communication session), which will be referenced herein. Figure 6 Further description.

[0085] In some implementations, rescaling techniques may include selective scaling based on the scatter points to be rendered in an effort to prevent gaps in the rendering. For example, the scatter points used for the torso may be disk-shaped or elliptical and are typically large, so gaps can be further separated and made more sparse. The rescaling technique may then selectively scale the torso scatter points in a different manner than the head / face scatter points (e.g., scaling when the scatter points are very large) (e.g., the torso might only require 10k scatter points, but the facial region might require 50k scatter points to achieve rendering quality). Selective scaling techniques can be used to determine trade-offs between quality and rendering efficiency and can be implemented in different levels of detail (LODs).

[0086] Figure 6 An example is illustrated whereby a representation of a user (e.g., a character) is generated by rendering scatter parameter data using a scatter rendering method selected based on the viewing experience context, according to some specific implementations. Specifically, Figure 6 An example user representation process 600 is illustrated, which obtains registration data (e.g., registered images 612a, 612b, 612c) from the transmitter's transmitting device and the viewpoint at the viewer's receiving device from the registration process 610, in order to generate a user representation view using Gaussian sputtering techniques. Furthermore, Figure 6 The example rendering process of the user representation process 600 illustrates the data stream processing between the transmitting device (e.g., transmitter stage 602) and the receiving device (e.g., receiver stage 604), as illustrated by the transmit / receive timeline marker 605.

[0087] The registration process 610 examples illustrate the user (e.g., Figure 1The registration process 610 may include user registration (e.g., pre-registration of registration data) and acquisition of sensor data (e.g., live data registration). As illustrated in Figure 611, user registration may include the user (e.g., user 110) acquiring a full-view image of his or her face using an external sensor on device 105, and thus removing and orienting device 105 (e.g., HMD) toward his or her face during the registration process. A registration avatar may be generated when the system acquires image data (e.g., RGB image) of the user's face while the user is providing different facial expressions. For example, the user may be instructed to "raise your eyebrows," "smile," "frown," etc., to provide the system with a range of facial features for the registration process. A preview of the registration avatar may be shown to the user as the user provides a registration image to visualize the state of the registration process. Registration image data 610 may include registration avatars with different user expressions and from different viewpoints (e.g., a front view in registration image 612a, a right view in registration image 612b, and a left view in registration image 612c). In some examples, more or fewer different facial expressions and / or viewpoints can be used to obtain enough data for the registration process (e.g., such as...). Figure 5 As illustrated, the 3DGS process begins in the registration phase, which can use merging of specific groupings / hierarchies of scatter points based on 3DGS parameters to generate multiple sets of 3DGS data based on different LODs.

[0088] In some implementations, the transformation from registered image data to feature data 622 can occur via a transformation (e.g., via a transformer) as part of the feature data process 620. For example, feature data 622 may include learned feature information obtained by user 110 from the registered image, such as skin, color, and other semantic information per pixel. Feature data 622 may include a list of locations for each feature value (e.g., 14 feature channels). Then, as part of a Gaussian mapping process 630, feature data 622 may be decoded by a decoder to generate a 3D Gaussian mapping for each feature. The 3D points of the feature data may be mapped to Gaussian parameters of a UV mapping (e.g., 3D points + Gaussian parameters) to generate multiple 3D Gaussian mappings as 3D Gaussian mapping data 632 to be rendered by a specific scatter point generation technique (e.g., a model / parameter specifying the total number of scatter points, scatter point size attributes, etc., for different resolutions). For example, the UV mapping stores the x, y, z positions for scatter plot parameters (e.g., orientation and angle information (visibility cone information), color (view-dependent / harmonic information), covariance, α / transparency, orientation, opacity, range in each axis, rotation, scaling, and semantic information (e.g., skin, hair, cheeks, nose, lips, eyebrows, etc.). In other words, the Gaussian mapping process 630 obtains 3D point information that includes sufficient information to generate scatter plots (e.g., 3D vector projection).

[0089] In some implementations, the scatter projection process obtains data from one or more representations and projects Gaussian scatter points onto a 2D height field or RGBDA texture. For example, the scatter points can be a 3D Gaussian distribution in a 2D space with color / density (e.g., parameterization), where a human face can be used as the height field, and multiple scatter points can be determined based on rays leaving the face / height field. The height field / parameterization of the scatter distribution provides higher quality data, allows for faster training of machine learning models, and enables faster (e.g., real-time) rasterization. In some implementations, 3D mapping information (e.g., identifying x, y, z positions corresponding to UV coordinates of a UV map) can be generated at registration (e.g., a Gaussian UV map).

[0090] In some implementations, process 600 continues after 3D Gaussian map data 632 is generated from Gaussian UV mapping process 640 (e.g., at the transmitter's device after registration). The system (e.g., at the viewer's device) can obtain 3D Gaussian map data 632 (e.g., feature data mapped to Gaussian parameters via UV mapping) and project Gaussian data using the current viewpoint (e.g., viewpoint data 644) to determine 2D Gaussian map data 642 (e.g., 2D points + Gaussian parameters) for Gaussian mapping process 630 at the receiving device. In other words, the receiving device can obtain multiple sets of scatter points comprising 3D Gaussian data that can be rendered based on a selected scatter rendering method (e.g., rendering mode).

[0091] In some specific implementations, the 3D Gaussian mapping process 640 may include a method based on... Figure 5 The 3DGS LOD process 500 determines the different LODs of the 3DGS data 560. For example, Gaussian map data 642 may include multiple sets of data based on different LODs, which are based on a specific grouping / hierarchy of scatter points merged according to 3DGS parameters, but merging regions that may move during a communication session (e.g., eye / mouth regions). In other words, Gaussian map data 642 may be based on LOD UV grouping data 532, 534, 538 from LOD UV grouping process 530, which is visualized in 3D space as LOD 3D data 542, 544, 546 from LOD 3D visualization process 540, and merged into LOD merged data 552, 554, 556 from LOD merging process 550 to generate different LODs of the 3DGS data 560.

[0092] In some implementations, a scatter rendering method (e.g., rendering mode) is selected during the rendering selection process 650. For example, a context analyzer 652 may identify a scatter rendering model (e.g., high-resolution versus low-resolution scatter models) based on sensor data 652 of the viewing environment (which may include the viewer's viewpoint / gaze data 644) and transmitter data 654. For example, user / scene understanding information may be obtained to determine different contexts, such as identifying viewpoint distances, determining whether the user is capturing images / videos, speaking to specific characters, and stating the user's name as represented by characters in the user's field of view. In some implementations, the viewing experience may be the experience of a view in which a representation of at least a portion of objects is rendered based on a viewpoint in a 3D environment. For example, scene understanding of the physical environment may be obtained or determined by device 105 based on depth data and light intensity image data. For example, user 110 may scan his or her environment and, based on image data (e.g., light intensity and / or depth), a scene understanding instruction set (e.g., Figure 6The context analyzer 652 can determine information about the physical environment (e.g., objects, people within the environment, etc.) to track object-based data based on generating a scene understanding of the environment. Additionally, user / scene understanding can be used to determine the context of the experience and / or environment (e.g., creating scene understanding to determine objects or people in the content or environment, where the user is, what the user is watching, etc.). The transmitter data 654 may include additional contextual data that the context analyzer 652 can use to determine a scatter rendering method (e.g., the transmitter or the transmitter's device may transmit information requiring higher resolution scatter rendering, i.e., taking a photograph or having higher resolution for a character the user is interacting with / speaking to during a communication session (among other characters in the environment)).

[0093] In some implementations, the context analyzer 652 determines the context of the viewing experience based on one or more different criteria and can continuously update the determined context, allowing for frame-by-frame updates to the scatter rendering method. In some implementations, the context of the viewing experience may be based on distance to the user (e.g., the user sees more character details closer than at a distance, and those characters are more likely to interact with the user). In some implementations, the context of the viewing experience may be based on detected user speech (e.g., the user saying a character's name, which would cause that character to be rendered at high resolution). In some implementations, the context of the viewing experience may be based on detected speech from a character speaking to the user (e.g., a character saying the user's name and starting to talk about something, and the rendering method will change to a higher resolution). In some implementations, the context of the viewing experience may be based on detected user intent communicating with a character and / or sensor fusion (e.g., gaze, spoken words, and gestures). For example, if a character is not interacting with the user, that character may be switched to rendering at a lower resolution, but when the character says the name of a user indicating that he / she wants to interact with the user, the character can still be rendered at a lower resolution until the user gazes at the character who said the user's name. Therefore, in an exemplary implementation, the resolution modification that is continuously updated during the viewing experience based on one or more 3D Gaussian-specific parameters (e.g., variance, size, color, etc.) can be modified.

[0094] After rendering selection process 650 and selecting a scatter rendering method (e.g., a high-resolution (e.g., 64k or LOD-0) versus a low-resolution (e.g., 32k, 16k, LOD-1, LOD-2, etc.) scatter rendering model), Gaussian sputtering can then be used for rendering to generate a view for the user representation generated in user representation generation process 660. For example, the user representation for user representation generation process 660 may generate a high-resolution user representation 662, a medium-resolution user representation 664, or a low-resolution user representation 666, depending on the selected scatter rendering method and the determined context of the viewing experience. For example, 3D Gaussian sputtering techniques use 2D points from a mapping (e.g., UV mapping) and associated Gaussian parameters to render an image using Gaussian sputtering based on the selected scatter rendering method and the viewer's current viewpoint (e.g., viewpoint and gaze data 644). Each rendering representation of 662, 664, 666 is instantiated as a user representation (e.g., a character), but the Gaussian sputtering technique described herein can be used for any 3D object or 3D scene reconstruction.

[0095] In some specific implementations, the scatter rendering methods used to generate different user representations 662, 664, and 666 can vary and can be adjusted based on resolution, different scatter sizes, different colors or fewer color gamuts, different numbers of Gaussian volumes rendered along the viewing ray 330, or combinations thereof. For example, as... Figure 3C As illustrated, high-resolution scatter rendering techniques can render more Gaussian scatter points along the viewing ray 330, while lower resolution can render only the first and / or second scatter points. In some specific implementations, lower-resolution scatter rendering techniques can use proxy rendering. For example, a proxy representation process can be used to generate a 3D proxy representation that is easier to render (e.g., a textured 3D mesh, a textured height field, etc.), and use that 3D proxy representation to render additional views needed to provide the view at a faster rate (e.g., second and third frames needed for a view that provides 30Hz content at 90Hz).

[0096] The various specific implementations discussed in this paper for each scatter rendering technique for different versions of the quality used to generate representations (e.g., low resolution to high resolution) are merely exemplary and are not intended to be limiting. For example, other implementations may have more or fewer than three different resolution levels (e.g., low, medium, high, LOD-1, LOD-2, LOD-3, etc.), and some implementations may not have discrete levels and may have a continuous slider approach that can be adjusted based on the context of the viewing experience, as the context of the viewing experience can change frame by frame.

[0097] Figure 7Exemplary electronic devices, according to some specific implementations, operate in different physical environments during communication sessions with multiple users, the communication sessions including views of other users' 3D representations using a selected scatter rendering method. For example, Figure 7 Exemplary operating environments 700 are illustrated for electronic devices 715, 725, 735, and 745 operating in different physical environments 710, 720, 730, and 740 during a communication session (e.g., when electronic devices 715, 725, 735, and 745 share information with each other or with intermediate devices (such as a communication session system / server (e.g., information system 790))). Similar to... Figure 2 The communication session instantiated in Figure 7 The example illustrates a communication session between users 712, 722, 732, and 742 when wearing devices 715, 725, 735, and 745 (e.g., HMD).

[0098] exist Figure 7 In this example, each physical environment 710, 720, 730, 740 is a room including the user and one or more objects when each user is viewing the environment on each corresponding device. Electronic devices 715, 725, 735, and 745 may include one or more cameras, microphones, depth sensors, or other sensors that can be used to capture and evaluate information about each physical environment and its objects, as well as information about each user on each electronic device. Information about the physical environment and / or the user can be used to provide (e.g., for user representation) visual content and (e.g., for audible speech or text transcription) audio content during a communication session. For example, a communication session may provide views of one or more participants to a 3D environment generated based on camera images and / or depth camera images of the physical environment, and provide one or more user representation views for each user of the communication session.

[0099] Figure 7 An example of a view 750 of user 742 in a virtual environment (e.g., 3D environment 770) at device 745 is further illustrated. Specifically, user 742's view 750 provides a representation 752 of table 748 and a representation 754 of plant 746 from environment 740. Furthermore, user 742's view 750 provides user representations 772, 774, and 776 of users 712, 722, and 732 respectively positioned around the representation 752 of table 748 (e.g., the role of each user in a communication session). In other words, Figure 7An example is illustrated with multiple people at different physical locations, each wearing an HMD and registered with their roles, and thus having a 3D Gaussian representation. All of these users are then in a shared visual space (such as a virtual meeting room), and the user representations are positioned around a virtual table (e.g., representation 752) and exist in their own respective physical environments (but appear to be in the same virtual or mixed reality environment) with many Gaussian representations of those wearing head-mounted devices. Specifically, each user representation 772, 774, 776 is generated in real-time at a receiving device (e.g., device 745) using a selected scatter rendering method. For example, as discussed herein, the scatter rendering method may be selected between a higher resolution representation and a lower resolution scatter rendering model based on one or more criteria (e.g., the context of the viewing experience). Gaussian sputtering can be used to render views to generate each user representation 772, 774, 776. For example, scatter rendering techniques (e.g., Figure 6 The user representation generation process 660 generates user representations to provide views of high-resolution user representation 772, medium-resolution user representation 774, or low-resolution user representation 776, depending on the selected scatter rendering method and the determined context of the viewing experience.

[0100] In some specific implementations, Figure 6The context analyzer 652 determines the context of the viewing experience based on one or more different criteria and can continuously update the determined context, allowing for frame-by-frame updates to the scatter rendering method. In some implementations, the context of the viewing experience may be based on distance to the user (e.g., a user can see more character details closer than at a distance, and those characters are more likely to interact with the user). For example, representation 772 is shown with higher detail / resolution because it is closer to the view of user 742. In some implementations, the context of the viewing experience may be based on detected user speech (e.g., a user saying a character's name, which would cause that character to be rendered at a high resolution). For example, representation 772 can be shown with higher detail / resolution because the user is speaking to user 742 (e.g., as illustrated, user 732 is standing, speaking, and moving his or her hand, so representation 772 is updated accordingly). In some implementations, the context of the viewing experience may be based on detected speech of a character speaking to the user (e.g., a character may say the user's name and begin talking about something, and the rendering method will change to a higher resolution). In some implementations, the context of the viewing experience can be based on detected user intent communicating with a character and / or sensor fusion (e.g., gazes, spoken words, and gestures). For example, if a character is not interacting with the user, that character can be switched to be rendered at a lower resolution, but when the character speaks the name of a user indicating that he / she wants to interact with the user, that character can still be rendered at a lower resolution until the user gazes at the character who spoke the user's name. For example, representation 774 may have already called the user's name 742, and then once the viewer (e.g., user 742) gazes at representation 774, the system can increase the resolution of representation 774 (e.g., to a higher resolution, such as representation 774). Thus, in exemplary implementations, the resolution modification, which is continuously updated based on context during the viewing experience, can be modified based on one or more 3D Gaussian-specific parameters (e.g., variance, size, color, etc.).

[0101] Additionally, in Figure 7In the example, the 3D environment 770 is an XR environment based on a common coordinate system that can be shared with other users (e.g., a virtual room for a role in a multi-user communication session). In other words, the common coordinate system of the 3D environment 770 is different from the coordinate system of the physical environment. For example, a common reference point can be used to align the coordinate system. In some implementations, the common reference point can be a virtual object within the 3D environment that each user can visualize within their respective view. For example, a common center piece table around which a user representation (e.g., a user's role) is positioned within the 3D environment. Alternatively, the common reference point is not visible within each view. For example, the common coordinate system of the 3D environment can use the common reference point to position each corresponding user representation (e.g., around a table / desk, such as representation 752). Therefore, if the common reference point is visible, each view of the device will be able to visualize the "center" of the 3D environment for perspective when viewing other user representations. The visualization of the common reference point can become more relevant to the multi-user communication session, allowing each user's view to add perspective to each other's fixed point during the communication session.

[0102] Figure 8 Examples are given based on some specific implementations. Figure 1 An example view 800 of the 3D environment 802 provided by a device (e.g., device 105). Specifically, Figure 8 A stereoscopic view 800 of a 3D environment 802 is illustrated, comprising three distinct characters 820, 830, and 840, each positioned at a different 3D fixed point relative to the user's viewpoint provided by a gaze cone 810 on device 105. The gaze cone 810 represents an example foveal region within the user's FoV relative to the gaze vector of the 3D environment 802 (e.g., the user's gaze). In other words, the foveal region is the gaze cone 810 surrounding the gaze vector. Furthermore, the gaze cone 810 illustrates 3D locations (e.g., the 3D fixed point of each character 820, 830, 840 as seen from the user's viewpoint) where objects can be depicted within the 3D environment 802 relative to the user's gaze vector. Specifically, the gaze cone 810 illustrates the user's foveal region of the user's FoV based on the gaze vector, such that the area outside the gaze cone 810 but within the user's view (the foveal region) is the peripheral region.

[0103] exist Figure 8In the illustrated example, character 820 is located within the concave region of the gaze cone 810 and can therefore be rendered at a higher level of detail (e.g., 64k scatter points). In contrast, characters 830 and 840 are each located outside the gaze cone 810 but within the user's field of view (e.g., at the periphery), thus characters 830 and 840 can be rendered at a lower level of detail (e.g., 16k scatter points) by leveraging the reduced sensitivity of the human visual system to peripheral changes. In other words, the system dynamically adjusts the level of detail for each character based on its relevance to the user's current focus (such as gaze direction or field of view (FoV)). This approach enables efficient use of GPU resources, improves overall frame rate, and increases battery life without noticeable loss of visual quality at the periphery. Figure 9 Figure 9 Examples are given based on viewpoint characteristics and specific implementations. Figure 8 Examples of different regions associated with the gaze cone 810. Specifically, Figure 9 This paper illustrates how a gaze cone 810 (e.g., the foveal region of the user's field of view) can be used to provide a dynamic, context-aware system for rendering 3D representations using 3D Gaussian scatter plots, where advanced LOD transitions are based on distance, gaze, and field of view. This approach minimizes visual artifacts during LOD changes, optimizes rendering efficiency, and maintains high visual fidelity in the region of interest.

[0104] For example, the foveal region (e.g., gaze cone 810) includes regions 920, 922, 924, and 926, each of which can be triggered to transition between different Levels of Detail (LODs) based on a distance associated with a 3D location at which at least a portion of an object representation (e.g., character 820) is to be depicted within the 3D environment. For instance, as character 820 moves further away from the user's viewpoint, the system can use merging techniques to gradually transition between LOD levels before switching to a lower level of detail representation. In other words, regions can define transition zones (e.g., switching from region 920 to region 922—starting from 2 meters to 2.5 meters), in which scattered points are merged into groups, thereby reducing the total number of scattered points (e.g., from 64k to 16k) before switching to a lower LOD character. For example, characters in region 920 can be rendered at LOD-0, characters in region 922 can be rendered at LOD-0 and / or LOD-1, characters in region 924 can be rendered at LOD-1 and / or LOD-2, and characters in region 926 can be rendered at LOD-2. This transition merging process can be continuous and can be adjusted based on a distance threshold.

[0105] Furthermore, for transitional merging, intermediate LOD states (e.g., a combination of LOD-0 and LOD-1) can be utilized. For example, the exact proportion of each LOD state varies with distance; more LOD-0 indicates closer proximity to the camera, and more LOD-1 indicates further distance. The runtime algorithm that generates LOD states operates by merging scatter points into larger scatter points. In some implementations, the algorithm does not necessarily need to output a fixed number of merged scatter points; instead, a variable number of scatter points can be selected for merging (randomly or in a predetermined manner), and this variable number can be controlled via a knob, such as the distance or angle to the camera.

[0106] In some implementations, this same process can be based on an angle threshold, as the user can change his or her field of view and / or gaze direction. For example, when a character transitions to the periphery (e.g., from one of the foveal regions 920, 922, 924, 926 to one of the peripheral regions 930, 932, 940, 942). In some implementations, the reduction in scatter point size can be based on the angle of gaze. Furthermore, the process can combine both an angle threshold (gaze direction) and a character distance threshold. For example, when transitioning from region 924 to region 940, both the angle threshold and the distance threshold can be applied progressively before transitioning to LOD-2. In other words, the merging of scatter points across characters may be uneven. Instead, the system uses gaze tracking to determine which regions are in the viewer's focus and to maintain higher detail in those regions while more aggressively merging scatter points in peripheral or less focused regions. The angle with the gaze center and the field of view are used to guide the merging process.

[0107] In some implementations, the outer regions 930, 932, 940, and 942 (e.g., outside the central concave region) may be prohibited from rendering at the highest resolution level (e.g., LOD-0). In other words, the central concave region may allow rendering of a first set of LOD-based regions (e.g., LOD-0, LOD-1, and LOD-2), and the outer regions may allow rendering of a second set of LOD-based regions (e.g., only LOD-1 and LOD-2), which is different from the first set.

[0108] In some implementations, the system can switch to a sparser UV representation in the peripheral region (outside the fovea region), allowing for more aggressive LOD reduction without perceptible loss of detail. Furthermore, the number of scatter points can be dynamically increased or decreased (e.g., between 8k and 64k) as the viewer's distance or gaze changes, providing a seamless user experience. In some implementations, to prevent flickering during transitions / merging, rescaling of the scatter points can be used to prevent gaps in the scatter points. Additionally or alternatively, in some implementations, to prevent flickering during transitions / merging, the size and / or threshold level of the gaze cone 810 and / or the region within the gaze cone 810 (fovea region) or the region outside the gaze cone 810 (e.g., in the periphery) can be increased or decreased based on LOD, the context of the viewing experience, and / or viewpoint characteristics associated with the current view.

[0109] In some implementations, as the rendered view is continuously updated between LODs, the transitional merging of scatter points involves deterministically calculating the merge ratio. For example, the system may begin merging groups of scatter points by deterministically calculating the merge ratio when scatter points cross distance boundaries (e.g., from region 922 to 924). In some implementations, scatter points from all groups may not be merged immediately. In some implementations, the merging threshold may become increasingly lower as a character crosses distance and / or angle boundaries. This merging of scatter points allows the system to create a large number of intermediate LOD states, which helps the system smoothly transition to a fixed LOD state (e.g., from merging all scatter point groups, such as LOD-0, LOD-1, and LOD-2).

[0110] Additionally or alternatively, in some implementations, random merging of scatter points can be utilized. For example, a mask (created via a random distribution) can be used for transition merging, where the system determines to merge scatter points in a group if the mask has a value above a certain threshold. Using a low-dispersion random distribution can further help to fade and hide any noticeable "pop-up" effects to the viewer during the transition. This partial merging of scatter points allows the system to create a large number of intermediate LOD states, which helps the system smoothly transition to a fixed LOD state (e.g., from merging all scatter point groups, such as LOD-0, LOD-1, and LOD-2).

[0111] Figures 10A to 10C Examples are given of different rendering levels of detail for user characters based on viewpoint characteristics (e.g., the user's current focus, such as gaze direction or field of view) in specific implementations. Specifically, Figures 10A to 10CAn example is illustrated of a process for representing instruction set 1030, which dynamically provides context-based selection and transitions between Levels of Detail (LOD) for rendering character 1020. Its focus is on minimizing perceptible transition artifacts while simultaneously optimizing distance, gaze, and field-of-view variations associated with the viewpoint of user 110 relative to the 3D position of the rendered character 1020. In other words, whether character 1020 is inside or outside the user's central field of view (e.g., the foveal region or periphery) can be determined based on the user's viewpoint, gaze, and / or the position of character 1020 relative to the centerline of the field of view. For example, Figure 10A This illustrates the process for generating characters at a higher level of detail (e.g., LOD-0 (approximately 64k scatter points)). Figure 10B The process for generating characters at an intermediate level of detail (e.g., LOD-1 (approximately 32k scatter points)) is illustrated, and Figure 10C An example is given of a process for generating characters at a lower level of detail (e.g., LOD-2 (16k scatter points or less)).

[0112] like Figure 10A As illustrated, the user 110 of the wearable device 105 has a field of view that includes a view of the character 1020 within a foveal region (e.g., the cone of gaze 1010). Then, the representation instruction set 1030 may utilize one or more techniques described herein for generating views to generate representation data 1032 to render a view 1040 of the 3D environment 1042. Figure 10A As illustrated, character 1020 in view 1040 is centered in the concave (focus) region of user 110 relative to the background (e.g., mountains) in order to generate a higher-fidelity view of the character (e.g., higher quality—more scattered points—with perceived stereo depth).

[0113] like Figure 10B As illustrated, user 110 of wearable device 105 has a foveal region (e.g., gaze cone 1010) within the field of view, which includes the view of character 1020 in the peripheral region—where character 1020's position is to the right relative to the center line of the field of view. In other words, user 110's gaze and / or head position has shifted relative to character 1020's 3D position, and / or character 1020's 3D position has shifted relative to user 110's gaze and / or head position in the view of the 3D environment, such that character 1020's 3D position is now mostly outside user 110's field of view (to its right), however, a portion of character 1020's 3D position is slightly inside user 110's foveal region, as depicted by gaze cone 1010. In other words, when user 105 moves his or her head (or gaze direction) to the left, a shift may occur from... Figure 10A LOD-0 in Figure 10BThe transition to LOD-1 in the process. Therefore, the representation instruction set 1030 can then utilize one or more techniques, as described herein, for generating transition views (merging) to generate representation data 1032 for rendering view 1060 of 3D environment 1062. Figure 10B As illustrated, the character 1020 in view 1060 is mostly outside the concave region (view cone 1010) of the user 110, but a portion remains slightly within the view cone 1010 relative to the background (e.g., now looking to the left of the mountain). Therefore, the system can determine the transition and generate an intermediate-fidelity view of the character (e.g., transitioning / merging from LOD-0 64k scatter points to LOD-1 32k scatter points).

[0114] like Figure 10C As illustrated, user 110 of wearable device 105 has a field of view that includes a view of character 1020 in the peripheral region—where the position of character 1020 is to the left of the center line of the field of view relative to the foveal region (gaze cone 1010). In other words, user 110's gaze and / or head position has moved relative to the 3D position of character 1020, and / or the 3D position of character 1020 has moved relative to user 110's gaze and / or head position in the view of the 3D environment, such that the 3D position of character 1020 is now outside (to its left) and further away from user 110's field of view, as depicted by gaze cone 1010. Then, representation instruction set 1030 may utilize one or more techniques described herein for generating transitional / merging views to generate representation data 1032 to transition to LOD-2 and render view 1050 of 3D environment 1052. Figure 10C As illustrated, the character 1020 in view 1050 is positioned relative to the background at the periphery of the central concave region of the user 110 (e.g., now looking to the right of the mountain) in order to generate a lower-fidelity view of the character (e.g., rendered with less scattering quality, since the human visual system is less sensitive to changes in the periphery and / or image quality).

[0115] Figure 11 This is a flowchart illustrating exemplary method 1100. In some specific implementations, the device (e.g., Figure 1Device 105) performs the techniques of method 1100 according to some specific embodiments to render scatter parameter data to provide a represented view by using a scatter rendering method selected based on the context of the viewing experience. In some specific embodiments, the techniques of method 1100 are performed on a mobile device, desktop computer, laptop computer, HMD, or server device. In some specific embodiments, method 1100 is performed on a processing logic component (including hardware, firmware, software, or a combination thereof). In some specific embodiments, method 1100 is performed on a processor executing code stored in a non-transitory computer-readable medium (e.g., memory). In some specific embodiments, method 1100 is implemented at the processor of a device (such as a viewing device) that renders a user representation (e.g., Figure 2 Device 210 renders a 3D representation 240 (character) of user 260 based on data obtained from device 265.

[0116] At box 1110, method 1100 obtains representation data representing at least a portion of an object at the device's processor. This representation data includes data for multiple sets of scatter points generated based on a scatter point generation technique, whereby the data for each of these multiple sets of scatter points includes 3D Gaussian data. For example, the scatter point generation technique (e.g., a model / parameter specifying the total number of scatter points, scatter point size attributes, etc.) may include multiple sets of scatter point parameter data (e.g., sets of scatter point parameter data at different resolutions / sizes ranging from 16k to 64k scatter points, or different LODs based on merged grouping of static scatter points, etc.). In some implementations, the representation data may be rendered using Gaussian sputtering, a technique in which individual 3D points are represented as a Gaussian distribution with color values ​​that change according to the viewpoint (e.g., scatter point parameter data, also referred to as "scatter points"). This technique uses spherical harmonics to model this view-dependent color variation and enables real-time rendering of high-quality, realistic scenes from sparse image sets. For example, each point has a color calculated based on its position relative to the camera, thus allowing for realistic shading effects across different viewpoints.

[0117] In exemplary embodiments, the device is a viewing device that renders a representation of an object, such as a user representation (e.g., a character). In some embodiments, at least a portion of the object (e.g., a person) includes the user's face and additional portions (e.g., the user's head, neck, clothing, hair, body, etc., on the transmitting device). In other words, method 1100 provides a view of the transmitter's representation at the viewer's device based on data from the transmitter's device. In exemplary embodiments, a user representation of the transmitter (e.g., a character) is provided for viewing; however, the techniques described herein (e.g., Gaussian sputtering culling) can render any type of object or scene for reconstructing the view at the viewer's device.

[0118] In some implementations, user representation data is based on UV mappings and 3D point cloud points associated with distributed data, which defines the size and shape of the 3D point cloud points used to render them as scatter points corresponding to each point of the UV mapping (e.g., a 3D Gaussian mapping). For example, as... Figure 6 As illustrated, a user representation 662 can be generated for a specific viewpoint based on the points rendered from the Gaussian UV map 642 after the rendering selection process 650.

[0119] In some specific implementations, scatter plot parameter data includes 3D Gaussian parameters. For example, scatter plot parameter data may include location information (e.g., 3D location), orientation and angle information (e.g., visibility cone), color information, covariance information, transparency information, orientation, opacity information, range information in each axis, rotation data, scaling, and / or semantic information (e.g., skin, hair, cheeks, nose, lips, eyebrows, etc.). For example, as... Figure 6 As illustrated, the 3D Gaussian mapping process 630 generates 3D Gaussian mapping data 632 (e.g., multiple sets of scatter parameter data) by transforming registered image data into multiple sets of feature data. These multiple sets of feature data may include user-learned feature information obtained from the registered image, such as skin, color, and other semantic information for each pixel. For example, scatter parameter location may identify where the scatter points are located based on xyz coordinates, scatter parameter covariance may identify how the scatter points are stretched / scaled (e.g., a 3×3 matrix), scatter parameter color may identify RGB colors, and scatter parameter alpha (α) may identify the transparency of the scatter points. In some specific implementations, user representation data includes 3D point cloud points associated with distribution data that defines the size and shape used to render the 3D point cloud points as scatter points. For example, the scatter model generated using Gaussian scattering includes texture / color, location, and scatter point shape, etc. In some specific implementations, the user representation data includes 3D mapping information, which includes feature values ​​and position information for each mapping point (e.g., identifying the x, y, z positions corresponding to the UV coordinates of the UV mapping).

[0120] In some implementations, the scatter point generation technique includes a scatter point criterion threshold (e.g., a quality scatter point threshold) associated with the resolution of each scatter point associated with each of the plurality of sets of scatter points. In some implementations, the scatter point generation technique includes a scatter point criterion (e.g., a quantity scatter point threshold such as 32k vs. 64k) associated with the number of scatter points associated with each of the plurality of sets of scatter points. In some implementations, the scatter point generation technique includes a scatter point criterion (e.g., a size scatter point threshold) associated with the size of each scatter point associated with each of the plurality of sets of scatter points.

[0121] In exemplary embodiments, the multiple sets of scatter points may be determined based on different levels of detail (LODs), which are determined based on a specific static grouping / hierarchy for merging scatter points according to 3DGS parameters, but avoid merging dynamic regions that may move during a communication session (e.g., eye / mouth regions). Therefore, for method 1100, in some embodiments, each of the multiple sets of scatter points is based on a different LOD. In some embodiments, the scatter point generation technique includes scatter point criteria associated with merging multiple subsets of scatter points based on the scatter point parameter data. In some embodiments, the multiple subsets of scatter points are selected based on determining one or more static regions associated with representing at least a portion of the object. For example, the multiple sets may be based on different LODs, which are based on a specific static grouping / hierarchy for merging scatter points according to 3DGS parameters, but avoid merging regions that move during a communication session (e.g., eye regions 522 and mouth regions 524 identified by the static / dynamic scatter point identification process 520). In some implementations, the merging of multiple subsets of scatter points is determined (e.g., generated and / or updated) at the rendering device (e.g., a viewer's device that renders the transmitter's character after receiving 3D Gaussian data). Additionally or alternatively, in some implementations, the merging of multiple subsets of scatter points is determined (e.g., generated and / or updated) during the registration process (e.g., the transmitter's device generates a 3D Gaussian LOD and then transmits the LOD of the 3D Gaussian data to the viewer's device, which renders the transmitter's character based on the received 3D Gaussian data and the different LODs based on the live data).

[0122] At box 1120, method 1100 determines the context of the viewing experience. For example, user / scene understanding information may be obtained to determine different contexts, such as identifying viewpoint distance and determining whether the user is capturing images / videos. In some implementations, the viewing experience may be the experience of a view in which a representation of at least a portion of an object is rendered based on a viewpoint in a 3D environment. For example, scene understanding of the physical environment may be obtained or determined by device 105 based on depth data and light intensity image data. For example, user 110 may scan his or her environment and, based on image data (e.g., light intensity and / or depth), a scene understanding instruction set (e.g., ... Figure 6 The context analyzer 652 can determine information about the physical environment (e.g., objects, people within the environment, etc.) to track object-based data based on the generated scene understanding of the environment. For example, scene understanding can be used to more accurately track the movement of user 110 relative to the environment. Additionally, scene understanding can be used to determine the context of the experience and / or environment (e.g., creating scene understanding to determine objects or people in the content or environment, where the user is, what the user is watching, etc.).

[0123] In some implementations, the context of the viewing experience includes the viewpoint of the represented view, and the scatter rendering method used to provide the view is selected based on a threshold distance associated with the viewpoint. For example, when the viewpoint exceeds the threshold distance, a lower-quality scatter set generated using a 32K scatter model or based on LOD-1, LOD-2, or similar partitioning / merging groups is used, and when the viewpoint is less than the threshold, a higher-quality scatter set generated using a 64K scatter model (e.g., scatter points that have not yet been partitioned / merged into groups, e.g., LOD-0) is used. Additionally or alternatively, in some implementations, the context of the viewing experience includes at least one frame recording content associated with the viewpoint of the represented view, and the scatter rendering method used to provide the view is selected based on that at least one frame recording the content. For example, a 64K scatter model is used where the viewer is capturing images / videos.

[0124] At box 1130, method 1100 selects a scatter rendering method (e.g., rendering mode) for providing a view that includes at least a portion of the object, based on the determined context of the viewing experience. For example, when the viewpoint exceeds a threshold distance, a lower-resolution scatter model (e.g., a 32K scatter model, or LODs that have been partitioned / merged, such as LOD-1, LOD-2, etc.) is used, or when the viewpoint is less than the threshold, a set of scatter points generated using a higher-resolution scatter model (e.g., a 64K scatter model, or LODs that have not yet been partitioned / merged, such as LOD-0) is used. For example, the threshold distance may be based on perceptual distance (e.g., the apparent distance from the view to the representation during view rendering). For example, if the perceptual distance of the representation to be rendered is greater than the perceptual distance threshold (e.g., in the background of the current view), a lower-quality rendering of the representation may be selected. On the other hand, if the perceptual distance of the representation to be rendered is closer than the perceptual distance threshold (e.g., a user representation (role) during a communication session), a higher-quality rendering of the representation may be selected. Additionally, the defined context of the viewing experience may indicate that the viewer is taking pictures of the viewing experience or recording videos of the viewing experience, thus expecting to render a higher quality representation.

[0125] In some implementations, the context analyzer determines the context of the viewing experience based on one or more different criteria and can continuously update the determined context, allowing for frame-by-frame updates to the scatter rendering method. In some implementations, the context of the viewing experience can be based on distance to the user (e.g., a user can see more character details closer than at a distance, and those characters are more likely to interact with the user). For example, representation 772 might be shown with higher detail / resolution because it is closer to user 742's view. In some implementations, the context of the viewing experience might be based on detected user speech (e.g., a user saying a character's name, which would cause that character to be rendered at high resolution). For example, representation 772 might be shown with higher detail / resolution because user 732 is speaking to user 742 (e.g., when user 732 is speaking and moving his or her hand, such as during group rendering, representation 772 is shown facing the viewer, i.e., user 742).

[0126] In some implementations, the context of the viewing experience may be based on detected speech from a character speaking to the user (e.g., the character may say the user's name and begin talking about something, and the rendering method will change to a higher resolution). In some implementations, the context of the viewing experience may be based on detected user intent communicating with a character and / or sensor fusion (e.g., gazes and spoken words and gestures). For example, if a character is not interacting with the user, that character may be switched to be rendered at a lower resolution, but when the character says the user's name indicating that he / she wants to interact with the user, that character may still be rendered at a lower resolution until the user gazes at the character who said the user's name. For example, representation 774 may have already called the user's name 742, and then once the viewer (e.g., user 742) gazes at representation 774, the system may increase the resolution of representation 774 (e.g., to a higher resolution, such as representation 774). Thus, in exemplary implementations, the resolution modification, which is continuously updated during the viewing experience based on context, may be modified based on one or more 3D Gaussian-specific parameters (e.g., variance, size, color, etc.).

[0127] At box 1140, method 1100 provides a view of a representation of at least a portion of the object based on a selected scatter rendering method, providing the view including rendering a subset of scatter points based on the scatter parameter data. For example, the selected scatter rendering method (e.g., selecting a rendering mode, such as high-resolution rendering versus low-resolution rendering) can be used to render a view of a user-presented object (e.g., a character), or to render another object or scene. Furthermore, 3D Gaussian sputtering can be used to avoid or fill holes, and body pose data can be applied to additional areas including the user (e.g., the neck / shoulder area). For example, as... Figure 6 As illustrated, user representations 662, 664, 666, etc., are generated for a specific viewpoint based on points rendered from a Gaussian map 642, which combines scatter parameters obtained from registration data, and these user representations are updated for each frame based on one or more marker points (e.g., a set of semantic points associated with facial features or other regions corresponding to the transmitters of the rendered user representations 662, 664, 666, etc.).

[0128] In some implementations, selecting the scatter plot rendering method involves choosing between a first method for rendering a first set of a plurality of scatter plots and a second method for rendering a second set of a plurality of scatter plots, wherein the first set of scatter plots has more scatter plots than the second set of scatter plots. For example, the system may determine whether to select LOD-0 (64k scatter plots) or LOD-1 (32k scatter plots).

[0129] In some implementations, selecting the scatter rendering method includes choosing a method to reduce the rendering level of detail (LOD) by switching to a lower LOD representation. Rendering the lower LOD representation includes merging a (second) subset of scatter points in a transition zone based on the location of the at least part of the object to be depicted within the 3D environment. For example, based on the distance between the viewer and the character, and selecting the rendering method before switching to a different LOD representation. In some implementations, during the merging of the (second) subset of scatter points in the transition zone, the number of scatter points rendered decreases as the distance associated with the location of the at least part of the object to be depicted increases. For example, the number of scatter points is gradually decreased (or increased) before switching to a lower LOD representation. In some implementations, merging the (second) subset of scatter points is performed in groups, where each group corresponds to a range of distances associated with the location of the at least part of the object to be depicted. For example, as... Figure 8 and Figure 9 As illustrated, transitions between LODs can be performed gradually without abruptly changing the number of rendered scatter points.

[0130] In some implementations, intermediate LOD states (e.g., a combination of LOD-0 and LOD-1) can be utilized for transitional merging. For example, the exact proportion of each LOD state varies with distance; more LOD-0 indicates closer proximity to the camera, and more LOD-1 indicates further distance. The runtime algorithm that generates LOD states operates by merging scatter points into larger scatter points. In some implementations, the algorithm does not necessarily need to output a fixed number of merged scatter points; instead, a variable number of scatter points can be randomly selected for merging, and this variable number can be controlled via a knob, such as the distance or angle to the camera.

[0131] In some implementations, merging scatter points can be based on determining whether viewpoint characteristics meet a criterion, where the viewpoint characteristics are determined based on the viewpoint and the location of at least a portion of the object to be depicted within the 3D environment. For example, as Figures 10A to 10C As illustrated, the techniques described herein can analyze a user's gaze direction, field of view, head pose, and the 3D anchor points of a rendered character to determine the viewpoint characteristics of the current view of a rendered virtual object (such as a character). Figures 10A to 10C As illustrated, character 1020 can be rendered at a higher level of detail (e.g., LOD-0 (approximately 64k scatter points)), a medium level of detail (e.g., LOD-1 (approximately 32k scatter points)), or a lower level of detail (e.g., LOD-2 (16k scatter points or less)).

[0132] In some implementations, the criterion is based on a threshold angular distance from the gaze center. For example, the criterion could be based on the character's position relative to the centerline of the field of view (e.g., associated with gaze cone 1010). In some implementations, the criterion is based on a threshold distance associated with the position of at least a portion of the object to be depicted within the 3D environment. For example, the criterion could be based on the character's position relative to a distance perceived by the user—e.g., near the periphery, but insufficient to trigger a periphery / gaze cone angle threshold, but still triggering an adjusted LOD level and merging between LOD levels if the character moves away from that distance.

[0133] In some implementations, merging scatter points in the outer regions is done using sparse UV mapping techniques. For example, this allows for a direct switch to a lower LOD in those regions.

[0134] In some implementations, the selection of scatter points for merging is based at least in part on a noise function. For example, the system can use a noise function to select scatter points for merging, thereby further reducing the perceptibility of LOD transitions.

[0135] In some implementations, the number of scatter points rendered to provide a representation of at least a portion of the object is adjusted (continuously) between a minimum and a maximum value based on at least one of the following: distance, gaze, or field of view associated with how the at least portion of the object should be depicted within the viewing experience. For example, the system provides a dynamic, context-aware process for rendering a 3D representation using 3D Gaussian scatter points, where advanced LOD transitions are performed based on distance, gaze, and / or field of view.

[0136] In various implementations, if the characteristics (e.g., viewpoint) do not change, the selected scatter rendering method can be reused (e.g., re-rendered) in subsequent frames. In some implementations, where the view of the representation is provided for a first frame of a plurality of frames for a first viewpoint, method 1100 further includes: determining that the second viewpoint is equivalent to the first viewpoint for a second viewpoint of a second frame of the plurality of frames; and (e.g., based on the selected scatter method from the first frame) reusing the view of the representation of at least a portion of the object from the first frame to the second frame. Additionally or alternatively, in some implementations, the view of the representation is provided for a first frame of a plurality of frames for a first viewpoint, and method 1100 further includes: determining that the second viewpoint is different from the first viewpoint for a second viewpoint of a second frame of the plurality of frames; selecting a different scatter rendering method (i.e., improving or reducing the quality of rendering) based on the context of the viewing experience associated with the second frame (e.g., a change in the viewer's perceptual distance, a change in another property of the context, etc.); and updating the view of the representation of the at least part of the object based on the different scatter rendering method.

[0137] In various embodiments, method 1100 further includes modifying the user representation data based on obtaining a set of sensor data acquired after the registration process. For example, the scatter parameter data may be modified based on live user data. For instance, the scatter parameter data may be obtained from a transmission device, such as from the registration process, and the modification of the registration scatter parameter data may be based on obtaining live sensors of the transmitter to determine the transmitter's live representation (e.g., a live view of a realistic character for a communication session). In some embodiments, the modification of the user representation data generates 3D Gaussian scatter points based on the image data of at least a portion of the user, where these Gaussian scatter points include texture, location, and scatter point shape. For example, a 3D Gaussian distribution in a 2D space with color / density (e.g., parameterized), where the face is represented as a grid, and the system determines multiple scatter points according to rays departing from the face / grid. In some embodiments, the scatter point generation technique generates the user representation via a machine learning model trained using training data obtained via one or more sensors in one or more environments. For example, machine learning models interpret image data and / or other sensor data captured during registration.

[0138] In various implementations, user representation data may be modified for the face instead of the body, for the body instead of the face, or for both the body and the face, and / or may be modified during the registration phase, on the transmitter-side device, and / or on the receiver-side device. In other words, the user representation (role) may be continuously updated to represent a live view of the current user's head / face and / or upper body movement, and may be modified at different stages during the communication session. In some implementations, a scatter rendering method associated with the user representation data is modified based on body pose data acquired during the registration process, during a communication session with another device, or a combination thereof. In some implementations, the device is the viewer's device, and the scatter rendering method associated with the user representation data is modified based on an additional set of sensor data acquired during a communication session with the transmitter's device associated with the user representation. Alternatively, in some implementations, a scatter rendering method associated with the user representation data is generated and updated during the registration process based on images of the user's face captured while the user is expressing multiple different facial expressions (e.g., registration images of the face when the user is smiling, raising eyebrows, puffing out cheeks, etc.).

[0139] In some implementations, a second set of sensor data acquired after a registration process performed by a device (e.g., a viewer's device) includes a sequence of frames for Gaussian UV mapping and corresponding marker points. The sequence of frames for Gaussian UV mapping and corresponding marker points may be obtained from a second device (e.g., a transmitter's device) during a communication session. The device (e.g., the viewer's device) uses one or more sputtering techniques described herein, based on the sequence of frames for Gaussian UV mapping and corresponding marker points, to render an animated depiction of the user (e.g., the transmitter). For example, the set of Gaussian UV mappings and corresponding marker points transmitted during a communication session with the second device may be used to render a view of the user's (transmitter's) face (and upper body). Additionally or alternatively, consecutive frames of facial data (the appearance of the user's face at different points in time) and body tracking data may be transmitted and used to display a live 3D video-like depiction of the user (e.g., a "live" character).

[0140] In some implementations, the second user representation is based on second image data obtained via a second set of sensors in a second physical environment with second lighting conditions (e.g., lighting conditions different from the first physical environment). For example, during the registration process, user representation data is acquired in a specific environment (also referred to herein as the "registration environment") that includes some lighting condition information (e.g., brightness values ​​and other lighting attributes), which may be lighting data different from live lighting data (e.g., two different physical environments between registration and during character generation based on "live" sensor data). In some implementations, the view of the representation of at least a portion of the object, provided based on a selected scatter rendering method, includes displaying the representation in a 3D environment (such as an extended reality (XR) environment).

[0141] In some implementations, method 1100 further includes adjusting the user representation to modify the view of the user representation by means of at least one color attribute of a plurality of color attributes of the environment, at least one light attribute of a plurality of light attributes of the environment, or a combination thereof. For example, adjusting the color or lighting on the user representation (such as hair, face, and clothing) based on the color and / or light associated with the viewer's environment and / or with the transmitter's environment. In other words, the lighting and / or color of the 3D representation (e.g., a character) can be changed to match the lighting and / or color of the viewer's environment (e.g., a pale red light emanating from the viewer's room will be reflected in the 3D representation). Alternatively, the lighting and / or color of the 3D representation (e.g., a character) can be changed to match the lighting and / or color of the transmitter's environment (e.g., even if the registration data does not reflect a pale green light emanating from the transmitter's room, that pale green light will still be reflected to the viewer in the 3D representation).

[0142] In some implementations, at least a portion of the user representation data obtained during the registration process is based on images captured of the user's face in different poses and / or while the user is expressing multiple different facial expressions. For example, these images are registration images of the user's face when facing the camera, to the left of the camera, and to the right of the camera, and / or when the user is smiling, raising their eyebrows, puffing out their cheeks, etc. For example, as by... Figure 6 The registration process 610, as illustrated, allows for the acquisition of registration images 612 of the face from different viewpoints while the user is smiling, raising their eyebrows, puffing out their cheeks, etc. In some implementations, a first set of sensor data corresponds to only a first region of the user (e.g., a portion not obstructed by the device, such as an HMD), and a second set of sensor data corresponds to a second region, which includes a third region different from the first region. For example, the second region may include portions of the portion of the HMD that is obscured when the user wears it. For instance, during the registration process, a larger portion of the user may be captured by image data compared to a live communication session with a user wearing an HMD (e.g., without an HMD).

[0143] In some specific implementations, such as Figure 2 As illustrated, rendering occurs during a communication session in which a second device (e.g., device 265) captures sensor data (e.g., image data of a portion of user 260 and environment 250) and provides a sequence of frame-specific 3D representations corresponding to multiple moments within a time period based on the sensor data. For example, second device 265 provides / transmits the sequence of frame-specific 3D representations to device 210, and device 210 generates a combined 3D representation to display a live 3D video-like facial depiction of user 260 (e.g., a realistic moving character) (e.g., representation 240 of user 260). Alternatively, in some embodiments, the second device provides a 3D representation of the user (e.g., representation 140 of user 260) (e.g., a realistic moving character) during the communication session. For example, the combined representation is determined at device 265 and transmitted to device 210. In some embodiments, a view of the combined 3D representation is displayed on the device (e.g., device 210) in real time relative to multiple moments within a time period. For example, the user's depiction is displayed in real time and based on live lighting data (e.g., a character shown to the second user on the display of the second user's second device).

[0144] In some implementations, the user-represented view may include sufficient data to achieve a stereoscopic view for the user (e.g., left-eye / right-eye view) that allows for a certain depth perception of the face. In one implementation, the depiction of the face includes a 3D model of the face, and generating representations from the left-eye and right-eye positions to provide a stereoscopic view of the face.

[0145] In some implementations, certain parts of the face (such as the eyes and mouth) that may be important for conveying a realistic or truthful appearance may be generated (e.g., based on marker points) in a different manner than other parts of the face. For example, parts of the face that may be important for conveying a realistic or truthful appearance may be based on current camera data, while other parts of the face may be based on previously acquired (e.g., registered) facial data.

[0146] In some implementations, a facial representation is generated using the texture, color, and / or geometry of various facial features. This facial representation identifies with confidence how well the generation technique, per data frame, accurately corresponds to the true texture, color, and / or geometry of those facial features based on depth and appearance values. In some implementations, the depiction is a 3D character. For example, the representation might represent a user (e.g., Figure 1 3D model of user 110.

[0147] In some implementations, a first set and / or a second set of sensor data (e.g., live data, such as video content including light intensity (RGB) and depth data) are associated with points in time, such as images from inward-facing / downward-facing sensors when a user is wearing the HMD (associated with frames). In some implementations, the sensor data includes depth data (e.g., infrared, time-of-flight, etc.) and light intensity image data acquired during the scanning process.

[0148] In some specific implementations, the first set of sensor data obtained during the registration process may include registration sensor data (e.g., data obtained from the device in multiple configurations corresponding to the user's facial features (e.g., texture, muscle activation, shape, depth, etc.)). Figure 6The registration data is 612. In some embodiments, the first set of data includes unobstructed image data of the user's face. For example, images of the face may be captured when the user smiles, raises eyebrows, puffs out cheeks, etc. In some embodiments, registration data may be obtained by the user removing the device (e.g., HMD) and capturing images without the device obscuring the face, or by using another device (e.g., mobile device) without the device (e.g., HMD) obscuring the face. In some embodiments, registration data (e.g., the first set of data) is acquired from light intensity images (e.g., RGB images). Registration data may include most (if not all) of the texture, muscle activation, etc., of the user's face. In some embodiments, registration data may be captured when the user is given different instructions for capturing different poses of the user's face. For example, a user interface guide may instruct the user to "raise your eyebrows," "smile," "frown," etc., to provide the system with a range of facial features for the registration process.

[0149] In some implementations, method 1100 may be repeated for each frame captured during each moment / frame of a live communication session or other experience. For example, for each iteration, while the user is using the device (e.g., a wearable HMD), method 1100 may involve continuously acquiring live sensor data (e.g., face tracking data and body tracking, etc.), and for each frame, updating a selected subset of scatter parameter data based on updated viewing characteristics of that frame (e.g., FoV, gaze, occlusion, etc.), and updating the displayed portion of the user representation based on updated Gaussian data. For example, for each new frame, the system may update the display of a 3D character based on the new data.

[0150] Figure 12 This is a flowchart illustrating an exemplary method 1200. In some specific implementations, the device (e.g., Figure 1 Device 105) performs method 1200 according to some specific embodiments to render scatter points to provide a view of the representation by merging a subset of representation data. In some specific embodiments, the techniques of method 1200 are performed on a mobile device, desktop computer, laptop computer, HMD, or server device. In some specific embodiments, method 1200 is performed on a processing logic component (including hardware, firmware, software, or a combination thereof). In some specific embodiments, method 1200 is performed on a processor executing code stored in a non-transitory computer-readable medium (e.g., memory). In some specific embodiments, method 1200 is implemented at a processor of a device (such as a viewing device) that renders a user representation (e.g., Figure 2 Device 210 renders a 3D representation 240 (character) of user 260 based on data obtained from device 265.

[0151] An exemplary implementation of Method 1200 provides a dynamic, context-based selection and transition process between Levels of Detail (LODs) used for 3DGS rendering, focusing on minimizing perceptible transition artifacts while simultaneously optimizing distance, gaze, and field of view. For example, Method 1200 can provide dynamic, context-aware rendering of the 3D representation using Gaussian scatter points, where advanced LOD transitions are based on whether the character is inside or outside the user's field of view (e.g., the fovea region or the periphery) (which can be determined based on the user's viewpoint, gaze, and / or the character's position relative to the field of view). This method achieves smooth, continuous LOD transitions, minimizes visual artifacts, and optimizes rendering efficiency and fidelity in the region of interest. For example, instead of abrupt LOD switching (e.g., from 64k scatter points to 16k scatter points), the system introduces a smooth, continuous reduction of scatter points. As the viewer moves away from the character, the scatter points gradually merge, creating a transition zone that avoids a noticeable "pop-up" effect. This merging can be adjusted and incorporated into noise-based selections for further optimization. This method minimizes visual artifacts during LOD changes, optimizes rendering efficiency, and maintains high visual fidelity in regions of interest.

[0152] At box 1210, method 1200 obtains representation data representing at least a portion of an object at the device's processor, wherein the representation data includes Gaussian data for rendering multiple scatter points for a first set of 3D points to represent the 3D appearance of the object. For example, scatter point generation techniques (e.g., models / parameters specifying the total number of scatter points, scatter point size attributes, etc.) may include generating multiple sets of scatter point parameter data (e.g., sets of scatter point parameter data at different resolutions / sizes ranging from 16k to 64k scatter points, or different LODs based on merged grouping of static scatter points, etc.). In some specific implementations, the representation data may be rendered using Gaussian sputtering, a technique in which individual 3D points are represented as a Gaussian distribution with color values ​​that change according to the viewpoint (e.g., scatter point parameter data, also referred to as "scatter points"). This technique uses spherical harmonics to model this view-dependent color variation and enables real-time rendering of high-quality, realistic scenes from sparse image sets. For example, each point has a color calculated based on its position relative to the camera, thus allowing for realistic shading effects across different viewpoints.

[0153] At box 1220, method 1200 determines viewpoint characteristics based on the viewpoint and the location of the at least portion of the object to be depicted within the 3D environment. For example, as Figures 10A to 10C As illustrated, the techniques described herein can analyze the user's gaze direction, field of view, head pose, and 3D anchor points of the rendered character to determine the viewpoint characteristics of the current view of the rendered virtual object.

[0154] At box 1230, method 1200 generates modified representation data for the region by merging a subset of the representation data to represent a second set of 3D points, based on the viewpoint characteristics meeting the criteria. This second set of 3D points may be smaller in number than the first set of 3D points. For example, one or more scatter points in transition zones LOD-1 to LOD-0 are dynamically merged based on distance, such that the number of rendered scatter points decreases continuously with increasing distance before switching to a lower LOD representation. Alternatively, the merging of scatter points is guided based on the gaze direction, such that regions within a threshold angle of the gaze center retain higher detail, and regions outside the threshold are more aggressively merged.

[0155] At box 1240, method 1200 renders scatter plots based on the modified representation data to provide a view of a representation of at least a portion of the object. For example... Figures 10A to 10C As illustrated, character 1020 can be rendered at a higher level of detail (e.g., LOD-0 (approximately 64k scatter points)), an intermediate level of detail (e.g., LOD-1 (approximately 32k scatter points)), or a lower level of detail (e.g., LOD-2 (16k scatter points or less)). When transitioning between different LODs, the system provides a view of a subset of scatter points to minimize visual artifacts during LOD changes, optimize rendering efficiency, and maintain high visual fidelity in the region of interest.

[0156] In some implementations, during the merging of the subset of representation data to represent a second set of 3D points, the number of rendered scatter points decreases as the distance associated with the location of the at least part of the object to be depicted increases. For example, the number of scatter points is gradually reduced (or increased) before switching to a lower LOD representation. In some implementations, merging the subset of representation data to represent a second set of 3D points is performed in groups, where each group corresponds to a range of distances associated with the location of the at least part of the object to be depicted. For example, as... Figure 8 and Figure 9 As illustrated, transitions between LODs can be performed gradually without abruptly changing the number of rendered scatter points.

[0157] In some implementations, intermediate LOD states (e.g., a combination of LOD-0 and LOD-1) can be utilized for transitional merging. For example, the exact proportion of each LOD state varies with distance; more LOD-0 indicates closer proximity to the camera, and more LOD-1 indicates further distance. The runtime algorithm that generates LOD states operates by merging scatter points into larger scatter points. In some implementations, the algorithm does not necessarily need to output a fixed number of merged scatter points; instead, a variable number of scatter points can be randomly selected for merging, and this variable number can be controlled via a knob, such as the distance or angle to the camera.

[0158] In some specific implementations, merging the subset of the representation data to represent a second set of 3D points is based on determining whether the viewpoint characteristics meet the criterion, and the viewpoint characteristics are determined based on the viewpoint and the location of at least a portion of the object to be depicted within the 3D environment. For example, as Figures 10A to 10C As illustrated, the techniques described herein can analyze a user's gaze direction, field of view, head pose, and the 3D anchor points of a rendered character to determine the viewpoint characteristics of the current view of a rendered virtual object (such as a character). Figures 10A to 10C As illustrated, character 1020 can be rendered at a higher level of detail (e.g., LOD-0 (approximately 64k scatter points)), a medium level of detail (e.g., LOD-1 (approximately 32k scatter points)), or a lower level of detail (e.g., LOD-2 (16k scatter points or less)).

[0159] In some implementations, the criterion is based on a threshold angular distance from the gaze center. For example, the criterion could be based on the character's position relative to the centerline of the field of view (e.g., associated with gaze cone 1010). In some implementations, the criterion is based on a threshold distance associated with the position of at least a portion of the object to be depicted within the 3D environment. For example, the criterion could be based on the character's position relative to a distance perceived by the user—e.g., near the periphery, but insufficient to trigger a periphery / gaze cone angle threshold, but still triggering an adjusted LOD level and merging between LOD levels if the character moves away from that distance.

[0160] In some implementations, merging scatter points in peripheral regions is performed using sparse UV mapping techniques. For example, this allows for a direct switch to a lower LOD in those regions. In some implementations, the selection of scatter points for merging is based at least in part on a noise function. For example, the system can use a noise function to select scatter points for merging, thereby further reducing the perceptibility of LOD transitions.

[0161] In some implementations, the number of scatter points rendered to provide a representation of at least a portion of the object is adjusted (continuously) between a minimum and a maximum value based on at least one of the following: distance, gaze, or field of view associated with how the at least portion of the object should be depicted within the viewing experience. For example, the system provides a dynamic, context-aware process for rendering a 3D representation using 3D Gaussian scatter points, where advanced LOD transitions are performed based on distance, gaze, and / or field of view.

[0162] Figure 13This is a block diagram of example device 1300. Device 1300 illustrates an exemplary device configuration for devices described herein (e.g., devices 105, 210, 265, 745, etc.). Although certain specific features are illustrated, those skilled in the art will recognize from this disclosure that various other features are not illustrated for the sake of brevity and so as not to obscure further relevant aspects of the specific embodiments disclosed herein. Therefore, as a non-limiting example, in some specific implementations, device 1300 includes one or more processing units 1302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, and / or processing cores, etc.), one or more input / output (I / O) devices and sensors 1306, one or more communication interfaces 1308 (e.g., USB, Firewire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, Bluetooth, ZigBee, SPI, I2C, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 1310, one or more displays 1312, one or more internal and / or external image sensor systems 1314, memory 1320, and one or more communication buses 1304 for interconnecting these components and various other components.

[0163] In some embodiments, one or more communication buses 1304 include circuitry that interconnects system components and controls communication between system components. In some embodiments, one or more I / O devices and sensors 1306 include at least one of the following: an inertial measurement unit (IMU), an accelerometer, a magnetometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, and / or one or more depth sensors (e.g., structured light or time-of-flight, etc.).

[0164] In some embodiments, one or more displays 1312 are configured to present a view of a physical or graphical environment to a user. In some embodiments, one or more displays 1312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conducting electron emission display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical systems (MEMS), and / or similar display types. In some embodiments, one or more displays 1312 correspond to waveguide displays such as diffraction, reflection, polarization, and holography. In one example, device 1300 includes a single display. In another example, device 1300 includes displays for each of the user's eyes.

[0165] In some embodiments, one or more image sensor systems 1314 are configured to acquire image data corresponding to at least a portion of the physical environment 102. For example, one or more image sensor systems 1314 may include one or more RGB cameras (e.g., having a complementary metal-oxide-semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), monochrome cameras, IR cameras, depth cameras, and / or event-based cameras, etc. In various embodiments, one or more image sensor systems 1314 may also include an illumination source emitting light, such as a flash. In various embodiments, one or more image sensor systems 1314 may also include an on-camera image signal processor (ISP) configured to perform multiple processing operations on the image data.

[0166] Memory 1320 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 1320 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 1320 optionally includes one or more storage devices remotely located to one or more processing units 1302. Memory 1320 includes non-transitory computer-readable storage media.

[0167] In some embodiments, memory 1320 or a non-transitory computer-readable storage medium of memory 1320 stores an optional operating system 1330 and one or more instruction sets 1340. Operating system 1330 includes procedures for handling various basic system services and for performing hardware-related tasks. In some embodiments, instruction set 1340 includes executable software defined by binary information stored in charge. In some embodiments, instruction set 1340 is software executable by one or more processing units 1302 to implement one or more of the techniques described herein.

[0168] Instruction set 1340 includes registration instruction set 1342, scenario understanding instruction set 1344, representation instruction set 1346, and communication session instruction set 1348. Instruction set 1340 can be represented as a single software executable file or multiple software executable files.

[0169] In some implementations, the registration instruction set 1342 can be executed by the processing unit 1302 to generate registration data from image data. The registration instruction set 1342 can be configured to provide instructions to the user to obtain image information to generate a registration avatar (e.g., registration image data 612) and to determine whether additional image information is needed to generate an accurate registration avatar to be used by the character display process. For these purposes, in various implementations, the instructions include instructions and / or logic for those instructions, as well as heuristics and metadata for those heuristics.

[0170] In some implementations, the scene understanding instruction set 1344 can be executed by the processing unit 1302 to determine the context of the experience and / or environment, or the context of the user's viewing environment (e.g., user understanding). The scene understanding instruction set 1344 can create scene understanding to determine objects or people in the content or environment, where the user is, what the user is watching, etc., using one or more of the techniques discussed herein (e.g., object detection, facial recognition, etc.) or other techniques that may be appropriate in other ways. For these purposes, in various implementations, the instructions include instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.

[0171] In some embodiments, the representation instruction set 1346 can be executed by the processing unit 1302 to generate a representation of an object, such as a user representation (e.g., rendered via Gaussian sputtering), using one or more of the techniques discussed herein or other techniques that may be appropriate in other ways. In some embodiments, the representation instruction set 1346 can be executed by the processing unit 1302 to analyze and select a scatter rendering method (e.g., rendering mode) for providing a view including the representation, based on a determined context as discussed herein or other appropriate for the viewing experience. For these purposes, in various embodiments, the instruction includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.

[0172] In some specific implementations, the communication session instruction set 1348 may be executed by the processing unit 1302 to facilitate two or more electronic devices (e.g., such as) using one or more of the techniques discussed herein or otherwise suitable. Figure 2 The communication session between the illustrated devices 210 and 265. For these purposes, in various specific implementations, the instruction includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.

[0173] Although instruction set 1340 is shown as residing on a single device, it should be understood that in other specific implementations, any combination of elements may reside on separate computing devices. Furthermore, Figure 13 This is intended more as a functional description of various features present in a particular implementation than as a structural diagram of the specific implementation described herein. As will be appreciated by those skilled in the art, the items shown individually can be combined, and some items can be separated. The actual number of instruction sets and how features are allocated therein will vary depending on the specific implementation and may depend in part on the specific combination of hardware, software, and / or firmware chosen for that particular implementation.

[0174] Figure 13 A block diagram illustrating an exemplary head-mounted device 1300 according to some specific embodiments is shown. The head-mounted device 1300 includes a housing 1301 (or shell) housing various components of the head-mounted device 1300. The housing 1301 includes (or is coupled to) eye pads (not shown) disposed at a proximal end of the housing 1301 (relative to the user 25). In various specific embodiments, the eye pads are plastic or rubber components that comfortably and snugly hold the head-mounted device 1300 in a proper position on the face of the user 25 (e.g., around the eyes 35 of the user 25).

[0175] The housing 1301 houses a display 1310 for displaying images, thereby emitting light toward or onto the eyes of the user 25. In various embodiments, the display 1310 emits light through an eyepiece having one or more optical elements 1305 that refract the light emitted by the display 1310, so that the display appears to the user 25 at a virtual distance greater than the actual distance from the eye to the display 1310. For example, the optical elements 1305 may include one or more lenses, waveguides, and other diffractive optical elements (DOEs). In order for the user 25 to focus on the display 1310, in various embodiments, the virtual distance is at least greater than the minimum focal length of the eye (e.g., 7 cm). Furthermore, to provide a better user experience, in various embodiments, the virtual distance is greater than 1 meter.

[0176] The housing 1301 also houses a tracking system including one or more light sources 1322, cameras 1324, 1332, 1334, and a controller 1380. One or more light sources 1322 emit light onto the eyes of user 25, which is reflected as a light pattern (e.g., a flash) detectable by camera 1324. Based on this light pattern, controller 1380 can determine the eye-tracking characteristics of user 25. For example, controller 1380 can determine the gaze direction and / or blinking state (open or closed eyes) of user 25. Also, controller 1380 can determine the pupil center, pupil size, or point of focus. Thus, in various embodiments, light is emitted by one or more light sources 1322, reflected from the eyes of user 25, and detected by camera 1324. In various embodiments, light from the eyes of user 25 is reflected from a hot mirror or passes through an eyepiece before reaching camera 1324.

[0177] Display 1310 emits light within a first wavelength range, and one or more light sources 1322 emit light within a second wavelength range. Similarly, camera 1324 detects light within the second wavelength range. In various specific embodiments, the first wavelength range is the visible wavelength range (e.g., a wavelength range of approximately 400 nm to 700 nm within the visible spectrum), and the second wavelength range is the near-infrared wavelength range (e.g., a wavelength range of approximately 700 nm to 1400 nm within the near-infrared spectrum).

[0178] In various implementations, eye tracking (or specifically, a defined gaze direction) is used to enable user interaction (e.g., user 25 selects an option by looking at display 1310), to provide foveated rendering (e.g., rendering a higher resolution in an area of ​​display 1310 that user 25 is viewing and a lower resolution elsewhere on display 1310), or to correct distortion (e.g., for an image to be presented on display 1310). In various implementations, one or more light sources 1322 emit light toward user 25's eye 35, which is reflected in the form of multiple flashes.

[0179] In various embodiments, camera 1324 is a frame / shutter-based camera that generates images of user 25's eye 35 at specific time points or multiple time points at a certain frame rate. Each image includes a matrix of pixel values ​​corresponding to pixels in the image, where each pixel corresponds to a fixed point in the camera's light sensor matrix. In specific embodiments, each image is used to measure or track pupil dilation by measuring changes in pixel intensity associated with one or both of the user's pupils.

[0180] In various specific implementations, camera 1324 is an event camera that includes multiple light sensors (e.g., a light sensor matrix) at multiple corresponding fixed points, which generates an event message indicating a specific fixed point of a specific light sensor in response to a specific light sensor detecting a change in light intensity.

[0181] In various specific implementations, cameras 1332 and 1334 are frame / shutter-based cameras capable of generating images of the user 25's face at specific or multiple time points at a certain frame rate. For example, camera 1332 captures an image of the user's face below the eyes, and camera 1334 captures an image of the user's face above the eyes. The images captured by cameras 1332 and 1334 may include light intensity images (e.g., RGB) and / or depth image data (e.g., time-of-flight, infrared, etc.).

[0182] It should be understood that the specific embodiments described above are cited by way of example, and the invention is not limited to what has been specifically shown and described above. Rather, the scope includes both combinations and sub-combinations of the various features described above, as well as variations and modifications of the various features that would occur to those skilled in the art upon reading the foregoing description and that are not disclosed in the prior art.

[0183] As described above, one aspect of the present invention is the collection and use of physiological data to improve the user experience with electronic devices in interacting with electronic content. This disclosure envisions that, in some cases, the collected data may include personal information data that uniquely identifies a particular person or can be used to identify the interests, characteristics, or tendencies of a particular person. Such personal information data may include physiological data, demographic data, location-based data, telephone numbers, email addresses, home addresses, device characteristics of personal devices, or any other personal information.

[0184] This disclosure recognizes that the use of such personal information data in the present invention can benefit users. For example, personal information data can be used to improve the interactivity and controllability of electronic devices. Therefore, the use of such personal information data enables planned control of electronic devices. Furthermore, this disclosure also anticipates other uses of personal information data that benefit users.

[0185] This disclosure further envisions that entities responsible for the collection, analysis, disclosure, transmission, storage, or other use of such personal information and / or physiological data will comply with established privacy policies and / or privacy practices. Specifically, such entities should implement and maintain privacy policies and measures that are recognized as meeting or exceeding industry or governmental requirements for maintaining the privacy and security of personal information data. For example, personal information from users should be collected for legitimate and reasonable purposes of the entity and not shared or sold outside of these legitimate purposes. Furthermore, such collection should only be conducted after receiving informed consent from users. Additionally, such entities should take any necessary steps to safeguard and protect access to such personal information data and ensure that others with access to such personal information data comply with their privacy policies and procedures. Furthermore, such entities may subject themselves to third-party assessments to demonstrate their compliance with widely accepted privacy policies and privacy practices.

[0186] Regardless of the foregoing, this disclosure also contemplates specific implementations allowing users to selectively block the use or access to personal information data. That is, this disclosure contemplates providing hardware or software components to prevent or block access to such personal information data. For example, with regard to a content delivery service tailored to a user, the technology of this invention can be configured to allow a user to choose to "join" or "opt out" of the collection of personal information data during service registration. In another example, a user may choose not to provide personal information data for a target content delivery service. In yet another example, a user may choose not to provide personal information but allow the transmission of anonymous information for improving device functionality.

[0187] Therefore, while this disclosure broadly covers the use of personal information data to implement one or more of the various disclosed embodiments, it is also contemplated that various embodiments can be implemented without access to such personal information data. That is, various embodiments of the present invention will not become inoperable due to the absence of all or part of such personal information data. For example, preferences or settings can be inferred based on non-personal information data or an absolute minimum amount of personal information, such as content requested by a device associated with a user, other non-personal information available to the content delivery service, or publicly available information, thereby selecting content and delivering it to the user.

[0188] In some implementations, data is stored using a public / private key system that allows only the data owner to decrypt the stored data. In other implementations, data may be stored anonymously (e.g., without identification and / or without personal information about the user, such as legal name, username, time, or location data). This prevents other users, hackers, or third parties from identifying the user associated with the stored data. In some implementations, a user can access their stored data from a user device different from the device used to upload the stored data. In these cases, the user may need to provide login credentials to access their stored data.

[0189] This document sets forth numerous specific details to provide a comprehensive understanding of the claimed subject matter. However, those skilled in the art will understand that the claimed subject matter can be practiced without these specific details. In other instances, methods, apparatus, or systems known to a person of ordinary skill have not been described in detail so as not to obscure the claimed subject matter.

[0190] Unless otherwise specifically stated, it should be understood that throughout this specification, discussions using terms such as “processing,” “calculating,” “calculating,” “determining,” and “identifying” refer to the actions or processes of computing devices, such as one or more computers or similar electronic computing devices, which manipulate or convert data representing physical electronic or magnetic quantities within the memory, registers, or other information storage, transmission, or display devices of a computing platform.

[0191] The one or more systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement of components that provide results conditioned on one or more inputs. Suitable computing devices include computer systems based on multi-purpose microprocessors that access stored software that programs or configures the computing system from a general-purpose computing device to a special-purpose computing device that implements one or more specific embodiments of the subject matter of this invention. The teachings contained herein may be implemented in the software used for programming or configuring the computing device using any suitable programming, scripting, or other type of language or combination of languages.

[0192] Specific implementations of the methods disclosed herein can be performed in the operation of such computing devices. The order of the boxes presented in the above examples can be changed; for example, the boxes can be reordered, grouped, or divided into sub-boxes. Some boxes or procedures can be executed in parallel.

[0193] The use of "applies to" or "configured to" in this document implies open and inclusive language, which does not exclude applicability to or configuration to devices performing additional tasks or steps. Furthermore, the use of "based on" implies openness and inclusivity, as processes, steps, calculations, or other actions "based on" one or more of the stated conditions or values ​​may in practice be based on additional conditions or values ​​beyond those stated. The headings, lists, and numbering included herein are for illustrative purposes only and are not intended to be restrictive.

[0194] It will also be understood that while terms such as "first," "second," etc., may be used in this document to describe various objects, these objects should not be limited by these terms. These terms are merely used to distinguish one object from another. For example, a first node can be called a second node, and similarly, a second node can be called a first node, changing the meaning of the description, provided that all occurrences of "first node" are consistently renamed and all occurrences of "second node" are consistently renamed. Both a first node and a second node are nodes, but they are not the same node.

[0195] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the claims. As used in the description of these embodiments and in the appended claims, the singular forms “a,” “an,” and “the” are intended to also cover the plural forms unless the context clearly indicates otherwise. It will also be understood that the term “or” as used herein refers to and covers any and all possible combinations of one or more of the associated listed items. It will also be understood that the terms “comprising” or “including” as used in this specification specify the presence of the stated features, integers, steps, operations, objects, or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, objects, components, or groups thereof.

[0196] As used herein, the term "if" can be interpreted as meaning "when the prerequisite is true" or "when the prerequisite is true" or "in response to determination" or "according to determination" or "in response to detection" that the prerequisite is true, depending on the context. Similarly, the phrase "if it is determined [the prerequisite is true]" or "if [the prerequisite is true]" or "when [the prerequisite is true]" can be interpreted as meaning "when it is determined that the prerequisite is true" or "in response to determination" or "according to determination" that the prerequisite is true or "when it is detected that the prerequisite is true" or "in response to detection" that the prerequisite is true, depending on the context.

[0197] The foregoing description and summary of the invention should be understood as exemplary and illustrative in every respect, and not restrictive, and the scope of the invention disclosed herein is determined not only by the detailed description of the exemplary specific implementations, but also by the full extent permitted by patent law.

[0198] It should be understood that the specific embodiments shown and described herein are merely illustrative of the principles of the invention, and those skilled in the art can make various modifications without departing from the scope and spirit of the invention.

Claims

1. A method, the method comprising: At the device's processor: Obtain representation data for at least a portion of the representation object, wherein the representation data includes data for a plurality of sets of scatter points generated based on a scatter point generation technique, and the data for each of the plurality of sets of scatter points includes three-dimensional (3D) Gaussian data; Determine the context of the viewing experience; A scatter rendering method is selected based on the determined context of the viewing experience to provide a view that includes at least a portion of the object. as well as Provide a view of at least a portion of the object based on the selected scatter rendering method, wherein providing the view includes rendering a subset of the scatter points based on the 3D Gaussian data.

2. The method of claim 1, wherein the at least portion of the object comprises a user’s facial portion and additional portions.

3. The method of claim 1, wherein the scatter point generation technique includes a scatter point criterion associated with at least one of the following: (i) the resolution of each scatter point associated with each of the plurality of sets of scatter points, (ii) the number of scatter points associated with each of the plurality of sets of scatter points, and (iii) the size of each scatter point associated with each of the plurality of sets of scatter points.

4. The method of claim 1, wherein selecting the scatter rendering method includes selecting between a first method for rendering a first set of the plurality of sets of scatter points and a second method for rendering a second set of the plurality of sets of scatter points, wherein the first set of the plurality of sets of scatter points has more scatter points than the second set of the plurality of sets of scatter points.

5. The method of claim 1, wherein the scatter point generation technique includes a scatter point criterion associated with multiple subsets of scatter points merged based on the 3D Gaussian data.

6. The method of claim 5, wherein the plurality of subsets of the scatter points are selected based on determining one or more static regions associated with representing the at least part of the object.

7. The method according to claim 1, wherein selecting the scatter rendering method comprises: The method of reducing the rendering level of detail is selected by switching to rendering a lower level of detail (LOD) representation, which includes merging a subset of scatter points in the transition region based on the location of the at least part of the object to be depicted in the 3D environment.

8. The method according to claim 7, wherein, During the merging of the subset of scatter points in the transition zone, the number of scatter points rendered decreases as the distance associated with the location of the at least part of the object to be depicted increases.

9. The method of claim 7, wherein merging the subsets of scatter points is performed in groups, wherein each group corresponds to a range of distances associated with the location of the at least portion of the object to be depicted.

10. The method of claim 7, wherein merging scatter points is based on determining whether viewpoint characteristics meet a criterion, wherein the viewpoint characteristics are determined based on the viewpoint and the location of the at least portion of the object to be depicted within the 3D environment.

11. The method of claim 10, wherein the criterion is based on a threshold angular distance from the gaze direction.

12. The method of claim 10, wherein the criterion is based on a threshold distance associated with the location of the at least part of the object to be depicted within the 3D environment.

13. The method of claim 7, wherein merging scatter points in the peripheral region is performed using sparse UV mapping technology.

14. The method of claim 7, wherein the selection of scatter points for merging is based at least in part on a noise function.

15. The method of claim 1, wherein the number of scatter points of the view rendered to provide a representation of at least a portion of the object is adjusted between a minimum and a maximum value according to at least one of the following: distance, gaze, or field of view associated with how the at least portion of the object is to be depicted within the viewing experience.

16. The method of claim 1, wherein the context of the viewing experience includes the viewpoint of the represented view, and the scatter rendering method for providing the view is selected based on a threshold distance associated with the viewpoint.

17. The method of claim 1, wherein the context of the viewing experience includes recording at least one frame of content associated with the viewpoint of the represented view, and the scatter rendering method for providing the view is selected based on the at least one frame of the recorded content.

18. The method of claim 1, wherein the represented view is provided for a first frame of a plurality of frames for a first viewpoint, the method further comprising: For the second viewpoint used in the second frame of the plurality of frames: In response to determining that the second viewpoint is equivalent to the first viewpoint, the view represented by the at least a portion of the object is reused from the first frame for the second frame; as well as In response to determining that the second viewpoint is different from the first viewpoint, Selecting an additional set from the plurality of sets of scatter points based on the context of the viewing experience associated with the second frame, and The view of the representation of at least a portion of the object is updated based on an additional set selected from the plurality of sets of scatter points.

19. An apparatus comprising: Non-transitory computer-readable storage medium; and One or more processors coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium includes program instructions that, when executed on the one or more processors, cause the one or more processors to perform operations, the operations including: Obtain representation data for at least a portion of the representation object, wherein the representation data includes data for a plurality of sets of scatter points generated based on a scatter point generation technique, and the data for each of the plurality of sets of scatter points includes three-dimensional (3D) Gaussian data; Determine the context of the viewing experience; Based on the determined context of the viewing experience, a scatter rendering method is selected to provide a view that includes a representation of at least a portion of the object; and Provide a view of at least a portion of the object based on the selected scatter rendering method, wherein providing the view includes rendering a subset of the scatter points based on the 3D Gaussian data.

20. A non-transitory computer-readable storage medium storing program instructions executable on a device to perform operations including: Obtain representation data for at least a portion of the representation object, wherein the representation data includes data for a plurality of sets of scatter points generated based on a scatter point generation technique, and the data for each of the plurality of sets of scatter points includes three-dimensional (3D) Gaussian data; Determine the context of the viewing experience; A scatter rendering method is selected based on the determined context of the viewing experience to provide a view that includes at least a portion of the object. as well as Provide a view of at least a portion of the object based on the selected scatter rendering method, wherein providing the view includes rendering a subset of the scatter points based on the 3D Gaussian data.