Visual encoding and decoding of 3D gaussian splats
By projecting 3D Gaussian splat scenes into 2D representations and encoding with image and video codecs, along with metadata, the high data volume issue is addressed, achieving efficient compression and quality preservation in 3D scene rendering.
Patent Information
- Application Number
- PCT/EP2025/050444
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-13
- Filing Date
- 2025-01-09
- Publication Date
- 2025-08-21
AI Technical Summary
3D Gaussian Splatting techniques require large memory and bandwidth due to high data volumes, and existing compression methods compromise on quality by reducing spherical harmonics, which affects the representation's detail and specular effects.
Project 3D Gaussian splat scenes into 2D representations, storing different parameters with associated metadata, and encode these into bitstreams using various image and video codecs, along with metadata encoded as atlas metadata or SEI messages, to reduce data volume while maintaining quality.
Efficiently compresses 3D Gaussian splat scenes, reducing bandwidth requirements while preserving detailed and specular representations, enabling effective streaming and rendering of complex 3D scenes.
Smart Images

Figure EP2025050444_21082025_PF_FP_ABST
Abstract
Description
Visual Encoding and Decoding of 3D Gaussian SplatsTECHNICAL FIELD
[0001] Examples of embodiments herein relate generally to video encoding and decoding and, more specifically, relate to video encoding and decoding of 3D (three- dimensional) Gaussian Splats.BACKGROUND
[0002] A technique in 3D (three-dimensional) modeling and rendering that is becoming more prevalent involves 3D Gaussian Splatting. 3D Gaussian Splatting is a technique in computer graphics that creates 3D scenes by projecting points, or “splats”, from a point cloud onto a 3D space, using Gaussian functions for each splat. The term “splatting” is based on the sound a snowball makes as it hits and spreads across a window. This technique supports complex view-dependent visual effects and surpasses the quality of traditional point cloud rendering by producing dynamic and lifelike visualizations.
[0003] The idea behind Gaussian splatting originated in a 1991 doctorate thesis by Lee Alan Westover at the University of North Carolina at Chapel Hill. The hardware at the time could not efficiently run the algorithms, so this technique was not widely used until recently. 3D Gaussian Splatting still has issues, including large memory requirements because of high amounts of data that are required.BRIEF SUMMARY
[0004] This section is intended to include examples and is not intended to be limiting.
[0005] In an exemplary embodiment, a method is disclosed that includes projecting, by an encoder, a scene represented by three-dimensional gaussian splats into one or more two- dimensional representations, where the one or more two-dimensional representations store different parameters of the three-dimensional gaussian splats, along with associated metadata describing transformation from three-dimensional space into two-dimensional representations; forming, by the encoder, the one or more two-dimensional representations and the associated metadata into one or more bitstreams; and outputting, by the encoder, the one or more bitstreams.
[0006] An additional exemplary embodiment includes a computer program, comprising instructions for performing the method of the previous paragraph, when the computer program is run on an apparatus. The computer program according to this paragraph, wherein the computer program is a computer program product comprising a computer-readable medium bearing the instructions embodied therein for use with the apparatus. Another example is the computer program according to this paragraph, wherein the program is directly loadable into an internal memory of the apparatus.
[0007] An exemplary apparatus includes one or more processors and one or more memories storing instructions that, when executed by the one or more processors, cause the apparatus at least to perform: projecting, by an encoder, a scene represented by three- dimensional gaussian splats into one or more two-dimensional representations, where the one or more two-dimensional representations store different parameters of the three-dimensional gaussian splats, along with associated metadata describing transformation from three- dimensional space into two-dimensional representations; forming, by the encoder, the one or more two-dimensional representations and the associated metadata into one or more bitstreams; and outputting, by the encoder, the one or more bitstreams.
[0008] An exemplary computer program product includes a computer-readable storage medium bearing instructions that, when executed by an apparatus, cause the apparatus to perform at least the following: projecting, by an encoder, a scene represented by three- dimensional gaussian splats into one or more two-dimensional representations, where the one or more two-dimensional representations store different parameters of the three-dimensional gaussian splats, along with associated metadata describing transformation from three- dimensional space into two-dimensional representations; forming, by the encoder, the one or more two-dimensional representations and the associated metadata into one or more bitstreams; and outputting, by the encoder, the one or more bitstreams.
[0009] In another exemplary embodiment, an apparatus comprises means for performing: projecting, by an encoder, a scene represented by three-dimensional gaussian splats into one or more two-dimensional representations, where the one or more two-dimensional representations store different parameters of the three-dimensional gaussian splats, along with associated metadata describing transformation from three-dimensional space into two-dimensional representations; forming, by the encoder, the one or more two-dimensional representations and the associated metadata into one or more bitstreams; and outputting, by the encoder, the one or more bitstreams.
[0010] In an exemplary embodiment, a method is disclosed that includes receiving, by a decoder, one or more bitstreams encoding one or more two-dimensional representations and associated metadata, the one or more two-dimensional representations and associated metadata encoding a scene represented by three-dimensional gaussian splats, where the one or more two- dimensional representations store different parameters of the three-dimensional gaussian splats and the associated metadata stores transformation information from three-dimensional space into two-dimensional representations; applying, by the decoder, decoding to corresponding individual bitstreams to form the one or more two-dimensional representations, and the associated metadata; reprojecting, as described by the associated metadata and by the decoder, the one or more two-dimensional representations into the scene represented by the three-dimensional gaussian splats; and outputting, by the decoder, the scene for rendering, or render the scene, to a viewer.
[0011] An additional exemplary embodiment includes a computer program, comprising instructions for performing the method of the previous paragraph, when the computer program is run on an apparatus. The computer program according to this paragraph, wherein the computer program is a computer program product comprising a computer-readable medium bearing the instructions embodied therein for use with the apparatus. Another example is the computer program according to this paragraph, wherein the program is directly loadable into an internal memory of the apparatus.
[0012] An exemplary apparatus includes one or more processors and one or more memories storing instructions that, when executed by the one or more processors, cause the apparatus at least to perform: receiving, by a decoder, one or more bitstreams encoding one or more two-dimensional representations and associated metadata, the one or more two- dimensional representations and associated metadata encoding a scene represented by three- dimensional gaussian splats, where the one or more two-dimensional representations store different parameters of the three-dimensional gaussian splats and the associated metadata stores transformation information from three-dimensional space into two-dimensional representations;applying, by the decoder, decoding to corresponding individual bitstreams to form the one or more two-dimensional representations, and the associated metadata; reprojecting, as described by the associated metadata and by the decoder, the one or more two-dimensional representations into the scene represented by the three-dimensional gaussian splats; and outputting, by the decoder, the scene for rendering, or render the scene, to a viewer.
[0013] An exemplary computer program product includes a computer-readable storage medium bearing instructions that, when executed by an apparatus, cause the apparatus to perform at least the following: receiving, by a decoder, one or more bitstreams encoding one or more two-dimensional representations and associated metadata, the one or more two- dimensional representations and associated metadata encoding a scene represented by three- dimensional gaussian splats, where the one or more two-dimensional representations store different parameters of the three-dimensional gaussian splats and the associated metadata stores transformation information from three-dimensional space into two-dimensional representations; applying, by the decoder, decoding to corresponding individual bitstreams to form the one or more two-dimensional representations, and the associated metadata; reprojecting, as described by the associated metadata and by the decoder, the one or more two-dimensional representations into the scene represented by the three-dimensional gaussian splats; and outputting, by the decoder, the scene for rendering, or render the scene, to a viewer.
[0014] In another exemplary embodiment, an apparatus comprises means for performing: receiving, by a decoder, one or more bitstreams encoding one or more two- dimensional representations and associated metadata, the one or more two-dimensional representations and associated metadata encoding a scene represented by three-dimensional gaussian splats, where the one or more two-dimensional representations store different parameters of the three-dimensional gaussian splats and the associated metadata stores transformation information from three-dimensional space into two-dimensional representations; applying, by the decoder, decoding to corresponding individual bitstreams to form the one or more two-dimensional representations, and the associated metadata; reprojecting, as described by the associated metadata and by the decoder, the one or more two-dimensional representations into the scene represented by the three-dimensional gaussian splats; and outputting, by the decoder, the scene for rendering, or render the scene, to a viewer.BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In the attached drawings:
[0016] FIG. 1 is a block diagram illustrating a system in accordance with an example;
[0017] FIG. 2 is an illustration of three gaussian splats;
[0018] FIG. 3 is a logic flow diagram for visual encoding of 3D gaussian splats;
[0019] FIG. 3A is a logic flow diagram for visual decoding of 3D gaussian splats;
[0020] FIG. 4 is a block diagram illustrating an encoder performing projection-based coding of 3D gaussian splats;
[0021] FIG. 4A is the encoder of FIG. 4 with static image encoders instead of video encoders;
[0022] FIG. 5 is a block diagram illustrating a decoder performing decoding and reconstruction of video-coded 3D gaussian splat parameters;
[0023] FIG. 5 A is the decoder of FIG. 5 with static image decoders instead of video decoders; and
[0024] FIG. 6 is an example of a block diagram of an apparatus suitable for implementing any of the encoders or decoders described herein.DETAILED DESCRIPTION OF THE DRAWINGS
[0025] Abbreviations that may be found in the specification and / or the drawing figures are defined below, at the end of the detailed description section.
[0026] The word “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments. All of the embodiments described in this Detailed Description are exemplary embodiments provided to enable persons skilled in the art to make or use the invention and not to limit the scope of the invention which is defined by the claims.
[0027] When more than one drawing reference numeral, word, or acronym is used within this description withand in general as used within this description, the “ / ” may be interpreted as “or”, “and”, or “both”. As used herein, “at least one of the following: <a list oftwo or more elements>” and “at least one of ” and similar wording, where the list of two or more elements are joined by “and” or “or,” mean at least any one of the elements, or at least any two or more of the elements, or at least all the elements.
[0028] As used herein, the singular forms “a”, “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises”, “comprising”, “has”, “having”, “includes” and / or “including”, when used herein, specify the presence of stated features, elements, and / or components etc., but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof.
[0029] It is noted that capital and lowercase words or phrases are considered to be the same herein. For instance, the words Slice and slice are the same, as are the phrases Network Repository Function and network repository function.
[0030] Any flow diagram (such as FIGS. 3 and 3A) or signaling diagram herein is considered to be a logic flow diagram, and illustrates the operation of an exemplary method, results of execution of computer program instructions embodied on a computer readable memory, functions performed by logic implemented in hardware, and / or interconnected means for performing functions in accordance with an exemplary embodiment. Block diagrams (such as FIGS. 1, 4, 4A, 5, and 5A) also illustrate the operation of an exemplary method, results of execution of computer program instructions embodied on a computer readable memory, functions performed by logic implemented in hardware, and / or interconnected means for performing functions in accordance with an exemplary embodiment. For methods, flow diagrams, and signaling diagrams, the orders of method steps, blocks in the flow, or signaling are not critical and instead are examples.
[0031] Examples of embodiments herein describe techniques for visual encoding and decoding of 3D gaussian splats. Additional description of these techniques is presented after an introduction to parts of the technical area is presented.
[0032] Turning to FIG. 1, FIG. 1 is a block diagram illustrating a system 100 in accordance with an example. This provides an overview of video coding, transmission and reception, and subsequent decoding and presentation. In the example, the encoder 130 is used to encode video, at a viewpoint 10 from the scene 15 having a human being 20, and the encoder130 is implemented in a transmitting apparatus 120. The encoder 130 produces bitstreams 101 that are transmitted through a channel 170 by the transmitting apparatus 120 and received by the receiving apparatus 140, which implements a decoder 150. The channel 170 could be wired or wireless or both and may introduce errors in the bitstreams 101. The encoder 130 performs an encoding process (one partial example of which is process 300 in FIG. 3), and the decoding 150 performs a decoding process (one partial example of which is process 391 of FIG. 3A). It is noted the processes 300 and 391 are examples and illustrate certain possible implementations. Many other examples that could be used as part of the processes 300 and 391 are described below. The decoder 150 forms the video for the scene 15-1, and the receiving apparatus 140 would present this to the user, e.g., via a smartphone, television, or projector among many other options. In this example, there is a capture of 3D media from the volumetric capture at a viewpoint 10 of a scene 15, which includes a human being 20. The receiving apparatus 140 reproduces a version of the 3D media at a viewpoint 10-1 of a scene 15-1, which includes a human being 20-1.
[0033] One topic of interest is the following. A volumetric frame (such as to capture scene 15 of FIG. 1) can capture in various methods, and the capture data can be represented in numerous formats depending on the processing applied as well as the target application requirements. Consider the following.
[0034] 1) A volumetric frame can be represented as a point cloud. A point cloud is a set of points in space, where each point position has its set of Cartesian coordinates (X, Y, Z) and corresponding attributes (e.g., color information provided as RGB A value).
[0035] 2) A volumetric frame can be represented as video frames with or without depth. In other words, the frame may be represented by one or more view frames (where view represents a view 10 from capture camera with known position, orientation and viewport), each of view frames may be represented by number of components, e.g., a geometry picture, an attribute picture for each attribute, and occupancy picture, which may be part of the geometry picture or represented separately.
[0036] 3) A volumetric frame can be represented as a mesh. Mesh is a collection of vertices, edges and faces that defines the shape of object.
[0037] 4) A volumetric frame can be represented by 3D Gaussian splats.
[0038] 5) A volumetric frame can be represented by a neural model.
[0039] A sequence of visual volumetric frames can create a visual volumetric video, which, when uncompressed, requires a large amount of data, which create challenges for storage and transmission. To tackle the large data requirements, volumetric frames can be coded, e.g., by converting the 3D volumetric information into a collection of 2D images and associated data. The converted 2D images can be coded using widely available video and image coding specifications, such as ISO / IEC 14496-10, ISO / IEC 23008-2, ISO / IEC 23090-3 and the associated data.
[0040] The associated data can be coded with mechanisms specified in ISO / IEC 23090-5, in form of supplemental enhancement information messages ISO / IEC 23002-7.
[0041] The coded images and the associated data can be stored or transmitted to a client that can decode and reconstruct the 3D volumetric information, i.e., volumetric frame.
[0042] An additional topic concerns gaussians. 3D gaussian splatting is a method for digitizing and rendering real-world objects. With gaussian splats, one can represent a 3D scene generated from 2D images.
[0043] In contrary to traditional mesh-based 3D scenes, gaussian splat scenes are not represented by polygons and textures. Instead, gaussian splat scenes are made up of individual, unconnected blobs called splats.
[0044] A splat is a particle in 3D space represented by position, size, normal, rotation, color or spherical harmonics, and / or opacity, where the opacity has a gaussian falloff from the center of the splat to its edge. A black-and-white representation of gaussian splats 200- 1, 200-2, and 200-3 is provided in FIG. 2. Given dense enough collage of splats 200, they can represent physical objects and complex surfaces very well. They also allow representing specular reflections in the splatting data using full or lower order specular harmonics.
[0045] In practice when rendering gaussian splats, the color and optionally depth is derived from multiple gaussian splats. For example, this can be done by projecting the gaussians into 2D camera perspective, sorting the gaussians by depth, and iterating the gaussians for every pixel, blending them together using the color and opacity.
[0046] Visual Volumetric Video-based Coding (V3C) is another topic of interest. ISO / IEC 23090-5 specifies a generic syntax and mechanism for volumetric video coding. Thegeneric syntax can be used by applications targeting volumetric content, such as point clouds, immersive video with depth, and mesh representations of volumetric frames. The purpose of the specification is to define how to decode and interpret the associated data (atlas data in ISO / IEC 23090-5) which tells a Tenderer how to interpret 3D frames to reconstruct volumetric frame.
[0047] Three applications of V3C (ISO / IEC 23090-5) have been or all under development: V-PCC, i.e., point clouds, (ISO / IEC 23090-5), MTV, i.e., multi-view plus depth, (ISO / IEC 23090-12), V-DMC, i.e., dynamic meshes, (ISO / IEC 23090-29). Those applications use a number of V3C syntax elements with a slightly modified semantics.
[0048] To differentiate between applications of V3C bitstream, that allow a client to properly interpret the decoded data, V3C uses ptl_profile_toolset_idc.
[0049] While designing the V3C specification, it was envisaged that amendments or new editions can be created in the future. In order to ensure that the first implementations of V3C decoders are compatible with any future extension, a number of fields for future extensions to parameter sets were reserved.
[0050] With respect to video codecs, the advanced video coding standard (which may be abbreviated H.264, AVC or H.264 / AVC) was developed by the Joint Video Team (JVT) of the Video Coding Experts Group (VCEG) of the Telecommunications Standardization Sector of International Telecommunication Union (ITU-T) and the Moving Picture Experts Group (MPEG) of International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC). The H.264 / AVC standard is published by both parent standardization organizations, and it is referred to as ITU-T Recommendation H.264 and ISO / IEC International Standard 14496-10, also known as MPEG-4 Part 10 Advanced Video Coding (AVC). There have been multiple versions of the H.264 / AVC standard, each integrating new extensions or features to the specification. These extensions include Scalable Video Coding (SVC) and Multiview Video Coding (MVC).
[0051] The High Efficiency Video Coding standard (which may be abbreviated H.265, HEVC or H.265 / HEVC) was developed by the Joint Collaborative Team - Video Coding (JCT-VC) of VCEG and MPEG. The standard is published by both parent standardization organizations, and it is referred to as ITU-T Recommendation H.265 and ISO / IEC International Standard 23008-2, also known as MPEG-H Part 2 High Efficiency Video Coding (HEVC).Extensions to H.265 / HEVC include scalable, multiview, three-dimensional, and fidelity range extensions, which may be referred to as SHVC, MV-HEVC, 3D-HEVC, and REXT, respectively. The references in this description to H.265 / HEVC, SHVC, MV-HEVC, 3D-HEVC and REXT that have been made for the purpose of understanding definitions, structures or concepts of these standard specifications are to be understood to be references to the latest versions of these standards that were available before the date of this application, unless otherwise indicated.
[0052] Versatile Video Coding (which may be abbreviated WC, H.266, or H.266 / VVC) is a video compression standard developed as the successor to HEVC. WC is specified in ITU-T Recommendation H.266 and equivalently in ISO / IEC 23090-3, which is also referred to as MPEG-I Part 3.
[0053] A specification of the AVI bitstream format and decoding process were developed by the Alliance for Open Media (AOM). The AVI specification was published in 2018. AOM is reportedly working on the AV2 specification.
[0054] Information about layered video codecs is now described. Multi-layered coding is a concept wherein an un-encoded visual representation of a scene is, by processes such as transformation and filtering, mapped into multiple dependent or independent representations (called layers). One or more encoders are used to encode a layered visual representation. When the layers contain redundancies, the use of a single encoder can, by using inter-layer prediction techniques, encode with a significant gain in coding efficiency. Layered video coding is typically used to provide some form of scalability in services - e.g., quality scalability, spatial scalability, temporal scalability, and view scalability.
[0055] In modern video coding standards, such as High Efficiency Video Coding (HEVC) and Versatile Video Coding (WC), layers can be grouped into layer-sets which the combination of layers in a layer-set can produce a valid decodable bitstream. When layers are coded using inter-layer prediction techniques, the coded layers have a decoding hierarchy.
[0056] Tools in WC for layered video are now described. WC supports temporal scalability by including a temporal ID (identification) in the NAL (network abstraction layer) unit header. A hierarchical coding structure is defined since pictures of a particular temporal sublayer cannot be used as reference for inter-prediction by pictures of a lower temporal sub-layer.WC also supports spatial scalability. The enhancement layers can be predicted using singlelayer intra-prediction and inter-prediction. In addition to inter-layer prediction, resampled reconstructed video content from the reference base layer can be used for prediction.Furthermore, SNR scalability is achieved similarly as spatial scalability, but without changes in the spatial resolution. With WC, it is possible to extract a sub-bitstream extraction process, and the requirement that each sub-bitstream extraction output be a conforming bitstream.
[0057] Another topic concerns tools in HEVC for layered video. Similar to WC, HEVC also supports temporal scalability through temporal layering. Spatial scalability is provided by extensions, e.g., Multiview HEVC (MV-HEVC) and 3D-HEVC. In MV-HEVC and 3D-HEVC, a layer can represent texture, depth, or the like. There can be dependencies between layers that are exploited through inter-layer prediction. And as in WC, it is possible to extract sub-bitstreams.
[0058] VSEI is now described. ITU-T Recommendation H.274, which is equivalent to ISO / IEC 23002-7, may be called “versatile supplemental enhancement information messages for coded video bitstreams” and be referred to as “versatile supplemental enhancement information” or VSEI. The VSEI standard specifies the syntax and semantics of video usability information (VUI) parameters and supplemental enhancement information (SEI) messages. The VUI parameters and SEI messages defined in the VSEI standard are designed to be conveyed within coded video bitstreams in a manner specified in a video coding specification or to be conveyed by other means determined by the specifications for systems that make use of such coded video bitstreams. The VSEI standard is intended for use with WC coded video bitstreams, although it is drafted in a manner intended to be sufficiently generic that it may also be used with other types of coded video bitstreams. VUI parameters and SEI messages may, for example, assist in processes related to decoding, display or other purposes.
[0059] Now that overviews of technical areas have been provided, problems in associated areas are described. A 3D gaussian splat scene, when represented in a raw format, requires large amount of data. For example, using a PLY format, each splat is represented by 232 bytes:
[0060] 1) position 12 bytes,
[0061] 2) orientation 16 bytes,
[0062] 3) scale 12 bytes, and
[0063] 4) color / spherical harmonics and opacity 192 bytes.
[0064] Each scene could contain millions of splats, which means that a single uncompressed (e.g., uncoded) volumetric frame represented by a million 3D Gaussian splats would require approximately 232 MB of data. For a sequence of gaussian splats at 30 frames per second, this would mean 6.9 GB / s or 56 Gbit / s bandwidth requirements. It is therefore paramount to find tools for compressing 3D gaussian splat scene representations such as through encoding.
[0065] Existing compression schemes for 3D Gaussian splats mainly focus on reduction of spherical harmonic orders, by dropping higher order spherical harmonics. This means that the spherical harmonics are replaced by a 4-byte color / opacity component that no longer supports specular representations. However, the quality of the representation remains very detailed without the specular component. Reduction of spherical harmonics alone would result in 44 bytes per splat, which would still require Gigabit level bandwidth for streaming 3D Gaussian splat sequences.
[0066] Examples herein address these high data rate requirements. An overview is provided now, and further details are provided below. It is noted that the examples presented herein can be applied at different levels. One higher level is encoder-agnostic, which means any of the techniques presented above for generating coded video bitstreams can be used. In particular, the examples presented herein detail how one can create the type of format that can be efficiently encoded by 3D video codec, and many different corresponding messaging may be used to carry this information.
[0067] Additional overview is presented in part in reference to FIG. 3, which is a logic flow diagram for visual encoding of 3D gaussian splats. The encoding process 300 is an example and illustrates certain possible implementations. Many other examples that could be used as part of the process 300 are described below. In FIG. 3, blocks 305, 325, and 350 are the main blocks, and the other blocks form examples of how their corresponding main blocks might be performed. The method in FIG. 3 is performed by an encoder 130, e.g., of a transmitting apparatus 120.
[0068] An example includes an encoding process 300 for projecting (see block 305), by the encoder 130, a scene represented by 3D gaussian splats into one or more 3D representations, that store different parameters of the 3D gaussian splats, along with the associated metadata describing transformation from three-dimensional space into two- dimensional representations. The encoder 130 forms, in block 325, the one or more two- dimensional representations and the associated metadata into one or more bitstreams. And in block 350, the encoder 130 outputs the one or more bitstreams, e.g., to the receiving apparatus 140 of FIG. 1.
[0069] The examples herein can be performed per single frames of a 3D scene, or be performed on dynamic 3D scenes of multiple frames. Single frames are addressed in block 330, and dynamic 3D scenes are address in block 340.
[0070] In one embodiment, one or more 3D representations are encoded using a static image codec (e.g., JPEG, H.264, H.265, H.266, AVI, VP8, VP9, AV2, PNG). The list of the image codes partially illustrates the encoder-agnostic possibilities of the examples herein, and any 2D image codec may be used. The compression with image codecs allows encoding a static scene represented by 3D gaussian splats. This is illustrated by block 330, where the forming includes individually encoding the one or more 3D representations using a corresponding static image encoder, and by block 335 where an individual static image encoder encodes a corresponding two-dimensional representation into one of the following formats: JPEG, H.264, H.265, H.266, AVI, VP9, AV2, PNG.
[0071] In another embodiment, one or more sequences of 3D representations are encoded using a video codec (e.g., JPEG, H.264, H.265, H.266, AVI, VP8, VP9, AV2, PNG). As with the single frames example, the list of the image codes partially illustrates the encoderagnostic possibilities of the examples herein, any 2D video codec may be used. As an example of this, for JPEG, there could be JPEG1, JPEG2, JPEG3 that correspond to the sequence of a 2D representation. Compression with video codecs allows encoding a dynamic scene represented by 3D Gaussians splats. Consider block 340, where the encoder 130 encodes one or more sequences of two-dimensional representations using a video encoder, and block 345, where an individual video encoder encodes a corresponding sequence of a two-dimensional representation into one of the following formats: JPEG, H.264, H.265, H.266, AVI, VP9, AV2, PNG.
[0072] It is noted that different ones of the video encoders 450 or static image encoders 490 could be used for the parameters that are encoded in the 2D representations 430. That is, there should be flexibility to consider different codecs for different parameters. Thus, while each of these encoders 450 / 490 could be only one of JPEG, H.264... , one could instead use JPEG for one encoding (say of the 2D geometry map 430-1) and H.264 for another encoding (say of the 2D SH / color map 430-2).
[0073] In one embodiment, the associated metadata is encoded by V3C and stored as atlas metadata, e.g., as defined in ISO / IEC 23090-5 and its extensions. See block 310. In an alternative embodiment, the associated metadata is stored as SEI messages, which, e.g., can be video-codec-agnostic, for example defined in ISO / IEC 23002-7. See block 315. In one embodiment, two or more 3D scenes represented with 3D gaussian splats with time relations are mapped to two or more 3D images (e.g., video frames) and associated metadata that are stored in sequence creating a volumetric video. See block 320. This is intended to allow coding multiple 3D gaussian splat scenes in one collection of 2D representations.
[0074] FIG. 3 has encoder examples presented in FIGS. 4 and 4A. These are described after a decoding process is described.
[0075] Referring to FIG. 3A, this figure is a logic flow diagram for visual decoding of 3D gaussian splats. This example includes a decoding process 391, which is implemented by a receiving apparatus 140. The decoding process 391 is an example and illustrates certain possible implementations. Many other examples that could be used as part of the process 391 are described below. The main blocks are blocks 353, 365, 390, and 395. The other blocks are examples of how the main blocks might be performed.
[0076] In block 353, a decoder 150 receives one or more bitstreams encoding one or more two-dimensional representations and associated metadata. The one or more two- dimensional representations and associated metadata encode a scene represented by three- dimensional gaussian splats, where the one or more two-dimensional representations store different parameters of the three-dimensional gaussian splats. In block 365, the decoder 150 applies decoding to corresponding individual bitstreams to form one or more two-dimensional representations, and associated metadata, the associated metadata stores transformation information from three-dimensional space into two-dimensional representations. The decoder150 in block 390 reprojects, as described by the associated metadata and by the decoder, the one or more two-dimensional representations into the scene represented by the three-dimensional gaussian splats, and in block 395 outputs the scene for rendering, or renders the scene, to a viewer.
[0077] Blocks 355, 360, and 363 are examples of (e.g., part of) block 353. In block 355, the associated metadata is encoded by V3C stored as atlas metadata, and the decoder 150 extracts the associated metadata from the atlas metadata and decodes using V3C. The associated metadata is encoded by SEI messages in block 360, and the decoder 150 extracts the associated metadata from the SEI messages for decoding. In block 363, two or more three-dimensional scenes represented with gaussian splatting with time relations are mapped to two or more two- dimensional images and associated metadata that are stored in a sequence creating a volumetric video. The decoder extracts the two or more three-dimensional scenes represented by gaussian splatting from two or more two-dimensional images and associated metadata represented as a sequence of volumetric video.
[0078] The applied decoding in block 365 may include block 370 (for static image decoders) or block 380 (for video decoders). In block 370, the decoder 150 individually decodes the one or more 3D representations using a corresponding static image decoder. This may be performed where (see block 375) an individual static image decoder decodes a corresponding two-dimensional representation using one of the following formats: JPEG, H.264, H.265, H.266, AVI, VP8, VP9, AV2, PNG. Another example is block 380, where the decoder 150 decodes one or more sequences of two-dimensional representations using a video decoder. This may include (see block 385) where an individual video decoder decodes a corresponding sequence of a two- dimensional representation using one of the following formats: JPEG, H.264, H.265, H.266, AVI, VP8, VP9, AV2, PNG. Block 386 corresponds to block 363 and the applying decoding performs extracting the two or more three-dimensional scenes represented by gaussian splatting from two or more two-dimensional images and associated metadata represented as a sequence of volumetric video.
[0079] FIG. 3A has corresponding decoder examples presented in FIGS. 5 and 5A.These are described after video-based coding corresponding to FIG. 3 is described in reference to FIGS. 4 and 4A.
[0080] Now that an overview has been provided, further details are provided. One example of a general design of video-based coding by an encoder 130 of 3D Gaussian splats is described in FIG. 4. As input, the system takes a sequence of volumetric frames represented by 3D Gaussian splats. See the 3D scene 410. The scene 410 is projected (by the projection-based transformation 420) into one or more 2D representations 430 characterizing different properties of the 3D gaussian splats, and into the associated metadata. This example has the associated metadata, shown as the per-frame metadata 440, as being encoded into its own bitstream. This is but one example, and the per-frame metadata can be multiplexed into a video bitstream. This is one case, as an example, where the metadata would be encoded as SEI messages and for example stored in the geometry / position video bitstream. The associated per-frame metadata describes the projection of 3D gaussian splats into 3D representations, i.e., maps. The metadata defines what parameters 425 (e.g., geometry; SH / color, including diffuse color; opacity; scaling; orientation) are present and how they are arranged into different maps. Furthermore, the metadata describes how different maps may be multiplexed into shared maps. This is useful for example in cases, when parameters 425 require little data per pixel in a map. Combining such parameters 425 in a shared map may reduce the number of separate encoded bitstreams. This is illustrated by block 475, where two or more of the parameters are combined into a shared two-dimensional representation.
[0081] This example has the following 3D representations 430: 3D geometry map 430-1; 3D SH / color map 430-2; 3D opacity map 430-3; 3D scaling map 430-4; and 3D orientation map 430-5. Other maps may be present, e.g., normal map, rotation map, and the like. Multiple parameters may be present in a shared map. As an example, scaling map and opacity map can be combined into a single shared map. Moreover, all details of a parameter 425 may not fit in a single representation. That is, a map may require 16 channels, in which case one would want to split the map into multiple representations. The sequence of 3D representations can be encoded using a (corresponding) video codec 450-1, 450-2, 450-3, 450-4, and 450-5 into one or more video bitstreams 470 (e.g., part of bitstream 101 of FIG. 1): geometry video bitstream 470-1; 3D SH / color video bitstream 470-2; opacity bitstream 470-3; scaling video bitstream 470-4; and orientation video bitstream 470-5. Additionally other video bitstreams may be present which carry parameters such as normal, or rotation. The per-frame metadata 440 canbe encoded by a metadata encoder 460 into a metadata bitstream 480 (e.g., as part of bitstreams 101 of FIG. 1).
[0082] Referring to FIG. 4A, this figure is the encoder 130 of FIG. 4 with static image encoders 490 instead of video encoders 450. That is, the 3D representations 430-1 through 430-5 are encoded with individual static image encoders 490-1 through 490-5 in this example.
[0083] FIGS. 4 and 4A describe how each parameter (geometry, color, opacity, scaling, and orientation in these examples, though other parameters such as spherical harmonics, diffuse color; rotation; normals; and / or orientation may be used) is projected into different 3D maps (referred to as 3D representations 430) and subsequently encoded as a separate video bitstream 470. In practice, multiple 3D maps (a.k.a. 3D representations 430) may be combined in the different channels of the same 3D representation 430 and thus encoded with a single video codec (also referred to as a video encoder) 450 or static image codec (also referred to as static image encoder) 490. What this means is that, for example, the opacity map 30-3 and the geometry map 430-1 could be combined and operated on by one video encoder 450 or one static image encoder 490. The per-frame metadata 440 describes how the parameters are packed in the 3D representations 430.
[0084] FIG. 5 describes an example of a decoder 150 performing the decoding and reconstruction process for video-coded 3D gaussian splat parameters. The input to the system is one or more encoded bitstreams of video / image data. This example has the following bitstreams (from bitstreams 101): geometry video bitstream 570-1; 3D SH / color video bitstream 570-2; opacity bitstream 570-3; scaling video bitstream 570-4; orientation video bitstream 570-5; and metadata bitstream 580. These bitstreams 570, 580 should be equivalent to corresponding bitstreams 470, 480, except for any errors caused by the transmission medium. Individual bitstreams 570-1 to 570-5 can be decoded by corresponding video decoders 550-1 to 550-5 to acquire 3D parameter maps 530-1 to 530-5 (which are examples of 3D representations, also referred to as 530) representing the 3D gaussian splat scene 510. The parameters 525 include one or more the following in this example: geometry; SH / color; opacity; scaling; and orientation, though other parameters such as spherical harmonics, diffuse color; rotation; normals; and / or orientation may be used. The metadata bitstream 580 is decoded by the metadata decoder 560 to retrieve the per-frame metadata 540. Block 575 corresponds to block 475 of FIG. 4, and thedecoder 150 decodes a shared two-dimensional representation having two or more of the parameters that have been combined into the shared two-dimensional representation. Together with the (decoded) per-frame metadata 540, the 3D parameter maps 530 can be reprojected via the re-projection 520 of the 3D parameter maps 530 to form the, e.g., original 3D scene 510 represented by 3D gaussian splats and subsequently rendered 585 to the viewer.
[0085] FIG. 5 A is the decoder 140 of FIG. 5 with static image decoders 590 instead of video decoders 550. Otherwise, operation is similar.
[0086] The following are additional examples, mainly for the encoding process. While the encoding process is used, the same or similar techniques could be used for the decoding process.
[0087] A (e.g., an original) 3D scene 410 may include multiple 3D gaussian splats, where each 3D gaussian splat includes one or more of the following parameters: position, size (e.g., scale), normal, rotation, color or spherical harmonics, or opacity. The scene can be projected into one or more 2D representations and associated metadata.
[0088] General metadata (e.g., per-frame metadata 440) is used to indicate which of the parameters are stored in the one or more 2D representations. As such, it is useful to establish relationships with the parameters and the 2D representations.
[0089] 1) In one embodiment, the relationship is established by allocating indices to each parameter and 2D representation.
[0090] 2) In another embodiment, the relationship is established by with a discrete mapping or a lookup-table.
[0091] Furthermore, the values of the parameters may be specific to a sub-region of the 2D representations 430 or the original 3D representation (from 3D scene 410). The subregions may contain higher level information for the parameters that are needed to efficiently store the parameters in the 2D representations. The possible division into sub-regions and the layout of sub-regions in the 2D representations should be indicated in the associated metadata.
[0092] In particular, the video components can be partitioned into tiles, where tiles correspond to regions of the 3D scene, for example obtained by an octree partitioning or any volume segmentation of the scene. Associated metadata may then describe the 3D scenepartitioning and the correspondence between the video component tiles and the partitions of the 3D scene.
[0093] The relationship between parameters and 2D representations as well as the division into sub-regions may be encoded in one or more parameter sets, SEI message or another implicit structure signaled in or along the 2D representations.
[0094] Position parameter (more commonly known as depth or geometry) of a 3D gaussian splat may be stored as a pixel value in one or more planes of the 2D representations 430 (for example luma or all planes as in colorized depth). The pixel value indicates the distance of a gaussian splat from the projection plane. In the 2D representation, the u- and v-offsets of the pixel indicate the transformation of the gaussian in tangential and co-tangential direction of the projection plane. The u- and v-offsets may be represented relative to smaller regions (e.g., patches or tiles) inside the projection plane. Consider the following.
[0095] 1) Associated metadata defines the projection plane position in 3D space either as orthogonal model, perspective model, equirectangular model or another mathematical construct.
[0096] 2) Associated metadata may provide information about how the projection plane is divided into smaller 2D regions (e.g., patches or tiles) and their offsets relative to the overall projection plane.
[0097] 3) Associated metadata should provide information how the 2D representations or regions of 2D representations are packed in large 2D frames (e.g., atlases in V3C context).
[0098] 4) Associated metadata may provide information how to interpret the pixel value, i.e., how to derive the 3D position from the pixel value. This information should contain information on the following:
[0099] a) the quantization range (e.g., 2d video coding bit-depth);
[0100] b) the minimum and maximum values (e.g., minimum and maximum depth range); and
[0101] c) a distribution function (e.g., linear, inverse, normalized, colorized, and the like).
[0102] Orientation of a 3D gaussian splat may be stored as a pixel value in one or more planes of the 2D representation, so that the orientation can be mapped to a gaussian in the corresponding texture coordinates of the geometry map. 3D orientation as a quaternion typically consists of 4 components (x,y,z,w). One of the components may be unknown as it can be derived from the three other components.
[0103] Associated metadata describes how to interpret the orientation pixel value, i.e., its quantization, distribution, and representation format. Possibilities include the following.
[0104] 1) Quantization range may depend on the codec used to compress the 2D representation. Typical quantization ranges allow 8- or 10-bit quantization. But more advanced codecs allow higher quantization ranges.
[0105] 2) Distribution defines how the values are allocated on the quantization range.Linear normalized distribution may be applicable in most scenarios because the components of the orientation don’t have necessary preferred distribution curve. Optionally other distributions modes could be supported.
[0106] 3) Representation type indicates how the orientation is encoded in the planes of the one or more 2D representations.
[0107] a) In one embodiment, the representation for the orientation can be stored as a normal map in the tangent or inverse tangent space of the projection plane. Only orientations for the gaussians that are pointing towards the projection plane can be encoded in the normal map. The normal map value for a 3D gaussian splat normally represents a unit normal vector consisting of three components (x,y,z), where each component can be normalized to a range from -1 to 1. These values can be then mapped to corresponding channels of the 2D representation.
[0108] b) In another embodiment, three components of the quaternion can be directly quantized, normalized and stored in the different channels of the 2D representation. This would allow encoding orientation for gaussians that are not necessarily pointing towards the projection plane. These values can then be mapped to the corresponding channels of the 2D representation.
[0109] Scaling (also known as radius or size) parameter for a 3D Gaussian splat is stored as pixel value in one or more planes of the 2D representation 430. The value of the scaling parameter would include the type of scaling, range of scaling, and the distribution.
[0110] 1) The type of scaling could indicate if the scaling is uniform, i.e., if a single scaling factor is applied to scale the splat equally in every dimension (x,y,z), or if the scaling is non-uniform and the scaling factor should be indicated for each dimension separately. There could also be a 2D scaling type, which would explain how the splat is transformed in the tangential space of the 3D gaussian.
[0111] 2) The range of scaling could indicate what is the maximum allowed range for the scaling values so that scaling values can be normalized appropriately. The scaling range would ideally inform the size of the smallest and largest 3D gaussian splat in the scene.
[0112] 3) The distribution could explain how the pixel values are allocated in the scaling range.
[0113] 4) Quantization range could be indicated by the maximum bit-depth of a codec for compressing the 2D representation. But the range could also be allocated smaller, if multiple 3D gaussian parameters were stored in the same channel of the 2D representation.
[0114] 5) Packing type informs how the scaling values are packed in the channels of the 2D representation.
[0115] A spherical harmonics parameter for a 3D Gaussian splat may be stored as pixel value in one or more planes of one or more 2D representations 430. Spherical harmonics have a large amount of data and can be divided into multiple “orders” or “levels”. The number of coefficients per layer is defined by a simple formula (e.g., num coefficents = 2*num_of_layer + 1). As such, the 0 (zero) level consists of one coefficient and the first level consists of three coefficients. The lower-level coefficients are generally needed to derive the spherical harmonics of the higher levels and as such they form a decoding dependency. Up to seven levels (including the 0-level) are needed to fully represent the spherical harmonics of a given 3D Gaussian splat, which means that there are 48 coefficients altogether. Natively, each coefficient may be represented by a floating-point value, which poses constraints for representing the data as pixel values of 2D representations. The floating-point values should be normalized and quantized before the values can be stored as texture data.
[0116] In another embodiment, the 45+3 spherical harmonics components are filtered down to a smaller set of spherical harmonics through dimensionality reduction, such as Principal1Component Analysis, or another reduction, and the total number of video frames for each triplet of transformed spherical harmonics is set to N, with N smaller than 16.
[0117] As is common, it is possible to reduce the order of spherical harmonics into just three components that represent the first level. The spherical harmonics parameter for 3D gaussian splats would therefore consist of the quantization range, distribution mode, storage mode, and the level or order of spherical harmonics.
[0118] 1) The level or order of spherical harmonics allows one to reconstruct the original spherical harmonics representation. The level or order also allows to select only lower levels of the spherical harmonics, for cases when higher levels cannot be efficiently encoded or streamed due to performance or bandwidth limitations.
[0119] 2) Quantization range requires storing min-and-max (minimum and maximum) values for the coefficients. It may be beneficial to store quantization range for each level or order implicitly.
[0120] 3) Distribution mode describes how the pixel values are allocated for the coefficients in the quantization range. This could be, for example, linearly allocated.
[0121] 4) The storage mode indicates if temporal interleaving is used to store spherical harmonic coefficients corresponding to one temporal instance in one or more subsequent 2D frame representation.
[0122] An opacity parameter for a 3D Gaussian splat may be stored as a pixel value in one or more planes of one or more 2D representations. The opacity value can be normalized and quantized before storing the value as texture data. It is expected that the opacity is quantized in a range 0... 1, the values can be distributed by any distribution curve, typically linear.
[0123] As a result of the above storage mechanisms, the 3D Gaussian splats for a scene can be encoded as a set of 2D representations and associated metadata. Consider the following.
[0124] 1) In one embodiment, an image codec such as JPEG, PNG, HEVC, or the like can be used to compress the 2D representations for a static scene having 3D Gaussian splats.
[0125] 2) In another embodiment, a video codec such as AVC, HEVC, WC, AVI,VP9, or the like can be used to compress a sequence of 2D representations for a dynamic scene having 3D Gaussian splats.
[0126] In one embodiment, one or more 2D representations 430 of the parameters for 3D gaussian splats may be combined into different planes of the same 2D representation. For example, opacity and scale, or scale and position, may be encoded in the different planes of the same 2D representation. In this case, the associated metadata stores the relationship between the multiple parameters and the one or more 2D representations.
[0127] In one embodiment, the associated metadata is according to V3C and its extension may include the following:
[0128] 1) a new profile related to gaussian splatting representation is defined;
[0129] 2) new video components such as orientation, scale and spherical harmonics are added and identified by V3C unit types and VPS; and / or
[0130] 3) atlas metadata provide information on how to interpret pixels from each video components to ensure proper reverse procedure from 2D to representation of 3D scene.
[0131] As described above, associated metadata may be stored as SEI messages. The SEI messages are fairly generic and some codecs define their own SEI messages. Meanwhile, VSEI is more specific standard that is intended for versatile SEI messages for video codecs. VSEI is an example of what the SEI messages could be in block 360 of FIG. 3, but the SEI messages could also be defined in some other standard. If VSEI is used, the associated metadata is according to VSEI and layered video codec, which may include the following:
[0132] 1) a new profile related to gaussian splatting representation is defined;
[0133] 2) video components carrying different data types are stored at different layers of the video bitstream;
[0134] 3) parameter sets of video codec are defined to indicate how to interpret the different layers;
[0135] 4) VSEI message provides information on how to interpret the layers; and / or
[0136] 5) VSEI message provides information on how to interpret the pixels of each video component to ensure proper reverse procedure from 2D to representation of 3D scene.
[0137] In another embodiment, the associated metadata can be signaled in or along the encoded 2D representations in any another explicitly defined structures.
[0138] In one embodiment, the spherical harmonic components can be parameterized by, for example, a color gradient defined on the sphere, or a discrete number of gaussians on thesphere. Spherical harmonics are generic and may represent a wide variety of visual anisotropy that is only fully exploited for specific parts of content with transparency or reflectivity properties. In one specific embodiment, the gaussian splats are segmented into different sub streams based on their content properties, and their encoding is then chosen specifically: for example, a color gradient for the mostly diffuse parts of the scene, and spherical harmonics for transpar ent / reflective surfaces.
[0139] Turning to FIG. 6, this figure is an example of a block diagram of an apparatus 680 suitable for implementing any of the encoders 130 (as transmitting apparatus 120) or decoders 140 (as receiving apparatus 140) described herein. The apparatus 680 includes circuitry comprising one or more processors 620, one or more memories 625, one or more transceivers 630, one or more network (N / W) interface(s) (I / F(s)) 655 and user interface (UI) circuitry and elements 657, interconnected through one or more buses 627. Depending on implementation, some apparatus may not have all of the circuitry. For example, an apparatus 680 might not have UI circuitry and elements 657. An apparatus may have additional circuitry, not described here. FIG. 6 is presented merely as an example.
[0140] Each of the one or more transceivers 630 includes a receiver, Rx, 632 and a transmitter, Tx, 633. The one or more buses 627 may be address, data, and / or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment, and the like. The one or more transceivers 630 are connected to one or more antennas 605, and may communicate using wireless link 611.
[0141] The one or more memories 625 include computer program code 623. The apparatus 680 includes a program 640, comprising one of or both parts 640-1 and / or 640-2. The program 640 may implement an encoder 130, a decoder 140, or a codec, which implements both encoding and decoding. The program itself may be implemented in a number of ways. The program 640 may be implemented in hardware as program 640-1, such as being implemented as part of the one or more processors 620. The program 640-1 may be implemented also as an integrated circuit or through other hardware such as a programmable gate array. In another example, the program 640 may be implemented as program 640-2, which is implemented as computer program code (having corresponding instructions) 623 and is executed by the one ormore processors 620. For instance, the one or more memories 625 store instructions (e.g., in the computer program code 623) that, when executed by the one or more processors 620, cause the apparatus 680 to perform one or more of the operations as described herein. Furthermore, the one or more processors 620, one or more memories 625, and example algorithms (e.g., as flowcharts and / or signaling diagrams), encoded as instructions, programs, or code, are means for causing performance of the operations described herein.
[0142] The network interface(s) (N / W I / F(s)) 655 are wired interfaces communicating using link(s) 656, which could be fiber optic or other wired interfaces. The apparatus 680 could include only wireless transceiver(s) 630, only N / W I / Fs 655, or both wireless transceiver(s) 630 and N / W I / Fs 655.
[0143] The apparatus 680 may or may not include UI circuitry and elements 657. These could include a display such as a touchscreen, speakers, or interface elements such as for headsets. For instance, an apparatus 680 of a smartphone would typically include at least a touchscreen and speakers. The UI circuitry and elements 657 may also include circuity to communicate with external UI elements (not shown) such as displays, keyboards, mice, headsets, and the like.
[0144] The computer readable memories 625 may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, flash memory, firmware, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The computer readable memories 625 may be means for performing storage functions. The processors 620 may be of any type suitable to the local technical environment, and may include one or more of general-purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs) and processors based on a multi-core processor architecture, as nonlimiting examples. The processors 620 may be means for performing functions, such as controlling the apparatus 680, and other functions as described herein.
[0145] Without in any way limiting the scope, interpretation, or application of the claims appearing below, a technical effect and / or advantage of one or more of the example embodiments disclosed herein is that it allows efficient compression of static and dynamic scenes consisting of 3D gaussian splats by projecting 3D gaussian splats into 2D representations,which can be coded with existing video or image codecs, and associated metadata, which can be coded using, e.g., V3C or SEI messages, describing the conversion from 3D into 2D space.Another technical effect and / or advantage of one or more of the example embodiments disclosed herein is that existing hardware intended for 2D video coding may be exploited.
[0146] The following are additional examples.
[0147] Example 1. A method, comprising: projecting, by an encoder, a scene represented by three-dimensional gaussian splats into one or more two-dimensional representations, where the one or more two-dimensional representations store different parameters of the three-dimensional gaussian splats, along with associated metadata describing transformation from three-dimensional space into two-dimensional representations; forming, by the encoder, the one or more two-dimensional representations and the associated metadata into one or more bitstreams; and outputting, by the encoder, the one or more bitstreams.
[0148] Example 2. The method according to example 1, wherein forming comprises individually encoding the one or more 2D representations using a corresponding static image encoder.
[0149] Example 3. The method according to example 1, wherein forming comprises encoding one or more sequences of two-dimensional representations using a video encoder.
[0150] Example 4. The method according to example 1, wherein the scene is an original three-dimensional scene, wherein values of the different parameters are specific to one or more sub-regions of the one or more two-dimensional representations and the original three- dimensional scene, and where the one or more sub-regions of the one or more two-dimensional representations are indicated in the associated metadata.
[0151] Example 5. The method according to any one of examples 1 to 4, wherein the associated metadata is encoded by V3C and stored as atlas metadata.
[0152] Example 6. The method according to any one of examples 1 to 4, wherein the associated metadata is stored as SEI messages.
[0153] Example 7. The method according to any one of examples 1 to 6, wherein projecting comprises mapping two or more three-dimensional scenes represented with 3D gaussian splats with time relations to two or more two-dimensional images and associated metadata that are stored in a sequence creating a volumetric video.
[0154] Example 8. The method according to any one of examples 1 to 7, wherein the two-dimensional representations comprise one or more of the following parameters: geometry; spherical harmonics, diffuse color; opacity; scaling; rotation; normals; or orientation.
[0155] Example 9. The method according to example 8, wherein two or more of the parameters are combined into a shared two-dimensional representation.
[0156] Example 10. A method, comprising: receiving, by a decoder, one or more bitstreams encoding one or more two-dimensional representations and associated metadata, the one or more two-dimensional representations and associated metadata encoding a scene represented by three-dimensional gaussian splats, where the one or more two-dimensional representations store different parameters of the three-dimensional gaussian splats and the associated metadata stores transformation information from three-dimensional space into two- dimensional representations; applying, by the decoder, decoding to corresponding individual bitstreams to form the one or more two-dimensional representations, and the associated metadata; reprojecting, as described by the associated metadata and by the decoder, the one or more two-dimensional representations into the scene represented by the three-dimensional gaussian splats; and outputting, by the decoder, the scene for rendering, or render the scene, to a viewer.
[0157] Example 11. The method according to example 10, wherein applying comprises individually decoding the one or more two-dimensional representations using a corresponding static image decoder.
[0158] Example 12. The method according to example 10, wherein applying comprises decoding one or more sequences of two-dimensional representations using a video decoder.
[0159] Example 13. The method according to example 10, wherein the scene is an original three-dimensional scene, wherein values of the different parameters are specific to one or more sub-regions of the one or more two-dimensional representations and the original three- dimensional scene, and where the one or more sub-regions of the one or more two-dimensional representations are indicated in the associated metadata.
[0160] Example 14. The method according to any one of examples 10 to 13, wherein the associated metadata is encoded by V3C stored as atlas metadata, and wherein the methodfurther comprises extracting the associated metadata from the atlas metadata and decoding using V3C.
[0161] Example 15. The method according to any one of examples 10 to 13, wherein the associated metadata are stored as SEI messages, and wherein the method further comprises extracting the associated metadata from the SEI messages for decoding.
[0162] Example 16. The method according to any one of examples 10 to 15, wherein two or more three-dimensional scenes represented with gaussian splatting with time relations are mapped to two or more two-dimensional images and associated metadata that are stored in a sequence creating a volumetric video, and wherein applying decoding comprises extracting the two or more three-dimensional scenes represented by gaussian splatting from two or more two- dimensional images and associated metadata represented as a sequence of volumetric video.
[0163] Example 17. The method according to any one of examples 10 to 16, wherein the two-dimensional representations comprise one or more of the following parameters: geometry; spherical harmonics, diffuse color; opacity; scaling; rotation; normals; or orientation.
[0164] Example 18. The method according to example 16, wherein the applying comprises decoding a shared two-dimensional representation having two or more of the parameters that have been combined into the shared two-dimensional representation.
[0165] Example 19. A computer program, comprising instructions for performing the methods of any of examples 1 to 18, when the computer program is run on an apparatus.
[0166] Example 20. The computer program according to example 19, wherein the computer program is a computer program product comprising a computer-readable medium bearing instructions embodied therein for use with the apparatus.
[0167] Example 21. The computer program according to example 19, wherein the computer program is directly loadable into an internal memory of the apparatus.
[0168] Example 22. An apparatus, comprising means for performing: projecting, by an encoder, a scene represented by three-dimensional gaussian splats into one or more two- dimensional representations, where the one or more two-dimensional representations store different parameters of the three-dimensional gaussian splats, along with associated metadata describing transformation from three-dimensional space into two-dimensional representations;forming, by the encoder, the one or more two-dimensional representations and the associated metadata into one or more bitstreams; and outputting, by the encoder, the one or more bitstreams.
[0169] Example 23. The apparatus according to example 22, wherein forming comprises individually encoding the one or more 2D representations using a corresponding static image encoder.
[0170] Example 24. The apparatus according to example 22, wherein forming comprises encoding one or more sequences of two-dimensional representations using a video encoder.
[0171] Example 25. The apparatus according to example 22, wherein the scene is an original three-dimensional scene, wherein values of the different parameters are specific to one or more sub-regions of the two-dimensional representations and the original three-dimensional scene, and where the one or more sub-regions of the 2D representations are indicated in the associated metadata.
[0172] Example 26. The apparatus according to any one of examples 22 to 25, wherein the associated metadata is encoded by V3C and stored as atlas metadata.
[0173] Example 27. The apparatus according to any one of examples 22 to 25, wherein the associated metadata is stored as SEI messages.
[0174] Example 28. The apparatus according to any one of examples 22 to 27, wherein projecting comprises mapping two or more three-dimensional scenes represented with 3D gaussian splats with time relations to two or more two-dimensional images and associated metadata that are stored in a sequence creating a volumetric video.
[0175] Example 29. The apparatus according to any one of examples 22 to 28, wherein the two-dimensional representations comprise one or more of the following parameters: geometry; spherical harmonics, diffuse color; opacity; scaling; rotation; normals; or orientation.
[0176] Example 30. The apparatus according to example 29, wherein two or more of the parameters are combined into a shared two-dimensional representation.
[0177] Example 31. An apparatus, comprising means for performing: receiving, by a decoder, one or more bitstreams encoding one or more two-dimensional representations and associated metadata, the one or more two-dimensional representations and associated metadata encoding a scene represented by three-dimensional gaussian splats, where the one or more two-dimensional representations store different parameters of the three-dimensional gaussian splats and the associated metadata stores transformation information from three-dimensional space into two-dimensional representations; applying, by the decoder, decoding to corresponding individual bitstreams to form the one or more two-dimensional representations, and the associated metadata; reprojecting, as described by the associated metadata and by the decoder, the one or more two-dimensional representations into the scene represented by the three-dimensional gaussian splats; and outputting, by the decoder, the scene for rendering, or render the scene, to a viewer.
[0178] Example 32. The apparatus according to example 31, wherein applying comprises individually decoding the one or more two-dimensional representations using a corresponding static image decoder.
[0179] Example 33. The apparatus according to example 31, wherein applying comprises decoding one or more sequences of two-dimensional representations using a video decoder.
[0180] Example 34. The apparatus according to example 31, wherein values of the different parameters are specific to one or more sub-regions of the one or more two-dimensional representations and the original three-dimensional scene, and where the one or more sub-regions of the one or more two-dimensional representations are indicated in the associated metadata.
[0181] Example 35. The apparatus according to any one of examples 31 to 34, wherein the associated metadata is encoded by V3C stored as atlas metadata, and wherein the means are further configured for performing extracting the associated metadata from the atlas metadata and decoding using V3C.
[0182] Example 36. The apparatus according to any one of examples 27 to 34, wherein the associated metadata are stored as SEI messages, and wherein the means are further configured for performing extracting the associated metadata from the SEI messages for decoding.
[0183] Example 37. The apparatus according to any one of examples 27 to 36, wherein two or more three-dimensional scenes represented with gaussian splatting with time relations are mapped to two or more two-dimensional images and associated metadata that are stored in a sequence creating a volumetric video, and wherein applying decoding comprisesextracting the two or more three-dimensional scenes represented by gaussian splatting from two or more two-dimensional images and associated metadata represented as a sequence of volumetric video.
[0184] Example 38. The apparatus according to any one of examples 27 to 37, wherein the two-dimensional representations comprise one or more of the following parameters: geometry; spherical harmonics, diffuse color; opacity; scaling; rotation; normals; or orientation.
[0185] Example 39. The apparatus according to example 38, wherein the applying comprises decoding a shared two-dimensional representation having two or more of the parameters that have been combined into the shared two-dimensional representation.
[0186] Example 40. The apparatus of any preceding apparatus example, wherein the means comprises: at least one processor; and at least one memory storing instructions that, when executed by at least one processor, cause the performance of the apparatus.
[0187] Example 41. An apparatus, comprising: one or more processors; and one or more memories storing instructions that, when executed by the one or more processors, cause the apparatus at least to perform: projecting, by an encoder, a scene represented by three- dimensional gaussian splats into one or more two-dimensional representations, where the one or more two-dimensional representations store different parameters of the three-dimensional gaussian splats, along with associated metadata describing transformation from three- dimensional space into two-dimensional representations; forming, by the encoder, the one or more two-dimensional representations and the associated metadata into one or more bitstreams; and outputting, by the encoder, the one or more bitstreams.
[0188] Example 42. The apparatus according to example 41, wherein forming comprises individually encoding the one or more 2D representations using a corresponding static image encoder.
[0189] Example 43. The apparatus according to example 41, wherein forming comprises encoding one or more sequences of two-dimensional representations using a video encoder.
[0190] Example 44. The apparatus according to example 41, wherein the scene is an original three-dimensional scene, wherein values of the different parameters are specific to one or more sub-regions of the one or more two-dimensional representations and the original three-dimensional scene, and where the one or more sub-regions of the one or more two-dimensional representations are indicated in the associated metadata.
[0191] Example 45. The apparatus according to any one of examples 41 to 44, wherein the associated metadata is encoded by V3C and stored as atlas metadata.
[0192] Example 46. The apparatus according to any one of examples 41 to 44, wherein the associated metadata is stored as SEI messages.
[0193] Example 47. The apparatus according to any one of examples 41 to 46, wherein projecting comprises mapping two or more three-dimensional scenes represented with 3D gaussian splats with time relations to two or more two-dimensional images and associated metadata that are stored in a sequence creating a volumetric video.
[0194] Example 48. The apparatus according to any one of examples 41 to 47, wherein the two-dimensional representations comprise one or more of the following parameters: geometry; spherical harmonics, diffuse color; opacity; scaling; rotation; normals; or orientation.
[0195] Example 49. The apparatus according to example 48, wherein two or more of the parameters are combined into a shared two-dimensional representation.
[0196] Example 50. An apparatus, comprising: one or more processors; and one or more memories storing instructions that, when executed by the one or more processors, cause the apparatus at least to perform: receiving, by a decoder, one or more bitstreams encoding one or more two-dimensional representations and associated metadata, the one or more two- dimensional representations and associated metadata encoding a scene represented by three- dimensional gaussian splats, where the one or more two-dimensional representations store different parameters of the three-dimensional gaussian splats and the associated metadata stores transformation information from three-dimensional space into two-dimensional representations; applying, by the decoder, decoding to corresponding individual bitstreams to form the one or more two-dimensional representations, and the associated metadata; reprojecting, as described by the associated metadata and by the decoder, the one or more two-dimensional representations into the scene represented by the three-dimensional gaussian splats; and outputting, by the decoder, the scene for rendering, or render the scene, to a viewer.
[0197] Example 51. The apparatus according to example 50, wherein applying comprises individually decoding the one or more two-dimensional representations using a corresponding static image decoder.
[0198] Example 52. The apparatus according to example 50, wherein applying comprises decoding one or more sequences of two-dimensional representations using a video decoder.
[0199] Example 53. The apparatus according to example 50, wherein the scene is an original three-dimensional scene, wherein values of the different parameters are specific to one or more sub-regions of the two-dimensional representations and the original three-dimensional scene, and where the one or more sub-regions of the 2D representations are indicated in the associated metadata.
[0200] Example 54. The apparatus according to any one of examples 50 to 53, wherein the associated metadata is encoded by V3C stored as atlas metadata, and wherein the one or more memories further store instructions that, when executed by the one or more processors, cause the apparatus at least to perform extracting the associated metadata from the atlas metadata and decoding using V3C.
[0201] Example 55. The apparatus according to any one of examples 50 to 53, wherein the associated metadata are stored as SEI messages, and wherein the one or more memories further store instructions that, when executed by the one or more processors, cause the apparatus at least to perform extracting the associated metadata from the SEI messages for decoding.
[0202] Example 56. The apparatus according to any one of examples 27 to 55, wherein two or more three-dimensional scenes represented with gaussian splatting with time relations are mapped to two or more two-dimensional images and associated metadata that are stored in a sequence creating a volumetric video, and wherein applying decoding comprises extracting the two or more three-dimensional scenes represented by gaussian splatting from two or more two-dimensional images and associated metadata represented as a sequence of volumetric video.
[0203] Example 57. The apparatus according to any one of examples 50 to 56, wherein the two-dimensional representations comprise one or more of the following parameters: geometry; spherical harmonics, diffuse color; opacity; scaling; rotation; normals; or orientation.
[0204] Example 58. The apparatus according to example 57, wherein the applying comprises decoding a shared two-dimensional representation having two or more of the parameters that have been combined into the shared two-dimensional representation.
[0205] As used in this application, the term “circuitry” may refer to one or more or all of the following:
[0206] (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry) and
[0207] (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions) and
[0208] (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.
[0209] This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.
[0210] Embodiments herein may be implemented in software (executed by one or more processors), hardware (e.g., an application specific integrated circuit), or a combination of software and hardware. In an example embodiment, the software (e.g., application logic, an instruction set) is maintained on any one of various conventional computer-readable media. Inthe context of this document, a “computer-readable medium” may be any media or means that can contain, store, communicate, propagate or transport the instructions for use by or in connection with an instruction execution system, apparatus, or device, such as a computer, with one example of a computer described and depicted, e.g., in FIG. 6. A computer-readable medium may comprise a computer-readable storage medium (e.g., memories 625 or other device) that may be any media or means that can contain, store, and / or transport the instructions for use by or in connection with an instruction execution system, apparatus, or device, such as a computer. A computer-readable storage medium does not comprise propagating signals, and therefore may be considered to be non-transitory. The term “non-transitory”, as used herein, is a limitation of the medium itself (i.e., tangible, not a signal) as opposed to a limitation on data storage persistency (e.g., RAM, random access memory, versus ROM, read-only memory).
[0211] If desired, the different functions discussed herein may be performed in a different order and / or concurrently with each other. Furthermore, if desired, one or more of the above-described functions may be optional or may be combined.
[0212] Although various aspects of the invention are set out in the independent claims, other aspects of the invention comprise other combinations of features from the described embodiments and / or the dependent claims with the features of the independent claims, and not solely the combinations explicitly set out in the claims.
[0213] It is also noted herein that while the above describes example embodiments of the invention, these descriptions should not be viewed in a limiting sense. Rather, there are several variations and modifications which may be made without departing from the scope of the present invention as defined in the appended claims.
[0214] The following abbreviations that may be found in the specification and / or the drawing figures are defined as follows:
[0215] 2D two dimensional
[0216] 3D three dimensional
[0217] 3 GPP third generation partnership project
[0218] a.k.a. also known as
[0219] AOM Alliance for Open Media
[0220] HEIF High Efficiency Image File Format
[0221] HEVC High Efficiency Video Coding
[0222] ID identification
[0223] IEC International Electrotechnical Commission
[0224] ISO International Organization for Standardization
[0225] ISOBMFF ISO base media file format
[0226] ITU-T Telecommunications Standardization Sector of International Telecommunication Union
[0227] JCT-VC Joint Collaborative Team - Video Coding
[0228] JPEG Joint Photographic Experts Group
[0229] JVT joint video team
[0230] MP4 MPEG 4
[0231] MPEG Moving Picture Experts Group
[0232] MV-HEVC Multiview-HEV C
[0233] NAL network abstraction layer
[0234] PLY format Polygon File Format
[0235] PNG Portable Network Graphics
[0236] SEI supplemental enhancement information
[0237] SH spherical harmonics
[0238] SNR signal-to-noise ratio
[0239] V3C Visual Volumetric Video-based Coding
[0240] VCEG Video Coding Experts Group
[0241] VPS V3C parameter set
[0242] VSEI versatile supplemental enhancement information
[0243] VUI video usability information
[0244] WC Versatile Video Coding
[0245] XML extensible Markup Language
Claims
What is claimed is:
1. A method, comprising: projecting, by an encoder, a scene represented by three-dimensional gaussian splats into one or more two-dimensional representations, where the one or more two- dimensional representations store different parameters of the three-dimensional gaussian splats, along with associated metadata describing transformation from three-dimensional space into two-dimensional representations; forming, by the encoder, the one or more two-dimensional representations and the associated metadata into one or more bitstreams; and outputting, by the encoder, the one or more bitstreams.
2. The method according to claim 1, wherein forming comprises individually encoding the one or more two-dimensional representations using a corresponding static image encoder.
3. The method according to claim 1, wherein forming comprises encoding one or more sequences of two-dimensional representations using a video encoder.
4. The method according to claim 1, wherein the scene is an original three-dimensional scene, wherein values of the different parameters are specific to one or more sub-regions of the one or more two-dimensional representations and the original three-dimensional scene, and where the one or more sub-regions of the one or more two-dimensional representations are indicated in the associated metadata.
5. The method according to any one of claims 1 to 4, wherein the associated metadata is encoded by V3C and stored as atlas metadata.
6. The method according to any one of claims 1 to 4, wherein the associated metadata is stored as SEI messages.
7. The method according to any one of claims 1 to 6, wherein projecting comprises mapping two or more three-dimensional scenes represented with 3D gaussian splats with time relations to two or more two-dimensional images and associated metadata that are stored in a sequence creating a volumetric video.
8. The method according to any one of claims 1 to 7, wherein the two-dimensional representations comprise one or more of the following parameters: geometry; spherical harmonics, diffuse color; opacity; scaling; rotation; normals; or orientation.
9. The method according to claim 8, wherein two or more of the parameters are combined into a shared two-dimensional representation.
10. A method, comprising: receiving, by a decoder, one or more bitstreams encoding one or more two-dimensional representations and associated metadata, the one or more two-dimensional representations and associated metadata encoding a scene represented by three- dimensional gaussian splats, where the one or more two-dimensional representations store different parameters of the three-dimensional gaussian splats and the associated metadata stores transformation information from three- dimensional space into two-dimensional representations; applying, by the decoder, decoding to corresponding individual bitstreams to form the one or more two-dimensional representations, and the associated metadata; reprojecting, as described by the associated metadata and by the decoder, the one or more two-dimensional representations into the scene represented by the three- dimensional gaussian splats; and outputting, by the decoder, the scene for rendering, or render the scene, to a viewer.
11. The method according to claim 10, wherein applying comprises individually decoding the one or more two-dimensional representations using a corresponding static image decoder.
12. The method according to claim 10, wherein applying comprises decoding one or more sequences of two-dimensional representations using a video decoder.
13. The method according to claim 10, wherein the scene is an original three-dimensional scene, wherein values of the different parameters are specific to one or more sub-regions of the one or more two-dimensional representations and the original three-dimensional scene, and where the one or more sub-regions of the one or more two-dimensional representations are indicated in the associated metadata.
14. The method according to any one of claims 10 to 13, wherein the associated metadata is encoded by V3C stored as atlas metadata, and wherein the method further comprises extracting the associated metadata from the atlas metadata and decoding using V3C.
15. The method according to any one of claims 10 to 13, wherein the associated metadata are stored as SEI messages, and wherein the method further comprises extracting the associated metadata from the SEI messages for decoding.
16. The method according to any one of claims 10 to 15, wherein two or more three- dimensional scenes represented with gaussian splatting with time relations are mapped to two or more two-dimensional images and associated metadata that are stored in a sequence creating a volumetric video, and wherein applying decoding comprises extracting the two or more three-dimensional scenes represented by gaussian splatting from two or more two-dimensional images and associated metadata represented as a sequence of volumetric video.
17. The method according to any one of claims 10 to 16, wherein the two-dimensional representations comprise one or more of the following parameters: geometry; spherical harmonics, diffuse color; opacity; scaling; rotation; normals; or orientation.
18. The method according to claim 17, wherein the applying comprises decoding a shared two-dimensional representation having two or more of the parameters that have been combined into the shared two-dimensional representation.
19. A computer program, comprising instructions for performing the methods of any of claims 1 to 18, when the computer program is run on an apparatus.
20. The computer program according to claim 19, wherein the computer program is a computer program product comprising a computer-readable medium bearing instructions embodied therein for use with the apparatus.
21. The computer program according to claim 19, wherein the computer program is directly loadable into an internal memory of the apparatus.
22. An apparatus, comprising means for performing: projecting, by an encoder, a scene represented by three-dimensional gaussian splats into one or more two-dimensional representations, where the one or more two- dimensional representations store different parameters of the three-dimensional gaussian splats, along with associated metadata describing transformation from three-dimensional space into two-dimensional representations; forming, by the encoder, the one or more two-dimensional representations and the associated metadata into one or more bitstreams; and outputting, by the encoder, the one or more bitstreams.
23. The apparatus according to claim 22, wherein forming comprises individually encoding the one or more 2D representations using a corresponding static image encoder.
24. The apparatus according to claim 22, wherein forming comprises encoding one or more sequences of two-dimensional representations using a video encoder.
25. The apparatus according to claim 22, wherein the scene is an original three-dimensional scene, wherein values of the different parameters are specific to one or more sub-regions of the two-dimensional representations and the original three-dimensional scene, and where the one or more sub-regions of the 2D representations are indicated in the associated metadata.
26. The apparatus according to any one of claims 22 to 25, wherein the associated metadata is encoded by V3C and stored as atlas metadata.
27. The apparatus according to any one of claims 22 to 25, wherein the associated metadata is stored as SEI messages.
28. The apparatus according to any one of claims 22 to 27, wherein projecting comprises mapping two or more three-dimensional scenes represented with 3D gaussian splats with time relations to two or more two-dimensional images and associated metadata that are stored in a sequence creating a volumetric video.
29. The apparatus according to any one of claims 22 to 28, wherein the two-dimensional representations comprise one or more of the following parameters: geometry; spherical harmonics, diffuse color; opacity; scaling; rotation; normals; or orientation.
30. The apparatus according to claim 29, wherein two or more of the parameters are combined into a shared two-dimensional representation.
31. An apparatus, comprising means for performing: receiving, by a decoder, one or more bitstreams encoding one or more two-dimensional representations and associated metadata, the one or more two-dimensional representations and associated metadata encoding a scene represented by three- dimensional gaussian splats, where the one or more two-dimensional representations store different parameters of the three-dimensional gaussian splats and the associated metadata stores transformation information from three- dimensional space into two-dimensional representations; applying, by the decoder, decoding to corresponding individual bitstreams to form the one or more two-dimensional representations, and the associated metadata; reprojecting, as described by the associated metadata and by the decoder, the one or more two-dimensional representations into the scene represented by the three- dimensional gaussian splats; and outputting, by the decoder, the scene for rendering, or render the scene, to a viewer.
32. The apparatus according to claim 31, wherein applying comprises individually decoding the one or more two-dimensional representations using a corresponding static image decoder.
33. The apparatus according to claim 31, wherein applying comprises decoding one or more sequences of two-dimensional representations using a video decoder.
34. The apparatus according to claim 31, wherein values of the different parameters are specific to one or more sub-regions of the one or more two-dimensional representations and the original three-dimensional scene, and where the one or more sub-regions of the one or more two-dimensional representations are indicated in the associated metadata.
35. The apparatus according to any one of claims 31 to 34, wherein the associated metadata is encoded by V3C stored as atlas metadata, and wherein the means are furtherconfigured for performing extracting the associated metadata from the atlas metadata and decoding using V3C.
36. The apparatus according to any one of claims 10 to 34, wherein the associated metadata are stored as SEI messages, and wherein the means are further configured for performing extracting the associated metadata from the SEI messages for decoding.
37. The apparatus according to any one of claims 10 to 36, wherein two or more three- dimensional scenes represented with gaussian splatting with time relations are mapped to two or more two-dimensional images and associated metadata that are stored in a sequence creating a volumetric video, and wherein applying decoding comprises extracting the two or more three-dimensional scenes represented by gaussian splatting from two or more two-dimensional images and associated metadata represented as a sequence of volumetric video.
38. The apparatus according to any one of claims 10 to 37, wherein the two-dimensional representations comprise one or more of the following parameters: geometry; spherical harmonics, diffuse color; opacity; scaling; rotation; normals; or orientation.
39. The apparatus according to claim 38, wherein the applying comprises decoding a shared two-dimensional representation having two or more of the parameters that have been combined into the shared two-dimensional representation.
Citation Information
Cited By
3D Gaussian Splitting compression method
CN121193947A
Three-dimensional Gaussian splash reconstruction method for underwater scene
CN121639946A
Dynamic scene reconstruction method and device based on Gaussian point cloud
CN121904264A
V-DMC based coding of gaussian splats
WO2026098877A1