Method for encoding / decoding gaussian for three-dimensional space representation

WO2026205989A1PCT designated stage Publication Date: 2026-10-01ELECTRONICS & TELECOMM RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2026/004769
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-09-29
Filing Date
2026-03-25
Publication Date
2026-10-01

Smart Images

  • Figure KR2026004769_01102026_PF_FP_ABST
    Figure KR2026004769_01102026_PF_FP_ABST
Patent Text Reader

Abstract

A method for encoding Gaussian for three-dimensional space representation, according to the present disclosure, may comprise the steps of: converting attribute information of Gaussians into 2D images; and encoding the 2D images. Here, the attribute information includes rotation information, and information indicating a representation type of the rotation information may be encoded as metadata.
Need to check novelty before this filing date? Find Prior Art

Description

Method for encoding / decoding Gaussians for 3D spatial representation

[0001] The present disclosure relates to a method for encoding / decoding a Gaussian splat for a three-dimensional spatial representation and an apparatus for performing the same.

[0002] Virtual reality services are evolving in a direction that maximizes immersion and a sense of presence by generating omnidirectional video in the form of real-life images or computer graphics (CG) and playing it on HMDs, smartphones, etc. It is known that currently, to play natural and immersive omnidirectional video through an HMD, it must support 6 degrees of freedom (DoF). 6DoF video must provide free video in six directions through the HMD screen, including (1) left-right rotation, (2) up-down rotation, (3) left-right movement, and (4) up-down movement. However, most omnidirectional videos based on real-life images currently only support rotational movement. Accordingly, research is actively underway in fields such as the acquisition and reproduction technologies of 6DoF omnidirectional video.

[0003] The present disclosure aims to provide a method for encoding / decoding attribute information of Gaussians into a 2D image for three-dimensional spatial representation.

[0004] The present disclosure aims to provide a method for encoding / decoding Gaussians by similarly modifying the distribution of attribute information of Gaussians.

[0005] The present disclosure aims to provide a method for clustering Gaussians to increase the encoding / decoding efficiency of 2D images.

[0006] The technical problems to be solved in this disclosure are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this disclosure belongs from the description below.

[0007] A method for encoding a Gaussian for a three-dimensional spatial representation according to the present disclosure may include the step of converting attribute information of Gaussians into 2D images; and the step of encoding the 2D images. In this case, the attribute information may include rotation information, and information indicating the representation type of the rotation information may be encoded as metadata.

[0008] In a method for encoding a Gaussian for a three-dimensional space representation according to the present disclosure, the rotation information may include three rotation information parameters and one Wigner D parameter.

[0009] In a method for encoding a Gaussian for a three-dimensional space representation according to the present disclosure, the three Euler angle parameters can be expressed in the same coordinate axis order as the Wigner D parameters.

[0010] In a method for encoding a Gaussian for a three-dimensional space representation according to the present disclosure, Gaussians with identical rotation information are classified into one group, and for said group, said rotation information can be encoded.

[0011] In a method for encoding a Gaussian for a three-dimensional spatial representation according to the present disclosure, information indicating whether rotation compensation has been performed on the Gaussian may be encoded as metadata.

[0012] In a method for encoding a Gaussian for a three-dimensional spatial representation according to the present disclosure, the attribute information includes position information, and the upper bits of the position information may be encoded through a first 2D image, and the lower bits may be encoded through a second 2D image.

[0013] In a method for encoding a Gaussian for a three-dimensional spatial representation according to the present disclosure, the attribute information includes location information, and the location information may include a first location information indicating the location of a local coordinate system to which the Gaussian belongs and a second location information indicating the location of the Gaussian within the local coordinate system.

[0014] In a method for encoding a Gaussian for a three-dimensional spatial representation according to the present disclosure, the first position information may represent an index of the local coordinate system.

[0015] In a method for encoding a Gaussian for a three-dimensional spatial representation according to the present disclosure, a mapping table defining a mapping relationship between the index of each of the local coordinate systems and the coordinates of the local coordinate system within the world coordinate system may be encoded as metadata.

[0016] In a method for encoding a Gaussian for a three-dimensional spatial representation according to the present disclosure, the first position information may represent the coordinates of the local coordinate system within the world coordinate system.

[0017] In a method for encoding a Gaussian for a three-dimensional spatial representation according to the present disclosure, the 2D image is divided into a plurality of blocks, and Gaussians with high similarity of AC coefficients of spherical harmonic functions can be packed into one block.

[0018] In a method for encoding a Gaussian for a three-dimensional spatial representation according to the present disclosure, the similarity of the AC coefficients can be determined based on the result of performing clustering on the AC coefficients.

[0019] In a method for encoding a Gaussian for a three-dimensional spatial representation according to the present disclosure, offset information is encoded for each of the blocks, and samples belonging to the blocks may be encoded in a state of difference or addition by an offset specified by the offset information.

[0020] In a method for encoding a Gaussian for a three-dimensional spatial representation according to the present disclosure, the attribute information may further include at least one of position change information and duration information of the Gaussian.

[0021] In a method for encoding a Gaussian for a three-dimensional spatial representation according to the present disclosure, Gaussians determined to be outliers are packed in a specific region of the 2D image, and information indicating the number of outlier Gaussians belonging to the specific region may be encoded.

[0022] A method for decoding a Gaussian for a three-dimensional spatial representation according to the present disclosure may include: decoding 2D images; restoring attribute information of Gaussians from the decoded 2D images; and rendering a target viewpoint image based on the attribute information of the restored Gaussians. In this case, the attribute information may include rotation information, and information indicating the representation type of the rotation information may be decoded as metadata.

[0023] According to the present disclosure, a computer-readable recording medium may be provided that records instructions for executing an image encoding method / image decoding method.

[0024] The technical problems to be solved in this disclosure are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art from the description below.

[0025] According to the present disclosure, a method for encoding / decoding attribute information of Gaussians into a 2D image for three-dimensional spatial representation may be provided.

[0026] According to the present disclosure, the encoding / decoding efficiency of Gaussians can be improved by similarly changing the distribution of attribute information of Gaussians.

[0027] According to the present disclosure, by clustering Gaussians, the encoding / decoding efficiency of 2D images can be increased.

[0028] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs from the description below.

[0029] Figure 1 shows multiple images captured using cameras at different viewpoints.

[0030] Figure 2 illustrates a method for removing duplicate data between multiple viewpoint images.

[0031] Figure 3 illustrates an example of capturing an object in three-dimensional space using multiple cameras at different locations.

[0032] Figure 4 illustrates a unit cell.

[0033] Figure 5 shows the incidence pattern of rays on reference points.

[0034] Figure 6 shows cases where the distribution of information represented by a sphere differs depending on the order and degrees of freedom of the spherical harmonic function.

[0035] Figure 7 illustrates the process of rasterizing color information of vertices into a target viewpoint image when the sizes of vertices located in three-dimensional space are different.

[0036] FIG. 8 shows the configuration of a Gaussian encoder and a Gaussian decoder according to the present disclosure.

[0037] Figure 9 shows an example of converting the properties of 3D Gaussians into multiple 2D images.

[0038] Figure 10 illustrates an example in which a single Gaussian is represented by multiple attributes in three-dimensional space.

[0039] Figure 11 shows the spherical harmonic function coefficients of three Gaussians as a bar graph.

[0040] Figure 12 illustrates a spherical Gaussian before rotation and a spherical Gaussian after rotation.

[0041] Figure 13 shows an example where the distribution of spherical harmonic function coefficients among Gaussians is similarly adjusted through the Wigner D function.

[0042] Figure 14 shows an example of expressing rotational information through Euler angles.

[0043] Figure 15 shows an example of dividing the world coordinate system of a target scene into local coordinate systems.

[0044] Figure 16 illustrates an example of representing Gaussian position information as a 4-channel image or two 3-channel images.

[0045] Figure 17 shows an example of applying a clustering algorithm to AC coefficients.

[0046] Figure 18 illustrates the change in the Gaussian over time.

[0047] FIG. 19 is a flowchart of a process for encoding / decoding Gaussian attribute information according to one embodiment of the present disclosure.

[0048] The present disclosure is subject to various modifications and may have various embodiments, and specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present disclosure to specific embodiments, and it should be understood that it includes all modifications, equivalents, and substitutions that fall within the spirit and scope of the present disclosure. Similar reference numerals in the drawings refer to the same or similar functions across various aspects. The shapes and sizes of elements in the drawings may be exaggerated for clearer explanation. The detailed description of exemplary embodiments described below refers to the accompanying drawings, which illustrate specific embodiments as examples. These embodiments are described in sufficient detail to enable those skilled in the art to practice the embodiments. It should be understood that various embodiments are different but need not be mutually exclusive. For example, specific shapes, structures, and characteristics described herein may be implemented in other embodiments without departing from the spirit and scope of the present disclosure in relation to one embodiment. It should also be understood that the location or arrangement of individual components within each disclosed embodiment may be changed without departing from the spirit and scope of the embodiment. Accordingly, the following detailed description is not intended to be taken in a limiting sense, and the scope of exemplary embodiments is limited only by the appended claims, together with all equivalents to those claimed therein, provided they are properly described.

[0049] In this disclosure, terms such as first, second, etc. may be used to describe various components, but said components should not be limited by said terms. Such terms are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of this disclosure, the first component may be named the second component, and similarly, the second component may be named the first component. The term "and / or" includes a combination of a plurality of related described items or any of a plurality of related described items.

[0050] Where it is stated that any component of the present disclosure is "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, or that there may be other components in between. On the other hand, where it is stated that a component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.

[0051] The components shown in the embodiments of the present disclosure are depicted independently to represent different characteristic functions and do not imply that each component consists of separate hardware or a single software unit. That is, each component is listed and included as a separate component for convenience of explanation; however, at least two of the components may be combined to form a single component, or a single component may be divided into multiple components to perform a function, and such integrated and separated embodiments of each component are included within the scope of the rights of the present disclosure as long as they do not depart from the essence of the present disclosure.

[0052] The terms used in this disclosure are used merely to describe specific embodiments and are not intended to limit this disclosure. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this disclosure, terms such as "comprising" or "having" are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof. That is, the description in this disclosure that a specific configuration "comprising" does not exclude configurations other than that configuration, but means that additional configurations may be included within the scope of the practice or technical concept of this disclosure.

[0053] Some components of the present disclosure may not be essential components performing an essential function in the present disclosure, but may be optional components merely for enhancing performance. The present disclosure may be implemented by including only the components essential to embody the essence of the present disclosure, excluding components used merely for enhancing performance, and a structure including only the essential components, excluding optional components used merely for enhancing performance, is also included within the scope of the rights of the present disclosure.

[0054] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In describing the embodiments of this specification, if it is determined that a detailed description of related known configurations or functions may obscure the gist of this specification, such detailed description is omitted, and the same reference numerals are used for identical components in the drawings, and redundant descriptions of identical components are omitted.

[0055] Figure 1 shows multiple images captured using cameras at different viewpoints.

[0056] If ViewC1 is the central viewpoint, ViewL1 and ViewR1 represent the left viewpoint image and the right viewpoint image of the central viewpoint, respectively.

[0057] When creating a virtual viewpoint image ViewV between the center viewpoint image ViewC1 and the left viewpoint image ViewL1, there may be areas that are obscured and invisible in the center viewpoint image ViewC1 but visible in the left viewpoint image ViewL1. Accordingly, image synthesis for the virtual viewpoint image ViewV can be performed by referencing not only the center viewpoint image ViewC1 but also the left viewpoint image ViewL1.

[0058] Figure 2 illustrates a method for removing duplicate data between multiple viewpoint images.

[0059] Among multiple viewpoint images, a primary viewpoint is selected, and for images other than the primary viewpoint, duplicate data with the primary viewpoint is removed through pruning. For example, if the central viewpoint ViewC1 is designated as the primary viewpoint, the remaining views excluding ViewC1 become additional views used as reference images during synthesis. By utilizing the 3D geometry and depth information (depth map) of each viewpoint image, all pixels of the primary viewpoint image can be mapped to the locations of the additional viewpoint images. At this time, the mapping can be performed through a 3D view warping process.

[0060] For example, as shown in the example illustrated in FIG. 2, a first warped image (Warped-view C1 onto L2 position) is generated by mapping the basic viewpoint image ViewC1 to the position of the first left viewpoint image ViewL1, and a second warped image (Warped-view C1 onto L1 position) can be generated by mapping the basic viewpoint image ViewC1 to the position of the second left viewpoint image ViewL2.

[0061] At this time, areas that were not visible in the original viewpoint image ViewC1 due to parallax are treated as hole areas with no data in the warped image. Areas with data (i.e., color) excluding the hole areas may be areas that are also visible in the original viewpoint image ViewC1.

[0062] A pruning process to remove redundant pixels can be performed by going through a procedure to determine whether overlapping pixels between a base viewpoint and an additional viewpoint can be determined as redundancy. For example, as shown in the example illustrated in FIG. 2, a first residual image (Residual Image 1) can be generated through pruning between a first warped image and a first left viewpoint image, and a second residual image (Residual Image 2) can be generated through pruning between a second warped image and a second left viewpoint image. By reducing image data through the pruning process, compression efficiency can be increased during image encoding.

[0063] Meanwhile, the determination of overlapping pixels may be based on whether at least one of the color value difference and / or depth value difference for pixels at the same location is smaller than a threshold. For example, if at least one of the color value difference or the depth value difference is smaller than a threshold, the two pixels may be determined to be overlapping pixels.

[0064] At this time, due to issues such as color or depth value noise within the image, errors in camera calibration values, or errors in the judgment formula, there may be cases where pixels are judged as duplicate pixels even though they are not. Additionally, even among pixels at the same location, color values ​​may differ depending on the camera position used to capture the pixel due to the characteristics of the light source and various reflective surfaces within the scene. Consequently, even if the pruning process is very accurate, loss of information representing the scene may occur, which can cause image quality degradation when rendering the target viewpoint image in the decoder.

[0065] Figure 3 illustrates an example of capturing an object in three-dimensional space using multiple cameras at different locations.

[0066] In Fig. 3(a), it was assumed that each image is projected onto a two-dimensional image.

[0067] In FIG. 3(a), V1 to V6 represent viewpoint images captured by cameras with different shooting angles (poses). As illustrated in the example, depending on the position and shooting angle (pose) of the camera acquiring the object, the way it is projected into a 2D image differs even for the same point in 3D space. For example, when an arbitrary point (302) on an object is projected onto each of the viewpoint images V1 to V6, depending on the camera shooting angle, the pixel value corresponding to the arbitrary point (302) in the projected 2D image may not be the same between the viewpoint images and may differ.

[0068] Likewise, the object (301) shown in FIG. 3 (a) may also have different brightness depending on the viewpoint due to the characteristics of the light source and the reflective surface.

[0069] However, if pixels corresponding to any point (302) on an object within the viewpoint images are determined to be duplicate pixels, a pruning process is performed so that pixels in the basic viewpoint image are retained and pixels in the additional viewpoint image are removed. That is, even if pixels corresponding to any point (302) on an object within the viewpoint images have different brightness, if the difference in depth value (or color value) is less than or equal to a threshold value, they are determined to be duplicate pixels.

[0070] The pruning process improves data compression efficiency by removing data redundancy; however, as in the example above, it identifies pixels with different brightness levels as duplicates, causing a loss of information, which leads to image quality degradation during rendering in the decoder.

[0071] In particular, color values ​​on transparent objects where an incident light source is refracted, or on surfaces such as mirrors where an incident light source is totally reflected, rather than on diffusely reflective surfaces, may be identified as duplicate pixels and removed during the pruning process even though the color values ​​are completely different depending on the angle.

[0072] To reproduce the color values ​​of an actual mixed reflective surface that appear differently depending on the observer's viewing position and angle, one can consider a method of modeling the reflection characteristics of the mixed reflective surface so that it possesses information when viewed from all angles, not just a specific angle, or can estimate information when viewed from a specific angle.

[0073] Recently, there has been active discussion regarding deep learning-based image processing methods. For example, instead of the traditionally used depth map-based rendering, a technique that models the radiance field by receiving multiple images of a target scene or target 3D space as input is gaining attention. Here, the modeling of the radiance field or Gaussian splats can be performed by inputting multiple images into a neural network.

[0074] Meanwhile, a radiation field represents a function or data structure that describes the characteristics of light for every point in three-dimensional space. For example, the characteristics of light may be how incident light is reflected as it passes through each point.

[0075] The technology for modeling a radiation field in three-dimensional space can be called NeRF (Neural Radiance Field).

[0076] When using NeRF, it becomes possible to more realistically reproduce non-Lambertian regions, which are difficult to express with traditional image synthesis methods. In addition, existing complex mathematical formulas or algorithms can be replaced with deep learning processes, that is, neural network learning processes.

[0077] In addition to NeRF, explicit feature information can also be utilized. Specifically, a scene in the target space can be represented in a form such as a voxel grid, and feature information capable of representing the voxel grid can be computed through the model training process.

[0078] For example, FIG. 3(a) shows an example in which a 3D region of interest to which an object within a target scene belongs is represented as a 3D grid structure. In the example illustrated in FIG. 3(a), the space containing the object (301) is illustrated as being represented as a 3D grid structure expressed by the world coordinate system. Here, the 3D grid structure refers to a cluster of 3D vertices arranged at equal intervals, and for example, reference numeral 311 is an approximation of one of the 3D vertices into a sphere shape. As in the example illustrated in FIG. 3(b), any point (302) represents a 3D vertex corresponding to any intermediate position within the 3D grid structure. Meanwhile, each of the multiple vertices can be referred to as a voxel.

[0079] When the target space is configured as a 3D grid in voxel units, a feature vector containing color and density information of the corresponding area may be assigned to each vertex.

[0080] A feature vector for a 3D vertex at any location in 3D space can be calculated by trilinear interpolating the feature vectors of neighboring vertices (e.g., the 8 vertices that make up a hexahedron containing the point at any location). That is, color and density information for a 3D vertex at any location can be obtained through trilinear interpolation of the feature vectors of neighboring vertices.

[0081] NeRF technology using explicit feature information requires representing the target space as a 3D grid, which leads to an increase in data volume. This increase in data volume may result in constraints on model file storage and inference. Accordingly, a structure for encoding / decoding feature information of the target space using a server-client model can be considered.

[0082] Below, we will examine in detail embodiments for encoding / decoding feature information of a target space.

[0083] Figure 4 illustrates a unit cell.

[0084] In FIG. 4, it is illustrated that a ray is unprojected from each pixel of the viewpoint images (V1 to V6) to a vertex (401) constituting a unit grid. As illustrated in the example, camera calibration information corresponding to the viewpoint image can be used to unproject from the pixels of the viewpoint image to a vertex (401) constituting a unit grid in the form of a ray. At this time, if the color value of the ray projected from each of the viewpoint images to the vertex (401) is referenced, the color value of the vertex (401) for any viewpoint can be estimated.

[0085] Furthermore, if the color values ​​of the eight vertices constituting the grid (400) can be estimated by referencing the pixels within the viewpoint images (V1 to V6), then at least one value among the color value, brightness value, or opacity of any point (402) within the grid can also be estimated. That is, by using the eight vertices constituting the grid as reference vertices, information about the target point (402) within the grid can be estimated by using methods such as tri-linear interpolation, averaging, or weighting of the reference vertices.

[0086] Meanwhile, in the example illustrated in FIG. 4, each vertex defining the unit cell can be extended into a form having a three-dimensional volume. For example, it can be defined in a three-dimensional shape, such as a sphere or an ellipsoid, centered around a vertex. To do this, size information for each of the vertices constituting the cell may be required.

[0087] Here, when the scaling of each vertex constituting the grid is the same, information about the target point can be estimated through simple methods such as 3D linear interpolation. On the other hand, when the sizes of the vertices constituting the grid are different from each other, when the vertex (801) is projected onto the target viewport, the size component can be modeled as an additional parameter so that the occupied area (or space) when the color intensity information of the ray projected onto the target viewport is rasterized can be varied. Through this, the representation parameters can be optimized so that the target scene can be rendered with the highest quality.

[0088] That is, based on size information, weights are set based on the occupancy at the location where each vertex is rasterized by projecting it onto the target viewpoint image at its respective size, and by undergoing a process of optimizing the final determined pixel based on the loss function with the pixel value of the reference viewpoint image through weighting operations of the vertices projected to the same location, the component information of the target point (401) constituting the voxel can be estimated.

[0089] Meanwhile, the size information can basically represent the radius of a circle or a sphere. In this case, for a vertex existing in 3D space, the radius for each of the x-axis, y-axis, and z-axis may be set individually. If the radius of at least one of the x-axis, y-axis, and z-axis differs from the other axes, it indicates that the shape of the vertex is an ellipsoid. By applying the above method to a 3D grid cluster, a target object can be reconstructed at any given viewpoint. Meanwhile, the denser the spacing of the vertices constituting the 3D grid cluster surrounding the target object, the higher the resolution at which the target object can be reconstructed.

[0090] In order to reproduce the target object in the way described above, for all reference vertices constituting the 3D grid cluster for the target object, color values ​​according to the angle of incidence (i.e., shooting angle) of the ray back-projected from each camera must be known.

[0091] Meanwhile, the number of lines passing through a reference vertex in the form of a ray may vary depending on at least one of the number of cameras (i.e., viewpoint images), the resolution of the images, or the camera geometry.

[0092] The more varied the angles of the rays incident from the cameras (i.e., viewpoint images) to the reference point, the more accurately the color values ​​of the target point can be reproduced according to the angle of incidence or orientation. In other words, the more reflected light information there is (i.e., the more reflected light information of the light source is obtained from various angles) that represents the information when the light source reflected from the target point is projected onto each camera (i.e., each viewpoint image), the more realistically the target point can be reproduced from various viewpoints and orientations.

[0093] As shown in the example illustrated in FIG. 3(a), if the reference vertex is assumed to be in the form of a sphere (311) with radius r, the color value at the moment when the ray passes through the reference vertex and reflects can be stored as reflected light information. Meanwhile, the reflected light information can be stored according to the incident angle (direction) of the ray.

[0094] By utilizing reflected light information, when synthesizing images from an arbitrary viewpoint, it becomes possible to reproduce appropriate colors depending on the angle at which the corresponding reference vertex is observed.

[0095] Meanwhile, the number of rays incident on a reference point by back-projection from the viewpoint image may vary depending on the reference point.

[0096] Figure 5 shows the incidence pattern of rays on reference points.

[0097] In the example illustrated in FIG. 5(a), five rays are incident on the first reference point, and in FIG. 5(b), three rays are incident on the second reference point. Since the number of rays incident on the first reference point is greater than the number of rays incident on the second reference point, the incident light source information for the first reference point can be understood as being more diverse than the incident light source information for the second reference point. Accordingly, when synthesizing images at any viewpoint, the first reference point can be reproduced with colors of higher realism from a wider variety of angles compared to the second reference point.

[0098] However, the maximum number of light rays incident on each of the first reference point and the second reference point is limited by the number of viewpoint images (i.e., cameras), and accordingly, incident light source information can be obtained only for the incident angle corresponding to each of the viewpoint images. That is, since incident light source information is not obtained for all angles in the omnidirectional direction, information regarding any orientation (angle) where no light rays are incident can be obtained through an approximation using peripheral values.

[0099] That is, information such as the color of a ray reflected from a reference point is set as a peripheral value, and at least one peripheral value is used to estimate the color value in an orientation or space where the ray is not incident.

[0100] Meanwhile, if the reference vertex is assumed to be in the shape of a sphere, the distribution of reflected light intensity on the sphere can be approximated based on peripheral values ​​using Laplace's equation in a spherical coordinate system. For example, the distribution of reflected light intensity can be approximated using spherical harmonics.

[0101] The following mathematical formula 1 represents the spherical harmonic function.

[0102]

[0103] Y in mathematical formula 1 l,m represents a spherical harmonic function. θ is the angle with respect to the positive z-axis in the spherical coordinate system, and φ is the angle with respect to the positive x-axis with respect to the z-axis. Since the function is continuous, l is a non-negative integer, and m is an integer satisfying -l ≤ m ≤ l.

[0104] In mathematical formula 1, c l,m It can be derived according to the following mathematical formula 2.

[0105]

[0106] Also, in Equation 1, P l m represents Legendre polynomials.

[0107] A spherical harmonic function that approximates the distribution of reflected light components at a spherical target point When saying, is the spherical harmonic function (basis function) Y of the reference points as shown in the following mathematical equation 3. l,m It can be expressed as the weighted sum of.

[0108]

[0109] Figure 6 shows cases where the distribution of information represented by a sphere differs depending on the order and degrees of freedom of the spherical harmonic function.

[0110] In the example illustrated in FIG. 6, assuming the order of the spherical harmonic function is 2, the spherical harmonic function of the target point can be approximated by 9 basis functions, which are the sum of information that can be expressed at an order of 2 or less (1 when the order is 0, 3 when the order is 1, and 5 when the order is 2). When the order of the spherical harmonic function is 3, the spherical harmonic function of the target point can be approximated by 16 basis functions, and when the order is 4, by 25 basis functions. Here, the number of basis functions may be the same concept as the number of coefficients constituting the spherical harmonic function.

[0111] As the order of the spherical harmonic function increases, the information regarding reflected light components corresponding to a local region in the spherical coordinate system can be approximated while distinguishing them from other regions. In other words, as the order of the spherical harmonic function increases, the high-frequency components of the reflected light expressed in the local region in the spherical coordinate system are included.

[0112] By referring to the intensity of the light rays incident on a spherical target point, the intensity of the reflected light component information that the sphere can represent can be approximated by direction. Specifically, when the order of the spherical harmonic function is 2, the spherical harmonic function for the target point can be approximated by calculating a weight function for a total of 9 basis functions.

[0113] At this point, the basis function Y in mathematical equation 3 l,m If information regarding the coefficients corresponding to the weights is available, the target point can be approximated using this information, and accordingly, the reflected light information at any orientation can be restored. To approximate the intensity of the R, G, and B trichromatic circles, the weights of the basis functions must be calculated by individually referencing the intensity of each channel.

[0114] In the embodiment of FIG. 3, it was assumed that the size information of all vertices (i.e., radius r) is the same. However, during the process of learning the target 3D scene, the sizes of the vertices may be set differently.

[0115] Figure 7 illustrates the process of rasterizing color information of vertices into a target viewpoint image when the sizes of vertices located in three-dimensional space are different.

[0116] In the process where vertices are projected onto the target viewpoint image and synthesized, the color information of the vertices can be rasterized into the target viewpoint image.

[0117] Specifically, in FIG. 7, the first vertex (701) is exemplified as a spherical shape (711) with radius r. On the other hand, the second vertex (702) is exemplified as an ellipsoidal shape (712) with radii for each of the x-axis, y-axis, and z-axis individually set.

[0118] A three-dimensional spherical shape (711) can be modeled using the directional color information (e.g., spherical harmonic coefficients) and size information held by the first vertex (701). Similarly, an ellipsoid (712) can be modeled using the directional color information and size information held by the second vertex (702). Meanwhile, in the case of the ellipsoid, not only size information for each axis but also rotation information for each axis can be set. Accordingly, the ellipsoid can be modeled so that each axis faces a different orientation.

[0119] Vertices in the shape of a sphere or an ellipsoid that follow a Gaussian distribution can be referred to as Gaussian or Gaussian splats.

[0120] The spherical shape (711) and the ellipsoid (712) are geometrically modeled as a probability distribution within a three-dimensional shape (e.g., within a range set by radius r) of directional color information (e.g., spherical harmonic function coefficients) corresponding to each direction (azimuth angle) based on the center of the three-dimensional shape. Meanwhile, the directional color information may also be referred to as directional intensity information.

[0121] If the probability distribution follows a Gaussian distribution, the Gaussian distribution can be defined by the following mathematical formula 4.

[0122]

[0123] In the above mathematical equation 4, Σ represents the 3D covariance matrix for a 3D Gaussian distribution. A 3D Gaussian distribution can be referred to as a Gaussian or a Gaussian splat.

[0124] When a 3D Gaussian, such as a sphere (711) or an ellipsoid (712), is projected onto a target viewport image (700), it is rasterized to form a 2D circle (721) and a 2D ellipse (722) shape.

[0125] In other words, the color intensity information of the ray radiated from the center of the three-dimensional shape is modeled as a three-dimensional Gaussian, such as a sphere (711) or an ellipsoid (712), by a Gaussian probability distribution. Additionally, when the three-dimensional Gaussian is projected onto a target two-dimensional image (700), it is rasterized into a two-dimensional circle or a two-dimensional ellipse, and the target viewpoint image can be synthesized.

[0126] When different Gaussians are projected onto pixels at the same location, the value of the pixel within the corresponding overlapping area (730) can be derived through a weighted sum operation using the occupancy information (or opacity) of the Gaussians projected onto the same location as weights.

[0127]

[0128] Equation 5 above represents the process of calculating the covariance matrix in the camera coordinate system given the viewing transformation matrix W. Equation 5 is used when a 3D Gaussian is projected onto a 2D image.

[0129] J represents the Jacobian matrix obtained by affine approximation of the perspective transformation. A 2 x 2 variance matrix can be derived from the Jacobian matrix.

[0130] In mathematical equation 5, Σ is a 3-dimensional covariance matrix, and given the scaling matrix and rotation matrix, it can be derived through the following mathematical equation 6.

[0131]

[0132] In mathematical equation 6, the magnitude matrix can be expressed as a 3-dimensional vector, and the rotation matrix can be expressed as a quaternion.

[0133] When modeling the Gaussian probability distribution of a light source radiating to the center of a three-dimensional shape using the described mathematical formulas, attributes such as directional color information, magnitude information, and / or rotation information may be required.

[0134] According to Equation 6, once the covariance matrix in the camera coordinate system is calculated, the color value of each pixel in the target viewpoint image can be determined through alpha blending (α-blending) between Gaussians. Specifically, the color value of a specific pixel can be determined through alpha blending between Gaussians that are superimposed and projected onto the specific pixel. Equation 7 shows an example in which the color value C is derived through alpha blending between Gaussians.

[0135]

[0136] In mathematical equation 7, c i and α i Each can represent the color value and occupancy information of the i-th Gaussian. N represents the number of Gaussians superimposed on a specific pixel.

[0137] Here, the directional color information may be spherical harmonic function coefficients. The directional color information may be expressed in a vector format or a hash code, or in the form of a matrix or feature vector that constitutes a Multi-Layer Perceptron (MLP) neural network learned by a deep learning-based algorithm.

[0138] Size information and / or rotation information can also take the form of a matrix, feature vector, or MLP neural network matrix.

[0139] The position of a 3D vertex (i.e., a 3D Gaussian) corresponds to the center point of a 3D ellipsoid. Through size and rotation information, the shape of the 3D vertex is determined, and through a spherical harmonic function, the color information of the rays radiating from the center point of the 3D vertex can be restored according to direction.

[0140] That is, a 3D Gaussian can be defined by attributes such as position information, size information, rotation information, occupancy (or transparency) information, and spherical harmonic functions. Meanwhile, the attributes representing a 3D Gaussian can be referred to as Gaussian parameters or parameters.

[0141] Finding the optimal value of each parameter through the learning process requires a long computation time. Accordingly, the encoder can determine the optimal value of each parameter, encode the determined optimal value, and signal it to the decoder. Meanwhile, the attributes of the 3D Gaussian can be converted into a 2D image form for encoding / decoding. Alternatively, the attributes of the 3D Gaussian can be encoded / decoded in the form of metadata.

[0142] FIG. 8 shows the configuration of a Gaussian encoder and a Gaussian decoder according to the present disclosure.

[0143] Referring to FIG. 8, the Gaussian encoder includes a Gaussian information extraction unit (810), a Gaussian transformation unit (820), and an image encoding unit (830).

[0144] The Gaussian information extraction unit (810) can extract Gaussian information from the input image. The Gaussian information can represent the attributes of the Gaussians.

[0145] The Gaussian transform unit (820) can convert the attributes of the Gaussians into 2D image type or point cloud type data. Meanwhile, the Gaussian transform unit can generate information indicating the type of codec used to encode the attribute for each attribute.

[0146] Additionally, the Gaussian transform unit (820) can generate mapping information between the 2D image and the point cloud. The mapping information may define the mapping relationship between coordinates in the 2D image and vertices in the 3D image.

[0147] The image encoding unit (830) can encode 2D images, point clouds, and mapping information generated by the Gaussian transform unit (820).

[0148] The video encoding unit (830) may include a 2D image encoding unit for encoding a 2D image and a PCC encoding unit for encoding a point cloud.

[0149] Meanwhile, mapping information may be represented in the form of a 2D image and encoded through a 2D image encoding unit, or represented in the form of a point cloud and encoded through a PCC encoding unit.

[0150] The Gaussian decoder may include an image decoding unit (840), a Gaussian restoration unit (850), and a rendering unit (860).

[0151] The image decoding unit (840) decodes 2D images, point clouds, and mapping information. The image decoding unit (840) may include a 2D image decoding unit for decoding 2D images and a PCC decoding unit for decoding point clouds.

[0152] Meanwhile, mapping information may be represented in the form of a 2D image and decoded through a 2D image decoding unit, or represented in the form of a point cloud and decoded through a PCC decoding unit.

[0153] The Gaussian restoration unit (850) can restore the attributes of each Gaussian based on 2D images, point clouds, and mapping information.

[0154] For each attribute, information indicating which codec the attribute was decoded by can be decoded. Based on the above information, it can be determined whether the restored 2D image or the restored point cloud corresponds to which attribute.

[0155] The rendering unit (860) can render 3D space based on the restored Gaussians.

[0156] Directional color information, size information, rotation information, and occupancy information can be derived through gradient descent. Specifically, through gradient descent, the optimal value of each parameter can be derived by finding the point where the cost function with respect to the training data set for each parameter is minimized. The following Equation 8 is an example of a cost function combining L1 loss and the D-SSIM (Depth Structural Similarity Index) evaluation metric.

[0157]

[0158] In the above mathematical formula 8, the variable λ may be used to prevent overfitting or weight adjustment during the learning process. The value of the variable λ can be adjusted arbitrarily.

[0159] The position and magnitude information of a 3D Gaussian are each expressed by three variables (i.e., x, y, z), and the rotation information is expressed by four variables (i.e., quaternion form). In addition, the occupancy information of a 3D Gaussian can be expressed by one variable. Meanwhile, the occupancy information may also be referred to as transparency information.

[0160] Meanwhile, the number of variables for the spherical harmonic function varies depending on the order. For example, when the order of the spherical harmonic function is 2, 9 variables are used, and when the order of the spherical harmonic function is 3, 16 variables are required. On the other hand, to represent each of RGB, the coefficients of the spherical harmonic function must be expressed for R, G, and B respectively. For example, when the order of the spherical harmonic function is 3, a total of 48 variables are required for the RGB image.

[0161] Accordingly, a total of 59 variables are required to represent the above properties of a 3D Gaussian.

[0162] The number of Gaussians required to reconstruct a target scene can be determined during the process of learning the target scene. Assuming the number of Gaussians is 500,000, a total of 29.5 million (i.e., 500,000 * 59) variables are required to represent the 500,000 Gaussians. This is a data amount similar to that of a 2D image with a resolution of approximately 31 million pixels.

[0163] To efficiently compress a vast amount of data, each attribute information can be converted into a 2D image format.

[0164] Figure 9 shows an example of converting the properties of 3D Gaussians into multiple 2D images.

[0165] Figure 9(a) illustrates an example in which the coefficients of spherical harmonic functions are configured into a multi-channel 2D image format.

[0166] FIG. 9(b) illustrates an example in which occupancy information, size information, position information, and rotation information are configured in a 2D image format. Occupancy information may be configured as a 1-channel 2D image, and position information and size information may each be configured as a 3-channel 2D image. Rotation information may be configured as a 4-channel 2D image.

[0167] As shown in the example illustrated in Fig. 9, in order to convert Gaussian attribute information into a 2D image format for encoding / decoding, a preprocessing step may be required to convert the attribute information into a data structure with high video encoding efficiency. For example, preprocessing may be performed to enable prediction based on spatiotemporal axis similarity, or to remove visually unimportant data and / or represent highly redundant data with short bits.

[0168] Figure 10 illustrates an example in which a single Gaussian is represented by multiple attributes in three-dimensional space.

[0169] Figure 10 (a) visualizes a Gaussian expressed by spherical harmonic function coefficients in the form of a sphere. When magnitude information and rotation information are added to the Gaussian, the Gaussian can have an ellipsoidal shape facing any direction, as shown in the example illustrated in Figure 10 (b).

[0170] The coefficients of the spherical harmonic function approximate the intensity of the light source radiated in each direction relative to the center point of the Gaussian. Here, to increase the expressiveness of the target scene, size and rotation information can be utilized during the learning process.

[0171] For example, you can rotate the reference coordinate system of the Gaussian or set different magnitudes for each of the three axes.

[0172] FIG. 10(b) illustrates a Gaussian with a different magnitude compared to the Gaussian shown in FIG. 10(a), and FIG. 10(c) illustrates an example where the reference coordinate system of the Gaussian shown in FIG. 10(a) is rotated.

[0173] Figure 11 shows the spherical harmonic function coefficients of three Gaussians as a bar graph.

[0174] If the coefficients of the spherical harmonic functions are arranged in order of degree and coefficient index, a graph representing the strength of the coefficients of the spherical harmonic functions of each of the three Gaussians can be obtained, as shown in the example illustrated in (a) of Fig. 11.

[0175] Meanwhile, as shown in the example illustrated in FIG. 11 (a), each of the three Gaussians may have a different distribution. In this case, to increase the temporal or spatial similarity between the three Gaussians, the distributions of the second and third Gaussians can be adjusted to match the distribution of the first Gaussian. Specifically, an offset can be applied to the distributions of the second and third Gaussians, and the graph with the applied offset can be shifted.

[0176] For example, as shown in the example illustrated in FIG. 11 (b), the distribution of the second Gaussian without an offset applied can be shifted two steps to the left to make the distribution of the second Gaussian similar to the distribution of the first Gaussian. Additionally, the distribution of the third Gaussian with an offset applied can be shifted 14 steps to the left to make the distribution of the third Gaussian similar to the distribution of the first Gaussian.

[0177] In this way, compression efficiency can be increased by transforming the distribution of the attributes of Gaussians so that they have similarity among themselves.

[0178] Meanwhile, information regarding changes in the distribution can be encoded as metadata and signaled. Here, the change information may include at least one of offset or shift information. For example, in the example illustrated in FIG. 10 (b), the shift information of the first Gaussian may be 0, the shift information of the second Gaussian may be 2, and the shift information of the third Gaussian may be 14.

[0179] Figure 11 illustrates the coefficients of a spherical harmonic function, but the method of adjusting the distribution can also be applied to rotation information, size information, or position information.

[0180] Meanwhile, when rotation information is expressed as quaternions, four channels are used. On the other hand, when rotation information is expressed as Euler angles, three channels are used. In this case, by applying offsets between channels and / or redistributing sample values, the distribution between channels can be made similar.

[0181] Size information and location information each consist of three channels, and the distribution between channels can be made similar using the method described above.

[0182] By making the distribution between channels similar, encoding / decoding efficiency can be improved through inter-channel prediction.

[0183] As another example, the reference coordinate system can be rotated to ensure that the attribute information of Gaussians has spatiotemporal similarity.

[0184] For example, as shown in the examples illustrated in FIG. 10 (a) and (c), the Gaussian can be rotated around the center point of the Gaussian. As the reference coordinate system is rotated, the light source radiating from the sphere also rotates accordingly. In order to represent the light source radiating from the sphere in accordance with the rotated reference coordinate system, the coefficients of the basis functions responsible for the intensity in each direction must also be changed.

[0185] In other words, when a spherical coordinate system expressed by spherical harmonic functions is rotated, the rotational change of the signal / wave expressed through the sphere must also be calculated. This rotational change can be calculated using the WignerD function.

[0186] The Wigner D function is used to calculate the coefficients of modified spherical harmonic functions based on rotation of the spherical coordinate system. Equation 9 shows an example of how the coefficients of modified spherical harmonic functions are obtained based on the Wigner D matrix.

[0187]

[0188] In Equation 9, a represents the original spherical harmonic function coefficients, and D represents the Wigner D matrix. The Wigner D matrix may contain rotation information. Additionally, a' represents the spherical harmonic function coefficients after rotation.

[0189] Figure 12 illustrates a spherical Gaussian before rotation and a spherical Gaussian after rotation.

[0190] Figure 12 illustrates an example of a sphere having bright and dark band shapes for a specific direction, rotated 45 degrees around the Z-axis based on the Wigner D matrix.

[0191] Using the Wigner D function, you can calculate how the shape or distribution of a pattern changes when rotating a pattern on a sphere.

[0192] The rotation information of the Wigner D matrix can be expressed based on Euler angles. For example, the Wigner D matrix can be expressed as shown in the following mathematical equation 10.

[0193]

[0194] In the above mathematical formula 11, α represents Z-axis rotation, β represents Y-axis rotation, and γ represents Z-axis rotation. Also, e -ima represents the rotational phase generated when the Gaussian is rotated by α around the Z-axis. d l mm' (β) is a function that indicates how the basis of the spherical harmonic function mixes when the Gaussian is rotated by β around the Y-axis. The above function can be expressed as shown in the following mathematical equation 11.

[0195]

[0196] In Equation 11, l and m are used to define the spherical harmonic function. Specifically, l represents the degree, or complexity, of the function to be rotated, and m represents the order, or the index of the spherical harmonic function before rotation. That is, l represents the complexity of the oscillation pattern of the spherical harmonic function, and m can represent the number of times the spherical harmonic function oscillates in the latitude direction. And, m' represents the index of the spherical harmonic function after rotation.

[0197] In mathematical equation 11, when l is 1, d l (β) can be defined as shown in the following mathematical formula 12.

[0198]

[0199] In mathematical equation 10, E -im'γ represents the rotation phase that occurs when the Gaussian is rotated by γ around the Z-axis.

[0200] Rotation information of a Gaussian expressed as quaternions can be utilized as a parameter for the local alignment of SH coefficients. Specifically, by applying the Wigner-D function to the SH coefficients of a Gaussian aligned to a specific reference coordinate system, coefficients that reflect rotational characteristics (compensated coefficients) can be generated. That is, after deriving the coefficients of the spherical harmonic function based on a specific reference coordinate system, the coefficients of the spherical harmonic function in a rotated state can be derived through the Wigner-D function. The coefficients of the spherical harmonic function in a rotated state can be referred to as the compensated coefficients, and the process of rotating the Gaussian to obtain the compensated coefficients can be referred to as rotation transformation or rotation compensation.

[0201] If the geometric rotation of the Gaussian can be derived from positional information or the previous frame, the amount of encoded data can be reduced by omitting the encoding / decoding of explicit rotational information (i.e., quaternion data) and restoring the compensated SH coefficients using only Wigner-D-based rotational transformation. In other words, without Gaussian rotational information, the coefficients of the compensated spherical harmonic function of the Gaussian derived through Wigner-D-based rotational transformation, the magnitude information of the Gaussian, the occupancy information of the Gaussian, and the positional information of the Gaussian are encoded / decoded, and the Gaussian rotational information can be derived in the decoder.

[0202] Through the Wigner D function, the distribution of spherical harmonic function coefficients among Gaussians can also be adjusted to be similar.

[0203] Figure 13 shows an example where the distribution of spherical harmonic function coefficients among Gaussians is similarly adjusted through the Wigner D function.

[0204] The leftmost figure in Fig. 13 shows the strength of the spherical harmonic coefficients as a bar graph. The subsequent figures show the distribution of the spherical harmonic coefficients as the Gaussian rotation angle is increased by 15 degrees. Specifically, the second figure is an example when the rotation angle is 15 degrees, and the rightmost figure is an example when the rotation angle is 120 degrees.

[0205] As shown in the example illustrated in Fig. 13, during the process of increasing the rotation angle, there may be cases where the strength of each coefficient exists only within a specific range. The fact that the strengths of the coefficients fall within a specific range indicates that the strengths of the coefficients have high similarity.

[0206] In video compression techniques, compression efficiency increases as the similarity along the spatiotemporal axis increases; therefore, if the compensated spherical harmonic function coefficients in a state of high similarity are compressed, the encoding / decoding efficiency can be improved. Accordingly, the Gaussian can be rotated to increase the distribution of the spherical harmonic function coefficients, and then the rotated spherical harmonic function coefficients can be encoded / decoded.

[0207] Meanwhile, if spherical harmonic coefficients are converted into a 2D image format, they undergo a quantization process during 2D image format encoding. At this time, if the strengths of the spherical harmonic coefficients are modified to have similarity, the difference between the maximum and minimum values ​​is reduced, so the efficiency of the quantization process using the maximum and minimum values ​​can also be increased.

[0208] In order to recover the coefficients of a spherical harmonic function in a decoder, Gaussian rotation information is required. Accordingly, the encoder can encode and signal Gaussian rotation information as attribute information or metadata.

[0209] For example, it is assumed that for a specific Gaussian, the rotation angle is 80 degrees such that the coefficients of the spherical harmonic function have a highly similar distribution. In this case, the encoder can rotate the Gaussian by 80 degrees to encode the coefficients of the spherical harmonic function of the Gaussian, and also encode and signal information indicating that the rotation angle of the Gaussian is 80 degrees.

[0210] In the decoder, the coefficients of the spherical harmonic function of the Gaussian can be restored by rotating the restored Gaussian in reverse by 80 degrees.

[0211] Rotation information can be expressed in various ways. For example, rotation information can be expressed as a quaternion with four parameters.

[0212] Alternatively, rotation information can be expressed in three parameters through Euler angles.

[0213] Figure 14 shows an example of expressing rotational information through Euler angles.

[0214] As shown in the example illustrated in Fig. 14, Euler angles can represent any three-dimensional rotation as a combination of three coordinate axis rotations.

[0215] This can be expressed as a mathematical formula, as in Equation 13.

[0216]

[0217] If rotation information expressed by four parameters is expressed in Euler angles, the number of parameters expressing rotation information can be reduced by one. In this case, if the remaining parameter is used to express a rotation angle based on Wigner D, it is possible to encode / decode new attribute information while maintaining the number of parameters expressing rotation information at four.

[0218] In mathematical equation 13, as a method of decomposing rotation information R into three coordinate axis rotations, there can be 12 combinations depending on the order of the coordinate axes.

[0219] Meanwhile, the Wigner D function is defined based on the ZYZ order. Accordingly, when expressing the rotation information of each Gaussian in Euler angles, if it is expressed in the same ZYZ order as the Wigner D function, the Wigner D-based rotation information can also be used directly without conversion.

[0220] On the other hand, when rotation information is expressed using Euler angles rather than quaternions, there is a possibility of a Gimbal Lock problem occurring. To resolve this problem, rotation information can be encoded / decoded using the Rodrigues rotation representation. The Rodrigues representation can represent an arbitrary 3D rotation transformation using four values, as shown in the example of Equation 14. For example, as shown in Equation 14, the Rodrigues rotation representation is expressed in the form of rotating point p in 3D by θ about the rotation axis v.

[0221]

[0222] Equation 14 represents that a Gaussian p in three-dimensional space has been rotated by θ about the axis of rotation v. In Equation 14, the rotation angle θ can be expressed as in Equation 15 below.

[0223]

[0224] The Rodriguez rotation transform has a lower likelihood of gimbal lock compared to Euler angles. Accordingly, there are advantages to encoding rotation information by representing it with the Rodriguez rotation transform instead of Euler angles.

[0225] When expressing rotation information using the Wigner D function, instead of including rotation information as attribute information, rotation information can also be expressed globally. For example, when all Gaussians are reconstructed into one or two dimensions based on position information, a one-dimensional or two-dimensional index can be assigned to each Gaussian. For example, when Gaussians are expressed in one dimension, indices from 0 to N-1 can be assigned to the Gaussians. On the other hand, when Gaussians are expressed in two dimensions, indices according to the (x, y) image coordinate system can be assigned to the Gaussians.

[0226] After clustering Gaussians with the same rotation information, the rotation information can be encoded / decoded on a group basis. At this time, metadata mapping the rotation information to index information, or metadata mapping the Gaussian to the group to which the corresponding point belongs, can be additionally encoded / decoded.

[0227] In encoding rotation information based on the above Wigner D function, in addition to the rotation information, at least one of a flag indicating whether rotation compensation has been applied, information required for defining the Wigner D function, information indicating the order of coordinate axes of Euler angles, or information indicating the method of representing the rotation information (e.g., quaternions, Euler angles, or Rodriguez rotation transformation) may be encoded / decoded as metadata.

[0228] The position information of a Gaussian can be expressed by three parameters (i.e., x, y, z). In this case, each parameter can constitute an independent layer (i.e., an independent image). By configuring each of the three parameters as one channel, the position information can also be represented as a 3-channel image.

[0229] A 3D Gaussian position can represent the distance from the origin of the world coordinate system defined through the camera calibration process. Accordingly, when expressing Gaussian position information as pixel values ​​constituting a 2D image, the accuracy may vary depending on the bit depth of the image. For example, if the image bit depth is 8 bits, the position information can be expressed in 256 steps. Consequently, if the range of any one of the x, y, or z axes exceeds 256 steps, the resolution of the position information representation decreases.

[0230] Generally, position information in three dimensions is represented by floating-point numbers. Also, generally, floating-point numbers are represented by 32 bits.

[0231] To represent 32 bits using existing image standards, two 16-bit images can be used. Specifically, 32-bit position information can be represented in binary, with the upper 16 bits represented as the first image and the lower 16 bits represented as the second image.

[0232] To further reduce the number of bits per pixel, floating-point data can also be represented as a 16-bit short float.

[0233] In this case, the upper 8 bits of the 16 bits can be represented as the first image, and the lower 8 bits can be represented as the second image.

[0234] In the case of the first image representing the upper bit, the difference between the data is small compared to the second image representing the lower bit (i.e., the similarity is high), so the efficiency during encoding / decoding can be high.

[0235] The above method can be easily applied when Gaussian position information can be represented using the commonly used number of bits for images (i.e., 8, 10, 12, or 16 bits). On the other hand, when Gaussian position information is difficult to represent using the commonly used number of bits for images, it is necessary to undergo a linear or non-linear normalization process. However, if a normalization process is performed, a problem may arise where the accuracy of the position information decreases.

[0236] To resolve the above problem, one can consider dividing the world coordinates of the target scene into one or more local coordinates.

[0237] Figure 15 shows an example of dividing the world coordinate system of a target scene into local coordinate systems.

[0238] In Fig. 15, the origin is O wAn example in which the in-world coordinate system is divided into multiple local coordinate systems is illustrated. Additionally, in FIG. 15, the x-axis, y-axis, and z-axis are each divided into four segments. Since the x-axis, y-axis, and z-axis are each divided into four segments, the world coordinate system is divided into a total of 64 local coordinate systems.

[0239] Multiple Gaussians can exist within a local coordinate system. For example, the location of a specific Gaussian within the local coordinate system i j p (i.e., ( i j p x , i j p y , i j p z It can be expressed as )). Here, i represents the index of the local coordinate system within the world coordinate system, and j represents the index of the corresponding Gaussian within the local coordinate system. That is, i j p represents the position of the Gaussian at index j within the local coordinate system at index i.

[0240] Meanwhile, the Gaussian position within the world coordinate system i j It can be expressed as G.

[0241] The position of a Gaussian within a local coordinate system can be expressed as a combination of the position of the local coordinate system within the world coordinate system and the position of the Gaussian within the local coordinate system. For example, the position of the Gaussian i j G can be expressed as shown in the following mathematical equation 15.

[0242]

[0243] In mathematical formula 16, w i L represents the position of the origin of the local coordinate system relative to the origin Ow of the world coordinate system. i jp represents the position of the Gaussian within the local coordinate system.

[0244] When the world coordinate system of a target scene is divided into multiple local coordinate systems, a large-scale target scene can be represented more easily. Additionally, dividing the world coordinate system into multiple local coordinate systems can also provide advantages in generating 2D image specifications.

[0245] Figure 16 illustrates an example of representing Gaussian position information as a 4-channel image or two 3-channel images.

[0246] The position of the Gaussian within the local coordinate system is ( i j p x , i j p y , i j p z It can be expressed as ), and the position of the corresponding local coordinate system within the world coordinate system is ( w i L x , w i L y , w i L z It can be expressed as ).

[0247] In this case, instead of the position of the local coordinate system within the world coordinate system, the index of the local coordinate system can be encoded / decoded. For example, since the world coordinate system is divided into 64 local coordinate systems, the position of the local coordinate system can be in the range from 0 to 63.

[0248] In this way, if the world coordinate system is divided into multiple local coordinate systems and the position information of the local coordinate systems is converted into a separate 2D image format, the 2D image can be composed solely of the indices of the local coordinate systems to which the Gaussians belong. In this case, the 2D image contains a large amount of data with similar data, which results in the advantage of increased compression efficiency.

[0249] That is, as in the example illustrated in FIG. 16 (a), the position information of the Gaussians can be encoded / decoded through a 2D image comprising one channel indicating the index of the local coordinate system and three channels indicating the coordinates of the Gaussians within the local coordinate system.

[0250] At this time, based on a mapping table that defines the mapping relationship between the index of the local coordinate system and the origin position of the local coordinate system, the position of the local coordinate system within the world coordinate system ( w i L x , w i L y , w i L z It may also derive ). The corresponding mapping table may be encoded and transmitted as metadata. In addition, information indicating the minimum and maximum coordinate values ​​of the world coordinate system may be transmitted as separate metadata to determine the range of coordinate values ​​that the local coordinate system can represent.

[0251] As another example, the position in the local coordinate system is the x, y, and z axis coordinate values ​​in the actual world coordinate system (i.e., ( w i L x , w i L y , w i L z It can also be encoded / decoded based on ). In this case, as in the example shown in FIG. 16 (b), the position information of the Gaussian can be encoded / decoded through a 3-channel 2D image representing the position of the local coordinate system within the world coordinate system and a 3-channel 2D image representing the position of the Gaussian within the local coordinate system.

[0252] The position of the Gaussian within the corresponding local coordinate system ( i j p x , i j p y , ij p z In the case of ), the position of the corresponding Gaussian in the world coordinate system ( i j G x , i j G y , i j G z )Is, ( w i L x + i j p x , w i L y + i j p y , w i L z + i j p z It can be expressed as ).

[0253] If the location information of the local coordinate system is expressed as coordinates in the world coordinate system rather than an index, the location of the local coordinate system and Gaussian can be restored without a separate mapping table.

[0254] In the example illustrated in FIG. 15, the x-axis, y-axis, and z-axis are each divided into four segments. However, the number of divisions for each axis is not limited to the illustrated example. Depending on the distribution of Gaussians that optimally represent the target scene, the number of divisions for the x-axis, y-axis, and z-axis may differ. For example, in a specific scene, if the Gaussians are evenly distributed in the x-axis and y-axis directions, but are concentrated within a specific range in the z-axis direction (i.e., the variance is large in the x-axis and y-axis directions, while the variance is small in the z-axis direction), the number of divisions in the z-axis direction can be set smaller compared to the number of divisions in the x-axis and y-axis directions. That is, by setting the quantization intervals between the three-dimensional axes unevenly, the positional information resolution can be increased for a specific axis, while the positional information representation resolution can be lowered for another axis.

[0255] In the embodiments described above, the world coordinate system is exemplified as being equally divided for each of the three axes. Alternatively, the origins and ranges of multiple local coordinate systems can be adaptively determined based on the distribution of Gaussian clusters within the world coordinate system. In this case, there is an advantage in that the resolution of positional information of Gaussians located in a specific area within the target 3D space can be appropriately adjusted. Additionally, the index of the local coordinate system, positional information of the local coordinate system in the world coordinate system, minimum and maximum coordinate value information for each of the x-axis, y-axis, and z-axis of the local coordinate system, number of local coordinate systems, and minimum and maximum value information for normalizing the positional values ​​of each local coordinate system can be transmitted as separate metadata.

[0256] When expressing the position of a Gaussian by combining the position in a local coordinate system with the position of the Gaussian within that local coordinate system, it becomes possible to represent a scene on a larger scale than when not using a local coordinate system. Additionally, by increasing the positional resolution of the target scene, the position of the Gaussian can be expressed with greater detail.

[0257] Meanwhile, the bit depth of a 2D image for expressing position information of a local coordinate system (e.g., an index or coordinates of a local coordinate system within a world coordinate system) or position information of a Gaussian within a local coordinate system can be encoded and transmitted as separate metadata.

[0258] The rotation information of a Gaussian can also be expressed through the log-quaternion method. The log-quaternion method represents performing quaternion operations in a logarithm space. Specifically, the log-quaternion method may represent a quaternion on a non-linear manifold, Lie Group SO (3), as a 3-dimensional vector in a linear tangent space, Lie Algebra SO (3). Under the log-quaternion method, the rotation information can be expressed in the form of a scalar product of a rotation axis and a rotation angle.

[0259] Mathematical Equation 17 represents rotation information expressed in quaternion form.

[0260]

[0261] In the above mathematical formula 17, w represents the scalar part, and x, y, and z represent the imaginary part. Also, u represents the axis of rotation, and θ represents the angle of rotation (radians).

[0262] A logarithmic quaternion can be calculated as shown in the following mathematical formula 18.

[0263]

[0264] In mathematical equation 18, log(q) represents ω, the product of the rotation axis and the rotation angle. In other words, by converting the rotation information of a quaternion into a logarithmic quaternion, the existing 4-dimensional representation can be changed to a 3-dimensional representation. The mathematical basis for this dimensionality reduction is that the quaternion used for the Gaussian property is a unit quaternion (x 2 + y 2 + z 2 This is because it has a constraint condition of (= 1). Due to this unit norm characteristic, at least one of the four components of a quaternion contains inherent redundancy that can be mathematically derived from the remaining components. Accordingly, according to the present disclosure, by eliminating redundant dimensions using the dependency between the four components, the rotation information of a Gaussian can be represented as 3-dimensional data without loss of information. Accordingly, by reducing the rotation information of a Gaussian from 4 dimensions to 3 dimensions, the bit rate transmitted during encoding / decoding can be effectively lowered.

[0265] As a result of the above equation, it is expressed as ω, which is the product of the rotation axis and the rotation angle. Through this, by converting the quaternion into a logarithmic quaternion, the existing 4-dimensional representation can be converted into a 3-dimensional representation. Accordingly, since the rotation information among the attribute information of a 3-dimensional Gaussian can be represented as 3-dimensional data rather than 4-dimensional data, the bitrate can be lowered during encoding / decoding.

[0266] When a logarithmic quaternion is encoded, the decoder can convert the logarithmic quaternion into a quaternion through an exponential map. Specifically, based on mathematical formulas 19 and 20, a quaternion can be restored from the logarithmic quaternion.

[0267]

[0268]

[0269] When encoding / decoding rotation information based on logarithmic quaternions, information indicating whether the rotation information is expressed as a logarithmic quaternion and information required for the logarithmic quaternion transformation can be encoded / decoded as metadata.

[0270] When expressing Gaussian directional color (intensity) information as a 3rd-order spherical harmonic function, 16 coefficients are required for each of the R, G, and B channels. That is, 48 ​​coefficients are required for each Gaussian. Among the 16 coefficients, if a clustering algorithm (e.g., K-means, LBG, etc.) is applied to the remaining AC coefficients (i.e., coefficients from the 1st to the 15th) excluding the DC coefficient (i.e., coefficient 0) representing the global directional color component, the remaining coefficients are grouped to show similar distributions.

[0271] Figure 17 shows an example of applying a clustering algorithm to AC coefficients.

[0272] The first group and the second group each represent subgroups classified through a clustering algorithm. Figure 17(a) illustrates the distribution of coefficients of spherical harmonic functions of three randomly selected samples within each group. As shown in the example illustrated in Figure 17(a), when a clustering algorithm is applied to AC coefficients, it can be observed that the distribution of coefficients within a single group exhibits similarity.

[0273] When Gaussians with similar distributions are grouped and mapped onto a two-dimensional plane, similar data are clustered, as shown in the example illustrated in Fig. 17 (b). By clustering data with similar distributions, encoding / decoding efficiency can be improved during motion prediction, DCT transformation, residual prediction, or quantization processes.

[0274] FIG. 17(b) is illustrated as an image being divided into multiple blocks along a grid, and data having a similar distribution being packed into one block. In the example of FIG. 15(b), the blocks within the image may represent a macroblock, a coding unit (CU), or a transform unit (TU), or may have a size that is a multiple of N of any of these.

[0275] As shown in the example illustrated in FIG. 17(b), Gaussians with similar distributions can be grouped into blocks, and then the blocks can be clustered according to similarity. At this time, to increase the similarity between blocks, an offset can be set for each block. That is, the similarity between blocks can be increased by differentiating or adding the values ​​of samples within a block by the offset. Blocks can be rearranged to increase the similarity between blocks to which the offset has been applied, and blocks with high similarity can be merged. At this time, information indicating the offset of each block can be encoded / decoded as metadata.

[0276] Figure 17 illustrates an example of clustering Gaussians based on the coefficients of a spherical harmonic function. However, even if the similarity between the coefficients of the spherical harmonic function among Gaussians is high, other attributes do not necessarily have high similarity. Accordingly, multiple 2D images can be encoded / decoded by applying multiple alignment criteria within a single block. For example, the first alignment criterion may be the coefficients of a spherical harmonic function, and the second criterion may be position information. For example, 2D images corresponding to color attribute information may have Gaussians aligned according to the first alignment criterion, and 2D images corresponding to geometric attribute information may have Gaussians aligned according to the second alignment criterion.

[0277] In this case, a mapping table indicating the mapping relationship of Gaussians between the first alignment criterion and the second alignment criterion can be encoded / decoded as metadata.

[0278] Meanwhile, FIG. 17 illustrates that clustering is performed for the R, G, and B color channels. The described embodiment is not limited to the RGB color space. The embodiment may also be applied to the YUV color space other than RGB. Specifically, the coefficients of the spherical harmonic function in RGB color information can be converted into the coefficients of the spherical harmonic function based on YUV color information.

[0279] Among the coefficients of the YUV-based spherical harmonic function, the coefficients corresponding to the U and V components among the remaining coefficients (i.e., AC component coefficients), excluding the coefficient corresponding to the DC component (i.e., the zero-order coefficient), exhibit a distribution similar to that of the coefficients corresponding to the Y component. This is because, when expressing directional color information based on a spherical coordinate system, the intensity of each color channel is heavily dependent on the luminance component. By utilizing this property, the encoding / decoding of the AC component coefficients corresponding to the U and V components can be omitted, and the AC component coefficients corresponding to the U and V components can be restored using the AC component coefficients corresponding to the Y component.

[0280] In addition, when encoding / decoding the coefficients of the spherical harmonic function corresponding to the Y component, clustering can be performed on the Y component to group Gaussians with similar distributions. By mapping each group to a single block, Gaussians with high similarity can be included in a single block.

[0281] The above-described embodiments may also be used to reproduce a video showing temporal and spatial changes of a target scene.

[0282] Figure 18 illustrates the change in the Gaussian over time.

[0283] The top of Fig. 18 shows the position change of the Gaussian on the time axis, and the bottom of Fig. 18 shows the duration during which the state of the Gaussian is maintained and the state change of the Gaussian on the time axis.

[0284] For example, the example illustrated in FIG. 18 indicates that during the periods t0 and t2, a change in the movement of the Gaussian occurs, and the Gaussian is affecting the scene representation.

[0285] As the 3D Gaussian is extended to the time axis, information regarding the change in motion of the Gaussian with respect to the time axis and the duration of each Gaussian is required. Here, the change in motion can be information on motion per time unit or information on additional change relative to a specific reference frame (i.e., deformation).

[0286] Additionally, the duration can be expressed in time units of hours, minutes, or seconds, or in video frames.

[0287] Changes in motion and duration can be additionally defined as attribute information. Accordingly, changes in motion and duration can also be converted into 2D video specifications and encoded / decoded.

[0288] When attempting to reconstruct a three-dimensional space using a 3D or 4D Gaussian-based scene representation method (3DGS or 4DGS representation), static representation mode and dynamic representation mode can be distinguished as methods for restoring the target scene. In static representation mode, information regarding changes in motion and duration is not required. On the other hand, in dynamic representation mode, information regarding changes in motion and duration is required. Accordingly, a flag indicating whether the currently encoded bitstream is in static representation mode or dynamic representation mode can be encoded / decoded as metadata.

[0289] Principal Component Analysis (PCA) can be used to encode and decode Gaussian attribute information. Here, Principal Component Analysis can be used to transform high-dimensional data into orthogonal components with low statistical correlation. Through Principal Component Analysis, the original data can be projected into a low-dimensional space while preserving the variance of the original data as much as possible.

[0290] In particular, when applying principal component analysis to high-dimensional feature vectors such as the coefficients of spherical harmonic functions, components that significantly contribute to data variance, such as lighting directionality and color variation, are preferentially extracted, thereby enabling the preservation of important characteristics of the original signal.

[0291] In other words, through principal component analysis, it becomes possible to reduce the amount of data to be compressed while maintaining reproducibility quality.

[0292] In the process of principal component analysis, the Mahalanobis distance can be used to determine the specificity of each data point. The Mahalanobis distance is an indicator of how far the data deviates from the mean when considering the covariance structure. Additionally, for specific data (e.g., coefficients of a spherical harmonic function), a p-value can be calculated based on the Mahalanobis distance, and then, based on the p-value, it can be determined whether a specific Gaussian is an outlier.

[0293] A WignerD-based rotation transformation can be applied to the coefficients of the spherical harmonic function corresponding to outliers. Accordingly, the distribution of the coefficients of the spherical harmonic function corresponding to outliers can be adjusted, and as a result, they can acquire similarity to the coefficients of the spherical harmonic function of other Gaussians.

[0294] Specifically, principal component vectors can be extracted by performing a PCA transformation on all Gaussians or Gaussians belonging to a specific group. Subsequently, based on the principal component vectors of each Gaussian, the Mahalanobis distance or p-value can be derived, and then a group of Gaussians and a group of Gaussians belonging to a certain upper proportion of the Mahalanobis distance or p-value can be constructed. Here, the group of Gaussians belonging to the upper proportion can be referred to as the upper group, and the group of Gaussians belonging to the lower proportion can be referred to as the lower group.

[0295] For each group, a principal direction vector can be extracted based on the coefficients of a spherical harmonic function (e.g., a vector corresponding to the coefficient l=1). Specifically, the principal direction vector can be obtained based on the coefficients of a spherical harmonic function where l is 1. Subsequently, the relative rotation angle between the principal direction vector of the upper group and the principal direction vector of the lower group is calculated, and based on the WignerD function, the coefficients of the spherical harmonic function of the Gaussians belonging to the lower group can be transformed to be similar to the principal direction vector of the upper group.

[0296] Through the above process, the distribution of Gaussians can be adjusted to be similar. If PCA transformation is performed again on the Gaussians with adjusted distributions, the Gaussians previously identified as outliers become distributed closer to the principal components, thereby enabling improved compression efficiency and quality based on PCA transformation.

[0297] The above process can be performed during the encoding or encoding preprocessing stage. To recover the coefficients of the spherical harmonic functions of the Gaussians belonging to the subgroup in the decoder, the relative rotation angle and the location of the subgroup within the image can be encoded / decoded as metadata. Additionally, the coordinates of the block containing the Gaussians of the subgroup, the size of the block, and the number of Gaussians belonging to the subgroup can also be encoded / decoded as metadata.

[0298] Generally, due to positional constraints of the acquired camera and / or the user's movement radius, the (x, y) coordinates of the learned Gaussian primitives exhibit a relatively symmetric Gaussian distribution with respect to the center of the scene. However, the z-axis coordinates of the Gaussian primitives show a completely different pattern from the (x, y) coordinates depending on the perspective projection characteristics of the camera. Specifically, the z-axis coordinates of the Gaussian primitives form a skewed long-tail distribution with a long tail depending on the background depth.

[0299] In this case, if uniform quantization is performed based on a single minimum and maximum pair, the step size along the z-axis increases, leading to a decrease in precision at the center of the distribution. This results in reduced rate distortion and degraded rendering quality.

[0300] Specifically, the (x, y) positional error of Gaussian primitives manifests as a loss of pixel-level alignment, while the z-axis error distorts the depth order of the Gaussian, leading to incorrect alpha blending. This results in visually noticeable problems during image rendering, such as color distortion, depth collisions, and / or occlusion errors.

[0301] Therefore, it is not desirable to apply the same quantization method to all axes. In particular, since the z-axis has the greatest impact on rendering quality, it is necessary to apply an optimized quantization method tailored to the characteristics of the axis.

[0302] In a skewed long-tail distribution, it is necessary to detect Gaussians corresponding to outliers. Since Euclidean distance only considers data far from the mean of each axis, it fails to reflect the characteristics of the elliptical distribution, which creates a possibility of misidentifying outliers. To resolve this issue, Gaussians corresponding to outliers can be detected based on Mahalanobis distance.

[0303] Mahalanobis distance represents the distance from the center of an ellipse to a specific data point when data is distributed in an elliptical shape. In particular, the Mahalanobis squared distance of 3D position data follows a chi-square distribution with 3 degrees of freedom. Accordingly, the probability that a specific value is an outlier can be statistically determined.

[0304] When converting Gaussian attribute information into a 2D image, the Gaussians can be aligned using PLAS. After aligning the Gaussians using PLAS, the 2D image can be divided into blocks of a certain size. Specifically, the 2D image can be divided into one or more blocks. That is, the 2D image may contain only one block, or it may contain multiple blocks.

[0305] Afterwards, based on the mean and covariance of the Gaussians within each block, the Mahalanobis squared distance for each Gaussian can be calculated.

[0306] If the calculated Mahalanobis squared distance is greater than a predefined threshold, the Gaussian may be determined to be an outlier. Gaussians determined to be outliers may be added to the outlier list.

[0307] For each block, the above process can be repeated to add outlier Gaussians to the outlier list. A rectangular target area can be designated to contain the Gaussians identified as outliers within the 2D image, corresponding to the total number of Gaussians added to the outlier list. For example, the bottom-right area within the 2D image can be set as the target area. Alternatively, an area at a different location (e.g., the bottom-left area) can be set as the target area. The target area can be a rectangular (i.e., block-shaped) area.

[0308] Afterward, the average value of the positions of the Gaussians included in the outlier list and the inlier Gaussians located in 8 directions around the Gaussian is set as the reference point, and the location within the target area that is closest in Euclidean distance to the reference point is determined. The Gaussian included in the outlier list can be assigned to the location that is closest in Euclidean distance to the reference point.

[0309] The above process can be performed for each Gaussian included in the outlier list.

[0310] The Gaussians in the outlier list are sorted based on their brightness (i.e., luminance) values, and then placed in the target area according to the sorting order.

[0311] Through the above process, Gaussians identified as outliers can be placed in independent spaces within the 2D image. Additionally, by maintaining spatial consistency with normal Gaussians, distortions that may occur in block-based quantization can be reduced.

[0312] Through the above process, outlier Gaussians are placed in block form in the target region within the 2D image. Subsequently, quantization is applied independently to the group to which the normal Gaussians belong (i.e., the normal group) and the group to which the outlier Gaussians belong (i.e., the outlier group).

[0313] By independently configuring the outlier group, the normalization range (i.e., minimum-maximum) of each of the outlier and normal value groups is reduced. This results in an improvement in the quantization resolution for both groups.

[0314] To this end, the encoder can encode / decode information about the location of outlier groups within a 2D image (e.g., the top-left and / or bottom-right coordinates of the target area) and the minimum and maximum values ​​of each group as metadata.

[0315] In addition, information indicating the number of outlier Gaussians belonging to the target region within a 2D image can also be encoded / decoded as metadata.

[0316] Meanwhile, if the number of outlier Gaussians is smaller than the size of the target area, the areas within the target area where the outlier Gaussians are not packed remain as empty areas. For example, when the target area consists of 10,000 pixels and the number of outlier Gaussians is 9,850, 150 pixels will form an empty area.

[0317] Accordingly, once the size of the empty area within the target region is known, the number of outlier Gaussians can be inversely calculated. Consequently, instead of the number of outlier Gaussians, the number of pixels belonging to the empty area within the target region can be encoded / decoded.

[0318] Alternatively, for each Gaussian, information indicating whether the Gaussian is an outlier Gaussian can be encoded / decoded. In addition, information indicating the number of outlier Gaussians can also be encoded / decoded.

[0319] Quantization of outlier groups can also be performed based on the Mu-law companding method, which is one of the nonlinear scaling techniques. The companding / expandtion process can be expressed as shown in the following mathematical equations 21 and 22.

[0320]

[0321]

[0322] Equation 21 represents the compression process, and Equation 22 represents the reverse compression process.

[0323] Parameters required for Mu-law companding / expansion transformations can be encoded / decoded. For example, the μ value, information indicating the normalization range of attribute information (e.g., minimum and / or maximum values ​​of attribute information), and quantization bit depth can be encoded and signaled.

[0324] Meanwhile, due to environmental constraints of the encoding or transmission system, there may be cases where a target area for placing outlier groups within a 2D image cannot be set. In this case, based on the distribution characteristics of each Gaussian, the data distribution range may be scaled into multiple intervals, and for each interval, a uniform quantization technique based on minimum and maximum values ​​or a μ-law companding / expansion process based on non-linear scaling may be applied.

[0325] As another example, if the variance of the Gaussian distributions differs along the axes, the quantization interval or quantization technique may be set differently by distinguishing between axes with a large data distribution range and axes with a small range. For example, uniform quantization based on minimum and maximum values ​​may be applied to the x-axis and y-axis, while a μ-law companding / expansion process may be applied to the z-axis.

[0326] In the encoder, at least one of the following can be encoded as metadata and signaled: a quantization technique selected according to the data distribution range and variance of Gaussians mapped to a 2D image, information indicating the normalization range required for quantization (e.g., minimum and maximum values), the number of quantization bits, a μ value, information on the quantization technique applied per axis, or interval information. In the decoder, the Gaussian can be restored by referring to the above metadata.

[0327] FIG. 19 is a flowchart of a process for encoding / decoding Gaussian attribute information according to one embodiment of the present disclosure.

[0328] To represent a 3D target scene, the properties of each Gaussian can be learned. Through the above learning, information on each Gaussian property can be derived (S1910).

[0329] The attribute information may include geometry attribute information that defines the shape and position of the Gaussian and appearance attribute information that defines the appearance. Here, the geometry attribute information includes position information (x, y, z coordinates), rotation information (quaternion), and scale information, and the appearance attribute information may include directional color information (SH coefficients) and transparency information (opacity).

[0330] Subsequently, preprocessing of attribute information can be performed. Preprocessing may include a pruning process to remove redundant data or extra component information.

[0331] Attribute information can be converted into a standard for encoding / decoding (S1920). For example, to use a 2D image compression technique, a 2D image can be generated by projecting the attribute information onto a 2D plane.

[0332] The attribute information (i.e., 2D image) with the converted specifications can be encoded (i.e., compressed) (S1930). The compressed bitstream can be transmitted to a decoder through a communication network.

[0333] In the decoder, the bitstream is decoded to restore a 2D image (S1940). Afterwards, the attribute information of the Gaussian projected onto the restored 2D image (i.e., color component information and geometric component information) can be restored in a 3D data format (S1950). Meanwhile, if a preprocessing process is performed in the encoder, a postprocessing process corresponding to the preprocessing process can be performed.

[0334] Based on the restored color component information and geometric component information, a target viewpoint image can be rendered (S1960).

[0335] The names of the syntax elements introduced in the embodiments described above are merely temporary for the purpose of describing the embodiments according to the present disclosure. Syntax elements may be named with names different from those proposed in the present disclosure.

[0336] The components described in the exemplary embodiments of the present disclosure may be implemented by hardware elements. For example, the hardware elements may include at least one of a digital signal processor (DSP), a processor, a controller, an application-specific integrated circuit (ASIC), a programmable logic element such as an FPGA, a GPU, other electronic devices, or a combination thereof. At least some of the functions or processes described in the exemplary embodiments of the present disclosure may be implemented in software, and the software may be recorded on a recording medium. The components, functions, and processes described in the exemplary embodiments may be implemented by a combination of hardware and software.

[0337] A method according to one embodiment of the present disclosure may be implemented as a program that can be executed by a computer, and said computer program may be recorded on various recording media such as magnetic storage media, optical reading media, digital storage media, etc.

[0338] The various technologies described in this disclosure may be implemented as digital electronic circuits or computer hardware, firmware, software, or a combination thereof. The technologies may be implemented as computer program products, namely, computer programs tangibly implemented on information media or computer programs (e.g., machine-readable storage devices (e.g., computer-readable media) or data processing devices), or as computer programs implemented as signals processed by or propagated to perform operations of data processing devices (e.g., programmable processors, computers, or a plurality of computers).

[0339] Computer program(s) may be written in any form of programming language, including compiled or interpreted languages, and may be distributed in any form, including standalone programs or modules, components, subroutines, or other units suitable for use in a computing environment. Computer programs may be executed on a single computer, or by multiple computers distributed across one site or multiple sites and interconnected by a communication network.

[0340] Examples of processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, and one or more processors of a digital computer. Generally, a processor receives instructions and data from read-only memory or random access memory, or both. Components of a computer may include at least one processor for executing instructions and one or more memory devices for storing instructions and data. Additionally, the computer may include one or more mass storage devices for storing data, such as magnetic, magneto-optical disks, or optical disks, or may be connected to said mass storage devices to receive and / or transmit data. Examples of information media suitable for implementing computer program instructions and data include semiconductor memory devices (magnetic media such as hard disks, floppy disks, and magnetic tapes), optical media such as compact disc read-only memory (CD-ROM) and digital video discs (DVD), magneto-optical media such as floptical disks, and Read Only Memory (ROM), Random Access Memory (RAM), flash memory, Erasable Programmable ROM (EPROM), Electrically Erasable Programmable ROM (EEPROM), and other known computer-readable media. Processors and memory may be complemented or integrated by special-purpose logic circuits.

[0341] A processor may execute an operating system (OS) and one or more software applications running on the OS. A processor unit may also access, store, manipulate, process, and generate data in response to software execution. For simplification, a processor unit is described in the singular; however, those skilled in the art will understand that the processor unit may include multiple processing elements and / or various types of processing elements. For example, a processor unit may include multiple processors or a processor and a controller. It may also constitute different processing structures, such as parallel processors. Furthermore, a computer-readable medium means any medium accessible to a computer and may include both computer storage media and transmission media.

[0342] The present disclosure includes detailed descriptions of various detailed embodiments, but such details are not intended to limit the invention or claims proposed in the present disclosure and should be understood as describing the features of specific exemplary embodiments.

[0343] Features individually described in exemplary embodiments in this disclosure may be implemented by a single exemplary embodiment. Conversely, various features described with respect to a single exemplary embodiment in this disclosure may be implemented by a combination of multiple exemplary embodiments or a suitable sub-combination. Furthermore, in this disclosure, said features may operate by a specific combination and may be described as said combination first claimed, but in some cases, one or more features may be excluded from the claimed combination, or the claimed combination may be changed into a sub-combination or a modified form of a sub-combination.

[0344] Likewise, even if operations are described in a specific order in the drawings, it should not be understood that it is necessary to execute the operations in a specific sequence or order, or that all operations must be performed, in order to obtain the desired result. In certain cases, multitasking and parallel processing may be useful. Furthermore, it should not be understood that the various device components in the exemplary embodiments of all embodiments must be separated, and the aforementioned program components and devices may be packaged into a single software product or multiple software products.

[0345] The exemplary embodiments disclosed in this specification are merely illustrative and are not intended to limit the scope of this disclosure. Those skilled in the art will recognize that various modifications to the exemplary embodiments may be made without departing from the spirit and scope of the claims and their equivalents.

[0346] Accordingly, the present disclosure shall be deemed to include all other substitutions, modifications, and changes falling within the scope of the following claims.

[0347] The embodiments included in the present disclosure may be applied to an electronic device capable of encoding / decoding images.

Claims

1. A step of converting attribute information of Gaussians into 2D images; and The method includes the step of encoding the above 2D images, The above attribute information includes rotation information, A method for encoding a Gaussian for a three-dimensional space representation, characterized in that information indicating the representation type of the rotation information is encoded as metadata.

2. In Paragraph 1, A method for encoding a Gaussian for a three-dimensional space representation, characterized in that the rotation information includes three rotation information parameters and one Wigner D parameter.

3. In Paragraph 2, A method for encoding a Gaussian for a three-dimensional space representation, characterized in that the three Euler angle parameters are expressed in the same coordinate axis order as the Wigner D parameters.

4. In Paragraph 1, Gaussians with identical rotation information are classified into one group, and A method for encoding a Gaussian for a three-dimensional space representation, characterized in that the rotation information is encoded for the above group.

5. In Paragraph 1, A method for encoding a Gaussian for a three-dimensional space representation, characterized in that information indicating whether rotation compensation has been performed on the above Gaussian is encoded as metadata.

6. In Paragraph 1, The above attribute information includes location information, and A method for encoding a Gaussian for a three-dimensional space representation, characterized in that the upper bits of the above position information are encoded through a first 2D image and the lower bits are encoded through a second 2D image.

7. In Paragraph 1, The above attribute information includes location information, and A method for encoding a Gaussian for a three-dimensional spatial representation, characterized in that the above-mentioned location information includes a first location information indicating the location of a local coordinate system to which the Gaussian belongs, and a second location information indicating the location of the Gaussian within the local coordinate system.

8. In Paragraph 7, A method for encoding a Gaussian for a three-dimensional space representation, characterized in that the first position information represents the index of the local coordinate system.

9. In Paragraph 8, A method for encoding a Gaussian for a 3D spatial representation, characterized in that a mapping table defining the mapping relationship between the index of each of the local coordinate systems and the coordinates of the local coordinate system within the world coordinate system is encoded as metadata.

10. In Paragraph 7, A method for encoding a Gaussian for a three-dimensional space representation, characterized in that the first position information represents the coordinates of the local coordinate system within the world coordinate system.

11. In Paragraph 1, The above 2D image is divided into a plurality of blocks, and A method for encoding Gaussians for a three-dimensional spatial representation, characterized in that Gaussians with high similarity of AC coefficients of spherical harmonic functions are packed into a single block.

12. In Paragraph 11, A method for encoding a Gaussian for a three-dimensional spatial representation, characterized in that the similarity of the above AC coefficients is determined based on the result of performing clustering on the above AC coefficients.

13. In Paragraph 11, For each block, offset information is encoded, and A method for encoding a Gaussian for a three-dimensional spatial representation, characterized in that samples belonging to a block are encoded in a state that is differentiated or added by an offset specified by offset information.

14. In Paragraph 1, A method for encoding a Gaussian for a three-dimensional spatial representation, characterized in that the above attribute information further includes at least one of position change information and duration information of the Gaussian.

15. In Paragraph 1, Gaussians determined to be outliers are packed into specific regions of the above 2D image, and A method for encoding a Gaussian for a three-dimensional space representation, characterized in that information indicating the number of outlier Gaussians belonging to the above specific region is encoded. Step of decoding 16.2D images; A step of restoring attribute information of Gaussians from decoded 2D images; and It includes a step of rendering a target viewpoint image based on the attribute information of the restored Gaussians, The above attribute information includes rotation information, A method for decoding a Gaussian for a three-dimensional space representation, characterized in that information indicating the representation type of the above-mentioned rotation information is decoded as metadata.