Encoding and decoding point cloud using patches for in-between samples
Patent Information
- Application Number
- JP2024136990
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-09-24
- Filing Date
- 2024-08-16
- Publication Date
- 2025-05-12
- Estimated Expiration
- 2040-09-15
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] At least one of the present embodiments relates generally to processing of point clouds. In particular, encoding / decoding of attributes of 3D samples in / from a separate video stream is disclosed. [Background technology]
[0002] SUMMARY OF THE DISCLOSURE This section is intended to introduce the reader to various aspects of art that may be related to various aspects of at least one of the present embodiments described and / or claimed below. This discussion is believed to be helpful in providing the reader with background information to facilitate a better understanding of the various aspects of at least one embodiment.
[0003] Point clouds can be used for various purposes such as cultural heritage / architecture, where objects like statues or buildings are scanned in 3D to share the spatial configuration of the object without sending or visiting it. Also, if the object may be destroyed, for example a temple may be destroyed by an earthquake, the point cloud is a way to ensure that knowledge of the object is preserved. Such point clouds are usually static, color-coded and huge.
[0004] Another use case is in topography and cartography, where the use of 3D representations allows maps that are not limited to flat surfaces and may include relief. Google Maps is currently a good example of a 3D map, but it uses meshes instead of point clouds. Nevertheless, point clouds may be the preferred data format for 3D maps, and such point clouds are usually static, color-coded, and large.
[0005] The automotive industry and autonomous vehicles are also areas where point clouds can be used. Autonomous vehicles need to be able to "explore" their environment and make good driving decisions based on the reality of their immediate neighborhood. Typical sensors like LIDAR (Light Detection and Ranging) generate dynamic point clouds that are used by decision engines. These point clouds are not intended for humans to see, they are usually small, not necessarily color-coded, and dynamic with a high capture frequency. These point clouds may have other attributes, such as reflectivity provided by LIDAR, when this attribute provides good information about the material of the detected object, which can help make decisions.
[0006] Virtual reality and immersive worlds have been a hot topic lately and are predicted by many to be the future of 2D flat video. The basic idea is to immerse the viewer in the environment that surrounds them, as opposed to standard TV, where the viewer can only see the virtual world in front of him / her. There are several degrees of immersion, depending on the viewer's degrees of freedom in the environment. Point clouds are a good candidate format for delivering Virtual Reality (VR) worlds.
[0007] In many applications, it is important to be able to deliver dynamic point clouds to end users (or store them in a server) while consuming only a reasonable amount of bitrate (or storage space for storage applications) while maintaining an acceptable (or preferably very good) quality of experience. Efficient compression of these dynamic point clouds is a key aspect for the practical implementation of many immersive world delivery networks.
[0008] At least one embodiment has been devised with the above in mind. Summary of the Invention
[0009] The following presents a simplified summary of at least one of the present embodiments in order to provide a basic understanding of some aspects of the present disclosure. This summary is not an extensive overview of the embodiments. It is not intended to identify key or essential elements of the embodiments. The following summary merely presents some aspects of at least one of the present embodiments in a simplified form as a prelude to the more detailed description provided elsewhere in this document.
[0010] According to a general aspect of at least one embodiment, there is provided a method for encoding attributes of an orthogonally projected 3D sample, wherein attributes of the orthogonally projected 3D sample are encoded as at least one first attribute patch of a 2D sample of an image and attributes of an intermediate 3D sample located between two orthogonally projected 3D samples along the same projection line are encoded as at least one second attribute patch of a 2D sample in the image, the method comprising encoding information indicating whether the at least one first attribute patch of the 2D sample and the at least one second attribute patch of the 2D sample are stored in separate images.
[0011] According to an embodiment, the video stream is hierarchically organized into picture level, frame level and patch level groups, and the information is available at either the picture level, frame level, atlas level or patch level groups.
[0012] According to an embodiment, the information is a first flag indicating whether at least one first attribute patch of the 2D sample is stored in a first image or whether at least one first attribute patch of the 2D sample is stored in a second image together with other attribute patches of the 2D sample, and a second flag indicating whether at least one second attribute patch of the 2D sample is stored in a third image or whether at least one second attribute patch of the 2D sample is stored in a second image together with other attribute patches of the 2D sample.
[0013] According to an embodiment, the method further comprises encoding further information indicating how the separate images are compressed.
[0014] According to a general aspect of at least one embodiment, there is provided a method for decoding attributes of a 3D sample, in which attributes of the 3D sample are decoded from at least one first attribute patch of a 2D sample of an image and attributes of an intermediate 3D sample located between two 3D samples along the same projection line are decoded as at least one second attribute patch of a 2D sample in the image, the method comprising decoding information indicative of whether the at least one first attribute patch of the 2D sample and the at least one second attribute patch of the 2D sample are stored in separate images.
[0015] According to an embodiment, the video stream is hierarchically organized into picture level, frame level and patch level groups, and the information is available at either the picture level, frame level, atlas level or patch level group.
[0016] According to an embodiment, the information is a first flag indicating whether at least one first attribute patch of the 2D sample is stored in a first image or whether at least one first attribute patch of the 2D sample is stored in a second image together with other attribute patches of the 2D sample, and a second flag indicating whether at least one second attribute patch of the 2D sample is stored in a third image or whether at least one second attribute patch of the 2D sample is stored in a second image together with other attribute patches of the 2D sample.
[0017] According to an embodiment, the method further comprises encoding further information indicating how the separate images are compressed.
[0018] One or more of at least one embodiment also provide an apparatus, a bitstream, a computer program product, and a non-transitory computer readable medium.
[0019] The specificity of at least one of the present embodiments, as well as other objects, advantages, features, and uses of at least one of the present embodiments, will become apparent from the following description of examples taken in conjunction with the accompanying drawings. [Brief description of the drawings]
[0020] Some example embodiments are illustrated in the drawings. [Figure 1] 1 illustrates a schematic block diagram of an example of a two-layer based point cloud coding structure according to at least one of the present embodiments. [Diagram 2] 1 illustrates a schematic block diagram of an example of a two-layer based point cloud decoding structure according to at least one of the present embodiments. [Diagram 3] 1 illustrates a schematic block diagram of an example image-based point cloud encoder according to at least one of the present embodiments. [Figure 3a] Illustrates an example canvas containing two patches and their 2D bounding boxes. [Figure 3b] 1 illustrates an example of two intermediate 3D samples located between two 3D samples along a projection line. [Figure 4] 1 illustrates a schematic block diagram of an example image-based point cloud decoder according to at least one of the present embodiments. [Diagram 5] 10 illustrates a schematic example of a syntax for a bitstream representing a base layer BL according to at least one of the present embodiments. [Figure 6] 1 illustrates a schematic block diagram of an example system in which various aspects and embodiments may be implemented. [Figure 7] 1 illustrates an example of a flowchart of a method for encoding orthogonal 3D samples of a point cloud frame in accordance with at least one embodiment. [Figure 7a]1 illustrates an example of a flowchart of a method for encoding orthogonal 3D samples of a point cloud frame in accordance with at least one embodiment. [Figure 8] 1 illustrates examples of syntax elements that embed information INF, according to at least one embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0021] At least one of the present embodiments will be described more fully below with reference to the accompanying drawings, in which at least one example of the present embodiments is shown. However, the embodiments may be embodied in many alternative forms and should not be construed as being limited to the examples set forth herein. It should therefore be understood that there is no intention to limit the embodiments to the particular forms disclosed. On the contrary, the present disclosure is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present application.
[0022] Where a diagram is presented as a flow diagram, it should be understood that the diagram also provides a block diagram of the corresponding apparatus. Similarly, where a diagram is presented as a block diagram, it should be understood that the diagram also provides a flow diagram of the corresponding method / process.
[0023] Similar or identical figure elements are referred to with the same reference numbers.
[0024] Some figures represent syntax tables widely used in V-PCC to define the structure of bitstreams according to V-PCC. In those syntax tables, the term "..." indicates the unchanged parts of the syntax with respect to the original definition given in V-PCC and the parts removed in the figures for ease of reading. The bolded terms in the figures indicate that the value of this term is obtained by parsing the bitstream. The right column of the syntax table indicates the number of bits for encoding the data of the syntax element. For example, u(4) indicates that 4 bits are used to encode the data, u(8) indicates that 8 bits are used to encode the data, and ae(v) indicates a context-adaptive arithmetic entropy coded syntax element.
[0025] The aspects described and contemplated below can be implemented in many different forms. In Figures 1-8 below, several embodiments are provided, however, other embodiments are contemplated and the discussion of Figures 1-8 is not intended to limit the breadth of implementations.
[0026] At least one of the aspects relates generally to point cloud encoding and decoding, and at least one other aspect relates generally to transmission of the generated or encoded bitstream.
[0027] More precisely, the various methods and other aspects described herein may be used to modify modules, for example the image-based encoder 3000 and decoder 4000, as shown in Figures 1-8.
[0028] Furthermore, the aspects are not limited to MPEG standards such as MPEG-I Part 5 related to point cloud compression, but may be applied, for example, to other standards and recommendations, whether existing or developed in the future, and extensions of any such standards and recommendations, including MPEG-I Part 5. Unless otherwise indicated or technically precluded, the aspects described in this application may be used individually or in combination.
[0029] In the following, image data refers to data, e.g. one or several arrays of 2D samples in a particular image / video format. A particular image / video format may specify information related to pixel values of an image (or video). A particular image / video format may also specify information that may be used by a display and / or any other device to, e.g., visualize and / or decode the image (or video). An image typically includes a first component, in the form of a first array of 2D samples, that usually represents the luminance (or luma) of the image. An image may also include a second and a third component, in the form of another array of 2D samples, that usually represents the chromaticity (or chroma) of the image. Some embodiments represent the same information using a set of 2D color sample arrays, such as a conventional three-color RGB representation.
[0030] A pixel value is represented in one or more embodiments by a vector of C values, where C is the number of components. Each value in the vector is typically represented using a number of bits that can define the dynamic range of the pixel value.
[0031] An image block refers to a set of pixels that belong to an image. The pixel values of an image block (or image block data) refer to the values of the pixels that belong to this image block. Image blocks can have any shape, but rectangular shapes are common.
[0032] A point cloud may be represented by a dataset of 3D samples in a 3D volumetric space, each having unique coordinates and may also have one or more attributes.
[0033] A 3D sample may include information defining the geometry of a 3D point of the point cloud, which may be represented by X, Y, and Z coordinates in 3D space. A 3D sample may also include information defining one or more associated attributes, such as, for example, color expressed in RGB or YUV color space, transparency, reflectance, a two-component normal vector, or any feature that represents a characteristic of this sample. For example, a 3D sample may include information defining six components (X, Y, Z, R, G, B), or in other words (X, Y, Z, y, U, V), where (X, Y, Z) define the coordinates of a 3D point in 3D space, and (R, G, B) or (y, U, V) define the color of this 3D point. Attributes of the same type may be present multiple times. For example, multiple color attributes may provide color information from different viewpoints.
[0034] A 2D sample may include information defining the geometry of an orthogonally projected 3D sample, which may be represented by three coordinates (u, v, Z), where (u, v) are the coordinates in the orthogonally projected 3D 2D space, and Z is the Euclidean distance between the 3D sample and the projection plane onto which it is orthogonally projected. Z is typically represented as a depth value. A 3D sample may also include information defining one or more associated attributes, such as, for example, color represented in RGB or YUV color space, transparency, reflectance, two component normal vectors, or any feature that characterizes this orthogonally projected 3D sample.
[0035] Thus, the 2D sample may contain information that defines the geometry and attributes of the orthogonally projected 3D sample in terms of (u,v,Z,R,G,B) or in other words (u,v,Z,y,U,V).
[0036] A point cloud can be static or dynamic, depending on whether the cloud changes over time. Instances of static or dynamic point clouds are typically denoted as point cloud frames. Note that in the case of a dynamic point cloud, the number of points is typically not constant, but in contrast typically changes over time. More generally, a point cloud can be considered dynamic if something changes over time, e.g., the number of points, the location of one or more points, or any attribute of any point.
[0037] FIG. 1 illustrates a schematic block diagram of an example of a two-layer based point cloud encoding structure 1000 according to at least one of the present embodiments.
[0038] The two-layer based point cloud encoding structure 1000 may provide a bitstream B representing an input point cloud frame IPCF, possibly representing a frame of a dynamic point cloud, which may then be encoded by the two-layer based point cloud encoding structure 1000.
[0039] A video stream for representing the dynamic point cloud may then be obtained by fully combining the bitstreams representing each frame of the dynamic point cloud.
[0040] Essentially, the two-layer based point cloud code structure 1000 may provide the ability to structure the bitstream B as a base layer BL and an enhancement layer EL. The base layer BL may provide a lossy representation of the input point cloud frame IPCF, and the enhancement layer EL may provide a higher quality (possibly lossless) representation by encoding isolated points not represented by the base layer BL.
[0041] The base layer BL may be provided by an image-based encoder 3000 as illustrated in Fig. 3, which may provide a geometry / attribute image representing the geometry / attributes of the 3D samples of the input point cloud frame IPCF, which may enable discarding orphaned 3D samples. The base layer BL may be decoded by an image-based decoder 4000 as illustrated in Fig. 4, which may provide an intermediate reconstructed point cloud frame IRPCF.
[0042] Returning now to the two-layer based point cloud encoding 1000 of FIG. 1, a comparator COMP may compare the 3D samples of the input point cloud frame IPCF with the 3D samples of the intermediate reconstructed point cloud frame IRPCF to detect / locate missed / orphaned 3D samples. Then, an encoder ENC may encode the missed 3D samples and provide an enhancement layer EL. Finally, the base layer BL and the enhancement layer EL may be multiplexed together by a multiplexing device MUX to generate a bitstream B.
[0043] According to an embodiment, the encoder ENC may include a detector capable of detecting 3D reference samples of the intermediate reconstructed point cloud frame IRPCF and associating them with the missed 3D sample M.
[0044] For example, the 3D reference sample R associated with a missed 3D sample M may be the closest neighbor of M according to a given metric.
[0045] According to an embodiment, the encoder ENC may then encode the spatial positions of the missed 3D samples M, and their attributes, as differences determined according to the spatial positions and attributes of said 3D reference sample R.
[0046] In a variant, the differences can be coded separately.
[0047] For example, for a missed 3D sample M, using spatial coordinates x(M), y(M), and z(M), the x-coordinate position difference Dx(M), the y-coordinate position difference Dy(M), the z-coordinate position difference Dz(M), the R-attribute component difference Dr(M), the G-attribute component difference Dg(M), and the B-attribute component difference Db(M) may be calculated as follows: Dx(M)=x(M)-x(R), where x(M) is the x-coordinate of the 3D sample M in the geometry image given by FIG. 3, and similarly for R, respectively; Dy(M) = y(M) - y(R) where y(M) is the y-coordinate of the 3D sample M in the geometry image given by FIG. 3, and similarly for R, respectively; Dz(M) = z(M) - z(R) where z(M) is the z-coordinate of the 3D sample M in the geometry image given by FIG. 3, and similarly for R, respectively. Dr(M) = R(M) - R(R). where R(M), R(R) are the r-color components of the color attributes of the 3D samples M and R, respectively; Dg(M) = G(M) - G(R). where G(M), G(R) are the g-color components of the color attributes of the 3D samples M and R, respectively; Db(M) = B(M) - B(R). where B(M), B(R) are the b-color components of the color attributes of 3D samples M, R, respectively.
[0048] FIG. 2 illustrates a schematic block diagram of an example of a two-layer based point cloud decoding structure 2000 according to at least one of the present embodiments.
[0049] The operation of the two-layer based point cloud decoding structure 2000 depends on its capabilities.
[0050] A two-layer based point cloud decoding structure 2000 with limited capabilities may access only the base layer BL from the bitstream B by using a demultiplexing device DMUX, and then provide a faithful (but lossy) version IRPCF of the input point cloud frame IPCF by decoding the base layer BL by a point cloud decoder 4000, as illustrated in FIG. 4.
[0051] The fully capable two-layer based point cloud decoding structure 2000 may access both the base layer BL and the enhancement layer EL from the bitstream B by using a demultiplexing device DMUX. The point cloud decoder 4000 may determine an intermediate reconstructed point cloud frame IRPCF from the base layer BL as illustrated in Fig. 4. The decoder DEC may determine a complementary point cloud frame CPCF from the enhancement layer EL. The combiner COMB may then combine the intermediate reconstructed point cloud frame IRPCF and the complementary point cloud frame CPCF together, thus providing a higher quality (possibly lossless) representation (reconstruction) CRPCF of the input point cloud frame IPCF.
[0052] FIG. 3 illustrates a schematic block diagram of an example image-based point cloud encoder 3000 according to at least one of the present embodiments.
[0053] The image-based point cloud encoder 3000 leverages existing video codecs to compress the geometry and attribute information of the 3D samples of the input dynamic point cloud using different video streams.
[0054] In a particular embodiment, two video streams, one for capturing geometric information of the 3D samples of the input point cloud and another for capturing attribute information of these 3D samples, may be generated and compressed using an existing video codec, such as the HEVC Main Profile Encoder / Decoder (ITU-T H.265 ITU Telecommunication Standardization Sector (02 / 2018), Series H, i.e., Audiovisual and Multimedia Systems, Infrastructure for Audiovisual Services - Coding of Video Moving Pictures, High Efficiency Video Coding, Recommendation ITU-T H.265).
[0055] Additional metadata used to interpret the two video streams is also typically generated and compressed separately, such additional metadata including, for example, occupancy maps OM and / or auxiliary patch information PI.
[0056] The generated video stream and metadata may then be multiplexed together to generate a composite stream.
[0057] Note that metadata typically represents a small amount of the overall information, the majority of which is in the video stream.
[0058] An example of such a point cloud coding / decoding process is given by the Test Model Category 2 algorithm (also denoted V-PCC) implementing the MPEG draft standard as defined in ISO / IEC JTC1 / SC29 / WG11, Information technology-Coded Representation of Immersive Media-Part 5: Video-based Point Cloud Compression, CD stage, SCD_d39, ISO / IEC 23090-5.
[0059] In step 3100, the module PGM may generate at least one patch of 2D samples by orthogonally projecting 3D samples of the frame IPCF of the input point cloud frame onto 2D samples on a projection plane using a strategy that provides the best compression.
[0060] A patch of 2D samples may be defined as a set of 2D samples that share a common property.
[0061] For example, in V-PCC, the normals for each 3D sample are first estimated, as described for example in Hoppe et al. (Hugues Hoppe, Tony DeRose, Tom Duchamp, John McDonald, Werner Stuetzle, Surface reconstruction from unorganized points, ACM SIGGRAPH 1992 Proceedings, 71-78). An initial clustering of the 3D samples is then obtained by associating each 3D sample with one of six orientation planes of a 3D bounding box that encloses the 3D sample. More precisely, each 3D sample is clustered and associated with the orientation plane that has the closest normal (maximizing the dot product of the point normal and the surface normal). The 3D samples are then projected orthogonally onto their associated plane (projection plane). A set of 3D samples that form a connected region in their plane is called a connected component. A connected component is thus a set of at least one 3D sample with similar normal and the same associated orientation plane. The initial clustering is then refined by iteratively updating the cluster associated with each 3D sample based on its normal and the clusters of its nearest neighbors. The final step consists of generating one patch of 2D samples from each connected component by projecting the 3D samples of each connected component onto the orientation plane associated with that connected component.
[0062] The 2D samples of a patch of 2D samples then share the same normal and the same orientation plane, and they are located close to each other.
[0063] A patch of 2D samples is associated with auxiliary patch information PI that represents auxiliary patch information used to interpret the 2D sample geometry / attributes of this patch of 2D samples.
[0064] In a V-PCC, for example, the auxiliary patch information PI includes information such as: 1) information indicating one of six orientation faces of a 3D bounding box that encloses the 3D sample of the connected component; 2) information about the face normals; 3) information determining the 3D position of the connected component with respect to the patch expressed in terms of depth, tangent shift, and both tangent shifts; and 4) coordinates in the projection plane (u0, v0, u1, v1) that define a 2D bounding box that encloses the patch.
[0065] In step 3200, the patch packing module PPM may map (arrange) at least one generated patch of 2D samples onto a 2D grid (also denoted as canvas or atlas), typically without any overlap, in a manner that minimizes unused space, and may ensure that every TxT (e.g., 16x16) block of the 2D grid is associated with a unique patch. A given minimum block size TxT of the 2D grid may specify a minimum distance between distinct patches of 2D samples when arranged on this 2D grid. The resolution of the 2D grid may depend on the input point cloud frame size, and its width W and height H as well as the block size T may be transmitted to the decoder as metadata.
[0066] The auxiliary patch information PI may further include information regarding the association between blocks of the 2D grid and patches of 2D samples.
[0067] Figure 3a illustrates an example of a canvas C that includes two patches of 2D samples P1 and P2 and their associated 2D bounding boxes B1 and B2. The two bounding boxes may overlap in the canvas C as illustrated in Figure 3a. Although the 2D grid (canvas partitions) is only represented within the bounding boxes, canvas partitions also occur outside of those bounding boxes. The bounding boxes associated with the patches can be partitioned into TxT blocks, typically T=16.
[0068] A TxT block containing a 2D sample belonging to a patch of 2D samples may be considered an occupied block. Each occupied block of the canvas is represented in the occupancy map OM by a certain pixel value (e.g., 1) and each unoccupied block of the canvas is represented by another certain value, e.g., 0. The pixel values of the occupancy map OM may then indicate whether a TxT block of the canvas is occupied or not, i.e. whether it contains at least one 2D sample belonging to a patch of 2D samples or not.
[0069] In Fig. 3a, occupied blocks are represented by white blocks and light grey blocks represent unoccupied blocks. The image generation process (steps 3300 and 3400) exploits the mapping of at least one generated patch of a 2D sample onto the 2D grid computed during step 3200 to store the geometry and attributes of the 3D sample as an image.
[0070] In step 3300, a geometric image generator GIG may generate at least one geometric image GI from at least one patch of 2D samples, the occupancy map OM, and the auxiliary patch information PI.
[0071] The geometry image GI may represent the geometry of at least one patch of 2D samples and may be, for example, a monochrome image of WxH pixels represented in YUV420-8bit format.
[0072] The geometric image generator GIG may exploit the occupancy map information to detect (locate) occupied blocks of the 2D grid in which at least one patch of the 2D sample is defined and thus non-empty pixels in the geometric image GI.
[0073] To better handle the case where multiple 3D samples are projected (mapped) to the same coordinates of the projection plane (along the same projection direction line), multiple layers can be generated. Thus, different depth values D1, ..., Dn can be obtained and associated with 2D samples of the same patch of 2D samples. Then, multiple geometry images GI1, ..., GIN can be generated, each for a specific depth value of the patch of 2D samples.
[0074] In V-PCC, the 2D samples of a patch may be projected onto two layers: the first layer, also called the near layer, may for example store the depth value D0 associated with the 2D sample with the lowest depth, and the second layer, called the far layer, may for example store the depth value D1 associated with the 2D sample with the highest depth.
[0075] According to an embodiment of step 3300, the geometry of at least one patch of 2D samples (the geometry of at least one orthogonally projected 3D point) is coded according to a regular geometry coding mode RGCM, which outputs at least one regular geometry patch of 2D samples RG2DP from the geometry of said at least one patch of 2D samples.
[0076] According to an embodiment, the regular geometry coding mode RGCM may code (derive) the depth values associated with 2D samples of a patch of 2D samples for a layer (first or second or both) as luma components g(u,v) given by (u,v)=δ(u,v)-δ0. Note that using this relationship, the 3D sample positions δ(0,s0,r0) may be reconstructed from the reconstructed geometry image g(u,v) using the associated auxiliary patch information PI.
[0077] According to an embodiment of step 3300, the geometry of at least one patch of 2D samples (the geometry of at least one orthogonally projected 3D point) is coded according to a first geometry coding mode FGCM, which outputs at least one first geometry patch of 2D samples FG2DP from the geometry of said at least one patch of 2D samples.
[0078] According to an embodiment, the first geometry coding mode FGCM directly codes the geometry of a 2D sample of at least one patch of said 2D samples as pixel values of a geometry image.
[0079] For example, when a geometric shape is represented by three coordinates (u, v, Z), three consecutive pixels of the image are used, one to encode the u coordinate, another to encode the v coordinate, and another to encode the Z coordinate.
[0080] According to an embodiment of step 3300, the geometry of the at least one intermediate 3D sample is coded according to a second geometry coding mode SGCM, which outputs at least one second geometry patch SG2DP of 2D samples from the geometry of the at least one intermediate 3D sample.
[0081] The intermediate 3D sample may be between a first orthogonally projected 3D sample and a second orthogonally projected 3D sample along the same projection line, the intermediate 3D sample and the first and second orthogonally projected 3D samples having the same coordinates on the projection plane but different depth values.
[0082] In a variant, the intermediate 3D sample may be defined from a single orthogonal projected 3D sample and the length of the EOM codeword. The depth value of the "virtual" second orthogonal projected 3D sample is then equal to the depth value of the first orthogonal projected 3D sample and the length value of that EOM codeword. The first and "virtual" orthogonal projected 3D samples have the same coordinates on the projection plane but different depth values.
[0083] In some cases, the length of the EOM codeword is embedded in a syntax element in the bitstream.
[0084] In the following, the intermediate 3D sample is considered to exist between a first orthogonal projected 3D sample and a second orthogonal projected 3D sample, even if the second orthogonal projected 3D sample is "virtual".
[0085] Furthermore, the intermediate 3D sample has a depth value that is greater than the depth value of the first orthogonal projected 3D sample and less than the depth value of the second orthogonal projected 3D sample.
[0086] A number of intermediate 3D samples may exist between the first and second orthogonally projected 3D samples, and a designated bit of the codeword may be set for each of the intermediate 3D samples to indicate when the intermediate 3D sample exists (or does not exist) at a particular distance (a particular spatial location along the projection line) from one of the two orthogonally projected 3D samples.
[0087] FIG. 3b shows two intermediate 3D samples P located between the two 3D samples P0 and P1 along the projection line PL. i1 and P i2 3D samples P0 and P1 have depth values equal to D0 and D1, respectively. Two intermediate 3D samples P i1 and P i2 Depth value D i1 and D. i2 are greater than D0 and less than D1, respectively.
[0088] Then, all the designated bits along the projection line may be concatenated to form a codeword, hereafter denoted as an Enhanced-Occupancy map (EOM) codeword. Assuming an EOM codeword of length 8 bits, as illustrated in FIG. 3b, with 2 bits equal to 1, two 3D samples P i1 and P i2 Indicates the location.
[0089] According to an embodiment of the second geometric coding mode SGCM, all EOM codewords are packed together to form at least one second geometric patch SG2DP of 2D samples.
[0090] At least one second geometric patch SG2DP of said 2D sample belongs to an image, the coordinates of pixels in said image indicate two of the three coordinates of the intermediate 3D samples (when these pixels point to EOM codewords) and the values of said pixels indicate the third coordinate of these intermediate 3D samples.
[0091] According to an embodiment, at least one second geometric patch SG2DP of said 2D sample belongs to the occupancy map OM.
[0092] In step 3400, the attribute image generator TIG may generate at least one attribute image TI from at least one patch of 2D samples, the occupancy map OM, the auxiliary patch information PI and at least one decoded geometry image DGI, i.e., the geometry of a 3D sample derived from the output of the video decoder VDEC (step 4200 in FIG. 4).
[0093] The attribute image TI may represent attributes of a 3D sample and may be, for example, an image of WxH pixels represented in YUV420-8 bit format.
[0094] The attribute image generator TG may exploit the occupancy map information to detect (locate) the occupied blocks of the 2D grid in which at least one patch of said 2D samples is defined and thus the non-empty pixels in the attribute image TI.
[0095] The attribute image generator TIG may be adapted to generate an attribute image TI and to associate the attribute image TI with each geometry image DGI.
[0096] Then, multiple attribute images TI1, ..., TIn may be generated, each for a particular depth value of the patch of 2D samples (for each geometry image).
[0097] According to an embodiment of step 3400, the attributes of at least one patch of 2D samples (attributes of the orthogonally projected 3D samples) are coded according to a regular attribute coding mode RACM, which outputs at least one regular attribute patch of 2D samples RA2DP from the attributes of the at least one patch of 2D samples.
[0098] According to an embodiment, the regular attribute coding mode RACM may code (store) attributes T0 associated with 2D samples of a patch of 2D samples as pixel values of a first attribute image TI0 for a first layer and attribute values T1 associated with 2D samples of a patch of 2D samples as pixel values of a second attribute image TI1 for a second layer.
[0099] Alternatively, the attribute image generation module TIG may code (store) attribute values T1 associated with the 2D samples of the patch of 2D samples as pixel values of a first attribute image TI0 for the second layer, and attribute values T0 associated with the 2D samples of the patch of 2D samples as pixel values of a second attribute image TI1 for the first layer.
[0100] For example, the color of the 3D sample may be obtained as described in sections 2.2.3, 2.2.4, 2.2.5, 2.2.8, or 2.5 of the V-PCC.
[0101] According to an embodiment of step 3400, the attributes of at least one patch of 2D samples (attributes of the orthogonally projected 3D samples) are coded according to a first attribute coding mode FACM, which outputs at least one first attribute patch of 2D samples FA2DP from the attributes of the at least one patch of 2D samples.
[0102] According to an embodiment, the first attribute coding mode FACM directly codes attributes of 2D samples of at least one patch of said 2D samples as pixel values of the image.
[0103] According to an embodiment, at least one patch of said 2D samples belongs to an attribute image.
[0104] For example, the attributes of the orthogonally projected 3D samples may be obtained as described in sections 2.2.3, 2.2.4, 2.2.5, 2.2.8, or 2.5 of the V-PCC.
[0105] According to an embodiment of step 3400, the attributes of the intermediate 3D sample are coded according to a second attribute coding mode SACM, which outputs at least one second attribute patch SA2DP of 2D samples from the attributes of said at least one intermediate 3D sample.
[0106] The attribute values of the intermediate 3D samples cannot be directly stored as pixel values of the attribute image because their pixel locations correspond to occupied blocks that are already used to store attribute values of other 2D samples, as illustrated in Fig. 3b.
[0107] According to an embodiment of the second attribute coding mode SACM, the attribute values of the intermediate 3D samples are fully packed to form at least one second attribute patch of the 2D sample.
[0108] According to an embodiment, at least one second attribute patch of said 2D sample belongs to an attribute image.
[0109] In the V-PCC, the location of at least one second attribute patch of the 2D sample is procedurally defined (Section 9.4.5 of the V-PCC). Briefly, the process involves determining the location of unoccupied blocks in an attribute image and storing attribute values associated with intermediate 3D samples as pixel values of the unoccupied blocks of the attribute image, thereby avoiding overlap between occupied blocks and second attribute patches of 2D samples.
[0110] At step 3500, the video encoder VENC may encode the generated images / layers TI and GI.
[0111] In step 3600, the encoder OMENC may encode the occupancy map as an image, for example as detailed in section 2.2.2 of V-PCC. Lossy or lossless encoding may be used.
[0112] According to an embodiment, the video encoders ENC and / or OMENC may be HEVC-based encoders.
[0113] In step 3700, the encoder PIENC may encode the auxiliary patch information PI and possibly additional metadata, such as the block size T, width W, and height H of the geometry / attribute image.
[0114] According to an embodiment, the auxiliary patch information may be differentially encoded (eg, as defined in section 2.4.1 of the V-PCC).
[0115] In step 3800, a multiplexing device may be applied to the generated outputs of steps 3500, 3600, and 3700, so that these outputs may be multiplexed together to generate a composite stream representing the base layer BL. Note that the metadata information represents a small percentage of the overall bitstream.
[0116] The encoder 3000 may also be used to encode a dynamic point cloud, where each frame of the point cloud is then repeatedly encoded. Then, at least one geometry image (step 3300), at least one attribute image (step 3400), an occupancy map (step 3600), and auxiliary patch information (step 3700) are generated for each frame. The generated geometry images for all frames of the point cloud may then be fully combined to form a video stream, attribute images to form another video stream, and an occupancy map to form another video stream. The auxiliary patch information may be added to generate a video stream, or all of the auxiliary patch information may be fully packed to form another video stream. Then, all of these video streams may be multiplexed (step 3800) to form a single bitstream BL.
[0117] FIG. 4 illustrates a schematic block diagram of an example image-based point cloud decoder 4000 according to at least one of the present embodiments.
[0118] The decoder 4000 may be used to decode point cloud frames from a bitstream that includes multiple image streams (at least one geometry image stream, at least one attribute image stream, an occupancy map stream, and an auxiliary patch information image stream). However, the decoder 4000 may also be used to decode dynamic point clouds that include multiple frames, where each frame of the dynamic point cloud is decoded by extracting information from the video streams (geometry video stream, attribute video stream, occupancy map video stream, auxiliary patch information video stream) embedded in the bitstream.
[0119] In step 4100, a demultiplexing device DMUX may be applied to demultiplex the coded information of the bitstream representing the base layer BL.
[0120] In step 4200, the video decoder VDEC may decode the encoded information to derive at least one decoded geometry image DGI and at least one decoded attribute image DTI for decoding of the 3D samples of the point cloud frame.
[0121] In step 4300, the decoder OMDEC may decode the encoded information to derive a decoded occupancy map DOM for decoding the 3D samples.
[0122] According to an embodiment, the video decoders VDEC and / or OMDEC may be HEVC-based decoders.
[0123] In step 4400, the decoder PIDEC may decode the encoded information to derive auxiliary patch information DPI for decoding of the 3D samples.
[0124] In some cases, metadata may also be derived from the bitstream BL.
[0125] In step 4500, the geometry generation module GGM may derive a geometry RG of a 3D sample of the point cloud frame IRPCF from at least one decoded geometry image DGI, the decoded occupancy map DOM, the decoded auxiliary patch information DPI and possibly additional metadata.
[0126] The geometry generation module GGM may exploit the decoded occupancy map information DOM in order to locate non-empty pixels in at least one decoded geometry image DGI.
[0127] The non-empty pixel belongs to either the occupied block or the EOM reference block, depending on the pixel value of the decoded occupancy information DOM and the values of D1 to D0 described above.
[0128] According to an embodiment of step 4500, when the non-empty pixel in question belongs to an occupied block, the geometry of the 3D sample is decoded according to a regular geometry decoding mode RGDM.
[0129] According to an embodiment, the regular geometry decoding mode RGDM derives the 3D coordinates of the 3D samples from the coordinates of the non-empty pixels, the values of said non-empty pixels of at least one decoded geometry image DGI, from the decoded auxiliary patch information and possibly from additional metadata.
[0130] The use of non-empty pixels is based on the relationship of the 2D pixels to the 3D samples. For example, with this projection in the V-PCC, the 3D coordinates of the reconstructed 3D sample in terms of depth δ(u,v), tangent shift s(u,v), and bi-tangent shift r(u,v) can be expressed as follows:
[0131] δ(u,v)=δ0+g(u,v) s(u,v)=s0-u0+u r(u,v)=r0-v0+v where g(u,v) is the luma component of the decoded geometry image DGI, (u,v) is the pixel associated with the reconstructed 3D sample, (δ0,s0,r0) is the 3D position of the connected component to which the reconstructed 3D sample belongs, and (u0,v0,u1,v1) are the coordinates in the projection plane that define the 2D bounding box that contains the projection of the patch associated with that connected component.
[0132] According to an embodiment of step 4500, the geometry of the 3D samples is decoded according to a first geometric decoding mode FGDM.
[0133] According to an embodiment, a first geometry decoding mode FGDM decodes the geometry of the 3D samples directly from the pixel values of the decoded geometry image DGI.
[0134] For example, when a geometric shape is represented by three coordinates (u, v, Z), three consecutive pixels of the image are used, where the u coordinate is equal to the value of one pixel in the geometric shape image, v is equal to the value of another pixel in the geometric shape image, and Z is equal to the value of another pixel in the geometric shape image.
[0135] According to an embodiment of step 4500, the geometry of the at least one intermediate 3D sample is decoded according to a decoding mode SGDM of the second geometry.
[0136] According to an embodiment, the coding mode SGCM of the second geometry may derive two of the 3D coordinates of the intermediate 3D sample from the coordinates of the non-empty pixel and from the third one of the bit values of the EOM codeword.
[0137] For example, according to the example of FIG. 3b, the EOM codeword EOMC is i1 and P i2 are used to determine the 3D coordinates of the intermediate 3D sample P i1 The third coordinate is, for example, D0xD. i1= D0 + 3, the reconstructed 3D sample P i2 The third coordinate is, for example, D0xD. i2 =D0+5. The offset value (3 or 5) is the number of intervals between D0 and D1 along the projection line.
[0138] In step 4600, the attribute generation module TGM may derive attributes of a 3D sample of the reconstructed point cloud frame IRPCF from the geometry RG of said 3D sample and from the at least one decoded attribute image DTI.
[0139] According to an embodiment of step 4600, the attributes of the 3D samples whose geometry has been decoded by the regular geometry decoding mode RGDM are decoded according to the regular attribute decoding mode RADM. The first attribute decoding mode RADM may decode the attributes of the 3D samples from pixel values of the attribute image.
[0140] According to an embodiment of step 4600, the attributes of the 3D samples whose geometry has been decoded by the first geometry decoding mode FGDM are decoded according to a first attribute decoding mode FADM, which may decode the attributes of the 3D samples from pixel values of an attribute image.
[0141] According to an embodiment of step 4600, the attributes of the intermediate 3D samples are decoded according to a second attribute decoding mode SADM.
[0142] According to an embodiment, the second attribute decoding mode SADM may derive attributes of the intermediate 3D samples from second attribute patches SA2DP of the 2D samples.
[0143] According to an embodiment, at least one second attribute patch of said 2D sample belongs to an attribute image.
[0144] In the V-PCC, the location of at least one second attribute patch of the 2D sample is procedurally defined (Section 9.4.10 of the V-PCC). Briefly, the process involves determining the location of unoccupied blocks in an attribute image and deriving attribute values associated with the intermediate 3D sample from pixel values of the unoccupied blocks of the attribute image.
[0145] FIG. 5 illustrates generally an example of a syntax of a bitstream representing a base layer BL according to at least one of the present embodiments.
[0146] The bitstream includes a bitstream header SH and at least one frame stream group GOFS.
[0147] The frame stream group GOFS includes a header HS, at least one syntax element OMS representing an occupancy map OM, at least one syntax element GVS representing at least one geometric shape image (or video), at least one syntax element TVS representing at least one attribute image (or video), and at least one syntax element PIS representing auxiliary patch information, as well as other additional metadata.
[0148] In a variant, the frame stream group GOFS comprises at least one frame stream.
[0149] FIG. 6 shows a schematic block diagram illustrating an example of a system in which various aspects and embodiments may be implemented.
[0150] System 6000 may be embodied as one or more devices including various components described below and configured to perform one or more of the aspects described in this document. Examples of devices that may form all or part of system 6000 include personal computers, laptops, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, connected vehicles and their associated processing systems, head mounted display devices (HMDs, see-through glasses), projectors (beamers), "immersive virtual reality experience caves" (systems including multiple displays), servers, video encoders, video decoders, post-processing output from video decoders, pre-processors providing input to video encoders, web servers, set-top boxes, and any other devices for processing point clouds, videos, or images, or other communication devices. Elements of system 6000 may be embodied, alone or in combination, in a single integrated circuit, multiple ICs, and / or individual components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 6000 may be distributed across multiple ICs and / or separate components. In various embodiments, system 6000 may be communicatively coupled to other similar systems or to other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various embodiments, system 6000 may be configured to implement one or more of the aspects described in this document.
[0151] The system 6000 may include at least one processor 6010 configured to execute instructions loaded therein, for example, to implement various aspects described herein. The processor 6010 may include embedded memory, input / output interfaces, and various other circuits known in the art. The system 6000 may include at least one memory 6020 (e.g., a volatile memory device and / or a non-volatile memory device). The system 6000 may include a storage device 6040, which may include non-volatile and / or volatile memory, including, but not limited to, Electrically Erasable Programmable Read-Only Memory (EEPROM), Read-Only Memory (ROM), Programmable Read-Only Memory (PROM), Random Access Memory (RAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Flash, magnetic disk devices, and / or optical disk devices. The storage device 6040 may include, as non-limiting examples, internal storage devices, attached storage devices, and / or network-accessible storage devices.
[0152] The system 6000 may include an encoder / decoder module 6030 configured to process data to provide encoded or decoded data, for example, and the encoder / decoder module 6030 may include its own processor and memory. The encoder / decoder module 6030 may represent a module that may be included within a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding and decoding module. Additionally, the encoder / decoder module 6030 may be implemented as a separate element of the system 6000 or may be incorporated within the processor 6010 as a combination of hardware and software, as is known to those skilled in the art.
[0153] Program code loaded into the processor 6010 or the encoder / decoder 6030 for implementing various aspects described herein may be stored in the storage device 6040 and then loaded onto the memory 6020 for execution by the processor 6010. According to various embodiments, one or more of the processor 6010, the memory 6020, the storage device 6040, and the encoder / decoder module 6030 may store one or more of various items during the performance of the processes described herein. Such stored items may include, but are not limited to, point cloud frames, encoded / decoded geometry / attribute videos / images or portions of encoded / decoded geometry / attribute videos / images, bitstreams, matrices, variables, and intermediate or final results from the processing of mathematical expressions, formulas, operations, and computational logic.
[0154] In some embodiments, memory within the processor 6010 and / or the encoder / decoder module 6030 may be used to store instructions and provide working memory for processes that may be performed during encoding or decoding.
[0155] However, in other embodiments, memory external to the processing device (e.g., the processing device can be either the processor 6010 or the encoder / decoder module 6030) may be used for one or more of these functions. The external memory may be the memory 6020 and / or the storage device 6040, e.g., dynamic volatile memory and / or non-volatile flash memory. In some embodiments, the external non-volatile flash memory may be used to store the television's operating system. In at least one embodiment, high-speed external dynamic volatile memory such as RAM may be used as working memory for video coding and decoding operations, such as for MPEG-2 Part 2 (also known as ITU-T Recommendation H,262 and ISO / IEC 13818-2, also known as MPEG-2 Video), HEVC (High Efficiency Video coding), or VVC (Versatille Video Coding).
[0156] Input to the elements of system 6000 may be provided through various input devices, as shown in block 6130. Such input devices include, but are not limited to, (i) an RF section that may receive RF signals transmitted over the air, for example by a broadcast station, (ii) a composite input, (iii) a USB input, and / or (iv) an HDMI input.
[0157] In various embodiments, the input devices of block 6130 may have associated respective input processing elements known in the art. For example, the RF section may be associated with elements necessary to (i) select a desired frequency (also referred to as selecting a signal or band-limiting a signal to a frequency band), (ii) down-convert the selected signal, (iii) band-limit again to a narrower frequency band to select a signal frequency band, which in certain embodiments may be referred to as a channel (for example), (iv) demodulate the down-converted and band-limited signal, (v) perform error correction, and (vi) demultiplex to select a desired stream of data packets. The RF section of various embodiments may include one or more elements that perform these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexing device. The RF section may include a tuner that performs these various functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or a near baseband frequency) or to baseband.
[0158] In one set-top box embodiment, the RF section and its associated input processing elements may receive an RF signal transmitted over a wired (e.g., cable) medium. The RF section may then perform frequency selection by filtering, down-converting, and filtering again to a desired frequency band.
[0159] Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions.
[0160] Adding elements may include inserting elements between existing elements, such as inserting amplifiers, analog to digital converters, etc. In various embodiments, the RF section may include an antenna.
[0161] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting the system 6000 to other electronic devices across USB and / or HDMI connections. It should be appreciated that various aspects of the input processing, e.g., Reed-Solomon error correction, may be implemented, for example, in a separate input processing IC or, if desired, in the processor 6010. Similarly, aspects of the USB or HDMI interface processing may be implemented in a separate interface IC or, if desired, in the processor 6010. The demodulated, error corrected, and demultiplexed stream may be provided to various processing elements, including, for example, the processor 6010 and an encoder / decoder 6030 operating in combination with memory and storage elements to process the data stream for presentation to an output device, if desired.
[0162] The various elements of the system 6000 may be provided in a unified housing in which the various elements may be interconnected and transmit data between each other using suitable connection arrangements 6140, such as internal buses known in the art, including an I2C bus, wiring, and printed circuit boards.
[0163] The system 6000 may include a communication interface 6050, which enables communication with other devices over a communication channel 6060. The communication interface 6050 may include, but is not limited to, a transceiver configured to transmit and receive data over the communication channel 6060. The communication interface 6050 may include, but is not limited to, a modem or a network card, and the communication channel 6060 may be implemented in a wired and / or wireless medium, for example.
[0164] In various embodiments, data may be streamed to the system 6000 using a Wi-Fi network, such as IEEE 802.11. The Wi-Fi signal in these embodiments may be received via a communication channel 6060 and communication interface 6050 adapted for Wi-Fi communication. The communication channel 6060 in these embodiments may typically be connected to an access point or router that provides access to external networks, including the Internet, to enable streaming applications and other over-the-top communications.
[0165] Other embodiments may provide streamed data to the system 6000 using a set-top box that delivers data via an HDMI connection in the input block 6130.
[0166] Yet other embodiments may use the RF connection of input block 6130 to provide streamed data to system 6000.
[0167] The streamed data may be used as a method for signaling information for use by the system 5000. The signaling information may include the information INF described above.
[0168] It should be appreciated that signaling may be accomplished in various ways, for example, in various embodiments, one or more syntax elements, flags, etc. may be used to signal information to a corresponding decoder.
[0169] System 6000 may provide output signals to a variety of output devices, including a display 6100, speakers 6110, and other peripheral devices 6120, which in various example embodiments may include one or more of a stand-alone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functionality based on the output of system 3000.
[0170] In various embodiments, control signals may be communicated between the system 6000 and the display 6100, speaker 6110, or other peripheral device 6120 using signaling methods such as AV.Link (Audio / Video Link), CEC (Consumer Electronics Control), or other communication protocols that enable inter-device control with or without user intervention.
[0171] Output devices may be communicatively coupled to system 6000 via dedicated connections through respective interfaces 6070, 6080, and 6090.
[0172] Alternatively, output devices may be connected to the system 6000 via the communication interface 6050 using a communication channel 6060. The display 6100 and speakers 6110 may be integrated into a single unit along with other components of the system 6000 in an electronic device such as a television.
[0173] In various embodiments, the display interface 6070 may include a display driver, such as a timing controller (TCon) chip.
[0174] The display 6100 and speakers 6110 may alternatively be separate from one or more of the other components, for example, if the RF portion of the input 6130 is part of a separate set-top box. In various embodiments in which the display 6100 and speakers 6110 may be external components, the output signal may be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.
[0175] In V-PCC the regular geometric patch of the 2D sample and the first geometric patch of the 2D sample are stored in a geometry image GI, the second geometric patch of the 2D sample are stored in an occupancy map OM, and the regular first and second attribute patches of the 2D sample are stored in an attribute image TI at locations indicated by syntax elements.
[0176] Regular first and second attribute patches of 2D samples stored together in the same video stream may have the advantage of requiring only one video encoder / decoder to encode / decode the attribute information of the 3D samples.
[0177] By design, the coding modes of the regular first and second geometric shapes / attributes respond to different needs. Thus, their attribute formats, scan orders, sizes, and shapes are different. Therefore, the methods for reducing information (or encoding / coding / compression) may require different methods, especially when reducing spatial "intra" redundancy. Moreover, while the geometry / attributes (EOM codewords) of intermediate 3D samples may be aimed to be coded with a lossless video codec, the geometry / attributes of sparse 3D samples (which are less likely to represent the focus of a typical point cloud scene in any way) are typically coded using the coding mode of the first geometric shape / attribute and may be subject to lossy coding or vice versa according to their application.
[0178] As a result, encoding the geometry / attributes of 3D samples requires both of these regular first and second coding modes, which require a specific (non-generic) profiled video codec that contains processes adapted to each of these two coding modes, and additional signaling to indicate to the decoder that one of these coding modes should be used to decode the geometry and attributes of the 3D samples.
[0179] Therefore, allowing multiple coding modes for encoding the 3D samples of a point cloud frame can impact performance (reduced encoder / decoder complexity and bandwidth).
[0180] Generally speaking, at least one of the embodiments provides a method for encoding attributes of an orthogonal projected 3D sample and an intermediate 3D sample, wherein information INF indicates whether at least one first attribute patch of a 2D sample obtained by encoding an attribute of the at least one orthogonal projected 3D sample according to a first attribute coding mode and at least one second attribute patch of a 2D sample of an image obtained by encoding an attribute of the at least one intermediate 3D sample are stored in separate images.
[0181] Distinct images may mean images of the same video stream, but also different times or images of different video streams (not belonging to the same video stream).
[0182] The information INF allows more flexibility in terms of the use of video codecs and also allows better compression performance, since the video codec can be adjusted to code the attributes of the orthogonal or intermediate 3D samples. The video codec adjustment can then take into account the characteristics of the attributes of these 3D samples and / or meet the constraints and requirements of the application. For example, the first and / or second attribute patches of 2D samples can be discarded according to the capabilities of the application or decoder / renderer (for example, some 3D samples may not be useful for point clouds rendered on a small display with an associated low-end SoC for decoding).
[0183] Furthermore, the video streams storing these attributes can be processed in a parallel manner.
[0184] FIG. 7 illustrates an example flowchart of a method for encoding orthogonal 3D samples of a point cloud frame in accordance with at least one embodiment.
[0185] In step 710, the information INF may be set to a first specific value (e.g. 1) to indicate that the at least one first attribute patch FA2DP of the 2D sample (step 3400) and the at least one second attribute patch SA2DP of the 2D sample (step 3400) are stored in separate images, and the information INF may be set to a second specific value (e.g. 0) to indicate that the at least one first attribute patch FA2DP of the 2D sample and the at least one second attribute patch SA2DP of the 2D sample are stored in the same image.
[0186] When the information INF is equal to the first specific value, in step 720, at least one first attribute patch FA2DP of the 2D sample is stored in a first image FAI and at least one second attribute patch SA2DP of the 2D sample is stored in a second image SAI.
[0187] When the information INF is equal to the second specific value, in step 730, at least one first attribute patch FA2DP of the 2D sample and at least one second attribute patch SA2DP of the 2D sample are stored in the same image AI.
[0188] According to an embodiment, as illustrated in FIG. 7a, the information INF comprises a first binary flag F1 and a second binary flag F2.
[0189] In step 740, the first flag F1 may be set to a first specific value (e.g., 1) to indicate that at least one first attribute patch FA2DP (step 3400) of the 2D sample is stored in the first image FA1, and may be set to a second specific value (e.g., 0) to indicate that at least one first attribute patch FA2DP (step 3400) of the 2D sample is stored in the image AI.
[0190] In step 750, the second flag F2 may be set to a first specific value (e.g., 1) to indicate that at least one second attribute patch SA2DP (step 3400) of the 2D sample is stored in the second image FA2, and may be set to a second specific value (e.g., 0) to indicate that at least one second attribute patch FA2DP (step 3400) of the 2D sample is stored in the image AI.
[0191] When the first flag FA and the second flag F2 are equal to 1, the first attribute patch FA2DP and the second attribute patch SA2DP of the 2D sample are stored in separate images.
[0192] When the first flag F1 is equal to 0 and the second flag F2 is equal to 1, the first attribute patch FA2DP of the 2D sample and the regular attribute 2D patch RA2DP are stored in the same image, and the second attribute patch SA2DP of the 2D sample is stored in a separate image.
[0193] When the first flag F1 is equal to 1 and the second flag F2 is equal to 0, the second attribute patch SA2DP of the 2D sample and the regular attribute 2D patch RA2DP are stored in the same image, and the first attribute patch FA2DP of the 2D sample is stored in a separate image.
[0194] When the first flag FA and the second flag F2 are equal to 0, the first attribute patch FA2DP and the second attribute patch SA2DP of the 2D sample are stored in the same image.
[0195] According to a variant, the information INF is valid for groups at picture level, frame / atlas level or patch level.
[0196] According to a variant, a syntax element is added to the bitstream to identify the video codec used to compress the video stream carrying the second attribute patch of 2D samples.
[0197] This codec may be identified via a component codec mapping SEI message or through means other than the V-PCC specification.
[0198] Such a syntax element, indicated by ai_eom_attribute_codec_id[atlas_id], may be added to the attribute information syntax structure conditional on a particular value of another syntax element, such as the syntax element vpcc_eom_patch_separate_video_present_flag of FIG.
[0199] FIG. 8 illustrates an example of a syntax element vpcc_eom_patch_separate_video_present_flag that embeds information INF according to at least one embodiment.
[0200] The syntax element vpcc_eom_patch_separate_video_present_flag may be coded in a parameter set, such as a Sequence Parameter Set (SPS), an Atlas Sequence Parameter Set (ASPS), or a Picture Parameter Set (PPS).
[0201] The elements in FIG. 8 have the following meaning:
[0202] vpcc_eom_patch_separate_video_present_flag[j] is equal to 1 and indicates that the second attribute patch of the 2D sample with index j may be stored in a separate video stream.
[0203] vpcc_eom_patch_separate_video_present_flag[j] is equal to 0, indicating that the second attribute patch of the 2D sample with index j shall not be stored in a separate video stream.
[0204] When vpcc_eom_patch_separate_video_present_flag[j] is not present, it is inferred to be equal to 0.
[0205] epdu_patch_in_eom_video_flag[p] specifies whether the attribute data associated with the second attribute patch of the 2D sample with index p in the current attra tile group is coded into a separate video compared to the attribute data of the intra- and inter-coded patch. If epdu_patch_in_eom_video_flag[p] is equal to 0, the attribute data associated with the second attribute patch of the 2D sample with index j in the current attra tile group is coded into the same video as the attribute data of the intra- and inter-coded patch. If epdu_patch_in_eom_video_flag[p] is equal to 1, the attribute data associated with the second attribute patch of the 2D sample with index j in the current attra tile group is coded into a separate video from the attribute data of the intra- and inter-coded patch. If epdu_patch_in_eom_video_flag[p] is not present, its value shall be inferred to be equal to 0.
[0206] In a variant, the syntax element vpcc_eom_patch_separate_video_present_flag[j] may be common to the first and second attribute patches of a 2D sample.
[0207] Then, when vpcc_separate_video_present_flag[j] is equal to 1, it indicates that the second attribute patch of the 2D sample, the first attribute patch of the 2D sample, and the first geometry patch of the 2D sample for the atlas with index j may be stored in separate images (separate video streams).
[0208] When vpcc_separate_video_present_flag[j] is equal to 0, it indicates that the second attribute patch of the 2D sample, the first attribute patch of the 2D sample, and the first geometry patch of the 2D sample for the atlas with index j shall not be stored in a separate video stream.
[0209] When vpcc_separate_video_present_flag[j] is not present, it is inferred to be equal to 0.
[0210] 1-8, various methods are described herein, each of which includes one or more steps or acts for achieving the described method. To the extent that a specific order of steps or acts is not required for the proper operation of the method, the order and / or use of certain steps and / or acts may be varied or combined.
[0211] Some examples are described with reference to block diagrams and operational flow charts. Each block represents a circuit element, module, or code portion that includes one or more executable instructions for implementing a specified logical function. It should also be noted that in other implementations, the functions noted in the blocks may occur out of the order shown. For example, two blocks shown in succession may in fact be executed substantially simultaneously, or the blocks may be executed in the reverse order, depending on the functionality involved.
[0212] The implementations and aspects described herein may be implemented, for example, as a method or process, an apparatus, a computer program, a data stream, a bit stream, or a signal. Even if discussed only in the context of a single implementation (e.g., discussed only as a method), the implementation of the discussed features may be implemented in other forms (e.g., an apparatus or a computer program).
[0213] The methods may be implemented, for example, in a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communications devices.
[0214] Additionally, the method may be implemented by instructions executed by a processor, and such instructions (and / or data values produced by the implementation) may be stored in a computer-readable storage medium. The computer-readable storage medium may take the form of a computer-readable program product embodied in one or more computer-readable media and embodied with computer-readable program code executable by a computer. As used herein, a computer-readable storage medium may be considered a non-transitory storage medium given the inherent capability of storing information therein as well as providing retrieval of information therefrom. The computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. It should be understood that the following, while providing more specific examples of computer-readable storage media to which the present embodiment may be applied, is merely illustrative and not an exhaustive list as would be readily understood by one of ordinary skill in the art: portable computer diskettes; hard disks; read-only memories (ROMs); erasable programmable read-only memories (EPROMs or flash memories); portable compact disc read-only memories (CD-ROMs); optical storage devices; magnetic storage devices; or any suitable combination of the foregoing.
[0215] The instructions may form an application program tangibly embodied on a processor-readable medium.
[0216] The instructions may be, for example, hardware, firmware, software, or a combination thereof. The instructions may be found, for example, in an operating system, a separate application, or a combination of the two. Thus, a processor may be characterized as, for example, both a device configured to perform a process and a device that includes a processor-readable medium (such as a storage device) having instructions for performing a process. Furthermore, the processor-readable medium may store data values produced by an implementation in addition to or in place of instructions.
[0217] The devices may be implemented, for example, with appropriate hardware, software, and firmware. Examples of such devices include personal computers, laptops, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, head-mounted display devices (HMDs, see-through glasses), projectors (beamers), "immersive virtual reality caves" (systems including multiple displays), servers, video encoders, video decoders, post-processing output from video decoders, pre-processors providing input to video encoders, web servers, set-top boxes, and any other devices for processing point clouds, videos, or images, or other communication devices. As should be clear, the equipment may be mobile, and even installed in a mobile vehicle.
[0218] The computer software may be implemented by the processor 6010, or by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments may also be implemented by one or more integrated circuits. The memory 6020 may be of any type suitable for the technological environment, and may be implemented using any suitable data storage technology, such as, as non-limiting examples, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memories, and removable memories. The processor 6010 may be of any type suitable for the technological environment, and may include, as non-limiting examples, one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.
[0219] As will be apparent to one of ordinary skill in the art, implementations may generate a variety of signals formatted to carry information that may be, for example, stored or transmitted. Information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal may be formatted to carry a bit stream of the described embodiments. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using a radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information that the signal carries may be, for example, analog information or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is known. The signal may be stored in a processor-readable medium.
[0220] The terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting. As used herein, the singular forms "a," "an," and "the" may be intended to include the plural unless the context clearly indicates otherwise. It will be further understood that as used herein, the terms "includes" and / or "including" may specify the presence of stated, e.g., features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Furthermore, when an element is referred to as "responsive to" or "connected" to another element, it may be directly responsive or connected to the other element, or intervening elements may be present. In contrast, when an element is referred to as "directly responsive to" or "directly connected" to another element, there are no intervening elements.
[0221] For example, in the case of "A / B," "A and / or B," and "at least one of A and B," it should be understood that use of any of the symbols / terms " / ," "and / or," and "at least one of" may be intended to include the selection of only the first enumerated alternative (A), or the selection of only the second enumerated alternative (B), or the selection of both alternatives (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C," such language is intended to include the selection of only the first enumerated alternative (A), or the selection of only the second enumerated alternative (B), or the selection of only the third enumerated alternative (C), or the selection of only the first and second enumerated alternatives (A and B), or the selection of only the first and third enumerated alternatives (A and C), or the selection of only the second and third enumerated alternatives (B and C), or the selection of all three alternatives (A and B and C). This may be expanded as many times as the number of items listed, as would be apparent to one of ordinary skill in the art of this and related arts.
[0222] Various numerical values may be used in this application. The specific values may be for illustrative purposes, and the described aspects are not limited to these specific values.
[0223] In this specification, terms such as first, second, etc. may be used to describe various elements, but it will be understood that these elements are not limited by these terms. These terms are used only to distinguish one element from another element. For example, a first element may be called a second element, and similarly, a second element may be called a first element, without departing from the teachings of this application. There is no order between a first element and a second element.
[0224] References to "one embodiment" or "embodiment" or "one implementation" or "implementation," as well as other variations thereof, are frequently used to convey that a particular feature, structure, characteristic, etc. (described in connection with the embodiment / implementation) is included in at least one embodiment / implementation. Thus, appearances of the phrases "in one embodiment" or "in an embodiment" or "in one implementation" or "in an implementation" appearing in various places throughout this specification, as well as any other variations thereof, do not necessarily all refer to the same embodiment.
[0225] Similarly, references herein to "according to an embodiment / example / implementation" or "in an embodiment / example / implementation," as well as other variations thereof, are frequently used to convey that a particular feature, structure, or characteristic (described in connection with an embodiment / example / implementation) may be included in at least one embodiment / example / implementation. Thus, appearances of the phrases "according to an embodiment / example / implementation" or "in an embodiment / example / implementation" in various places throughout this specification do not necessarily all refer to the same embodiment / example / implementation, nor do separate or alternative embodiments / examples / implementations necessarily exclude other embodiments / examples / implementations from one another.
[0226] Reference numerals appearing in the claims are merely exemplary and shall have no limiting effect on the scope of the claims. Although not expressly described, the present embodiments / examples and variants may be used in any combination or subcombination.
[0227] Where a diagram is presented as a flow diagram, it should be understood that the diagram also provides a block diagram of the corresponding apparatus. Similarly, where a diagram is presented as a block diagram, it should be understood that the diagram also provides a flow diagram of the corresponding method / process.
[0228] Some figures include arrows on communication paths to indicate a primary direction of communication, however, it should be understood that communication may occur in the opposite direction to that of the illustrated arrow.
[0229] Various implementations involve decoding. As used herein, "decoding" may encompass all or part of the processes performed, for example, on received point cloud frames (including, in some cases, a received bitstream encoding one or more point cloud frames) to generate a final output suitable for display or for further processing in a reconstructed point cloud domain. In various embodiments, such processes include one or more of the processes typically performed by an image-based decoder. In various embodiments, such processes also or alternatively include, for example, processes performed by a decoder of the various implementations described herein.
[0230] As a further example, in one embodiment, "decoding" may refer only to entropy decoding, in another embodiment, "decoding" may refer only to differential decoding, and in another embodiment, "decoding" may refer to a combination of entropy and differential decoding. Whether the phrase "decoding process" may be intended to refer specifically to a subset of operations or to the broader decoding process in general will be clear based on the context of a particular description and will be well understood by one of ordinary skill in the art.
[0231] Various implementations involve encoding. In a manner similar to the above discussion regarding "decoding," as used herein, "encoding" may encompass all or a portion of the processes performed on an input point cloud frame, for example, to generate an encoded bitstream. In various embodiments, such processes include one or more of the processes typically performed by an image-based decoder. In various embodiments, such processes also or alternatively include processes performed by an encoder of the various implementations described herein.
[0232] As a further example, in one embodiment, "encoding" may refer only to entropy encoding, in another embodiment, "encoding" may refer only to differential encoding, and in another embodiment, "encoding" may refer to a combination of differential and entropy encoding. Whether the phrase "encoding process" may be intended to refer specifically to a subset of operations or to a broader encoding process in general will be clear based on the context of a particular description and will be well understood by one of ordinary skill in the art.
[0233] It should be noted that as used herein, syntax elements are descriptive terms, and therefore they do not preclude the use of other syntax element names.
[0234] Various embodiments refer to rate-distortion optimization. In particular, during the encoding process, a balance or trade-off between rate and distortion is considered, which usually often gives a constraint on the computational complexity. Rate-distortion optimization may be formulated to minimize a rate-distortion function, which is usually a weighted sum of rate and distortion. There are different approaches to solving the rate-distortion optimization problem. For example, the approach may be based on an extensive test of all encoding options, including all considered modes or coding parameter values, and a full evaluation of their coding cost and associated distortion of the reconstructed signal after coding and decoding. Also, faster approaches may be used to reduce the coding complexity, in particular with an approximate distortion calculation based on a predicted or predicted residual signal, rather than on the reconstructed signal. Also, a mixture of these two approaches may be used, such as by using approximate distortion of only some of the possible encoding options, and full distortion of other encoding options. Other approaches evaluate only a subset of the possible encoding options. More generally, many approaches use any of a variety of techniques to perform the optimization, but the optimization is not necessarily a full evaluation of both the coding cost and the associated distortion.
[0235] Additionally, the application may refer to "determining" various information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from memory.
[0236] Additionally, the application may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0237] Additionally, the present application may refer to "receiving" various information. Receiving, like "accessing," is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from a memory). Furthermore, "receiving" typically involves in some way, an operation such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0238] Also, as used herein, the word "signaling" refers to, among other things, indicating something to a corresponding decoder. For example, in an embodiment, the encoder signals a specific information INF, which may be carried by the syntax element vpcc_eom_patch_separate_video_present_flag. In this way, in an embodiment, the same parameters may be used at both the encoder side and the decoder side. Thus, for example, the encoder may send (explicitly signal) a specific parameter to the decoder so that the decoder may use the same specific parameter. Conversely, if the decoder already has the specific parameter as well as other parameters, signaling may be used without sending (implicitly signaling) to simply enable the decoder to know and select the specific parameter. By avoiding the sending of any actual function, bit savings are realized in various embodiments. It should be understood that signaling may be achieved in various ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder. Although the above is about the verb form of the word "signaling", the word "signal" may also be used as a noun in this specification.
[0239] Multiple implementations have been described. Nevertheless, it will be understood that various modifications may be made. For example, elements of different implementations may be combined, supplemented, modified, or deleted to produce other implementations. Additionally, those skilled in the art will appreciate that other structures and processes may be substituted for those disclosed, such that the resulting implementations perform at least substantially the same functions, in at least substantially the same manner, to achieve at least substantially the same results as the disclosed implementations. Accordingly, these and other implementations are contemplated by the present application.
Claims
1. decoding data representing a point cloud from a first video stream, the data including at least one first attribute patch of a 2D sample, the at least one first attribute patch of the 2D sample including attributes of an orthogonally projected 3D sample of the point cloud, the 3D coordinates of the sample being derived from coordinates of a pixel in a geometry image and a value of the pixel in the geometry image; decoding a syntax element indicating whether at least one second attribute patch of a 2D sample, at least one geometry patch of a 2D sample, and at least one third attribute patch of a 2D sample can be stored in a second video stream separate from the first video stream; at least one geometric patch of the 2D sample includes 3D coordinates of at least one 3D sample of a point cloud, the 3D coordinates being derived from pixel values of the at least one geometric patch; at least one second attribute patch of the 2D sample includes attribute values of the at least one 3D sample whose 3D coordinates are derived from pixel values of the at least one geometric shape patch; at least one third attribute patch of the 2D samples includes attributes of at least one intermediate 3D sample of the point cloud, the 3D coordinates of which are derived from coordinates of a pixel in a geometry image and from one bit of a codeword indicating a position of the intermediate 3D sample along a projection line between two 3D samples of the point cloud projected orthogonally along the projection line; decoding at least one of at least one geometry patch of the 2D sample, at least one second attribute patch of the 2D sample, or at least one third attribute patch of the 2D sample from the first video stream or the second video stream based on the syntax element; A method comprising:
2. The method of claim 1 , wherein the syntax element is a flag common to the at least one second attribute patch and the at least one third attribute patch.
3. 2. The method of claim 1, further comprising: decoding a flag based on the syntax element, the flag indicating whether attribute data associated with at least one third attribute patch of the 2D sample is encoded in the first video stream or in the second video stream.
4. The method of claim 1 , wherein the syntax elements are decoded from one of a sequence parameter set, an atlas sequence parameter set, or a picture parameter set.
5. 2. The method of claim 1, further comprising: decoding another syntax element that identifies a video codec used to compress the video stream that carries at least one third attribute patch of the 2D sample.
6. 2. The method of claim 1, wherein the first video stream and the second video stream are hierarchically organized in picture-level, frame-level, and patch-level groups, and the syntax element is valid at either the picture-level, frame-level, atlas-level, or patch-level group.
7. The method of claim 1 , wherein the codewords are packed to form at least one other geometric patch of 2D samples that is stored in an occupancy map.
8. 1. An apparatus comprising one or more processors, the one or more processors comprising: decoding data representing a point cloud from a first video stream, the data including at least one first attribute patch of a 2D sample, the at least one first attribute patch of the 2D sample including attributes of an orthogonally projected 3D sample of the point cloud, the 3D coordinates of the sample being derived from coordinates of a pixel in a geometry image and a value of the pixel in the geometry image; decoding a syntax element indicating whether at least one second attribute patch of a 2D sample, at least one geometry patch of a 2D sample, and at least one third attribute patch of a 2D sample can be stored in a second video stream separate from the first video stream; at least one geometric patch of the 2D sample includes 3D coordinates of at least one 3D sample of a point cloud, the 3D coordinates being derived from pixel values of the at least one geometric patch; at least one second attribute patch of the 2D sample includes attribute values of the at least one 3D sample whose 3D coordinates are derived from pixel values of the at least one geometric shape patch; at least one third attribute patch of the 2D samples includes attributes of at least one intermediate 3D sample of the point cloud, the 3D coordinates of which are derived from coordinates of a pixel in a geometry image and from one bit of a codeword indicating a position of the intermediate 3D sample along a projection line between two 3D samples of the point cloud projected orthogonally along the projection line; decoding at least one of at least one geometry patch of the 2D sample, at least one second attribute patch of the 2D sample, or at least one third attribute patch of the 2D sample from the first video stream or the second video stream based on the syntax element; An apparatus configured to:
9. 1. An apparatus comprising one or more processors, the one or more processors comprising: encoding data representing a point cloud in a first video stream, the data including at least one first attribute patch of a 2D sample, the at least one first attribute patch of the 2D sample including attributes of an orthogonally projected 3D sample of the point cloud, the 3D coordinates of the sample being derived from coordinates of a pixel in a geometry image and a value of the pixel in the geometry image; encoding a syntax element indicating whether at least one second attribute patch of a 2D sample, at least one geometry patch of a 2D sample, and at least one third attribute patch of a 2D sample may be stored in a second video stream separate from the first video stream; at least one geometric patch of the 2D sample includes 3D coordinates of at least one 3D sample of a point cloud, the 3D coordinates being derived from pixel values of the at least one geometric patch; at least one second attribute patch of the 2D sample includes attribute values of the at least one 3D sample whose 3D coordinates are derived from pixel values of the at least one geometric shape patch; at least one third attribute patch of the 2D samples includes attributes of at least one intermediate 3D sample of the point cloud, the 3D coordinates of which are derived from coordinates of a pixel in a geometry image and from one bit of a codeword indicating a position of the intermediate 3D sample along a projection line between two 3D samples of the point cloud projected orthogonally along the projection line; encoding at least one of at least one geometry patch of the 2D sample, at least one second attribute patch of the 2D sample, or at least one third attribute patch of the 2D sample in the first video stream or in the second video stream based on the syntax element; An apparatus configured to:
10. encoding data representing a point cloud in a first video stream, the data including at least one first attribute patch of a 2D sample, the at least one first attribute patch of the 2D sample including attributes of an orthogonally projected 3D sample of the point cloud, the 3D coordinates of the sample being derived from coordinates of a pixel in a geometry image and a value of the pixel in the geometry image; encoding a syntax element indicating whether at least one second attribute patch of a 2D sample, at least one geometry patch of a 2D sample, and at least one second attribute patch of a 2D sample may be stored in a second video stream separate from the first video stream; at least one geometric patch of the 2D sample includes 3D coordinates of at least one 3D sample of a point cloud, the 3D coordinates being derived from pixel values of the at least one geometric patch; at least one second attribute patch of the 2D sample includes attribute values of the at least one 3D sample whose 3D coordinates are derived from pixel values of the at least one geometric shape patch; at least one third attribute patch of the 2D samples includes attributes of at least one intermediate 3D sample of the point cloud, the 3D coordinates of which are derived from coordinates of a pixel in a geometry image and from one bit of a codeword indicating a position of the intermediate 3D sample along a projection line between two 3D samples of the point cloud projected orthogonally along the projection line; encoding at least one of at least one geometry patch of the 2D sample, at least one second attribute patch of the 2D sample, or at least one third attribute patch of the 2D sample in the first video stream or in the second video stream based on the syntax element; A method comprising:
11. A non-transitory computer readable medium comprising instructions for causing one or more processors to perform the method of any one of claims 1-7 or 10.
12. 10. The apparatus of claim 8 or 9, wherein the syntax element is a flag common to the at least one second attribute patch and the at least one third attribute patch.
13. 9. The apparatus of claim 8, wherein the one or more processors are further configured to decode a flag based on the syntax element, the flag indicating whether attribute data associated with at least one third attribute patch of the 2D sample is encoded in the first video stream or in the second video stream.
14. The apparatus of claim 8 , wherein the syntax elements are decoded from one of a sequence parameter set, an atlas sequence parameter set, or a picture parameter set.
15. 10. The apparatus of claim 8, wherein the one or more processors are further configured to decode another syntax element that identifies a video codec used to compress the video stream that carries at least one third attribute patch of the 2D sample.
16. 10. The apparatus of claim 9, wherein the one or more processors are further configured to decode a flag based on the syntax element, the flag indicating whether attribute data associated with at least one third attribute patch of the 2D sample is encoded in the first video stream or in the second video stream.
17. 10. The apparatus of claim 9, wherein the one or more processors are further configured to decode another syntax element that identifies a video codec used to compress the video stream that carries at least one third attribute patch of the 2D sample.