Encoding and decoding of point clouds using patches of intermediate samples

The method addresses the challenge of encoding and decoding 3D sample attributes by structuring video streams with separate 2D attribute patches and encoding storage information, resulting in efficient compression and quality preservation for dynamic point clouds.

JP7690663B2Active Publication Date: 2025-06-10INTERDIGITALCE PATENT HLDG SAS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024136990
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-09-24
Filing Date
2024-08-16
Publication Date
2025-06-10
Estimated Expiration
2040-09-15

AI Technical Summary

Technical Problem

Existing technologies face challenges in efficiently encoding and decoding attributes of 3D samples within video streams, particularly for dynamic point clouds, which require effective compression to maintain quality while minimizing bitrate.

Method used

A method for encoding attributes of orthographically projected 3D samples as 2D attribute patches, where information indicating whether these patches are stored in separate images is encoded, allowing for flexible video stream structuring and compression.

Benefits of technology

This approach enables efficient compression of dynamic point clouds, allowing for better quality preservation and reduced bitrate, thereby facilitating the distribution of immersive world data formats.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007690663000001
    Figure 0007690663000001
  • Figure 0007690663000002
    Figure 0007690663000002
  • Figure 0007690663000003
    Figure 0007690663000003
Patent Text Reader

Abstract

To provide a method and device for encoding / decoding of attributes of orthogonally projected 3D samples and attributes of in-between 3D samples.SOLUTION: A method of encoding of attributes of orthogonally projected 3D samples comprises: encoding an attribute of orthogonally projected 3D samples, as at least one first attribute patch of 2D samples of an image; and encoding an attribute of in-between 3D samples located between two orthogonally projected 3D samples along the same projection line, as at least one second attribute patch of 2D samples in an image. The step of encoding the attribute of the orthogonally projected 3D samples includes encoding information indicating whether the at least one first attribute patch of 2D samples and the at least one second attribute patch of 2D samples are stored in separate images.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] At least one of the embodiments generally relates to the processing of point clouds. In particular, the encoding / decoding of attributes of 3D samples within / from separate video streams is disclosed.

Background Art

[0002] This section is intended to introduce the reader to various aspects of techniques that may be related to various aspects of at least one of the embodiments described and / or claimed below. This discussion is thought to be helpful in providing background information to the reader and facilitating a better understanding of the various aspects of at least one embodiment.

[0003] Point clouds can be used for various purposes such as cultural heritage / buildings, scanning objects such as statues or buildings therein in 3D, and sharing the spatial configuration of the objects without sending or visiting the objects. Also, when the object can be destroyed, for example, when a temple can be destroyed by an earthquake, the point cloud is a way to reliably preserve the knowledge of the object. Such point clouds are usually static, colored, and huge.

[0004] Another example of use is in topography and mapping methods, where the use of 3D representations enables maps that are not limited to a plane and can include undulations. Google Maps is currently a good example of a 3D map but uses a mesh instead of a point cloud. Nevertheless, point clouds can be a suitable data format for 3D maps, and such point clouds are usually static, colored, and huge.

[0005] The automotive industry and self-driving vehicles are also fields where point clouds can be used. Self-driving vehicles need to be able to "explore" their environment and make good driving decisions based on the reality of their immediate neighborhood. Typical sensors such as LIDAR (Light Detection and Ranging) generate dynamic point clouds that are used by the decision engine. These point clouds are not intended for human viewing; they are usually small, not necessarily color-coded, and dynamic with a high capture frequency. These point clouds can have other attributes such as the reflectivity provided by LIDAR when this attribute provides good information about the material of the detected object and can help in making decisions.

[0006] Virtual reality and immersive worlds have recently become a topic of discussion and are predicted by many as the future of 2D flat video. The basic idea is to immerse the viewer within an environment that surrounds the viewer, in contrast to standard TVs where the viewer can only look at the virtual world in front of the viewer. Depending on the degree of freedom of the viewer within the environment, there are several levels of immersion. Point clouds are a good candidate for a format to deliver Virtual Reality (VR) worlds.

[0007] In many applications, it is important that dynamic point clouds can be delivered to end-users (or stored within a server) by consuming only a reasonable amount of bitrate (or storage space for storage applications) while maintaining an acceptable (or preferably very good) quality of experience. Efficient compression of these dynamic point clouds is an important point for enabling many immersive world distribution networks.

[0008] At least one embodiment has been devised with the above in mind. SUMMARY OF THE INVENTION

[0009] The following presents a simplified overview of at least one of the embodiments in order to provide a basic understanding of some aspects of the present disclosure. This overview is not an extensive overview of the embodiments. It is not intended to identify key or essential elements of the embodiments. The following overview merely presents some aspects of at least one of the embodiments in a simplified form as a prelude to a more detailed description provided elsewhere in this document.

[0010] According to a general aspect of at least one embodiment, a method for encoding attributes of an orthographically projected 3D sample, wherein the attributes of the orthographically projected 3D sample are encoded as at least one first attribute patch of 2D samples of an image, and the attributes of an intermediate 3D sample located between two orthographically projected 3D samples along the same projection line are encoded as at least one second attribute patch of 2D samples in the image, the method including encoding information indicating whether at least one first attribute patch of the 2D samples and at least one second attribute patch of the 2D samples are stored in separate images.

[0011] According to an embodiment, a video stream is hierarchically structured in groups at the picture level, frame level, and patch level, and the information is valid in any of the groups at the picture level, frame level, atlas level, or patch level.

[0012] According to an embodiment, the information is a first flag indicating whether at least one first attribute patch of the 2D samples is stored in a first image or whether at least one first attribute patch of the 2D samples is stored together with other attribute patches of the 2D samples in a second image, and a second flag indicating whether at least one second attribute patch of the 2D samples is stored in a third image or whether at least one second attribute patch of the 2D samples is stored together with other attribute patches of the 2D samples in the second image.

[0013] According to an embodiment, the method further includes encoding additional information indicating how the separate images are compressed.

[0014] According to a general aspect of at least one embodiment, a method for decoding an attribute of a 3D sample, wherein the attribute of the 3D sample is decoded from at least one first attribute patch of 2D samples of an image, and an attribute of an intermediate 3D sample located between two 3D samples along the same projection line is decoded as at least one second attribute patch of 2D samples in the image, the method including decoding information indicating whether at least one first attribute patch of the 2D samples and at least one second attribute patch of the 2D samples are stored in separate images.

[0015] According to an embodiment, the video stream is hierarchically structured in groups at the picture level, frame level, and patch level, and the information is valid in any of the groups at the picture level, frame level, atlas level, or patch level.

[0016] According to an embodiment, the information is a first flag indicating whether at least one first attribute patch of the 2D samples is stored in a first image or whether at least one first attribute patch of the 2D samples is stored together with other attribute patches of the 2D samples in a second image, and a second flag indicating whether at least one second attribute patch of the 2D samples is stored in a third image or whether at least one second attribute patch of the 2D samples is stored together with other attribute patches of the 2D samples in the second image.

[0017] According to an embodiment, the method further includes encoding additional information indicating how the separate images are compressed.

[0018] One or more of at least one embodiment also provide an apparatus, a bitstream, a computer program product, and a non-transitory computer-readable medium.

[0019] At least one specificity of at least one of the present embodiments, as well as at least one other object, advantage, feature, and use of at least one of the present embodiments, will become apparent from the following description of the examples taken in conjunction with the accompanying drawings.

Brief Description of the Drawings

[0020] In the drawings, examples of several embodiments are illustrated. The drawings show the following.

Figure 1

Figure 2

Figure 3

Figure 3a

Figure 3b

Figure 4

Figure 5

Figure 6

Figure 7

Figure 7a

Figure 8

[0021] At least one of the present embodiments will be described in more detail below with reference to the accompanying drawings, in which examples of at least one of the present embodiments are shown. However, the embodiments can be embodied in many alternative forms and should not be construed as limited to the examples described herein. Therefore, it should be understood that there is no intention to limit the embodiments to the specific forms disclosed. In contrast, the present disclosure is intended to cover all modifications, equivalents, and alternatives within the spirit and scope of the present application.

[0022] It should be understood that when a figure is presented as a flowchart, the figure also provides a block diagram of the corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that the figure also provides a flowchart of the corresponding method / process.

[0023] Like or identical elements of the same figure are referenced using the same reference numbers.

[0024] Some figures represent syntax tables that are widely used in V-PCC to define the structure of a bitstream conforming to V-PCC. In those syntax tables, the term "..." indicates the unchanging part of the syntax with respect to the original definition given in V-PCC and the part removed in the figure for ease of reading. The terms in bold in the figure indicate that the value of this term can be obtained by analyzing the bitstream. The right column of the syntax table indicates the number of bits used to encode the data of the syntax element. For example, u(4) indicates that 4 bits are used to encode the data, u(8) indicates that 8 bits are used to encode the data, and ae(v) indicates a syntax element encoded with context-adaptive arithmetic entropy coding.

[0025] The aspects described and contemplated below can be implemented in many different forms. In the following FIGS. 1 to 8, several embodiments are provided, but other embodiments are contemplated, and the consideration of FIGS. 1 to 8 does not limit the scope of the implementation forms.

[0026] At least one of the aspects generally relates to point cloud encoding and decoding, and at least one other aspect generally relates to the transmission of the generated or encoded bitstream.

[0027] More precisely, the various methods and other aspects described herein can be used to modify modules, such as the image-based encoder 3000 and decoder 4000 as shown in FIGS. 1 to 8.

[0028] Furthermore, this aspect is not limited to MPEG standard specifications such as MPEG-I Part 5 related to point cloud compression. For example, it can be applied to other standard specifications and recommendations, whether existing or to be developed in the future, and extensions of any such standard specifications and recommendations (including MPEG-I Part 5). Unless otherwise indicated or technically excluded, the aspects described in this application can be used individually or in combination.

[0029] Hereinafter, image data refers to data, for example, one or several arrays of 2D samples in a specific image / video format. The specific image / video format may specify information related to the pixel values of the image (or video). The specific image / video format may also specify information that can be used by a display and / or any other device to visualize and / or decode, for example, the image (or video). An image typically has the shape of a first array of 2D samples and usually includes a first component representing the luminance (or luma) of the image. The image may also include a second component and a third component, typically in the shape of other arrays of 2D samples, representing the chrominance (or chroma) of the image. Some embodiments use a set of 2D color sample arrays, such as the conventional three-color RGB representation, to represent the same information.

[0030] Pixel values are represented in one or more embodiments by a vector of C values, where C is the number of components. Each value of the vector is typically represented using the number of bits that can define the dynamic range of the pixel value.

[0031] An image block means a set of pixels belonging to an image. The pixel values of an image block (or image block data) refer to the values of the pixels belonging to this image block. An image block can have any shape, but a rectangle is common.

[0032] A point cloud can be represented by a dataset of 3D samples in a 3D volume space, and the dataset of 3D samples has unique coordinates and may also have one or more attributes.

[0033] A 3D sample can include information defining the geometric shape of a point cloud of 3D points that can be represented by X, Y, and Z coordinates in 3D space. The 3D sample can also include information defining one or more associated attributes such as, for example, a color represented in an RGB or YUV color space, transparency, reflectivity, two component normal vectors, or any feature representing a feature of this sample. For example, a 3D sample can include information defining six components (X, Y, Z, R, G, B), or alternatively (X, Y, Z, y, U, V), where (X, Y, Z) define the coordinates of a 3D point in 3D space and (R, G, B) or (y, U, V) define the color of this 3D point. Attributes of the same type can be present multiple times. For example, multiple color attributes can provide color information from different viewpoints.

[0034] A 2D sample can include information defining the geometric shape of an orthogonally projected 3D sample that can be represented by three coordinates (u, v, Z), where (u, v) are coordinates in the 2D space of the orthogonally projected 3D and Z is the Euclidean distance between the 3D sample and the projection plane onto which the 3D sample is orthogonally projected. Z is typically represented as a depth value. The 3D sample can also include information defining one or more associated attributes such as, for example, a color represented in an RGB or YUV color space, transparency, reflectivity, two component normal vectors, or any feature representing a feature of this orthogonally projected 3D sample.

[0035] Accordingly, a 2D sample can include information defining the geometric shape and attributes of an orthogonally projected 3D sample by (u, v, Z, R, G, B) or alternatively (u, v, Z, y, U, V).

[0036] A point cloud can be static or dynamic depending on whether the group changes over time. Examples of static or dynamic point clouds are typically shown as point cloud frames. In the case of a dynamic point cloud, note that the number of points is generally not constant, but in contrast generally changes over time. More generally, a point cloud can be considered dynamic if something, such as, for example, the number of points, the position of one or more points, or any attribute of any point, changes over time.

[0037] FIG. 1 illustrates a schematic block diagram of an example of a two - layer - based point cloud encoding structure 1000 according to at least one of the present embodiments.

[0038] The two - layer - based point cloud encoding structure 1000 may provide a bitstream B representing an input point cloud frame IPCF. In some cases, the input point cloud frame IPCF represents a frame of a dynamic point cloud. Then, the frame of the dynamic point cloud may be encoded by the two - layer - based point cloud encoding structure 1000.

[0039] Then, by completely combining the bitstreams representing each frame of the dynamic point cloud, a video stream for representing the dynamic point cloud may be obtained.

[0040] Basically, the two - layer - based point cloud code structure 1000 may provide the ability to structure the bitstream B as a base layer BL and an enhancement layer EL. The base layer BL may provide an irreversible representation of the input point cloud frame IPCF, and the enhancement layer EL may provide a higher - quality (optionally reversible) representation by encoding the isolated points not represented by the base layer BL.

[0041] The base layer BL may be provided by an image - based encoder 3000 as illustrated in FIG. 3. The image - based encoder 3000 may provide a geometric shape / attribute image representing the geometric shape / attributes of the 3D samples of the input point cloud frame IPCF. It may be possible to discard the isolated 3D samples. The base layer BL may be decoded by an image - based decoder 4000 as illustrated in FIG. 4, and the image - based decoder may provide an intermediate reconstructed point cloud frame IRPCF.

[0042] Next, returning to the two-layer base point cloud encoding 1000 of FIG. 1, the comparator COMP can compare the 3D samples of the input point cloud frame IPCF with the 3D samples of the intermediate reconstructed point cloud frame IRPCF to detect / locate missed / isolated 3D samples. Next, the encoder ENC can encode the missed 3D samples and can provide an enhancement layer EL. Finally, the base layer BL and the enhancement layer EL can be multiplexed together by a multiplexing device MUX to generate a bitstream B.

[0043] According to an embodiment, the encoder ENC may include a detector that can detect 3D reference samples of the intermediate reconstructed point cloud frame IRPCF and associate them with the missed 3D samples M.

[0044] For example, the 3D reference sample R associated with the missed 3D sample M can be the one closest adjacent to M according to a given metric.

[0045] According to an embodiment, the encoder ENC can then encode the spatial positions and attributes of the missed 3D samples M as differences determined according to the spatial positions and attributes of the 3D reference sample R.

[0046] In a variant, those differences can be encoded separately.

[0047] For example, in the case of the missed 3D sample M, using the spatial coordinates x(M), y(M), and z(M), the x - coordinate position difference Dx(M), y - coordinate position difference Dy(M), z - coordinate position difference Dz(M), R - attribute component difference Dr(M), G - attribute component difference Dg(M), and B - attribute component difference Db(M) can be calculated as follows. Dx(M)=x(M)-x(R), where x(M) is the x - coordinate of the 3D sample M in the geometric shape image given by FIG. 3, and the same applies to R respectively, Dy(M)=y(M)-y(R) Where y(M) is the y - coordinate of the 3D sample M in the geometric shape image given by FIG. 3, and the same applies to R respectively. Dz(M)=z(M)-z(R) Where z(M) is the z - coordinate of the 3D sample M in the geometric shape image given by FIG. 3, and the same applies to R respectively. Dr(M)=R(M)-R(R). Where R(M) and R(R) are the r - color components of the color attributes of the 3D samples M and R respectively. Dg(M)=G(M)-G(R). Where G(M) and G(R) are the g - color components of the color attributes of the 3D samples M and R respectively. Db(M)=B(M)-B(R). Where B(M) and B(R) are the b - color components of the color attributes of the 3D samples M and R respectively.

[0048] FIG. 2 illustrates a schematic block diagram of an example of a two - layer - based point - group decoding structure 2000 according to at least one of the present embodiments.

[0049] The operation of the two - layer - based point - group decoding structure 2000 depends on its capabilities.

[0050] A two - layer - based point - group decoding structure 2000 with limited capabilities can access only the base layer BL from the bitstream B by using a multiplex separation device DMUX, and then, as illustrated in FIG. 4, by decoding the base layer BL by a point - group decoder 4000, a faithful (but irreversible) version IRPCF of the input point - group frame IPCF can be provided.

[0051] The two-layer-based point cloud decoding structure 2000 with full capabilities can access both the base layer BL and the enhancement layer EL from the bitstream B by using a multiplexing separation device DMUX. As illustrated in FIG. 4, the point cloud decoder 4000 can determine an intermediate reconstructed point cloud frame IRPCF from the base layer BL. The decoder DEC can determine a complementary point cloud frame CPCF from the enhancement layer EL. Then, the combiner COMB can combine the intermediate reconstructed point cloud frame IRPCF and the complementary point cloud frame CPCF together, and thus provide a higher-quality (possibly reversible) representation (reconstruction) CRPCF of the input point cloud frame IPCF.

[0052] FIG. 3 illustrates a schematic block diagram of an example of an image-based point cloud encoder 3000 according to at least one of the present embodiments.

[0053] The image-based point cloud encoder 3000 utilizes an existing video codec and compresses the geometric shapes and attribute information of the 3D samples of the input dynamic point cloud using different video streams.

[0054] In certain embodiments, two video streams, namely, one video stream for capturing the geometric information of the 3D samples of the input point cloud and another video stream for capturing the attribute information of these 3D samples, can be generated and compressed using an existing video codec. Examples of existing video codecs include the HEVC main profile encoder / decoder (ITU-T H.265, ITU Telecommunication Standardization Sector (02 / 2018), Series H, i.e., Audiovisual and Multimedia Systems, Infrastructure for Audiovisual Services - Coding of Moving Pictures, High Efficiency Video Coding, Recommendation ITU-T H.265).

[0055] Additional metadata used to interpret the two video streams is also typically generated and compressed separately. Such additional metadata includes, for example, an occupancy map OM and / or auxiliary patch information PI.

[0056] Subsequently, the generated video stream and metadata can be multiplexed together to generate a composite stream.

[0057] Note that metadata typically represents only a small amount of the overall information. Most of the information is within the video stream.

[0058] An example of such a point cloud coding / decoding process is given by the test model category 2 algorithm (also denoted as V-PCC) that implements the MPEG draft standard, as defined in ISO / IEC JTC1 / SC29 / WG11, Information technology - Coded Representation of Immersive Media - Part 5: Video-based Point Cloud Compression, CD stage, SCD_d39, ISO / IEC 23090-5.

[0059] In step 3100, the module PGM can generate at least one patch of 2D samples by orthogonally projecting the 3D samples of the frame IPCF of the input point cloud frame onto 2D samples on the projection plane, using a strategy that provides the best compression.

[0060] A patch of 2D samples can be defined as a set of 2D samples that share common characteristics.

[0061] For example, in V-PCC, as described, for example, in the report by Hoppe et al. (Hugues Hoppe, Tony DeRose, Tom Duchamp, John McDonald, Werner Stuetzle, Surface reconstruction from unorganized points, ACM SIGGRAPH 1992 Proceedings, 71-78), the normal for each 3D sample is first estimated. Next, an initial clustering of the 3D samples is obtained by associating each 3D sample with one of the six oriented surfaces of the 3D bounding box surrounding the 3D samples. More precisely, each 3D sample is clustered and associated with the oriented surface having the closest normal (maximizing the dot product of the point normal and the surface normal). The 3D samples are then projected orthogonally onto their associated planes (projection planes). A set of 3D samples forming a connected region within those planes is referred to as a connected component. Thus, a connected component is a set of at least one 3D sample having similar normals and the same associated oriented surface. The initial clustering is then refined by repeatedly updating the cluster associated with each 3D sample based on its normal and the clusters of its closest neighboring samples. The final step consists of generating one patch of 2D samples from each connected component, which is done by projecting the 3D samples of each connected component onto the oriented surface associated with that connected component.

[0062] The 2D samples of the patch of 2D samples then share the same normal and the same oriented surface and are positioned close to each other.

[0063] The patch of 2D samples is associated with auxiliary patch information PI representing auxiliary patch information used to interpret the geometric shape / attributes of the 2D samples of this patch of 2D samples.

[0064] In V-PCC, for example, the auxiliary patch information PI includes: 1) information indicating one of the six oriented surfaces of the 3D bounding box that surrounds the 3D samples of the connected components; 2) information regarding the surface normal; 3) information for determining the 3D position of the connected components with respect to the depth, tangent shift, and both tangent shifts converted and represented for the patch; and 4) information such as the coordinates (u0, v0, u1, v1) in the projection plane that defines the 2D bounding box surrounding the patch.

[0065] In step 3200, the patch packing module PPM can typically map (arrange) at least one generated patch of 2D samples onto a 2D grid (also shown as a canvas or atlas) without overlapping at all in a way that minimizes the unused space, and can ensure that it is associated with a unique patch for each block of TxT (e.g., 16x16) of the 2D grid. A given minimum block size TxT of the 2D grid can specify the minimum distance between separate patches of 2D samples when arranged on this 2D grid. The resolution of the 2D grid can depend on the input point cloud frame size, and its width W, height H, and block size T can be sent to the decoder as metadata.

[0066] The auxiliary patch information PI may further include information regarding the association between the blocks of the 2D grid and the patches of the 2D samples.

[0067] Figure 3a illustrates an example of a canvas C that includes two patches of 2D samples P1 and P2 and their associated 2D bounding boxes B1 and B2. The two bounding boxes can overlap within the canvas C as illustrated in Figure 3a. The 2D grid (division of the canvas) is represented only within the bounding boxes, but the division of the canvas also occurs outside of those bounding boxes. The bounding boxes associated with the patches can be divided into TxT blocks, typically with T = 16.

[0068] A TxT block containing 2D samples belonging to a patch of 2D samples can be regarded as an occupied block. Each occupied block of the canvas is represented in the occupancy map OM by a specific pixel value (e.g., 1), and each unoccupied block of the canvas is represented by a specific other value, e.g., 0. Subsequently, the pixel values of the occupancy map OM can indicate whether the TxT block of the canvas is occupied, i.e., whether it contains at least one 2D sample belonging to a patch of 2D samples.

[0069] In FIG. 3a, the occupied blocks are represented by white blocks, and the light gray blocks represent unoccupied blocks. The image generation process (steps 3300 and 3400) utilizes the mapping of at least one generated patch of 2D samples onto the 2D grid calculated during step 3200 to store the geometric shape and attributes of the 3D samples as an image.

[0070] In step 3300, the geometric shape image generator GIG can generate at least one geometric shape image GI from at least one patch of 2D samples, the occupancy map OM, and the auxiliary patch information PI.

[0071] The geometric shape image GI can represent the geometric shape of at least one patch of 2D samples and can be, for example, a monochromatic image of WxH pixels represented in the YUV420 - 8 - bit format.

[0072] The geometric shape image generator GIG can utilize the occupancy map information to detect (locate) the occupied blocks of the 2D grid in which at least one patch of the 2D samples is defined, and thus the non - empty pixels in the geometric shape image GI.

[0073] To better handle the case where multiple 3D samples are projected (mapped) to the same coordinates on the projection plane (along the same projection direction line), multiple layers can be generated. Thus, different depth values D1, ..., Dn can be obtained and associated with the 2D samples of the same patch of 2D samples. Then, multiple geometric shape images GI1, ..., GIN can be generated respectively for a specific depth value of the patch of 2D samples.

[0074] In V-PCC, the 2D samples of a patch can be projected onto two layers. The first layer, also called the near layer, can store, for example, the depth value D0 associated with the 2D sample having the lowest depth. The second layer, referred to as the far layer, can store, for example, the depth value D1 associated with the 2D sample having the highest depth.

[0075] According to an embodiment of step 3300, the geometric shape (the geometric shape of at least one orthogonally projected 3D point) of at least one patch of 2D samples is encoded according to the regular geometric shape coding mode RGCM. The regular geometric shape coding mode RGCM outputs at least one regular geometric shape patch RG2DP of the 2D samples from the geometric shape of at least one patch of the 2D samples.

[0076] According to an embodiment, the regular geometric shape coding mode RGCM can encode (derive) the depth value associated with the 2D samples of the patch of 2D samples as the luma component g(u, v) given by (u, v)=δ(u, v)-δ0 for the layer (the first or the second or both). Note that using this relationship, the 3D sample position δ(0, s0, r0) can be reconstructed from the reconstructed geometric shape image g(u, v) using the accompanying auxiliary patch information PI.

[0077] According to the embodiment of step 3300, the geometric shape of at least one patch of the 2D sample (the geometric shape of at least one orthogonally projected 3D point) is encoded according to the encoding mode FGCM of the first geometric shape. The encoding mode FGCM of the first geometric shape outputs at least one first geometric shape patch FG2DP of the 2D sample from the geometric shape of at least one patch of the 2D sample.

[0078] According to the embodiment, the encoding mode FGCM of the first geometric shape directly encodes the geometric shape of the 2D sample of at least one patch of the 2D sample as the pixel value of the geometric shape image.

[0079] For example, when the geometric shape is represented by three coordinates (u, v, Z), three consecutive pixels of the image are used, one for encoding the u coordinate, another for encoding the v coordinate, and another for encoding the Z coordinate.

[0080] According to the embodiment of step 3300, the geometric shape of at least one intermediate 3D sample is encoded according to the encoding mode SGCM of the second geometric shape. The encoding mode SGCM of the second geometric shape outputs at least one second geometric shape patch SG2DP of the 2D sample from the geometric shape of at least one intermediate 3D sample.

[0081] The intermediate 3D sample may exist between the first orthogonally projected 3D sample and the second orthogonally projected 3D sample along the same projection line. The intermediate 3D sample and the first and second orthogonally projected 3D samples have the same coordinates and different depth values on the projection plane.

[0082] In a modification, the intermediate 3D sample can be defined from a single orthogonally projected 3D sample and the length of the EOM codeword. The depth value of the "virtual" second orthogonally projected 3D sample is then equal to the depth value of the first orthogonally projected 3D sample and the length value of the EOM codeword. The first and "virtual" orthogonally projected 3D samples have the same coordinates on the projection plane and different depth values.

[0083] In some cases, the length of the EOM codeword is embedded in the syntax elements of the bitstream.

[0084] Hereinafter, the intermediate 3D sample is regarded as existing between the first orthogonally projected 3D sample and the second orthogonally projected 3D sample even when the second orthogonally projected 3D sample is "virtual".

[0085] Furthermore, the intermediate 3D sample has a depth value that is greater than the depth value of the first orthogonally projected 3D sample and lower than the depth value of the second orthogonally projected 3D sample.

[0086] A plurality of intermediate 3D samples can exist between the first orthogonally projected 3D sample and the second orthogonally projected 3D sample. Thus, the specified bits of the codeword can be set for each of the intermediate 3D samples to indicate whether the intermediate 3D sample exists (or does not exist) at a specific distance (a specific spatial position along the projection line) from one of the two orthogonally projected 3D samples.

[0087] FIG. 3b illustrates an example of two intermediate 3D samples P i1 and P i2 located between two 3D samples P0 and P1 along the projection line PL. The 3D samples P0 and P1 have depth values equal to D0 and D1, respectively. The depth values D i1 and D i2 of the two intermediate 3D samples P i1 and D i2 are each greater than D0 and lower than D1.

[0088] Next, all the specified bits along the projection line can be concatenated to form a codeword, hereinafter referred to as an Enhanced-Occupancy map (EOM) codeword. As illustrated in FIG. 3b, assuming an EOM codeword of 8-bit length, 2 bits become equal to 1, and two 3D samples P i1 and P i2 indicate the positions.

[0089] According to an embodiment of the second geometric shape coding mode SGCM, all EOM codewords are packed together to form at least one second geometric shape patch SG2DP of 2D samples.

[0090] At least one second geometric shape patch SG2DP of the 2D samples belongs to an image, and the coordinates of the pixels in the image indicate two of the three coordinates of the intermediate 3D samples (when these pixels refer to the EOM codeword), and the values of those pixels indicate the third coordinate of these intermediate 3D samples.

[0091] According to an embodiment, at least one second geometric shape patch SG2DP of the 2D samples belongs to an occupancy map OM.

[0092] In step 3400, the attribute image generator TIG can generate at least one attribute image TI from at least one patch of 2D samples, an occupancy map OM, auxiliary patch information PI, and at least one decoded geometric shape image DGI, i.e., the geometric shape of 3D samples derived from the output of the video decoder VDEC (step 4200 in FIG. 4).

[0093] The attribute image TI can represent the attributes of the 3D samples, for example, an image of WxH pixels represented in the YUV420-8-bit format.

[0094] The attribute image generator TG can effectively utilize the occupancy map information to detect (locate) the occupied blocks of the 2D grid in which at least one patch of the 2D sample is defined, and thus the non-empty pixels in the attribute image TI.

[0095] The attribute image generator TIG can be adapted to generate an attribute image TI and associate the attribute image TI with each geometric shape image DGI.

[0096] Next, a plurality of attribute images TI1,..., TIn can be generated respectively (for each geometric shape image) for a specific depth value of the patch of the 2D sample.

[0097] According to an embodiment of step 3400, the attributes of at least one patch of the 2D sample (attributes of the orthogonally projected 3D sample) are encoded according to the regular attribute coding mode RACM. The regular attribute coding mode RACM outputs at least one regular attribute patch RA2DP of the 2D sample from the attributes of at least one patch of the 2D sample.

[0098] According to an embodiment, the regular attribute coding mode RACM can code (store) the attribute T0 associated with the 2D sample of the patch of the 2D sample as the pixel value of the first attribute image TI0 for the first layer, and also code (store) the attribute value T1 associated with the 2D sample of the patch of the 2D sample as the pixel value of the second attribute image TI1 for the second layer.

[0099] Alternatively, the attribute image generation module TIG can code (store) the attribute value T1 associated with the 2D sample of the patch of the 2D sample as the pixel value of the first attribute image TI0 for the second layer, and also code (store) the attribute value T0 associated with the 2D sample of the patch of the 2D sample as the pixel value of the second attribute image TI1 for the first layer.

[0100] For example, the color of the 3D sample can be obtained as described in Sections 2.2.3, 2.2.4, 2.2.5, 2.2.8, or 2.5 of the V-PCC.

[0101] According to an embodiment of step 3400, the attribute of at least one patch of the 2D sample (the attribute of the orthogonally projected 3D sample) is encoded according to the coding mode FACM of the first attribute. The coding mode FACM of the first attribute outputs at least one first attribute patch FA2DP of the 2D sample from the attributes of at least one patch of the 2D sample.

[0102] According to an embodiment, the coding mode FACM of the first attribute directly encodes the attribute of the 2D sample of at least one patch of the 2D sample as the pixel value of the image.

[0103] According to an embodiment, at least one patch of the 2D sample belongs to the attribute image.

[0104] For example, the attribute of the orthogonally projected 3D sample can be obtained as described in Sections 2.2.3, 2.2.4, 2.2.5, 2.2.8, or 2.5 of the V-PCC.

[0105] According to an embodiment of step 3400, the attribute of the intermediate 3D sample is encoded according to the coding mode SACM of the second attribute. The coding mode SACM of the second attribute outputs at least one second attribute patch SA2DP of the 2D sample from the attributes of the at least one intermediate 3D sample.

[0106] Since the attribute values of the intermediate 3D samples correspond to the occupied blocks that have already been used to store the attribute values of other 2D samples as illustrated in FIG. 3b for the positions of their pixels, they cannot be directly stored as the pixel values of the attribute image.

[0107] According to an embodiment of the coding mode SACM of the second attribute, the attribute values of the intermediate 3D samples are fully packed to form at least one second attribute patch of the 2D samples.

[0108] According to an embodiment, at least one second attribute patch of the 2D sample belongs to an attribute image.

[0109] In V-PCC, the position of at least one second attribute patch of the 2D sample is defined as per the procedure (section 9.4.5 of V-PCC). Briefly, in this process, the location of the non-occupied blocks in the attribute image is determined, and the attribute values associated with the intermediate 3D samples are stored as the pixel values of the non-occupied blocks of the attribute image. This avoids the overlap between the occupied blocks and the second attribute patches of the 2D samples.

[0110] In step 3500, the video encoder VENC may encode the generated images / layers TI and GI.

[0111] In step 3600, the encoder OMENC may encode the occupancy map as an image, for example, as detailed in section 2.2.2 of V-PCC. Irreversible or reversible coding may be used.

[0112] According to an embodiment, the video encoder ENC and / or OMENC may be an HEVC-based encoder.

[0113] In step 3700, the encoder PIENC may encode the auxiliary patch information PI, as well as additional possible metadata such as the block size T, width W, and height H of the geometric shape / attribute image.

[0114] According to an embodiment, the auxiliary patch information may be encoded differentially (e.g., as defined in section 2.4.1 of V-PCC).

[0115] In step 3800, the multiplexing device can be applied to the generated outputs of steps 3500, 3600, and 3700, such that these outputs can be multiplexed together to generate a composite stream representing the base layer BL. It should be noted that the metadata information represents only a small percentage of the entire bitstream.

[0116] Encoder 3000 can also be used to encode the dynamic point cloud, and then each frame of this point cloud is encoded repeatedly. Then, at least one geometric shape image (step 3300), at least one attribute image (step 3400), occupancy map (step 3600), and auxiliary patch information (step 3700) are generated for each frame. Then, the geometric shape images generated for all frames of the point cloud can be fully combined to form a video stream, an attribute image for forming another video stream, and an occupancy map for forming another video stream. The auxiliary patch information can be added to generate a video stream, or all of the auxiliary patch information can be fully packed to form another video stream. Then, all of these video streams can be multiplexed (step 3800) to form a single bitstream BL.

[0117] FIG. 4 illustrates a schematic block diagram of an example of an image-based point cloud decoder 4000 according to at least one of the present embodiments.

[0118] Decoder 4000 can be used to decode a point cloud frame from a bitstream including a plurality of image streams (at least one geometric shape image stream, at least one attribute image stream, an occupancy map stream, and an auxiliary patch information image stream). However, decoder 4000 can also be used to decode a dynamic point cloud including a plurality of frames. In that case, each frame of the dynamic point cloud is decoded by extracting information from a video stream (geometric shape video stream, attribute video stream, occupancy video stream, auxiliary patch information video stream) embedded in the bitstream.

[0119] In step 4100, the demultiplexing device DMUX can be applied to demultiplex the encoded information of the bitstream representing the base layer BL.

[0120] In step 4200, the video decoder VDEC can decode the encoded information to derive at least one decoded geometric shape image DGI and at least one decoded attribute image DTI for the decoding of the 3D samples of the point cloud frame.

[0121] In step 4300, the decoder OMDEC can decode the encoded information to derive a decoded occupancy map DOM for the decoding of the 3D samples.

[0122] According to an embodiment, the video decoder VDEC and / or OMDEC can be an HEVC-based decoder.

[0123] In step 4400, the decoder PIDEC can decode the encoded information to derive auxiliary patch information DPI for the decoding of the 3D samples.

[0124] In some cases, metadata can also be derived from the bitstream BL.

[0125] In step 4500, the geometric shape generation module GGM may derive the geometric shape RG of the 3D samples of the point cloud frame IRPCF from at least one decoded geometric shape image DGI, the decoded occupancy map DOM, the decoded auxiliary patch information DPI, and possibly additional metadata.

[0126] The geometric shape generation module GGM may effectively utilize the decoded occupancy map information DOM to locate non-empty pixels within at least one decoded geometric shape image DGI.

[0127] The non-empty pixels belong to either an occupancy block or an EOM reference block according to the pixel values of the decoded occupancy information DOM and the values of D1 to D0 described above.

[0128] According to an embodiment of step 4500, when the non-empty pixel belongs to an occupancy block, the geometric shape of the 3D sample is decoded according to the decoded mode RGDM of the regular geometric shape.

[0129] According to an embodiment, the decoded mode RGDM of the regular geometric shape derives the 3D coordinates of the 3D sample from the coordinates of the non-empty pixel, the value of the non-empty pixel of at least one decoded geometric shape image DGI, the decoded auxiliary patch information, and possibly additional metadata.

[0130] The use of non-empty pixels is based on the relationship between the 2D pixels and the 3D samples. For example, using the projection within the V-PCC, the 3D coordinates of the reconstructed 3D samples can be converted into depth δ(u, v), tangent shift s(u, v), and both tangent shift r(u, v) and expressed as follows.

[0131] δ(u, v)=δ0+g(u, v) s(u, v)=s0 - u0 + u r(u, v)=r0 - v0 + v Where g(u, v) is the lumina component of the decoded geometric shape image DGI, (u, v) is the pixel associated with the reconstructed 3D sample, (δ0, s0, r0) is the 3D position of the connected component to which the reconstructed 3D sample belongs, and (u0, v0, u1, v1) are the coordinates in the projection plane that define the 2D bounding box encompassing the projection of the patch associated with the connected component.

[0132] According to an embodiment of step 4500, the geometric shape of the 3D sample is decoded according to the decoding mode FGDM of the first geometric shape.

[0133] According to an embodiment, the decoding mode FGDM of the first geometric shape directly decodes the geometric shape of the 3D sample from the pixel values of the decoded geometric shape image DGI.

[0134] For example, when the geometric shape is represented by three coordinates (u, v, Z), three consecutive pixels of the image are used, the u coordinate is equal to the value of one pixel of the geometric shape image, v is the value of another pixel of the geometric shape image, and Z is equal to the value of another pixel of the geometric shape image.

[0135] According to an embodiment of step 4500, the geometric shape of at least one intermediate 3D sample is decoded according to the decoding mode SGDM of the second geometric shape.

[0136] According to an embodiment, the coding mode SGCM of the second geometric shape can derive two of the 3D coordinates of the intermediate 3D sample from the coordinates of the non-empty pixels and the third bit value among the bit values of the EOM codeword.

[0137] For example, according to the example of FIG. 3b, the EOM codeword EOMC is used to determine the 3D coordinates of the intermediate 3D samples P i1 and P i2 . The third coordinate of the intermediate 3D sample P i1 is, for example, D0xD i1can be derived from D0 + 3, the reconstructed 3D sample P i2 The third coordinate of i2 is, for example, D0xD i2 can be derived from D0 + 5. The offset value (3 or 5) is the number of intervals between D0 and D1 along the projection line.

[0138] In step 4600, the attribute generation module TGM can derive the attributes of the 3D sample of the reconstructed point cloud frame IRPCF from the geometric shape RG of the 3D sample and at least one decoded attribute image DTI.

[0139] According to an embodiment of step 4600, the attributes of the 3D sample decoded by the decoding mode RGDM of the regular geometric shape are decoded according to the decoding mode RADM of the regular attributes. The first attribute decoding mode RADM can decode the attributes of the 3D sample from the pixel values of the attribute image.

[0140] According to an embodiment of step 4600, the attributes of the 3D sample decoded by the decoding mode FGDM of the first geometric shape are decoded according to the decoding mode FADM of the first attributes. The first attribute decoding mode FADM can decode the attributes of the 3D sample from the pixel values of the attribute image.

[0141] According to an embodiment of step 4600, the attributes of the intermediate 3D sample are decoded according to the second attribute decoding mode SADM.

[0142] According to an embodiment, the second attribute decoding mode SADM can derive the attributes of the intermediate 3D sample from the second attribute patch SA2DP of the 2D sample.

[0143] According to an embodiment, at least one second attribute patch of the 2D sample belongs to the attribute image.

[0144] In V-PCC, the position of at least one second attribute patch of the 2D sample is defined as per the procedure (section 9.4.10 of V-PCC). Briefly speaking, in this process, the location of non-occupied blocks in the attribute image is determined, and the attribute values associated with the intermediate 3D samples are derived from the pixel values of the non-occupied blocks of the attribute image.

[0145] Figure 5 schematically illustrates an example of the syntax of a bitstream representing the base layer BL according to at least one of the present embodiments.

[0146] The bitstream includes a bitstream header SH and at least one frame stream group GOFS.

[0147] The frame stream group GOFS includes a header HS, at least one syntax element OMS representing an occupancy map OM, at least one syntax element GVS representing at least one geometric shape image (or video), at least one syntax element TVS representing at least one attribute image (or video), and at least one syntax element PIS representing auxiliary patch information, as well as other additional metadata.

[0148] In a variant, the frame stream group GOFS includes at least one frame stream.

[0149] Figure 6 shows a schematic block diagram illustrating an example of a system in which various aspects and embodiments are implemented.

[0150] System 6000 can be embodied as one or more devices including the various components described below and is configured to implement one or more of the aspects described in this document. Examples of devices that can form all or part of System 6000 include personal computers, laptops, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, connected vehicles and their associated processing systems, head-mounted display devices (HMDs, see-through glasses), projectors (beamers), "cave automatic virtual environment" (CAVE) systems (systems including multiple displays), servers, video encoders, video decoders, post-processor processing output from video decoders, pre-processors providing input to video encoders, web servers, set-top boxes, and any other device for processing point clouds, videos, or images, or other communication devices. The elements of System 6000 can be embodied singly or in combination as a single integrated circuit, multiple ICs, and / or individual components. For example, in at least one embodiment, the processing and encoder / decoder elements of System 6000 can be distributed across multiple ICs and / or individual components. In various embodiments, System 6000 can be communicatively coupled to other similar systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various embodiments, System 6000 can be configured to implement one or more of the aspects described in this document.

[0151] System 6000 may include at least one processor 6010 configured to execute instructions loaded internally, for example, to implement various aspects described in this document. The processor 6010 may include an embedded memory, an input / output interface, and various other circuits known in the art. System 6000 may include at least one memory 6020 (e.g., a volatile memory device and / or a non-volatile memory device). System 6000 may include a storage device 6040, which may include non-volatile memory and / or volatile memory, and these memories may include, but are not limited to, electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash, magnetic disk devices, and / or optical disk devices. The storage device 6040 may include, by way of non-limiting example, an internal storage device, an attachable storage device, and / or a network-accessible storage device.

[0152] System 6000 may include, for example, an encoder / decoder module 6030 configured to process data to provide encoded data or decoded data, and the encoder / decoder module 6030 may include its own processor and memory. The encoder / decoder module 6030 may represent a module included within a device and capable of performing encoding and / or decoding functions. As is known, a device may include one or both of an encoding and a decoding module. Additionally, the encoder / decoder module 6030 may be implemented as a separate element of the system 6000 or may be incorporated within the processor 6010 as a combination of hardware and software, as is known to those skilled in the art.

[0153] The program code loaded into the processor 6010 or the encoder / decoder 6030 to implement the various aspects described in this document may be stored in the storage device 6040 and then loaded onto the memory 6020 for execution by the processor 6010. According to various embodiments, one or more of the processor 6010, the memory 6020, the storage device 6040, and the encoder / decoder module 6030 may store one or more of the various items during the implementation of the processes described in this document. Such stored items may include, but are not limited to, point cloud frames, encoded / decoded geometric shape / attribute videos / images or portions thereof, bitstreams, matrices, variables, and intermediate or final results from the processing of mathematical expressions, formulas, operations, and operation logic.

[0154] In some embodiments, the memory internal to the processor 6010 and / or the encoder / decoder module 6030 may be used to store instructions and provide a working memory for processes that may be performed during encoding or decoding.

[0155] However, in other embodiments, memory external to the processing device (e.g., the processing device can be either the processor 6010 or the encoder / decoder module 6030) can be used for one or more of these functions. The external memory can be the memory 6020 and / or the storage device 6040, e.g., dynamic volatile memory and / or non-volatile flash memory. In some embodiments, an external non-volatile flash memory can be used to store the television's operating system. In at least one embodiment, a high-speed external dynamic volatile memory such as RAM can be used as a working memory for video encoding and decoding operations such as for MPEG-2 Part 2 (also known as ITU-T Recommendation H.262 and ISO / IEC 13818-2 and also known as MPEG-2 video), HEVC (High Efficiency Video Coding), or VVC (Versatile Video Coding).

[0156] Inputs to the elements of the system 6000 can be provided through various input devices, as shown in block 6130. Such input devices include, but are not limited to, (i) an RF portion that can receive RF signals transmitted over the air, e.g., by a broadcast station, (ii) composite input terminals, (iii) USB input terminals, and / or (iv) HDMI input terminals.

[0157] In various embodiments, the input device of block 6130 may have respective input processing elements known in the art. For example, the RF portion may be associated with elements necessary for (i) selecting a desired frequency (also referred to as selecting a signal or band-limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) in certain embodiments, band-limiting to a narrower frequency band again to select a signal frequency band, which may be referred to as a channel for example, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF portion of various embodiments may include one or more elements for performing these functions, such as a frequency selector, signal selector, band limiter, channel selector, filter, down-converter, demodulator, error corrector, and demultiplexing device. The RF portion may include a tuner for performing these various functions, including for example down-converting a received signal to a lower frequency (such as an intermediate frequency or near baseband frequency) or to baseband.

[0158] In one embodiment of a set-top box, the RF portion and its associated input processing elements may receive an RF signal transmitted via a wired (e.g., cable) medium. The RF portion may then perform frequency selection by filtering to a desired frequency band, down-converting, and filtering again.

[0159] Various embodiments may rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions.

[0160] Adding elements may include inserting elements between existing elements, such as inserting an amplifier and an analog-to-digital converter for example. In various embodiments, the RF portion may include an antenna.

[0161] Additionally, the USB and / or HDMI terminals may each include an interface processor for connecting the system 6000 to other electronic devices across the USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or, if necessary, within the processor 6010. Similarly, aspects of USB or HDMI interface processing may be implemented within a separate interface IC or, if necessary, within the processor 6010. The demodulated, error-corrected, and demultiplexed streams may be provided to various processing elements, including, for example, the processor 6010, and, if necessary, an encoder / decoder 6030 that operates in combination with memory and storage elements for processing the data streams for presentation to an output device.

[0162] The various elements of the system 6000 may be provided within an integrated housing. Within the integrated housing, the various elements may be interconnected using a suitable connection arrangement 6140 and may transmit data between each other. Examples of such connection arrangements include internal buses known in the art, including I2C buses, wiring, and printed circuit boards.

[0163] The system 6000 may include a communication interface 6050 that enables communication with other devices via a communication channel 6060. The communication interface 6050 may include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 6060. The communication interface 6050 may include, but is not limited to, a modem or network card, and the communication channel 6060 may be implemented, for example, within wired and / or wireless media.

[0164] In various embodiments, data can be streamed to system 6000 using a Wi-Fi network such as IEEE 802.11. The Wi-Fi signals of these embodiments can be received via communication channel 6060 and communication interface 6050 adapted for Wi-Fi communication. The communication channel 6060 of these embodiments can typically be connected to an access point or router that provides access to an external network, which includes the Internet for enabling streaming applications and other over-the-top communications.

[0165] Other embodiments can provide the streamed data to system 6000 using a set-top box that delivers data via the HDMI connection of input block 6130.

[0166] Yet other embodiments can provide the streamed data to system 6000 using the RF connection of input block 6130.

[0167] The streamed data can be used as a way to signal information used by system 5000. The signaling information can include the information INF described above.

[0168] It should be understood that signaling can be achieved in various ways. For example, in various embodiments, one or more syntax elements, flags, etc. can be used to signal information to the corresponding decoder.

[0169] System 6000 can provide the output signal to various output devices including display 6100, speaker 6110, and other peripheral devices 6120. Other peripheral devices 6120 can include, in various examples of the embodiments, one or more of a stand-alone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functionality based on the output of system 3000.

[0170] In various embodiments, the control signal may be transmitted between the system 6000 and the display 6100, the speaker 6110, or other peripheral devices 6120 using a signaling scheme such as an AV Link (Audio / Video Link), CEC (Consumer Electronics Control), or other communication protocol that enables inter-device control regardless of the presence or absence of user intervention.

[0171] The output devices may be communicatively coupled to the system 6000 via dedicated connections through their respective interfaces 6070, 6080, and 6090.

[0172] Alternatively, the output devices may be connected to the system 6000 using the communication channel 6060 via the communication interface 6050. The display 6100 and the speaker 6110 may be integrated within a single unit together with other components of the system 6000, such as within an electronic device such as a television.

[0173] In various embodiments, the display interface 6070 may include a display driver, such as a timing controller (TCon) chip.

[0174] The display 6100 and the speaker 6110 may alternatively be separated from one or more of the other components, such as when the RF portion of the input 6130 is part of a separate set-top box. In various embodiments where the display 6100 and the speaker 6110 may be external components, the output signal may be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.

[0175] In V-PCC, the regular geometric shape patches of the 2D samples and the first geometric shape patches of the 2D samples are stored in the geometric shape image GI, the second geometric shape patches of the 2D samples are stored in the occupancy map OM, and the regular first and second attribute patches of the 2D samples are stored in the attribute image TI at the locations indicated by the syntax elements.

[0176] The regular first and second attribute patches of the 2D samples, which are stored together in the same video stream, may have the advantage of requiring only one video encoder / decoder for encoding / decoding the attribute information of the 3D samples.

[0177] By design, the encoding modes of the regular first and second geometric shapes / attributes respond to different needs. Therefore, their attribute formats, scan orders, sizes, and shapes are different. Therefore, methods for reducing information (or encoding / coding / compression) may require different methods, especially when reducing spatial "intra" redundancy. Moreover, while it may be aimed to encode the geometric shape / attributes (EOM codewords) of the intermediate 3D samples with a reversible video codec, the geometric shape / attributes of the sparse 3D samples (which typically represent the focus of a typical point cloud scene in some form less often) are typically encoded using the encoding mode of the first geometric shape / attributes and can be the subject of non-reversible encoding or vice versa according to their use.

[0178] As a result, encoding the geometric shape / attributes of the 3D samples requires both of these regular first and second encoding modes, and these modes require a specific (not general) profiled video codec that includes processes adapted to each of these two encoding modes, and additional signaling to indicate to the decoder that it is necessary to use either of these encoding modes for decoding the geometric shape and attributes of the 3D samples.

[0179] Thus, enabling multiple coding modes for encoding 3D samples of a point cloud frame can affect performance (reduction of encoder / decoder complexity and bandwidth).

[0180] Generally speaking, at least one of the embodiments provides a method for encoding attributes of orthogonally projected 3D samples and intermediate 3D samples, and the information INF indicates whether at least one first attribute patch of 2D samples obtained by encoding the attributes of the at least one orthogonally projected 3D sample according to a first attribute coding mode and at least one second attribute patch of 2D samples of an image obtained by encoding the attributes of the at least one intermediate 3D sample are stored in separate images.

[0181] Separate images can mean images of the same video stream, but can also mean different times or images of different video streams (not belonging to the same video stream).

[0182] The information INF enables more flexibility regarding the use of a video codec and also enables better compression performance since the video codec can be adjusted to encode the attributes of the orthogonal 3D samples or intermediate 3D samples. The adjustment of the video codec can then take into account the characteristics of the attributes of these 3D samples and / or satisfy the constraints and requirements of the application. For example, the first and / or second attribute patches of the 2D samples can be discarded according to the capabilities of the application or decoder / renderer (e.g., some 3D samples may not be useful for rendering a point cloud on a small display with a low-end SoC for decoding in some cases).

[0183] Furthermore, the video streams storing those attributes can be processed in parallel.

[0184] FIG. 7 shows an example of a flowchart of a method for encoding an orthogonal 3D sample of a point cloud frame according to at least one embodiment.

[0185] In step 710, the information INF can be set to a first specific value (e.g., 1) to indicate that at least one first attribute patch FA2DP (step 3400) of the 2D sample and at least one second attribute patch SA2DP (step 3400) of the 2D sample are stored in separate images. Also, the information INF can be set to a second specific value (e.g., 0) to indicate that at least one first attribute patch FA2DP of the 2D sample and at least one second attribute patch SA2DP of the 2D sample are stored in the same image.

[0186] When the information INF is equal to the first specific value, in step 720, at least one first attribute patch FA2DP of the 2D sample is stored in a first image FAI, and at least one second attribute patch SA2DP of the 2D sample is stored in a second image SAI.

[0187] When the information INF is equal to the second specific value, in step 730, at least one first attribute patch FA2DP of the 2D sample and at least one second attribute patch SA2DP of the 2D sample are stored in the same image AI.

[0188] According to an embodiment, as illustrated in FIG. 7a, the information INF includes a first binary flag F1 and a second binary flag F2.

[0189] In step 740, the first flag F1 can be set to a first specific value (e.g., 1) to indicate that at least one first attribute patch FA2DP of the 2D samples (step 3400) is stored in the first image FA1, and can be set to a second specific value (e.g., 0) to indicate that at least one first attribute patch FA2DP of the 2D samples (step 3400) is stored in the image AI.

[0190] In step 750, the second flag F2 can be set to a first specific value (e.g., 1) to indicate that at least one second attribute patch SA2DP of the 2D samples (step 3400) is stored in the second image FA2, and can be set to a second specific value (e.g., 0) to indicate that at least one second attribute patch FA2DP of the 2D samples (step 3400) is stored in the image AI.

[0191] When the first flag FA and the second flag F2 are equal to 1, the first attribute patch FA2DP and the second attribute patch SA2DP of the 2D samples are stored in separate images.

[0192] When the first flag F1 is equal to 0 and the second flag F2 is equal to 1, the first attribute patch FA2DP and the regular attribute 2D patch RA2DP of the 2D samples are stored in the same image, and the second attribute patch SA2DP of the 2D samples is stored in a separate image.

[0193] When the first flag F1 is equal to 1 and the second flag F2 is equal to 0, the second attribute patch SA2DP and the regular attribute 2D patch RA2DP of the 2D samples are stored in the same image, and the first attribute patch FA2DP of the 2D samples is stored in a separate image.

[0194] When the first flag FA and the second flag F2 are equal to 0, the first attribute patch FA2DP and the second attribute patch SA2DP of the 2D samples are stored in the same image.

[0195] According to the modification example, the information INF is valid in a group at the picture level, frame / atlas level, or patch level.

[0196] According to the modification example, syntax elements are added to the bitstream to identify a video codec used to compress a video stream that carries a second attribute patch of 2D samples.

[0197] This codec can be identified via a component codec mapping SEI message or through means other than the V-PCC specification.

[0198] Such a syntax element indicated by ai_eom_attribute_codec_id[atlas_id] can be conditionally added to the attribute information syntax structure based on a specific value of another syntax element such as the syntax element vpcc_eom_patch_separate_video_present_flag in FIG. 8.

[0199] FIG. 8 illustrates an example of the syntax element vpcc_eom_patch_separate_video_present_flag that embeds the information INF according to at least one embodiment.

[0200] The syntax element vpcc_eom_patch_separate_video_present_flag can be coded in a parameter set such as a Sequence Parameter Set (SPS), an Atlas Sequence Parameter Set (ASPS), or a Picture Parameter Set (PPS).

[0201] The elements in FIG. 8 have the following meanings.

[0202] vpcc_eom_patch_separate_video_present_flag[j] is equal to 1, indicating that the second attribute patch of the 2D sample with index j can be stored in a separate video stream.

[0203] vpcc_eom_patch_separate_video_present_flag[j] is equal to 0, indicating that the second attribute patch of the 2D sample with index j is not stored in a separate video stream.

[0204] When vpcc_eom_patch_separate_video_present_flag[j] does not exist, it is inferred to be equal to 0.

[0205] epdu_patch_in_eom_video_flag[p] specifies whether the attribute data associated with the second attribute patch of the 2D sample with index p within the current atlas style group is encoded in a separate video as compared to the attribute data of the intra- and inter-coded patches. When epdu_patch_in_eom_video_flag[p] is equal to 0, the attribute data associated with the second attribute patch of the 2D sample with index j within the current atlas style group is encoded in the same video as the attribute data of the intra- and inter-coded patches. When epdu_patch_in_eom_video_flag[p] is equal to 1, the attribute data associated with the second attribute patch of the 2D sample with index j within the current atlas style group is encoded in a separate video from the attribute data of the intra- and inter-coded patches. When epdu_patch_in_eom_video_flag[p] does not exist, its value is assumed to be inferred as equal to 0.

[0206] In a modification example, the syntax element vpcc_eom_patch_separate_video_present_flag[j] may be common to the first and second attribute patches of the 2D samples.

[0207] Next, when vpcc_separate_video_present_flag[j] is equal to 1, it indicates that the second attribute patch of the 2D sample, the first attribute patch of the 2D sample, and the first geometric shape patch of the 2D sample for the atlas having index j can be stored in separate images (separate video streams).

[0208] When vpcc_separate_video_present_flag[j] is equal to 0, it indicates that the second attribute patch of the 2D sample, the first attribute patch of the 2D sample, and the first geometric shape patch of the 2D sample for the atlas having index j are not stored in a separate video stream.

[0209] When vpcc_separate_video_present_flag[j] does not exist, it is inferred to be equal to 0.

[0210] In FIGS. 1 to 8, various methods are described herein, and each of the methods includes one or more steps or acts for achieving the described method. The order and / or use of the specific steps and / or acts can be changed or combined, provided that a specific order of steps or acts is not required for the proper operation of the method.

[0211] Several examples are described with respect to block diagrams and operation flowcharts. Each block represents a circuit element, module, or portion of code that includes one or more executable instructions for implementing the specified logical function. It should also be noted that in other implementations, the functions described in the blocks may occur from the order shown. For example, two blocks shown consecutively may actually be executed substantially simultaneously, or these blocks may be executed in the reverse order depending on the functionality involved.

[0212] The implementations and aspects described herein may be implemented, for example, as a method or process, apparatus, computer program, data stream, bit stream, or signal. Even if considered only in the context of a single form of implementation (e.g., considered only as a method), the implementation of the features considered may be implemented in other forms (e.g., apparatus or computer program).

[0213] The method may be implemented, for example, by a processor that generally refers to a processing device, and examples of the processor include a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device.

[0214] Additionally, the method may be implemented by instructions executed by a processor, and such instructions (and / or data values generated by the implementation) may be stored in a computer-readable storage medium. The computer-readable storage medium may be embodied in one or more computer-readable media and may take the form of a computer-readable program product having computer-readable program code embodied thereon that is executable by a computer. As used herein, a computer-readable storage medium may be regarded as a non-transitory storage medium that is provided with the inherent ability to store information therein and the inherent ability to provide retrieval of information therefrom. The computer-readable storage medium may be, for example, an electronic system, a magnetic system, an optical system, an electromagnetic system, an infrared system, or a semiconductor system, apparatus, or device, or any suitable combination of the foregoing, but is not limited thereto. The following provides more specific examples of computer-readable storage media to which the present embodiment may be applied, but is merely illustrative and not an exhaustive list as would be readily understood by one of ordinary skill in the art: portable computer diskette; hard disk; read-only memory (ROM); erasable programmable read-only memory (EPROM or flash memory); portable compact disc read-only memory (CD-ROM); optical storage device; magnetic storage device; or any suitable combination of the foregoing.

[0215] The instructions may form an application program embodied in a tangible manner on a processor-readable medium.

[0216] The instructions can be, for example, hardware, firmware, software, or a combination thereof. The instructions can be found, for example, in an operating system, a separate application, or a combination of the two. Thus, a processor can be characterized as both, for example, a device configured to perform a process and a device that includes a processor-readable medium (such as a storage device) having instructions for performing the process. Further, the processor-readable medium can store data values generated by an implementation, in addition to or instead of the instructions.

[0217] The apparatus can be implemented, for example, with appropriate hardware, software, and firmware. Examples of such apparatus include personal computers, laptops, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, head-mounted display devices (HMDs, see-through glasses), projectors (beamers), "immersive virtual reality experience devices (caves)" (systems including multiple displays), servers, video encoders, video decoders, post-processor processing output from a video decoder, preprocessors providing input to a video encoder, web servers, set-top boxes, and any other device for processing point clouds, videos, or images, or other communication devices. It should be made clear that the device can be mobile and can even be installed in a mobile vehicle.

[0218] Computer software can be implemented by processor 6010, or by hardware, or by a combination of hardware and software. As a non-limiting example, embodiments can also be implemented by one or more integrated circuits. Memory 6020 can be of any type suitable for the technical environment, and can be implemented using any suitable data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples. Processor 6010 can be of any type suitable for the technical environment, and can include, as non-limiting examples, one or more of a microprocessor, a general-purpose computer, a dedicated computer, and a processor based on a multi-core architecture.

[0219] As will be apparent to those skilled in the art, embodiments can generate various signals formatted to carry information that can be stored or transmitted, for example. The information can include, for example, instructions for executing a method or data generated by one of the described embodiments. For example, the signal can be formatted to carry a bitstream of the described embodiment. Such a signal can be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier using the encoded data stream. The information carried by the signal can be analog information or digital information, for example. The signal can be transmitted via various different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.

[0220] The terms used in this specification are for the purpose of describing particular embodiments only and are not intended to be limiting. As used herein, the singular forms "a", "an", and "the" may be intended to include the plural as well, unless the context clearly dictates otherwise. As used herein, the terms "includes" / "comprises" and / or "including" / "comprising" may specify the presence of the stated, for example, features, integers, steps, operations, elements, and / or components, but it will be further understood that they do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Further, when an element is referred to as "responding to" or "connected to" another element, it may directly respond to or be connected to the other element, or intervening elements may be present. In contrast, when an element is referred to as "directly responding to" or "directly connected to" another element, no intervening elements are present.

[0221] For example, in the case of "A / B", "A and / or B", and "at least one of A and B", it should be understood that the use of any of the symbols / terms " / " ", "and / or", and "at least one of" is intended to include the selection of only the first listed option (A), or only the second listed option (B), or the selection of both options (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C", such expressions are intended to include the selection of only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or the selection of only the first and second listed options (A and B), or the selection of only the first and third listed options (A and C), or the selection of only the second and third listed options (B and C), or the selection of all three options (A and B and C). As will be apparent to those skilled in the art of this technology and related technologies, this can be extended for as many listed items as there are.

[0222] In this application, various numerical values may be used. The specific values may be for illustrative purposes and the described embodiments are not limited to these specific values.

[0223] In this specification, terms such as first, second, etc. may be used to describe various elements, but it should be understood that these elements are not limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the teachings of this application, the first element may be referred to as the second element, and similarly, the second element may be referred to as the first element. There is no order between the first element and the second element.

[0224] References to "one embodiment" or "an embodiment" or "one implementation" or "an implementation", and other variations thereof, are frequently used to convey that certain features, structures, characteristics, etc. (described in relation to the embodiment / implementation) are included in at least one embodiment / implementation. Thus, the appearance of the phrases "in one embodiment" or "in an embodiment" or "in one implementation" or "in an implementation" in various places in this specification, and any other variations, do not necessarily all refer to the same embodiment.

[0225] Similarly, references in this specification to "in accordance with an embodiment / example / implementation" or "in an embodiment / example / implementation", and other variations thereof, are frequently used to convey that certain features, structures, or characteristics (described in relation to the embodiment / example / implementation) may be included in at least one embodiment / example / implementation. Thus, the appearance of the expressions "in accordance with an embodiment / example / implementation" or "in an embodiment / example / implementation" in various places in this specification does not necessarily all refer to the same embodiment / example / implementation, and also, separate or alternative embodiments / examples / implementations do not necessarily exclude each other.

[0226] The reference numbers appearing in the claims are for illustrative purposes only and shall not have a limiting effect on the claims. Although not explicitly stated, the present embodiments / examples and variations can be used in any combination or sub-combination.

[0227] When a figure is presented as a flowchart, it should be understood that the figure also provides a block diagram of the corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that the figure also provides a flowchart of the corresponding method / process.

[0228] Some figures include arrows on the communication path to indicate the main direction of communication, but it should be understood that communication can occur in the direction opposite to the illustrated arrows.

[0229] Various implementations involve decoding. As used in this application, "decoding" can include all or part of a process that is performed, for example, on a received point cloud frame (which may in some cases include a received bitstream encoding one or more point cloud frames), to generate a final output suitable for display or for further processing in a reconstructed point cloud region. In various embodiments, such a process typically includes one or more of the processes performed by an image-based decoder. In various embodiments, such a process may also or alternatively include, for example, the processes performed by the decoders of the various implementations described in this application.

[0230] As a further example, in one embodiment, "decoding" may refer only to entropy decoding, in another embodiment, "decoding" may refer only to differential decoding, and in another embodiment, "decoding" may refer to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally to a broader decoding process is apparent based on the context of the particular description and is considered to be fully understood by those skilled in the art.

[0231] Various implementations involve encoding. In a manner similar to the above considerations regarding "decoding", when used in this application, "encoding" can include, for example, all or part of the process performed on an input point cloud frame to generate an encoded bitstream. In various embodiments, such a process typically includes one or more of the processes performed by an image-based decoder. In various embodiments, such a process also includes, or alternatively, the processes performed by the encoders of the various implementations described in this application.

[0232] As a further example, in one embodiment, "encoding" can refer to only entropy encoding, in another embodiment, "encoding" can refer to only differential encoding, and in another embodiment, "encoding" can refer to a combination of differential encoding and entropy encoding. Whether the phrase "encoding process" is intended to specifically refer to a subset of operations or generally to a broader encoding process is clear based on the context of a particular description and is considered to be well understood by those skilled in the art.

[0233] Note that the syntax elements are used herein as illustrative terms. Thus, they do not preclude the use of other syntax element names.

[0234] Various embodiments refer to rate-distortion optimization. In particular, during the encoding process, the balance or trade-off between rate and distortion is typically considered to often impose constraints on the computational complexity. Rate-distortion optimization can typically be formulated to minimize a rate-distortion function that is a weighted sum of rate and distortion. There are different approaches to solving the rate-distortion optimization problem. For example, the approach can be based on an extensive test of all encoding options, including all considered modes or coded parameter values, and fully evaluate their coding cost and the associated distortion of the reconstructed signal after encoding and decoding. Also, a faster approach can be used to reduce the encoding complexity, especially by using an approximation of the distortion based on the predicted or prediction residual signal rather than the reconstructed signal. Also, a mixture of these two approaches can be used, such as by using the approximate distortion of only some of the possible encoding options and the full distortion of other encoding options. Other approaches evaluate only a subset of the possible encoding options. More generally, many approaches use any of various techniques for performing the optimization, but the optimization is not necessarily a full evaluation of both the coding cost and the associated distortion.

[0235] Additionally, this application may refer to "determining" various information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from memory.

[0236] Furthermore, this application may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0237] Additionally, the present application may refer to "receiving" various information. Receiving is intended to be a broad term, similar to "accessing". Receiving information may include, for example, one or more of accessing the information or retrieving the information (e.g., from a memory). Further, "receiving" typically involves, in some form, actions such as storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information, among others.

[0238] Also, as used herein, the word "signal" refers, among other things, to indicating something to the corresponding decoder. For example, in one embodiment, the encoder signals a particular piece of information INF that may be carried by the syntax element vpcc_eom_patch_separate_video_present_flag. In this way, in embodiments, the same parameters may be used on both the encoder side and the decoder side. Thus, for example, the encoder may transmit (explicitly signal) a particular parameter to the decoder so that the decoder may use the same particular parameter. Conversely, if the decoder already has a particular parameter as well as other parameters, signaling may be used without transmitting (implicitly signaling) in order to simply enable the decoder to know and select the particular parameter. Bit savings are achieved in various embodiments by avoiding the transmission of any actual functionality. It should be understood that signaling may be achieved in various ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to signal information to the corresponding decoder. The foregoing relates to the verb form of the word "signal", but the word "signal" may also be used as a noun herein.

[0239] Multiple implementations have been described. Nevertheless, it will be understood that various changes can be made. For example, elements of different implementations can be combined, supplemented, modified, or deleted to produce other implementations. Additionally, those skilled in the art will understand that other structures and processes can be substituted for those disclosed, and that the resulting implementations will perform at least substantially the same function in at least substantially the same way to achieve at least substantially the same result as the disclosed implementations. Accordingly, these and other implementations are contemplated by this application.

Claims

1. decoding data representing a point cloud from a first video stream, the data including at least one first attribute patch of a 2D sample, the at least one first attribute patch of the 2D sample including attributes of an orthogonally projected 3D sample of the point cloud, the 3D coordinates of the sample being derived from coordinates of a pixel in a geometry image and a value of the pixel in the geometry image; decoding a syntax element indicating whether at least one second attribute patch of a 2D sample, at least one geometry patch of a 2D sample, and at least one third attribute patch of a 2D sample can be stored in a second video stream separate from the first video stream; at least one geometric patch of the 2D sample includes 3D coordinates of at least one 3D sample of a point cloud, the 3D coordinates being derived from pixel values ​​of the at least one geometric patch; at least one second attribute patch of the 2D sample includes attribute values ​​of the at least one 3D sample whose 3D coordinates are derived from pixel values ​​of the at least one geometric shape patch; at least one third attribute patch of the 2D samples includes attributes of at least one intermediate 3D sample of the point cloud, the 3D coordinates of which are derived from coordinates of a pixel in a geometry image and from one bit of a codeword indicating a position of the intermediate 3D sample along a projection line between two 3D samples of the point cloud projected orthogonally along the projection line; decoding at least one of at least one geometry patch of the 2D sample, at least one second attribute patch of the 2D sample, or at least one third attribute patch of the 2D sample from the first video stream or the second video stream based on the syntax element; A method comprising:

2. The method of claim 1 , wherein the syntax element is a flag common to the at least one second attribute patch and the at least one third attribute patch.

3. 2. The method of claim 1, further comprising: decoding a flag based on the syntax element, the flag indicating whether attribute data associated with at least one third attribute patch of the 2D sample is encoded in the first video stream or in the second video stream.

4. The method of claim 1 , wherein the syntax elements are decoded from one of a sequence parameter set, an atlas sequence parameter set, or a picture parameter set.

5. 2. The method of claim 1, further comprising: decoding another syntax element that identifies a video codec used to compress the video stream that carries at least one third attribute patch of the 2D sample.

6. 2. The method of claim 1, wherein the first video stream and the second video stream are hierarchically organized in picture-level, frame-level, and patch-level groups, and the syntax element is valid at either the picture-level, frame-level, atlas-level, or patch-level group.

7. The method of claim 1 , wherein the codewords are packed to form at least one other geometric patch of 2D samples that is stored in an occupancy map.

8. 1. An apparatus comprising one or more processors, the one or more processors comprising: decoding data representing a point cloud from a first video stream, the data including at least one first attribute patch of a 2D sample, the at least one first attribute patch of the 2D sample including attributes of an orthogonally projected 3D sample of the point cloud, the 3D coordinates of the sample being derived from coordinates of a pixel in a geometry image and a value of the pixel in the geometry image; decoding a syntax element indicating whether at least one second attribute patch of a 2D sample, at least one geometry patch of a 2D sample, and at least one third attribute patch of a 2D sample can be stored in a second video stream separate from the first video stream; at least one geometric patch of the 2D sample includes 3D coordinates of at least one 3D sample of a point cloud, the 3D coordinates being derived from pixel values ​​of the at least one geometric patch; at least one second attribute patch of the 2D sample includes attribute values ​​of the at least one 3D sample whose 3D coordinates are derived from pixel values ​​of the at least one geometric shape patch; at least one third attribute patch of the 2D samples includes attributes of at least one intermediate 3D sample of the point cloud, the 3D coordinates of which are derived from coordinates of a pixel in a geometry image and from one bit of a codeword indicating a position of the intermediate 3D sample along a projection line between two 3D samples of the point cloud projected orthogonally along the projection line; decoding at least one of at least one geometry patch of the 2D sample, at least one second attribute patch of the 2D sample, or at least one third attribute patch of the 2D sample from the first video stream or the second video stream based on the syntax element; An apparatus configured to:

9. 1. An apparatus comprising one or more processors, the one or more processors comprising: encoding data representing a point cloud in a first video stream, the data including at least one first attribute patch of a 2D sample, the at least one first attribute patch of the 2D sample including attributes of an orthogonally projected 3D sample of the point cloud, the 3D coordinates of the sample being derived from coordinates of a pixel in a geometry image and a value of the pixel in the geometry image; encoding a syntax element indicating whether at least one second attribute patch of a 2D sample, at least one geometry patch of a 2D sample, and at least one third attribute patch of a 2D sample may be stored in a second video stream separate from the first video stream; at least one geometric patch of the 2D sample includes 3D coordinates of at least one 3D sample of a point cloud, the 3D coordinates being derived from pixel values ​​of the at least one geometric patch; at least one second attribute patch of the 2D sample includes attribute values ​​of the at least one 3D sample whose 3D coordinates are derived from pixel values ​​of the at least one geometric shape patch; at least one third attribute patch of the 2D samples includes attributes of at least one intermediate 3D sample of the point cloud, the 3D coordinates of which are derived from coordinates of a pixel in a geometry image and from one bit of a codeword indicating a position of the intermediate 3D sample along a projection line between two 3D samples of the point cloud projected orthogonally along the projection line; encoding at least one of at least one geometry patch of the 2D sample, at least one second attribute patch of the 2D sample, or at least one third attribute patch of the 2D sample in the first video stream or in the second video stream based on the syntax element; An apparatus configured to:

10. encoding data representing a point cloud in a first video stream, the data including at least one first attribute patch of a 2D sample, the at least one first attribute patch of the 2D sample including attributes of an orthogonally projected 3D sample of the point cloud, the 3D coordinates of the sample being derived from coordinates of a pixel in a geometry image and a value of the pixel in the geometry image; encoding a syntax element indicating whether at least one second attribute patch of a 2D sample, at least one geometry patch of a 2D sample, and at least one second attribute patch of a 2D sample may be stored in a second video stream separate from the first video stream; at least one geometric patch of the 2D sample includes 3D coordinates of at least one 3D sample of a point cloud, the 3D coordinates being derived from pixel values ​​of the at least one geometric patch; at least one second attribute patch of the 2D sample includes attribute values ​​of the at least one 3D sample whose 3D coordinates are derived from pixel values ​​of the at least one geometric shape patch; at least one third attribute patch of the 2D samples includes attributes of at least one intermediate 3D sample of the point cloud, the 3D coordinates of which are derived from coordinates of a pixel in a geometry image and from one bit of a codeword indicating a position of the intermediate 3D sample along a projection line between two 3D samples of the point cloud projected orthogonally along the projection line; encoding at least one of at least one geometry patch of the 2D sample, at least one second attribute patch of the 2D sample, or at least one third attribute patch of the 2D sample in the first video stream or in the second video stream based on the syntax element; A method comprising:

11. A non-transitory computer readable medium comprising instructions for causing one or more processors to perform the method of any one of claims 1-7 or 10.

12. 10. The apparatus of claim 8 or 9, wherein the syntax element is a flag common to the at least one second attribute patch and the at least one third attribute patch.

13. 9. The apparatus of claim 8, wherein the one or more processors are further configured to decode a flag based on the syntax element, the flag indicating whether attribute data associated with at least one third attribute patch of the 2D sample is encoded in the first video stream or in the second video stream.

14. The apparatus of claim 8 , wherein the syntax elements are decoded from one of a sequence parameter set, an atlas sequence parameter set, or a picture parameter set.

15. 10. The apparatus of claim 8, wherein the one or more processors are further configured to decode another syntax element that identifies a video codec used to compress the video stream that carries at least one third attribute patch of the 2D sample.

16. 10. The apparatus of claim 9, wherein the one or more processors are further configured to decode a flag based on the syntax element, the flag indicating whether attribute data associated with at least one third attribute patch of the 2D sample is encoded in the first video stream or in the second video stream.

17. 10. The apparatus of claim 9, wherein the one or more processors are further configured to decode another syntax element that identifies a video codec used to compress the video stream that carries at least one third attribute patch of the 2D sample.

Citation Information

Patent Citations

  • Point cloud compression

    US20190087979A1

  • Point cloud compression using hybrid transforms

    US20190122393A1