Encoding and decoding a point cloud using a patch of intermediate samples

By using layered processing and orthogonal projection, the geometric shape and attribute information of point cloud frames are encoded into attribute patches of 2D samples, solving the problem of efficient compression of dynamic point clouds and achieving high-quality reconstruction at a limited bit rate, which is suitable for immersive virtual reality and autonomous driving.

CN114503579BActive Publication Date: 2026-03-24INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-15
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently compress dynamic point clouds to provide an acceptable quality of experience on end-user devices while maintaining reasonable bit rate consumption.

Method used

A two-layer point cloud coding structure is adopted to process the geometric shape and attribute information of point cloud frames in layers. The base layer provides a lossy representation and the enhancement layer provides a higher quality lossless representation through combined coding of the base layer and the enhancement layer. The 3D sample attributes are encoded into attribute patches of 2D samples through orthogonal projection.

Benefits of technology

It achieves efficient compression of dynamic point clouds at a limited bit rate, providing higher quality reconstruction results, and is suitable for applications such as immersive virtual reality and autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114503579B_ABST
    Figure CN114503579B_ABST
Patent Text Reader

Abstract

At least one embodiment of the invention relates to providing a method of encoding / decoding properties of an orthogonal projected 3D sample and an intermediate 3D sample, wherein the information indicates whether at least one first property patch of a 2D sample obtained by encoding a property of at least one orthogonal projected 3D sample according to a first property encoding mode and at least one second property patch of a 2D sample of an image obtained by encoding a property of at least one intermediate 3D sample are stored in a separate image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] At least one of these embodiments substantially relates to point cloud processing. Specifically, the encoding / decoding of properties of 3D samples from / in a separate video stream is disclosed. Background Technology

[0002] This section is intended to introduce various aspects of the art that may relate to at least one aspect of the embodiments described below and / or claimed. This discussion is intended to help provide the reader with background information to facilitate a better understanding of the various aspects of at least one embodiment.

[0003] Point clouds can be used for various purposes, such as cultural heritage / buildings, where objects like sculptures or buildings are scanned in 3D to share their spatial configuration without sending or accessing them. Additionally, it's a way to ensure the preservation of knowledge about objects in case they are destroyed; for example, a temple collapsed due to an earthquake. Such point clouds are typically static, shaded, and massive.

[0004] Another use case is in topography and cartography, where using 3D representation allows for maps that are not limited to flat surfaces and can include terrain. Google Maps is now a good example of a 3D map, but it uses a grid instead of point clouds. However, point clouds can be a suitable data format for 3D maps, and such point clouds are typically static, shaded, and massive.

[0005] The automotive industry and autonomous vehicles are also areas where point clouds can be used. Autonomous vehicles should be able to "detect" their environment to make sound driving decisions based on the actual conditions of their neighbors. Typical sensors, such as light detection and ranging (LIDAR), generate dynamic point clouds used by decision engines. These point clouds are not intended for human viewing, and they are typically small, not necessarily colored, and dynamic at high capture frequencies. These point clouds may possess other properties, such as reflectivity provided by LIDAR, because this property provides good information about the material of the sensed object and can potentially aid in decision-making.

[0006] Virtual reality and immersive worlds have recently become hot topics, and many foresee a future for 2D flat video. The basic idea, compared to standard TV where viewers can only see a virtual world in front of them, is to immerse them in the environment surrounding them. Immersion is categorized into several levels depending on the viewer's degree of freedom within the environment. Point clouds are a good candidate format for distributing virtual reality (VR) worlds.

[0007] In many applications, it is important to be able to distribute dynamic point clouds to end users (or store them on servers) while consuming only a reasonable amount of bitrate (or storage space for the application) and maintaining an acceptable (or preferably very good) quality of experience. Efficient compression of these dynamic point clouds is key to making distribution chains useful for many immersive worlds.

[0008] In view of the foregoing, at least one implementation scheme has been designed. Summary of the Invention

[0009] The following presents a simplified overview of at least one of the embodiments to provide a basic understanding of some aspects of this disclosure. This summary is not a broad overview of the embodiments. It is not intended to identify key or essential elements of the embodiments. The following summary presents only some aspects of at least one of the embodiments in a simplified form as a prelude to a more detailed description provided elsewhere in this document.

[0010] According to a general aspect of at least one embodiment, a method is provided for encoding properties of orthographically projected 3D samples, wherein the properties of the orthographically projected 3D samples are encoded as at least one first attribute patch of a 2D sample of an image, and the properties of an intermediate 3D sample located between two orthographically projected 3D samples along the same projection line are encoded as at least one second attribute patch of a 2D sample in the image, wherein the method includes encoding information indicating whether the at least one first attribute patch and the at least one second attribute patch of the 2D sample are stored in separate images.

[0011] According to the implementation plan, the video stream is hierarchically structured at a set of image level, frame level, and tile level, and the information is valid at the set of image level, frame level, atlas level, or tile level.

[0012] According to the implementation scheme, the information is a first mark indicating whether at least one first attribute patch of the 2D sample is stored in a first image or whether the at least one first attribute patch of the 2D sample is stored in a second image together with other attribute patches of the 2D sample, and a second mark indicating whether at least one second attribute patch of the 2D sample is stored in a third image or whether the at least one second attribute patch of the 2D sample is stored in the second image together with other attribute patches of the 2D sample.

[0013] According to the implementation scheme, the method further includes encoding another piece of information indicating how to compress individual images.

[0014] According to a general aspect of at least one embodiment, a method for decoding the properties of a 3D sample is provided, wherein the properties of the 3D sample are decoded based on at least one second attribute patch of a 2D sample in an image, and the properties of an intermediate 3D sample located between two 3D samples along the same projection line are decoded as at least one second attribute patch of the 2D sample in the image, wherein the method includes decoding information indicating whether the at least one first attribute patch of the 2D sample and the at least one second attribute patch of the 2D sample are stored in separate images.

[0015] According to the implementation plan, the video stream is hierarchically structured at a set of image level, frame level, and tile level, and the information is valid at the set of image level, frame level, atlas level, or tile level.

[0016] According to the implementation scheme, the information is a first mark indicating whether at least one first attribute patch of the 2D sample is stored in a first image or whether the at least one first attribute patch of the 2D sample is stored in a second image together with other attribute patches of the 2D sample, and a second mark indicating whether at least one second attribute patch of the 2D sample is stored in a third image or whether the at least one second attribute patch of the 2D sample is stored in the second image together with other attribute patches of the 2D sample.

[0017] According to the implementation scheme, the method further includes encoding another piece of information indicating how to compress individual images.

[0018] One or more embodiments also provide devices, bit streams, computer program products, and non-transitory computer-readable media.

[0019] The specific nature of at least one of the embodiments and other objects, advantages, features and uses of the at least one of the embodiments will become apparent from the following description of examples in conjunction with the accompanying drawings. Attached Figure Description

[0020] Examples of several implementation schemes are shown in the accompanying drawings. (See attached drawings for details.)

[0021] - Figure 1 A schematic block diagram illustrating an example of a two-layer point cloud coding structure according to at least one of the embodiments herein;

[0022] - Figure 2 A schematic block diagram illustrating an example of a two-layer point cloud decoding structure according to at least one of the embodiments is shown.

[0023] - Figure 3 A schematic block diagram illustrating an example of an image-based point cloud encoder according to at least one of the present embodiments is shown.

[0024] - Figure 3a An example of a canvas including two patches and their 2D bounding boxes is shown;

[0025] - Figure 3b An example of an intermediate 3D sample located between two 3D samples along the projection line is shown;

[0026] - Figure 4 A schematic block diagram illustrating an example of an image-based point cloud decoder according to at least one of the present embodiments;

[0027] - Figure 5 An example of a syntax for representing a bitstream of a base layer BL according to at least one of the embodiments is illustrated schematically;

[0028] - Figure 6 A schematic block diagram showing an example of a system in which various aspects and implementation schemes are carried out;

[0029] - Figure 7 and Figure 7a An example flowchart illustrating a method for encoding orthogonal 3D samples of point cloud frames according to at least one embodiment;

[0030] - Figure 8 An example of a syntax element for embedding information INF according to at least one implementation scheme is shown; Detailed Implementation

[0031] At least one of the embodiments is described more fully below with reference to the accompanying drawings, which illustrate examples of at least one of the embodiments. However, embodiments may be embodied in many alternative forms and should not be construed as limited to the examples set forth herein. Therefore, it should be understood that the embodiments are not intended to be limited to the specific forms disclosed. Rather, this disclosure is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of this application.

[0032] When the accompanying drawings are presented as flowcharts, it should be understood that block diagrams of the corresponding devices are also provided. Similarly, when the accompanying drawings are presented as block diagrams, it should be understood that flowcharts of the corresponding methods / processes are also provided.

[0033] Similar or identical elements in the diagram are referenced using the same reference numerals.

[0034] Some diagrams represent syntax tables widely used in V-PCC to define bitstream structures consistent with V-PCC. In those syntax tables, the term '...' indicates the unchanged portion of the syntax related to the original definition given in V-PCC, and said unchanged portion is removed in the diagram for readability. Bold terms in the diagram indicate the value of the term obtained by parsing the bitstream. The right column of the syntax table indicates the number of bits used to encode the data of the syntax element. For example, u(4) indicates 4 bits used to encode the data, u(8) indicates 8 bits, and ae(v) indicates context-adaptive arithmetic entropy encoding of the syntax element.

[0035] The aspects described and envisioned below can be implemented in many different forms. Figures 1 to 8 Some implementation schemes are provided below, but other implementation schemes are also considered, and Figures 1 to 8 The discussion does not limit the breadth of implementation methods.

[0036] At least one of these aspects typically involves point cloud encoding and decoding, and at least one other aspect typically involves transmitting generated or encoded bit streams.

[0037] More precisely, the various methods and other aspects described herein can be used to modify modules, such as the image-based encoder 3000 and decoder 4000, as... Figures 1 to 8 As shown.

[0038] Furthermore, these aspects of the invention are not limited to MPEG standards, such as MPEG-I Part 5 related to point cloud compression, and can be applied to, for example, other standards and recommendations, whether pre-existing or future-developed, and any extensions of such standards and recommendations (including MPEG-I Part 5). Unless otherwise indicated or technically excluded, the aspects described in this application may be used alone or in combination.

[0039] In the following text, image data refers to reference data, such as one or more arrays of 2D samples in a particular image / video format. A particular image / video format may specify information relating to the pixel values ​​of an image (or video). A particular image / video format may also specify information that can be used by a display and / or any other device to, for example, visualize and / or decode the image (or video). An image typically contains a first component in the shape of a first 2D array of samples, typically representing the luminance (luma) of the image. An image may also contain second and third components in the shape of other 2D arrays of samples, typically representing the chrominance (chroma) of the image. Some embodiments represent the same information using a 2D array of a set of color samples (e.g., conventional three-color RGB representation).

[0040] Pixel values ​​are represented by C values ​​in one or more implementations, where C is the number of components. Each value of the vector is typically represented by a number of bits that can define the dynamic range of the pixel value.

[0041] An image patch is a set of pixels that belong to an image. The pixel value of an image patch (or image patch data) refers to the value of the pixels belonging to that image patch. Image patches can have any shape, although rectangles are common.

[0042] A point cloud can be represented by a dataset of 3D samples within a 3D volumetric space, which may have unique coordinates and may also have one or more attributes.

[0043] A 3D sample may include information defining the geometry of 3D points in a point cloud, which may be represented by X, Y, and Z coordinates in 3D space. It may also include information defining one or more associated attributes, such as colors represented in RGB or YUV color spaces, such as transparency, reflectivity, a two-component normal vector, or any feature representing a characteristic of the sample. For example, a 3D sample may include information defining six components (X, Y, Z, R, G, B) or equivalently (X, Y, Z, y, U, V), where (X, Y, Z) defines the coordinates of a 3D point in 3D space, and (R, G, B) or (y, U, V) defines the color of that 3D point. Attributes of the same type may exist multiple times. For example, multiple color attributes can provide color information from different viewpoints.

[0044] The 2D sample may include information defining the geometry of a 3D sample orthographically projected onto it, which may be represented by three coordinates (u, v, Z), where (u, v) are the coordinates in the 2D space of the orthographically projected 3D sample, and Z is the Euclidean distance between the 3D sample and the projection plane onto which the 3D sample is orthographically projected. Z is typically represented as a depth value. The 3D sample may also include information defining one or more associated properties, such as color represented in RGB or YUV color space, such as transparency, reflectivity, a two-component normal vector, or any feature representing the characteristics of this orthographically projected 3D sample.

[0045] Therefore, a 2D sample may include information defining the geometry and properties of a 3D sample by orthogonal projections of (u, v, Z, R, G, B) or equivalently (u, v, Z, y, U, V).

[0046] Point clouds can be static or dynamic, depending on whether the cloud changes relative to time. Instances of static or dynamic point clouds are typically represented as point cloud frames. It should be noted that in the case of dynamic point clouds, the number of points is usually not constant, but rather typically changes over time. More generally, a point cloud can be considered dynamic if anything (e.g., the number of points, the location of one or more points, or any property of any point) changes over time.

[0047] Figure 1 A schematic block diagram showing an example of a two-layer point cloud coding structure 1000 according to at least one of the embodiments herein.

[0048] A two-layer point cloud coding structure 1000 can provide a bitstream B representing an input point cloud frame IPCF. Possibly, the input point cloud frame IPCF represents a frame of a dynamic point cloud. This frame of the dynamic point cloud can then be encoded by the two-layer point cloud coding structure 1000.

[0049] Next, a video stream representing the dynamic point cloud can be obtained by combining the bitstreams representing each frame of the dynamic point cloud together.

[0050] Essentially, the two-layer point cloud coding structure 1000 provides the ability to structure a bitstream B into a base layer (BL) and an enhancement layer (EL). The base layer (BL) provides a lossy representation of the input point cloud frame (IPCF), while the enhancement layer (EL) can provide a higher-quality (potentially lossless) representation by encoding isolated points not represented by the base layer (BL).

[0051] The base layer BL can be provided by an image-based encoder 3000, such as... Figure 3 As shown. The image-based encoder 3000 can provide a geometry / attribute image representing the geometry / attributes of a 3D sample from an input point cloud frame (IPCF). This allows for the discarding of isolated 3D samples. (As shown...) Figure 4 As shown, the base layer BL can be decoded by an image-based decoder 4000, which can provide intermediate reconstructed point cloud frames IRPCF.

[0052] Next, return to Figure 1 The two-layer point cloud encoding 1000 uses a comparator COMP to compare 3D samples from the input point cloud frame IPCF with 3D samples from the intermediate reconstructed point cloud frame IRPCF to detect / locate missing / isolated 3D samples. Next, the encoder ENC encodes the missing 3D samples and provides an enhancement layer EL. Finally, the base layer BL and the enhancement layer EL are multiplexed together by the multiplexer MUX to generate the bitstream B.

[0053] According to the implementation scheme, the encoder ENC may include a detector that can detect a 3D reference sample R of the intermediate reconstructed point cloud frame IRPCF and associate it with the missing 3D sample M.

[0054] For example, based on a given metric, a 3D reference sample R associated with a missing 3D sample M can be its nearest neighbor.

[0055] According to the implementation scheme, the encoder ENC can then encode the spatial location and attributes of the missing 3D sample M as a difference determined based on the spatial location and attributes of the 3D reference sample R.

[0056] In the variants, those differences can be encoded individually.

[0057] For example, for a missing 3D sample M, the spatial coordinates x(M), y(M), and z(M), the x-coordinate position difference Dx(M), the y-coordinate position difference Dy(M), the z-coordinate position difference Dz(M), the R-attribute component difference Dr(M), the G-attribute component difference Dg(M), and the B-attribute component difference Db(M) can be calculated as follows:

[0058] Dx(M) = x(M) - x(R),

[0059] Where x(M) is the x-coordinate of the 3D sample M, respectively, derived from... Figure 3 The R coordinate in the provided geometric image,

[0060] Dy(M) = y(M) - y(R),

[0061] Where y(M) is the y-coordinate of the 3D sample M, respectively, derived from... Figure 3 The R coordinate in the provided geometric image,

[0062] Dz(M)=z(M)-z(R),

[0063] Where z(M) is the z-coordinate of the 3D sample M, respectively, derived from... Figure 3 The R coordinate in the provided geometric image,

[0064] Dr(M) = R(M) - R(R),

[0065] Where R(M) and R(R) are the r-color components of the color attribute of the 3D sample M, respectively, and R is R,

[0066] Dg(M)=G(M)-G(R),

[0067] Where G(M) and G(R) are the g-color components of the color attribute of the 3D sample M, and R is the color component of the sample M.

[0068] Db(M)=B(M)-B(R),

[0069] Where B(M) and B(R) are the b color components of the color attribute of the 3D sample M, and R is the b color component.

[0070] Figure 2 A schematic block diagram illustrating an example of a two-layer point cloud decoding structure 2000 according to at least one of the embodiments herein.

[0071] The behavior of the two-layer point cloud decoding structure 2000 depends on its capabilities.

[0072] A two-layer point cloud decoding architecture 2000 with limited capabilities can provide a faithful (but lossy) version of the input point cloud frame IPCF, such as the IRPCF, by using a demultiplexer DMUX to access only the base layer BL from the bitstream B, and then decoding the base layer BL by the point cloud decoder 4000. Figure 4 As shown.

[0073] The fully-capable two-layer point cloud decoding architecture 2000 can access both the base layer BL and the enhancement layer EL from the bitstream B using a demultiplexer DMUX. Figure 4 As shown, the point cloud decoder 4000 can determine the intermediate reconstructed point cloud frame IRPCF based on the base layer BL. The decoder DEC can determine the complementary point cloud frame CPCF based on the enhancement layer EL. Then, the combiner COMB can combine the intermediate reconstructed point cloud frame IRPCF with the complementary point cloud frame CPCF, thus providing a higher quality (potentially lossless) representation (reconstruction) of the input point cloud frame IPCF (CRPCF).

[0074] Figure 3 A schematic block diagram illustrating an example of an image-based point cloud encoder 3000 according to at least one of the embodiments herein is shown.

[0075] The image-based point cloud encoder 3000 utilizes existing video codecs to compress the geometry and attribute information of 3D samples of input dynamic point clouds using different video streams.

[0076] In a particular implementation, two video streams can be generated and compressed using an existing video encoder: one to capture the geometric information of 3D samples of the input point cloud, and the other to capture the attribute information of these 3D samples. An example of an existing video codec is the HEVC Master Profile Encoder / Decoder (ITU-T H.265, ITU Telecommunications Standardization Sector, (February 2018), H Series: Audiovisual and Multimedia Systems, Infrastructure for Audiovisual Services – Coding of Mobile Video, Efficient Video Coding, Recommendation ITU-T H.265).

[0077] Additional metadata used to interpret the two video streams is typically generated and compressed separately. This additional metadata includes, for example, occupancy map (OM) and / or auxiliary tile information (PI).

[0078] The generated video stream and metadata can then be reused together to generate a combined stream.

[0079] It should be noted that metadata typically represents a small amount of overall information. The majority of the information is in the video stream.

[0080] Examples of such point cloud encoding / decoding processes are given by the test model category 2 algorithm (also referred to as V-PCC) implementing the MPEG draft standard, as defined in Information Technology - Coding Representation of Immersive Media - Part 5: Video-based Point Cloud Compression, CD Grade, SCD_d39, ISO / IEC 23090-5, in ISO / IEC JTC1 / SC29 / WG11.

[0081] In step 3100, by orthogonally projecting the 3D sample of the frame IPCF of the input point cloud frame onto the 2D sample on the projection plane using a strategy that provides optimal compression, the module PGM can generate at least one patch of the 2D sample.

[0082] A patch of 2D samples can be defined as a group of 2D samples that share common characteristics.

[0083] For example, in V-PCC, the normal of each 3D sample is first estimated, as described, for example, by Hope et al. (Hughes Hope, Tony DeRoss, Tom Duchamp, John McDonald, Werner Stuzil: Surface Reconstruction from Unorganized Points. ACM SIGGR4PH 1992, Proceedings, pp. 71-78). Next, an initial cluster of the 3D samples is obtained by associating each 3D sample with one of the six orientation planes of the 3D bounding box surrounding the 3D sample. More precisely, each 3D sample is clustered and associated with the orientation plane having the closest normal (i.e., maximizing the dot product of the point normal and the plane finding). Then, the 3D sample is orthogonally projected onto its associated plane (projection plane). A group of 3D samples that form connecting regions in its plane is called a connecting assembly. Thus, a connecting assembly is a group of at least one 3D sample with similar normals and the same associated orientation plane. Next, the initial clusters are refined by iteratively updating the clusters associated with each 3D sample based on the normal of each 3D sample and the clusters of its nearest neighbor samples. The final step involves generating a patch of a 2D sample from each connector, which is accomplished by projecting the 3D sample of each connector onto an orientation plane associated with the connector.

[0084] Next, the 2D samples of the 2D sample patch share the same normal and the same orientation plane, and they are closely positioned to each other.

[0085] The patch of the 2D sample is associated with auxiliary patch information (PI), which represents auxiliary patch information used to interpret the geometry / attributes of the 2D sample for this patch.

[0086] In V-PCC, for example, the auxiliary patch information PI includes: 1) information indicating one of the six orientation planes of the 3D bounding box surrounding the 3D sample of the connector assembly; 2) information relative to the plane normal; 3) information determining the 3D position of the connector assembly relative to the patch in terms of depth, tangential displacement, and bitangential displacement; and 4) information such as coordinates (u0, v0, u1, v1) in the projection plane that defines the 2D bounding box surrounding the patch.

[0087] In step 3200, the Patch Filling Module (PPM) can map at least one generated patch of the 2D sample onto a 2D mesh (also represented as a canvas or atlas) without overlapping in a manner that typically minimizes unused space, and can guarantee that each T×T (e.g., 16×16) block of the 2D mesh is associated with a unique patch. A given minimum block size T×T of the 2D mesh can specify the minimum distance between different patches of the 2D sample placed on this 2D mesh. The 2D mesh resolution can depend on the input point cloud frame size as well as its width W and height H, and the block size T can be emitted to the decoder as metadata.

[0088] The auxiliary patch information (PI) can further include information about the relationship between the blocks relative to the 2D mesh and the patches of the 2D sample.

[0089] Figure 3a An example of a canvas C is shown, comprising two patches P1 and P2 of a 2D sample and their associated 2D bounding boxes B1 and B2. Note that the two bounding boxes can overlap in canvas C, as shown below. Figure 3a As shown. 2D meshes (canvas divisions) are represented only within bounding boxes, but canvas divisions also occur outside those bounding boxes. The bounding box associated with a tile can be divided into TxT blocks, typically T=16.

[0090] A T×T block containing a 2D sample belonging to a 2D sample patch can be considered an occupied block. Each occupied block of the canvas is represented by a specific pixel value (e.g., 1) in the occupancy map OM, and each unoccupied block of the canvas is represented by another specific value (e.g., 0). The pixel values ​​of the occupancy map OM can then indicate whether a T×T block of the canvas is occupied, i.e., whether it contains at least one 2D sample belonging to a 2D sample patch.

[0091] exist Figure 3aIn the image, occupied blocks are represented by white blocks, and light gray blocks represent unoccupied blocks. The image generation process (steps 3300 and 3400) utilizes the mapping of at least one generated patch of the 2D sample calculated during step 3200 onto the 2D mesh to store the geometry and properties of the 3D sample as an image.

[0092] In step 3300, the geometry image generator GIG can generate at least one geometry image GI from at least one patch of the 2D sample, the occupancy map OM, and the auxiliary patch information PI.

[0093] A geometric image (GI) can represent the geometry of at least one patch of a 2D sample and can be, for example, a monochrome image of WxH pixels represented in YUV420-8-bit format.

[0094] The geometry image generator GIG can utilize occupancy map information to detect (locate) occupancy blocks of a 2D grid that defines at least one patch of a 2D sample, and thus detect (locate) non-empty pixels in the geometry image GI.

[0095] To better handle situations where multiple 3D samples are projected (mapped) onto the same coordinates on the projection plane (along the same projection direction line), multiple layers can be generated. Therefore, different depth values ​​D1, ..., Dn can be obtained and associated with the same patch of the 2D sample. Then, multiple geometric images GI1, ..., GIn can be generated, each for a specific depth value of the 2D sample patch.

[0096] In V-PCC, the 2D sample of the patch can be projected onto two layers. The first layer (also called the near layer) can store, for example, a depth value D0 associated with a 2D sample having a smaller depth. The second layer, called the far layer, can store, for example, a depth value D1 associated with a 2D sample having a larger depth.

[0097] According to the implementation scheme of step 3300, the geometry (geometry of at least one orthogonally projected 3D point) of at least one patch of the 2D sample is encoded according to the Regular Geometry Encoding Pattern (RGCM). The Regular Geometry Encoding Pattern (RGCM) outputs at least one regular geometry patch of the 2D sample RG2DP based on the geometry of the at least one patch of the 2D sample.

[0098] According to the implementation scheme, the conventional geometric coding mode (RGCM) can encode (derive) the depth values ​​associated with a 2D sample relative to a 2D sample patch of a layer (first layer, second layer, or both) as a luminance component g(u, v), which is given by the following formula: g(u, v) = δ(u, v) - δ0. It should be noted that this relationship can be used to reconstruct the 3D sample position (δ0, s0, r0) from the reconstructed geometric image g(u, v) with attached auxiliary patch information PI.

[0099] According to the implementation scheme of step 3300, the geometry (geometry of at least one orthogonally projected 3D point) of at least one patch of the 2D sample is encoded according to the first geometric coding mode FGCM. The first geometric coding mode FGCM outputs at least one first geometric shape patch FG2DP of the 2D sample according to the geometry of the at least one patch of the 2D sample.

[0100] According to the implementation scheme, the first geometric coding mode (FGCM) directly encodes the geometry of the 2D sample of at least one patch of the 2D sample into pixel values ​​of the geometric image.

[0101] For example, when the geometry is represented by three coordinates (u, v, Z), then three consecutive pixels of the image are used: one to encode u, another to encode v, and another to encode the Z coordinate.

[0102] According to the implementation scheme of step 3300, the geometry of at least one intermediate 3D sample is encoded according to the second geometric coding mode SGCM. The second geometric coding mode SGCM outputs at least one second geometric shape patch SG2DP of the 2D sample according to the geometry of the at least one intermediate 3D sample.

[0103] The intermediate 3D sample can be located between the 3D sample in the first orthogonal projection and the 3D sample in the second orthogonal projection, along the same projection line. This intermediate 3D sample, as well as the 3D sample in the first orthogonal projection and the 3D sample in the second orthogonal projection, have the same coordinates and different depth values ​​on the projection plane.

[0104] In a variant, an intermediate 3D sample can be defined based on the length of the 3D sample from a single orthographic projection and the EOM codeword. Then, the depth value of the 3D sample from the "virtual" second orthographic projection is equal to the depth value of the 3D sample from the first orthographic projection plus the stated length value of the EOM codeword. The 3D sample from the first orthographic projection and the 3D sample from the "virtual" orthographic projection have the same coordinates and different depth values ​​on the projection plane.

[0105] Possibly, the length of the EOM codeword is embedded in the syntax elements of the bitstream.

[0106] In the following text, the intermediate 3D sample will be considered as being located between the 3D sample of the first orthographic projection and the 3D sample of the second orthographic projection, even if the 3D sample of the second orthographic projection is "virtual".

[0107] Furthermore, the depth value of the intermediate 3D sample is greater than the depth value of the 3D sample in the first orthogonal projection, but lower than the depth value of the 3D sample in the second orthogonal projection.

[0108] Multiple intermediate 3D samples may exist between the 3D sample of the first orthogonal projection and the 3D sample of the second orthogonal projection. Therefore, a designated bit of a codeword can be set for each of the intermediate 3D samples to indicate whether the intermediate 3D sample is present (or not present) at a specific distance from one of the two orthogonal projection 3D samples (at a specific spatial position along the projection line).

[0109] Figure 3b This shows the intermediate 3D sample P located between two 3D samples P0 and P1 along the projection line PL. i1 and P i2 Examples. 3D samples P0 and P1 have depth values ​​equal to D0 and D1, respectively. These correspond to two intermediate 3D samples P... i1 and P i2 Depth value D i1 and D i2 It is greater than D0 and less than D1.

[0110] Next, all the designated bits along the projection line can be concatenated to form a codeword, hereinafter referred to as an Enhanced Occupation Map (EOM) codeword. For example... Figure 3b As shown, assuming an 8-bit EOM codeword, 2 bits are equal to 1 to indicate the position P of the two 3D samples. i1 and P i2 The location.

[0111] According to the implementation scheme of the second geometric coding mode SGCM, all EOM codewords are filled together to form at least one second geometric shape patch SG2DP of the 2D sample.

[0112] The at least one second geometry patch SG2DP of the 2D sample belongs to an image, and the coordinates of the pixels in the image indicate two of the three coordinates of the intermediate 3D sample (when those pixels reference EOM codewords), and the values ​​of those pixels indicate the third coordinate of these intermediate 3D samples.

[0113] According to the implementation scheme, the at least one second geometry patch SG2DP of the 2D sample belongs to the occupied map OM.

[0114] In step 3400, the attribute image generator TIG can generate at least one attribute image TI based on at least one patch of the 2D sample, the occupancy map OM, the auxiliary patch information PI, and the geometry of the 3D sample derived from at least one decoded geometry image DGI, and the output of the video decoder VDEC. Figure 4 Step 4200 in the middle.

[0115] Attribute images (TI) can represent the attributes of a 3D sample and can be, for example, WxH pixel images represented in YUV420-8-bit format.

[0116] The attribute image generator TG can utilize occupancy map information to detect (locate) occupancy blocks of a 2D grid that defines at least one patch of a 2D sample, and thus detect (locate) non-empty pixels in the attribute image TI.

[0117] The Attribute Image Generator (TIG) can be adapted to generate attribute images and associate them with each geometric image (DGI).

[0118] Then, multiple attribute images TI1, ..., TIn can be generated, each attribute image representing a specific depth value of the patch on the 2D sample (for each geometry image).

[0119] According to the implementation scheme of step 3400, the attributes of at least one patch of the 2D sample (attributes of the orthographically projected 3D sample) are encoded according to the Regular Attribute Encoding Mode (RACM). The Regular Attribute Encoding Mode (RACM) outputs at least one regular attribute patch RA2DP of the 2D sample based on the attributes of the at least one patch of the 2D sample.

[0120] According to the implementation scheme, the conventional attribute encoding mode RACM can encode (store) the attribute T0 associated with the 2D sample of the patch relative to the 2D sample of the first layer as the pixel value of the first attribute image TI0, and encode (store) the attribute value T1 associated with the 2D sample of the patch relative to the 2D sample of the second layer as the pixel value of the second attribute image TI1.

[0121] Alternatively, the attribute image generation module TIG can encode (store) the attribute value T1 associated with the 2D sample of the patch relative to the 2D sample of the second layer as the pixel value of the first attribute image TI0, and encode (store) the attribute value T0 associated with the 2D sample of the patch relative to the 2D sample of the first layer as the pixel value of the second attribute image TI1.

[0122] For example, the color of a 3D sample can be obtained as described in Sections 2.2.3, 2.2.4, 2.2.5, 2.2.8, or 2.5 of the V-PCC.

[0123] According to the implementation scheme of step 3400, the attributes of at least one patch of the 2D sample (attributes of the orthographically projected 3D sample) are encoded according to the first attribute encoding mode FACM. The first attribute encoding mode FACM outputs at least one first attribute patch FA2DP of the 2D sample according to the attributes of the at least one patch of the 2D sample.

[0124] According to the implementation scheme, the first attribute encoding mode FACM directly encodes the attributes of the at least one patch of the 2D sample into the pixel values ​​of the image.

[0125] According to the implementation scheme, the at least one patch of the 2D sample belongs to an attribute image.

[0126] For example, the properties of orthographically projected 3D samples can be obtained as described in Sections 2.2.3, 2.2.4, 2.2.5, 2.2.8, or 2.5 of V-PCC.

[0127] According to the implementation scheme of step 3400, the attributes of the intermediate 3D sample are encoded according to the second attribute encoding mode SACM. The second attribute encoding mode SACM outputs at least one second attribute patch SA2DP of the 2D sample based on the attributes of the at least one intermediate 3D sample.

[0128] The attribute values ​​of intermediate 3D samples cannot be directly stored as pixel values ​​of the attribute image because the positions of those pixels correspond to occupied blocks already used to store attribute values ​​of other 2D samples, such as... Figure 3b As shown.

[0129] According to the implementation scheme of the second attribute encoding mode (SACM), the attribute values ​​in the intermediate 3D sample are filled together to form at least one second attribute patch of the 2D sample.

[0130] According to the implementation scheme, the at least one second attribute patch of the 2D sample belongs to the attribute image.

[0131] In V-PCC, the position of the at least one second attribute patch of the 2D sample is defined procedurally (Section 9.4.5 of V-PCC). In short, this process determines the position of the unoccupied block in the attribute image and stores the attribute value associated with the intermediate 3D sample as the pixel value of the unoccupied block in the attribute image. This avoids overlap between the occupied block and the second attribute patch of the 2D sample.

[0132] In step 3500, the video encoder VENC can encode the generated image / layer TI and GI.

[0133] In step 3600, the encoder OMENC can encode the occupancy map into an image, as detailed in section 2.2.2 of V-PCC, for example. Lossy or lossless encoding can be used.

[0134] According to the implementation plan, the video encoder ENC and / or OMENC can be HEVC-based encoders.

[0135] In step 3700, the encoder PIENC can encode the auxiliary patch information PI and possible additional metadata, such as the block size T, the width W and height H of the geometry / attribute image.

[0136] According to the implementation plan, auxiliary patch information can be differentially encoded (as defined in Section 2.4.1 of V-PCC).

[0137] In step 3800, a multiplexer can be applied to the outputs generated in steps 3500, 3600, and 3700, and thus these outputs can be multiplexed together to generate a combined stream representing the base layer BL. It should be noted that the metadata information represents a small portion of the overall bitstream.

[0138] The encoder 3000 can also be used to encode the dynamic point cloud: then iteratively encodes each frame of the point cloud. For each frame, at least one geometry image (step 3300), at least one attribute image (step 3400), an occupancy map (step 3600), and auxiliary patch information are then generated (step 3700). The geometry images generated for all frames of the point cloud can then be combined to form a video stream, the attribute images can be combined to form another video stream, and the occupancy maps can be combined to form yet another video stream. Auxiliary patch information can be added to generate the video stream, or all auxiliary patch information can be padded together to form another video stream. All these video streams can then be multiplexed (step 3800) to form a single bitstream BL.

[0139] Figure 4 A schematic block diagram illustrating an example of an image-based point cloud decoder 4000 according to at least one of the embodiments herein.

[0140] The decoder 4000 can be used to decode point cloud frames based on a bitstream comprising multiple image streams (at least one geometric image stream, at least one attribute image stream, occupancy map stream, and auxiliary patch information image stream). However, it can also be used to decode dynamic point clouds comprising multiple frames. In this case, each frame in the dynamic point cloud is decoded by extracting information from the video streams (geometric video stream, attribute video stream, occupancy video stream, and auxiliary patch information video stream) embedded in the bitstream.

[0141] In step 4100, the demultiplexer DMUX can be applied to demultiplex the encoded information representing the bit stream of the base layer BL.

[0142] In step 4200, relative to the decoding of the 3D sample of the point cloud frame, the video decoder VDEC can decode the encoded information to derive at least one decoded geometry image DGI and at least one decoded attribute image DTI.

[0143] In step 4300, the decoder OMDEC can decode the encoded information to derive a decoded occupancy map DOM relative to the decoding of the 3D sample.

[0144] According to the implementation plan, the video decoder VDEC and / or OMDEC can be HEVC-based decoders.

[0145] In step 4400, relative to the decoding of the 3D sample, the decoder PIDEC can decode the encoded information to derive the auxiliary patch information DPI.

[0146] Metadata may also be derived from the bitstream BL.

[0147] In step 4500, the geometry generation module GGM can derive the geometry RG of a 3D sample of a point cloud frame IRPCF from at least one decoded geometry image DGI, a decoded occupancy map DOM, decoded auxiliary patch information DPI, and possible additional metadata.

[0148] The geometry generation module (GGM) can utilize the decoded occupied map information (DOM) to locate non-empty pixels in at least one decoded geometry image (DGI).

[0149] As described above, depending on the pixel values ​​of the decoded occupancy information DOM and the values ​​of D1-D0, the non-empty pixel belongs to the occupancy block or the EOM reference block.

[0150] According to the implementation of step 4500, when the non-empty pixel belongs to the occupied block, the geometry of the 3D sample is decoded according to the conventional geometric decoding mode RGDM.

[0151] According to the implementation scheme, the conventional geometric decoding mode RGDM derives the 3D coordinates of the 3D sample from the coordinates of the non-empty pixels, the value of the non-empty pixel in at least one of the decoded geometric images DGI, decoded auxiliary patch information, and possibly from additional metadata.

[0152] The use of non-empty pixels is based on their relationship to the 2D pixels of the 3D sample. For example, in the case of the projection in V-PCC, the 3D coordinates of the reconstructed 3D sample can be represented by depth δ(u, v), tangential shift s(u, v), and bidirectional shift r(u, v), as follows:

[0153] δ(u, v) = δ0 + g(u, v)

[0154] s(u, v) = s0 - u0 + u

[0155] r(u, v) = r0 - v0 + v

[0156] Where g(u, v) is the luminance component of the decoded geometric image DGI, (u, v) is the pixel associated with the reconstructed 3D sample, (δ0, s0, r0) is the 3D position of the connection component to which the reconstructed 3D sample belongs, and (u0, v0, ul, v1) are the coordinates in the projection plane that defines the 2D bounding box surrounding the projection of the patch associated with the connection component.

[0157] According to the implementation scheme of step 4500, the geometry of the 3D sample is decoded according to the first geometric decoding mode FGDM.

[0158] According to the implementation plan, the first geometric decoding mode, FGDM, directly decodes the geometry of the 3D sample based on the pixel values ​​of the decoded geometric image DGI.

[0159] For example, when the geometry is represented by three coordinates (u, v, Z), then three consecutive pixels of the image are used: the u coordinate is equal to the value of one pixel of the geometry, the v coordinate is equal to the value of another pixel of the geometry, and the Z coordinate is equal to the value of another pixel of the geometry.

[0160] According to the implementation of step 4500, the geometry of at least one intermediate 3D sample is decoded according to the second geometric decoding mode SGDM.

[0161] According to the implementation plan, the second geometric coding mode SGCM can derive two 3D coordinates of the intermediate 3D sample from the coordinates of the non-empty pixels, and derive a third 3D coordinate from the bit value of the EOM codeword.

[0162] For example, according to Figure 3b For example, the EOM codeword EOMC is used to determine the intermediate 3D sample P. i1 and P i2 3D coordinates. Middle 3D sample PP i1 The third coordinate can be, for example, from D0 through D i1 =D0+3 is derived, and the 3D sample P is reconstructed. i2The third coordinate can be, for example, from D0 through D i2 =D0+5 is derived. The offset value (3 or 5) is the number of intervals along the projection line between D0 and D1.

[0163] In step 4600, the attribute generation module TGM can derive the attributes of the 3D sample from the geometry RG of the 3D sample and at least one decoded attribute image DTI to reconstruct the point cloud frame IRPCF.

[0164] According to the implementation scheme of step 4600, the attributes of the 3D sample whose geometry is decoded by the conventional geometry decoding mode RGDM are decoded according to the conventional attribute decoding mode RADM. The first attribute decoding mode RADM can decode the attributes of the 3D sample based on the pixel values ​​of the attribute image.

[0165] According to the implementation scheme of step 4600, the attributes of the 3D sample whose geometry is decoded by the first geometry decoding mode (FGDM) are decoded according to the first attribute decoding mode (FADM). The first attribute decoding mode (FADM) can decode the attributes of the 3D sample based on the pixel values ​​of the attribute image.

[0166] According to the implementation scheme of step 4600, the attributes of the intermediate 3D sample are decoded according to the second attribute decoding mode SADM.

[0167] According to the implementation plan, the second attribute decoding mode (SADM) can derive the attributes of the intermediate 3D sample from the second attribute patch (SA2DP) of the 2D sample.

[0168] According to the implementation scheme, the at least one second attribute patch of the 2D sample belongs to the attribute image.

[0169] In V-PCC, the position of the at least one second attribute patch of the 2D sample is defined by procedure (Section 9.4.10 of V-PCC). In short, this process determines the position of the unoccupied block in the attribute image and derives the attribute values ​​associated with the intermediate 3D sample from the pixel values ​​of the unoccupied block in the attribute image.

[0170] Figure 5 An example syntax for representing a bitstream of a base layer BL according to at least one of the embodiments is illustrated schematically.

[0171] The bitstream includes a bitstream header (SH) and at least one frame stream group (GOFS).

[0172] The Frame Stream Group GOFS includes a header HS, at least one syntax element OMS representing the occupied map OM, at least one syntax element GVS representing at least one geometric image (or video), at least one syntax element TVS representing at least one attribute image (or video), and at least one syntax element PIS representing auxiliary patch information and other additional metadata.

[0173] In the variant, the Frame Stream Group GOFS includes at least one frame stream.

[0174] Figure 6 A schematic block diagram illustrating examples of systems in which various aspects and implementation schemes are carried out.

[0175] System 6000 may be embodied as one or more devices comprising the various components described below and configured to perform one or more of the aspects described in this document. Examples of devices that may form all or part of System 6000 include personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, connected vehicles and their associated processing systems, head-mounted displays (HMDs, see-through glasses), projectors (beam emitters), "cave displays" (systems containing multiple displays), servers, video encoders, video decoders, post-processors that process the output from the video decoder, pre-processors that provide input to the video encoder, video encoder, web server, set-top box, and any other means for processing point clouds, video, or images, or other communication means. Elements of System 6000 may be embodied individually or in combination in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of System 6000 may be distributed across multiple ICs and / or discrete components. In various embodiments, system 6000 may be communicatively coupled to other similar systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 6000 may be configured to implement one or more of the aspects described in this document.

[0176] System 6000 may include at least one processor 6010 configured to execute instructions loaded therein for implementation of various aspects as described in this document. Processor 6010 may include embedded memory, input / output interfaces, and various other circuitry as known in the art. System 6000 may include at least one memory 6020 (e.g., a volatile memory device and / or a non-volatile memory device). System 6000 may include a storage device 6040, which may include non-volatile memory and / or volatile memory, including, but not limited to, electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 6040 may include internal storage, attached storage, and / or network-accessible storage.

[0177] System 6000 may include an encoder / decoder module 6030 configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 6030 may include its own processor and memory. The encoder / decoder module 6030 may represent a module that can be included in a device to perform encoding and / or decoding functions. It is well known that a device may include one or both of an encoding module and a decoding module. Furthermore, the encoder / decoder module 6030 may be implemented as a separate element of system 6000, or may be incorporated within processor 6010 as a combination of hardware and software known to those skilled in the art.

[0178] Program code to be loaded onto processor 6010 or encoder / decoder 6030 to execute the various aspects described in this document may be stored in storage device 6040 and subsequently loaded onto memory 6020 for execution by processor 6010. According to various embodiments, one or more of processor 6010, memory 6020, storage device 6040, and encoder / decoder module 6030 may store one or more items from various projects during execution of the processes described in this document. Such stored items may include, but are not limited to, point cloud frames, encoded / decoded geometry / attribute video / images or portions of encoded / decoded geometry / attribute video / images, bitstreams, matrices, variables, and intermediate or final results of processing equations, formulas, operations, and operational logic.

[0179] In several embodiments, the memory within the processor 6010 and / or encoder / decoder module 6030 may be used to store instructions and provide working memory for processing that is executable during encoding or decoding.

[0180] However, in other embodiments, external memory (e.g., the processing device may be processor 6010 or encoder / decoder module 6030) may be used for one or more of these functions. External memory may be memory 6020 and / or storage device 6040, such as dynamic volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory may be used to store the television's operating system. In at least one embodiment, fast external dynamic volatile memory, such as RAM, may be used as working memory for video encoding and decoding operations, such as for MPEG-2 Part 2 (also known as ITU-T Recommendation H.262 and ISO / IEC 13818-2, also known as MPEG-2 video), High Efficiency Video Coding (HEVC), or Multi-Function Video Coding (VVC).

[0181] As indicated in block 6130, inputs to the components of system 6000 can be provided through various input devices. Such input devices include, but are not limited to: (i) inputs that can receive signals, for example, from a broadcaster via the air.

[0182] The RF section of the transmitted RF signal, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.

[0183] In various embodiments, the input device of block 6130 has associated corresponding input processing elements as known in the art. For example, the RF section may be associated with elements required for: (i) selecting a desired frequency (also known as selecting a signal, or limiting a signal band to a band), (ii) down-converting the selected signal, (iii) re-band-limiting the signal to a narrower band to select (e.g.,) a signal band that may be referred to as a channel in some embodiments), (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired data packet stream. The RF section of various embodiments may include one or more elements for performing these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include tuners that perform various functions among these functions, including, for example, down-converting received signals to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or to baseband.

[0184] In one top-box implementation, the RF section and its associated input processing elements can receive RF signals transmitted via a wired (e.g., cable) medium. The RF section can then perform frequency selection by filtering, down-converting, and re-filtering to the desired frequency band.

[0185] Various implementation schemes rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions.

[0186] Adding elements may include inserting elements between existing elements, such as inserting amplifiers and analog-to-digital converters. In various implementations, the RF section may include an antenna.

[0187] Additionally, the USB and / or HDMI terminals may include corresponding interface processors for connecting the system 6000 to other electronic devices across USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) may be implemented as needed, for example, within a separate input processing IC or within the processor 6010. Similarly, various aspects of USB or HDMI interface processing may be implemented as needed, either within a separate interface IC or within the processor 6010. Demodulated streams, error-correcting streams, and demultiplexed streams can be provided to various processing elements, including, for example, the processor 6010, which operates in conjunction with memory and storage elements to process data streams as needed for presentation on an output device, and the encoder / decoder 6030.

[0188] Various components of the system 6000 can be housed within an integrated housing. Within the integrated housing, the various components can be interconnected using a suitable connection arrangement 6140 (e.g., an internal bus known in the art, including an I2C bus, wiring, and printed circuit boards) and data can be transferred between these components.

[0189] System 6000 may include a communication interface 6050 capable of communicating with other devices via a communication channel 6060. The communication interface 6050 may include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 6060. The communication interface 6050 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 6060 may be implemented, for example, within a wired and / or wireless medium.

[0190] In various implementations, data can be streamed to system 6000 using a Wi-Fi network such as IEEE 802.11. The Wi-Fi signal in these implementations can be received via a communication channel 6060 and a communication interface 6050 suitable for Wi-Fi communication. The communication channel 6060 in these implementations can typically be connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other cross-platform communications.

[0191] Other implementations may use a set-top box to provide streaming data to system 6000, the set-top box delivering data via an HDMI connection of input block 6130.

[0192] Other implementations may use the RF connection of input block 6130 to provide streaming data to system 6000.

[0193] Streaming data can be used as a method for sending signaling information in System 5000. Signaling information may include the information INF explained above.

[0194] It should be understood that signaling can be implemented in various ways. For example, in various implementation schemes, one or more syntax elements, flags, etc., are used to signal information to the corresponding decoder.

[0195] System 6000 can provide output signals to various output devices, including display 6100, speaker 6110, and other peripheral devices 6120. In various embodiments of the implementation, other peripheral devices 6120 may include one or more of a standalone DVR, disk player, stereo system, lighting system, and other devices that provide output functionality based on system 6000.

[0196] In various implementations, signaling such as audio / video link (AV.Link), consumer electronics control (CEC), or other communication protocols capable of device-to-device control with or without user intervention can be used to transmit control signals between system 6000 and display 6100, speaker 6110, or other peripheral devices 6120.

[0197] The output devices can be communicatively coupled to the system 6000 via dedicated connections through the corresponding interfaces 6070, 6080 and 6090.

[0198] Alternatively, the output device can be connected to the system 6000 via communication interface 6050 using communication channel 6060. The display 6100 and speaker 6110 can be integrated into a single unit with other components of the system 6000 in an electronic device (e.g., a television).

[0199] In various implementations, the display interface 6070 may include a display driver, such as, for example, a timing controller (TCon) chip.

[0200] Alternatively, if the RF portion of input 6130 is part of a separate set-top box, the display 6100 and speaker 6110 may be separated from one or more other components. In various embodiments where the display 6100 and speaker 6110 can be external components, the output signal may be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.

[0201] In V-PCC, the regular geometry patch and the first geometry patch of the 2D sample are stored in the geometry image GI, the second geometry patch of the 2D sample is stored in the occupancy map OM, and the regular, first and second attribute patches of the 2D sample are stored in the attribute image TI at the location indicated by the syntax element.

[0202] Storing the regular, first and second attribute patches of a 2D sample together in the same video stream has the advantage of requiring only one video encoder / decoder to encode / decode the attribute information of the 3D sample.

[0203] By design, conventional and first / second geometry / attribute encoding modes respond to different needs. Therefore, their attribute formats, scan orders, sizes, and shapes differ. Consequently, the way information is reduced (or encoded / compressed) may require different approaches, especially when it comes to reducing spatial “internal” redundancy. Furthermore, it may be of interest to encode the geometry / attributes of 3D samples (EOM codewords) using a lossless video encoder, while the geometry / attributes of sparse 3D samples (something less representative of a typical point cloud scene) typically encoded using the first geometry / attribute encoding mode may undergo lossy encoding, or vice versa, depending on the application.

[0204] Therefore, encoding the geometry / attributes of a 3D sample requires both these conventional and first- and second-order encoding modes, which require specific (uncommon) molded video codecs containing processes suitable for each of these two encoding modes, as well as additional signaling to the decoder indicating which of these encoding modes must be used to decode the geometry and attributes of the 3D sample.

[0205] Therefore, allowing the use of multiple encoding modes to encode 3D samples of point cloud frames may impact performance (causing encoder / decoder complexity and bandwidth reduction).

[0206] Generally, at least one of the embodiments provides a method for encoding the attributes of an orthographically projected 3D sample and an intermediate 3D sample, wherein information INF indicates whether at least one first attribute patch of a 2D sample obtained by encoding the attributes of the at least one orthographically projected 3D sample according to a first attribute encoding mode and at least one second attribute patch of a 2D sample of an image obtained by encoding the attributes of at least one intermediate 3D sample are stored in a separate image.

[0207] A single image can refer to an image in the same video stream but at different times, or an image from a different video stream (which does not belong to the same video stream).

[0208] The information INF allows for greater flexibility in the use of video codecs and also allows for better compression performance because the video codec can be tailored to encode the properties of orthogonal 3D samples or intermediate 3D samples. The video codec can then be tailored to take into account the characteristics of these 3D sample properties and / or to meet application constraints and requirements. For example, first and / or second attribute patches of 2D samples can be discarded depending on the application or decoder / renderer capabilities (e.g., some 3D samples may not be helpful for rendering point clouds on small displays and may have associated low-end SoCs for decoding).

[0209] Furthermore, video streams that store those attributes can be processed in parallel.

[0210] Figure 7 An example flowchart illustrating a method for encoding orthogonal 3D samples of point cloud frames according to at least one implementation scheme is provided.

[0211] In step 710, the information INF can be set to a first specific value (e.g., 1) to indicate that at least one first attribute patch FA2DP of the 2D sample (step 3400) and at least one second attribute patch SA2DP of the 2D sample (step 3400) are stored in separate images. Alternatively, the information INF can be set to a second specific value (e.g., 0) to indicate that the at least one first attribute patch FA2DP of the 2D sample and the at least one second attribute patch SA2DP of the 2D sample are stored in the same image.

[0212] When the information INF equals the first specific value, then in step 720, the at least one first attribute patch FA2DP of the 2D sample is stored in the first image FAI, and the at least one second attribute patch SA2DP of the 2D sample is stored in the second image SAI.

[0213] When the information INF equals the second specific value, then in step 730, the at least one first attribute patch FA2DP of the 2D sample and the at least one second attribute patch SA2DP of the 2D sample are stored in the same image AI.

[0214] According to the implementation plan, such as Figure 7a As shown, the information INF includes a first binary tag F1 and a second binary tag F2.

[0215] In step 740, the first marker F1 may be set to a first specific value (e.g., 1) to indicate that at least one first attribute patch FA2DP of the 2D sample is stored in the first image FA1 (step 3400), and the first marker may be set to a second specific value (e.g., 0) to indicate that at least one first attribute patch FA2DP of the 2D sample is stored in the image AI (step 3400).

[0216] In step 750, the second marker F2 may be set to a first specific value (e.g., 1) to indicate that at least one second attribute patch SA2DP of the 2D sample is stored in the second image FA2 (step 3400), and the second marker may be set to a second specific value (e.g., 0) to indicate that at least one second attribute patch FA2DP of the 2D sample is stored in the image AI (step 3400).

[0217] When the first marker FA and the second marker F2 are equal to 1, the first attribute patch and the second attribute patch FA2DP and SA2DP of the 2D sample are stored in a separate image.

[0218] When the first marker F1 equals 0 and the second marker F2 equals 1, the first attribute patch FA2DP and the regular attribute 2D patch RA2DP of the 2D sample are stored in the same image, and the second attribute patch SA2DP of the 2D sample is stored in a separate image.

[0219] When the first marker F1 equals 1 and the second marker F2 equals 0, the second attribute patch SA2DP and the regular attribute 2D patch RA2DP of the 2D sample are stored in the same image, and the first attribute patch FA2DP of the 2D sample is stored in a separate image.

[0220] When the first marker FA and the second marker F2 are equal to 0, the first attribute patch and the second attribute patch FA2DP and SA2DP of the 2D sample are stored in the same image.

[0221] Depending on the variant, the information INF is valid at a set of picture levels, frame / atlas levels, or patch levels.

[0222] Depending on the variant, syntax elements are added to the bitstream to identify the video codec used to compress the video stream carrying a second attribute patch of 2D samples.

[0223] This codec can be identified either by mapping SEI messages using component codecs or by means outside the V-PCC specification.

[0224] Syntax elements represented as ai_eom_attribute_codec_id[atlas_id] can be conditionally added to the attribute information syntax structure until a specific value of another syntax element is added, for example... Figure 8 The syntax element vpcc_eom_patch_separate_video_present_flag.

[0225] Figure 8 An instance of the syntax element vpcc_eom_patch_separate_video_present_flag of the embedding information INF according to at least one implementation is shown.

[0226] The syntax element vpcc_eom_patch_separate_video_present_flag can be encoded in a parameter set, such as a sequence parameter set (SPS), an atlas sequence parameter set (ASPS), or a picture parameter set (PPS).

[0227] Figure 8 The elements have the following semantics:

[0228] vpcc_eom_patch_separate_video_present_flag is set to 1 to indicate that the second attribute patch of the 2D sample with index j can be stored in a separate video stream.

[0229] vpcc_eom_patch_separate_video_present_flag[j] is equal to 0 to indicate that the second attribute patch of the 2D sample with index j should not be stored in a separate video stream.

[0230] If vpcc_eom_patch_separate_video_present_flag[j] does not exist, it is inferred that it is equal to 0.

[0231] `epdu_patch_in_eom_video_flag[p]` specifies whether the attribute data associated with the second attribute patch of the 2D sample with index p in the current atlas tile group is encoded in a separate video compared to the attribute data of the intra-frame and inter-frame encoded patches. If `epdu_patch_in_eom_video_flag[p]` equals 0, the attribute data associated with the second attribute patch of the 2D sample with index j in the current atlas tile group is encoded in the same video as the attribute data of the intra-frame and inter-frame encoded patches. If `epdu_patch_in_eom_video_flag[p]` equals 1, the attribute data associated with the second attribute patch of the 2D sample with index j in the current atlas tile group is encoded in a different video than the attribute data of the intra-frame and inter-frame encoded patches. If `epdu_patch_in_eom_video_flag[p]` does not exist, its value should be inferred to be 0.

[0232] In the variant, the syntax element vpcc_eom_patch_separate_video_present_flag[j] can be common for the first and second attribute patches of a 2D sample.

[0233] Next, when vpcc_separate_video_present_flag[j] equals 1, it indicates that for an atlas with index j, the second attribute patch of the 2D sample, the first attribute patch of the 2D sample, and the first geometry patch of the 2D sample can be stored in a separate image (separate video stream).

[0234] When vpcc_separate_video_present_flag[j] equals 0, it indicates that for an atlas with index j, the second attribute patch, the first attribute patch, and the first geometry patch of the 2D sample should not be stored in separate video streams.

[0235] If vpcc_separate_video_present_flag[j] does not exist, it is inferred that it is equal to 0.

[0236] exist Figures 1 to 8 This document describes various methods, and each method includes one or more steps or actions for implementing the described method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or purpose of a particular step and / or action may be modified or combined.

[0237] The block diagrams and flowcharts describe some examples. Each block represents a circuit element, module, or section of code, containing one or more executable instructions for implementing a specified logical function. It should also be noted that in other implementations, the functions marked in a block may not appear in the indicated order. For example, two blocks shown consecutively may actually execute substantially simultaneously, or these blocks may sometimes execute in reverse order depending on the functionality involved.

[0238] The implementations and aspects described herein may be implemented, for example, in a method or process, an apparatus, a computer program, a data stream, a bit stream, or a signal. Even if discussed only in the context of a single form of implementation (e.g., discussed only as a method), the implementation of the discussed features may also be implemented in other forms (e.g., an apparatus or a computer program).

[0239] These methods can be implemented, for example, in a processor, which typically refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. The processor also includes communication devices.

[0240] Additionally, the method can be implemented by instructions executed by a processor, and such instructions (and / or data values ​​generated by implementation) can be stored on a computer-readable storage medium. The computer-readable storage medium can take the form of a computer-readable program product embodied in one or more computer-readable media, on which computer-readable program code executable by a computer is embodied. Given its inherent ability to store and retrieve information therein, the computer-readable storage medium as used herein can be considered a non-transitory storage medium. A computer-readable storage medium can be, for example, but not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or apparatuses, or any suitable combination thereof. It should be understood that while providing more specific examples of computer-readable storage media to which this embodiment can be applied, the following are merely illustrative and not an exhaustive list readily understood by those skilled in the art: portable computer floppy disks; hard disks; read-only memory (ROM); erasable programmable read-only memory (EPROM or flash memory); portable optical disc read-only memory (CD-ROM); optical storage devices; magnetic storage devices; or any suitable combination thereof.

[0241] Instructions can form an application program, which is tangibly embodied in a processor-readable medium.

[0242] Instructions can be, for example, hardware, firmware, software, or a combination thereof. Instructions can be found, for example, in an operating system, a standalone application, or a combination of both. Therefore, a processor can be characterized as, for example, both a means configured to execute a process and a means containing a processor-readable medium (e.g., a storage device) with instructions for executing the process. Furthermore, in addition to or instead of instructions, the processor-readable medium can store data values ​​generated by the implementation.

[0243] The device can be implemented, for example, in appropriate hardware, software, and firmware. Examples of such devices include personal computers, laptops, smartphones, tablets, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, head-mounted displays (HMDs, see-through glasses), projectors (beam emitters), "cave displays" (systems containing multiple displays), servers, video encoders, video decoders, post-processors that process the output from video decoders, pre-processors that provide input to video encoders, video encoders, web servers, set-top boxes, and any other means for processing point clouds, video, or images, or other communication devices. It should be understood that the device can be mobile and can even be installed in mobile vehicles.

[0244] Computer software may be implemented via processor 6010, hardware, or a combination of hardware and software. As a non-limiting example, the implementation may be implemented using one or more integrated circuits. Memory 6020 may be of any type suitable for the technical environment, and as a non-limiting example, any suitable data storage technology may be used, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. Processor 6010 may be of any type suitable for the technical environment, and as a non-limiting example, may encompass one or more of microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.

[0245] It will be apparent to those skilled in the art that embodiments may produce various signals formatted to carry, for example, storable or transmissible information. The information may include, for example, instructions for performing a method or data generated by one of the embodiments. For example, signals may be formatted to carry a bitstream of the embodiments. Such signals may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or baseband signals. Formatting may include, for example, encoding the data stream and using a modulated carrier with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is known, signals can be transmitted via a variety of different wired or wireless links. Signals may be stored on a processor-readable medium.

[0246] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used herein, the singular forms “a / an” and “the” may also be intended to include the plural forms unless the context clearly indicates otherwise. It should be further understood that, when used in this specification, the terms “includes / comprises” and / or “including / comprising” may specify the stated presence, such as features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Furthermore, when an element is referred to as “responding” or “connected” to another element, it may directly respond to or be connected to the other element, or there may be intermediate elements present. Conversely, when an element is referred to as “directly responding” or “directly connected” to another element, there are no intermediate elements.

[0247] It should be understood that, for example, in the cases of “A / B,” “A and / or B,” and “at least one of A and B,” the use of any of the symbols / terms “ / ,” “and / or,” and “at least one” may be intended to cover selecting only the first listed option (A), or only the second listed option (B), or both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C,” such phrases are intended to cover selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A, B, and C). As will be apparent to those skilled in the art and related fields, this can be extended to as many items as possible listed.

[0248] Various numerical values ​​may be used in this application. Specific values ​​may be used, for example, for illustrative purposes, and the aspects described are not limited to these specific values.

[0249] It should be understood that although the terms first, second, etc., may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the teachings of this application. There is no implied order between the first element and the second element.

[0250] References to “an implementation,” “implementation,” “an implementation,” or “implementation,” and their other variations, are frequently used to convey that a particular feature, structure, characteristic, etc., described in connection with an implementation / implementation is included in at least one implementation / implementation. Therefore, the appearance of the phrase “in an implementation,” “in an implementation,” “in a specific implementation,” or “in a specific implementation,” and any other variations appearing throughout this application, do not necessarily refer to the same implementation.

[0251] Similarly, references to "according to an implementation scheme / example / implementation" or "in an implementation scheme / example / implementation" and their variations are frequently used herein to convey that a particular feature, structure, or characteristic (described in conjunction with an implementation scheme / example / implementation) may be included in at least one implementation scheme / example / implementation. Therefore, expressions "according to an implementation scheme / example / implementation" or "in an implementation scheme / example / implementation" appearing in various places in the specification do not necessarily refer to the same implementation scheme / example / implementation, and individual or alternative implementation schemes / examples / implementations are not necessarily mutually exclusive with other implementation schemes / examples / implementations.

[0252] Reference numerals appearing in the claims are for illustrative purposes only and should not be construed as limiting the scope of the claims. Although not explicitly described, embodiments / examples of the invention may be employed in any combination or sub-combination.

[0253] When an illustration is presented as a flowchart, it should be understood that a block diagram of the corresponding device is also provided. Similarly, when an illustration is presented as a block diagram, it should be understood that a flowchart of the corresponding method / process is also provided.

[0254] Although some diagrams include arrows along the communication path to show the main direction of communication, it should be understood that communication can occur in the opposite direction to the arrows depicted.

[0255] Various specific implementations involve decoding. As used herein, "decoding" can encompass all or part of the processes performed, for example, on received point cloud frames (containing possible received bitstreams encoding one or more point cloud frames) to produce a final output suitable for display or suitable for further processing in the reconstructed point cloud domain. In various embodiments, such processes include one or more processes typically performed by an image-based decoder. In various embodiments, such processes also include, or alternatively include, processes performed by the decoders of the various embodiments described herein.

[0256] As a further example, in one embodiment, "decoding" may refer only to entropy decoding; in another embodiment, "decoding" may refer only to differential decoding; and in yet another embodiment, "decoding" may refer to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" specifically refers to a subset of operations or broadly refers to a wider decoding process will be clear based on the specific context of the description and is believed to be well understood by those skilled in the art.

[0257] Various specific implementations involve encoding. In a manner similar to the discussion of “decoding” above, “encoding,” as used herein, can encompass all or part of the process performed, for example, on an input point cloud frame to produce an encoded bitstream. In various embodiments, such processes include one or more processes typically performed by an image-based decoder. In various embodiments, such processes also include, or alternatively include, processes performed by an encoder of the various specific implementations described in this application.

[0258] As a further example, in one embodiment, "encoding" may refer only to entropy encoding; in another embodiment, "encoding" may refer only to differential encoding; and in yet another embodiment, "encoding" may refer to a combination of differential encoding and entropy encoding. Whether the phrase "encoding process" specifically refers to a subset of operations or broadly refers to a wider encoding process will be clear based on the specific context of the description and is believed to be well understood by those skilled in the art.

[0259] Note that the grammatical elements used in this article are descriptive terms. Therefore, the use of other grammatical element names is not excluded.

[0260] Various implementation schemes refer to rate distortion optimization. Specifically, during the encoding process, a balance or trade-off between rate and distortion is typically considered, often taking into account computational complexity constraints. Rate distortion optimization can generally be formulated as minimizing a rate distortion function, which is a weighted sum of rate and distortion. Different approaches exist to solve the rate distortion optimization problem. For example, these methods may be based on extensive testing of all encoding options (including all considered modes or encoding parameter values) and a complete evaluation of their encoding costs and the associated distortion of the reconstructed signal after encoding and decoding. Faster methods can also be used to reduce encoding complexity, particularly for the computation of approximate distortion based on prediction or prediction of the residual signal rather than the reconstructed residual signal. A hybrid of these two approaches can also be used, for example by using approximate distortion only for some possible encoding options and full distortion for others. Other methods evaluate only a subset of possible encoding options. More generally, many methods employ any of a variety of techniques to perform optimization, but optimization is not necessarily a complete evaluation of both encoding costs and associated distortion.

[0261] Additionally, this application may involve "determining" various types of information. Determining information may include, for example, one or more of the following: estimation information, calculation information, prediction information, or information retrieved from memory.

[0262] Furthermore, this application may relate to "accessing" various types of information. Accessing information may include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information, or one or more of these.

[0263] Furthermore, this application may relate to "receiving" various types of information. Like "access," "receiving" is intended to be a broad term. Receiving information may include, for example, accessing information or retrieving information from memory, or one or more of these. Moreover, "receiving" typically involves one or more of the following during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0264] Moreover, as used herein, the term "signaling" refers to (among other things) instructing the corresponding decoder to do something. For example, in some implementations, the encoder signals specific information INF, which may be carried by the syntax element vpcc_eom_patch_separate_video_present_flag. In this way, in implementations, the same parameters are used on both the encoder and decoder sides. Thus, for example, the encoder may transmit (explicit signaling) specific parameters to the decoder so that the decoder can use the same specific parameters. Conversely, if the decoder already has specific parameters as well as other parameters, signaling may be used without transmission (implicit signaling) to simply allow the decoder to know and select specific parameters. Bit savings are achieved in various implementations by avoiding the transmission of any actual functionality. It should be understood that signaling can be implemented in many ways. For example, in various implementations, information is signaled to the corresponding decoder using one or more syntax elements, flags, etc. Although the verb form of the word "signal" was mentioned above, the word "signal" may also be used as a noun in this document.

[0265] Several embodiments have been described. However, it should be understood that various modifications can be made. For example, elements of different embodiments can be combined, supplemented, modified, or removed to produce other embodiments. Furthermore, those skilled in the art will understand that the disclosed structures and processes can be replaced with other structures and processes, and the resulting embodiments will perform at least substantially the same function in at least substantially the same manner to achieve at least substantially the same results as the disclosed embodiments. Therefore, these and other embodiments are considered in this application.

Claims

1. A method comprising encoding attributes of 3D samples of a point cloud, comprising: - encoding at least one attribute of a first 3D sample of the point cloud, wherein the at least one attribute of the first 3D sample is encoded as at least one first attribute patch of a 2D sample, wherein the first 3D sample is a 3D sample whose 3D coordinates are encoded as pixel values of a patch of a geometry image, - encoding at least one attribute of an intermediate 3D sample, the intermediate 3D sample being a 3D sample of the point cloud located between two orthogonally projected 3D samples of the point cloud along a same projection line, and the 3D coordinates of the intermediate 3D sample being encoded as coordinates of a pixel on a projection plane onto which the intermediate 3D sample is projected along the projection line, and a codeword being encoded in an occupancy image, wherein one bit of the codeword indicates a position of the intermediate 3D sample along the projection line, wherein the at least one attribute of the intermediate 3D sample is encoded as at least one second attribute patch of a 2D sample, and - encoding information indicating whether the at least one first attribute patch of a 2D sample and the at least one second attribute patch of a 2D sample are stored in separate images, wherein the information comprises: a first flag indicating whether the at least one first attribute patch of a 2D sample is stored in a first image or whether the at least one first attribute patch of a 2D sample is stored together with other attribute patches of a 2D sample in a second image, and a second flag indicating whether the at least one second attribute patch of a 2D sample is stored in a third image or whether the at least one second attribute patch of a 2D sample is stored together with other attribute patches of a 2D sample in the second image.

2. The method according to claim 1, wherein the method further comprises encoding a further information indicating how the separate images are compressed.

3. A method comprising decoding attributes of 3D samples of a point cloud, comprising: - decoding at least one attribute of a first 3D sample of the point cloud from at least one first attribute patch of a 2D sample, wherein the first 3D sample is a 3D sample whose 3D coordinates are encoded as pixel values of a patch of a geometry image, - decoding at least one attribute of an intermediate 3D sample, the intermediate 3D sample being a 3D sample of the point cloud located between two orthogonally projected 3D samples of the point cloud along a same projection line, and the 3D coordinates of the intermediate 3D sample being decoded as coordinates of a pixel on a projection plane onto which the intermediate 3D sample is projected along the projection line, and a codeword being decoded from an occupancy image, wherein one bit of the codeword indicates a position of the intermediate 3D sample along the projection line, wherein the at least one attribute of the intermediate 3D sample is decoded as at least one second attribute patch of a 2D sample, - decoding information indicating whether the at least one first attribute patch of a 2D sample and the at least one second attribute patch of a 2D sample are stored in separate images, wherein the information comprises: a first flag indicating whether the at least one first attribute patch of a 2D sample is stored in a first image or whether the at least one first attribute patch of a 2D sample is stored together with other attribute patches of a 2D sample in a second image, and a second flag indicating whether the at least one second attribute patch of a 2D sample is stored in a third image or whether the at least one second attribute patch of a 2D sample is stored together with other attribute patches of a 2D sample in the second image. a first flag indicating whether the at least one first attribute patch of the 2D sample is stored in a first image or whether the at least one first attribute patch of the 2D sample is stored together with other attribute patches of the 2D sample in a second image, and a second flag indicating whether the at least one second attribute patch of the 2D sample is stored in a third image or whether the at least one second attribute patch of the 2D sample is stored together with other attribute patches of the 2D sample in the second image.

4. The method of claim 1, wherein the at least one attribute of the first 3D sample, the at least one attribute of the intermediate 3D sample and the information are encoded in a video stream, the video stream being hierarchically structured at a group of picture level, a frame level and a patch level, and wherein the information is valid at the group of picture level, frame level, atlas level or patch level.

5. The method of claim 3, wherein the at least one attribute of the first 3D sample, the at least one attribute of the intermediate 3D sample and the information are decoded from a video stream, the video stream being hierarchically structured at a group of picture level, a frame level and a patch level, and wherein the information is valid at the group of picture level, frame level, atlas level or patch level.

6. The method of claim 3, wherein the method further comprises decoding a further information indicating how the individual images are compressed.

7. A device for encoding attributes of 3D samples of a point cloud, comprising one or more processors configured to: - encode at least one attribute of a first 3D sample of the point cloud, wherein the at least one attribute of the first 3D sample is encoded as at least one first attribute patch of a 2D sample, wherein the first 3D sample is a 3D sample whose 3D coordinates are encoded as pixel values of patches of a geometry image, - encode at least one attribute of an intermediate 3D sample, the intermediate 3D sample being a 3D sample of the point cloud located between two orthogonal projections of 3D samples of the point cloud along a same projection line and the 3D coordinates of the intermediate 3D sample being encoded as coordinates of a pixel on a projection plane onto which the intermediate 3D sample is projected along the projection line and a codeword being encoded in an occupancy image, wherein one bit of the codeword indicates a position of the intermediate 3D sample along the projection line, wherein the at least one attribute of the intermediate 3D sample is encoded as at least one second attribute patch of a 2D sample, - encode an information indicating whether the at least one first attribute patch of the 2D sample and the at least one second attribute patch of the 2D sample are stored in individual images, wherein the information comprises: a first flag indicating whether the at least one first attribute patch of the 2D sample is stored in a first image or whether the at least one first attribute patch of the 2D sample is stored together with other attribute patches of the 2D sample in a second image, and a second flag indicating whether the at least one second attribute patch of the 2D sample is stored in a third image or whether the at least one second attribute patch of the 2D sample is stored together with other attribute patches of the 2D sample in the second image. a second flag indicating whether the at least one first attribute patch of the 2D sample is stored in a first image or whether the at least one first attribute patch of the 2D sample is stored together with other attribute patches of the 2D sample in a second image, and 8. A device for decoding attributes of 3D samples of a point cloud, comprising one or more processors configured to: - decode at least one attribute of a first 3D sample of the point cloud from at least one first attribute patch of a 2D sample, wherein the first 3D sample is a 3D sample whose 3D coordinates are encoded as pixel values of patches of a geometry image, - decode at least one attribute of an intermediate 3D sample, the intermediate 3D sample being a 3D sample of the point cloud located between two orthogonally projected 3D samples of the point cloud along a same projection line, and the 3D coordinates of the intermediate 3D sample being decoded as coordinates of a pixel on a projection plane onto which the intermediate 3D sample is projected along the projection line, and a codeword being decoded from an occupancy image, wherein one bit of the codeword indicates a position of the intermediate 3D sample along the projection line, wherein the at least one attribute of the intermediate 3D sample is decoded as at least one second attribute patch of a 2D sample, - decode information indicating whether the at least one first attribute patch of the 2D sample and the at least one second attribute patch of the 2D sample are stored in separate images, wherein the information comprises: a first flag indicating whether the at least one first attribute patch of the 2D sample is stored in a first image or whether the at least one first attribute patch of the 2D sample is stored together with other attribute patches of the 2D sample in a second image, and a second flag indicating whether the at least one second attribute patch of the 2D sample is stored in a third image or whether the at least one second attribute patch of the 2D sample is stored together with other attribute patches of the 2D sample in the second image.

9. The device of claim 7, wherein the at least one attribute of the first 3D sample, the at least one attribute of the intermediate 3D sample and the information are encoded in a video stream, the video stream being hierarchically structured at a set of picture level, frame level and patch level, and the information being valid at the set of picture level, frame level, atlas level or patch level.

10. The device of claim 8, wherein the at least one attribute of the first 3D sample, the at least one attribute of the intermediate 3D sample and the information are decoded from a video stream, the video stream being hierarchically structured at a set of picture level, frame level and patch level, and the information being valid at the set of picture level, frame level, atlas level or patch level.

11. The device of claim 7 or 8, wherein the device further comprises means for encoding or decoding another information indicating how the separate images are compressed.

12. A non-transitory computer readable medium containing instructions for causing one or more processors to perform the method of claim 1 or 3.

13. The method of claim 1 or 3, wherein, the 3D coordinates of the orthogonally projected 3D samples are encoded as coordinates and pixel values of pixels on a projection plane onto which the orthogonally projected 3D samples are projected.

14. The method of claim 13, wherein the codewords are padded to form at least one second geometric patch of 2D samples stored in an occupancy image.

15. The device of claim 7 or 8, wherein, the 3D coordinates of the orthogonally projected 3D samples are encoded as coordinates and pixel values of pixels on a projection plane onto which the orthogonally projected 3D samples are projected.

16. The device of claim 15, wherein the codewords are padded to form at least one second geometric patch of 2D samples stored in an occupancy image.

Citation Information

Patent Citations

  • Point cloud compression using hybrid transforms

    US20190122393A1

  • An apparatus, a method and a computer program for volumetric video

    WO2019115867A1