Processing point cloud

Through the two-layer point cloud encoding structure and the point local reconstruction mode of signaling notification, the efficiency problem in dynamic point cloud data compression and distribution is solved, and bit rate reduction and quality maintenance are achieved.

CN113614786BActive Publication Date: 2025-07-08INTERDIGITAL VC HOLDINGS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080021646.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-03-15
Filing Date
2020-01-27
Publication Date
2025-07-08
Estimated Expiration
2040-01-27

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently compress and distribute dynamic point cloud data, especially to reduce bit consumption while maintaining an acceptable quality of experience.

Method used

Using a two-layer point cloud encoding structure, including the basic layer and the enhancement layer, the geometric and texture information of the point cloud is compressed by detecting and encoding isolated points, combining existing video codecs, and optimizing point local reconstruction mode using lookup tables and signaling notification methods to reduce computing resources and bit rate.

Benefits of technology

It realizes the significant reduction in bit rate consumption while maintaining point cloud quality, and improves the compression efficiency and distribution capabilities of point cloud data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113614786B_ABST
    Figure CN113614786B_ABST
Patent Text Reader

Abstract

At least one embodiment relates to a method for signaling a syntax element representing a point local reconstruction mode, where the point local reconstruction mode represents at least one parameter defining a mode for reconstructing at least one point of a point cloud frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] At least one of the present embodiments generally relates to the processing of point clouds. Background Art

[0002] This section is intended to introduce the reader to aspects of the field that may be related to aspects of at least one current embodiment described and / or claimed below. This discussion is considered to be helpful to provide the reader with background information to better understand aspects of at least one embodiment.

[0003] Point clouds can be used for various purposes, such as cultural heritage / architecture, where objects like statues or buildings are scanned in 3D to share the spatial configuration of the object without sending or accessing it. Additionally, this is a way to ensure the preservation of knowledge of an object in case the object may be destroyed; for example, a temple that may be destroyed by an earthquake. Such point clouds are typically static, colored, and huge.

[0004] Another use case is in terrain and cartography, where using 3D representation allows maps to not be limited to a plane and can include terrain. Google Maps is now a good example of a 3D map, but it uses meshes instead of point clouds. However, for 3D maps, point clouds may be a suitable data format, and such point clouds are typically static, colored, and huge.

[0005] The automotive industry and autonomous vehicles are also areas where point clouds can be used. Autonomous vehicles should be able to "sense" their environment and make good driving decisions based on the reality of their neighbors. Typical sensors like Light Detection And Ranging (LIDAR) produce dynamic point clouds for use by the decision-making engine. These point clouds are not intended to be seen by humans, they are typically small, not necessarily colored, and are captured dynamically at a high frequency. These point clouds can have other attributes, such as the reflectivity provided by LIDAR, as this attribute provides good information about the material of the sensed object and may help in making decisions.

[0006] Virtual reality and immersive worlds have recently become a hot topic and are foreseen by many as the future of 2D flat video. The basic idea is to immerse the viewer in the environment around the viewer, which is in contrast to a standard TV where the viewer can only look at the virtual world in front of the viewer. There are several levels of immersion, depending on the degree of freedom of the viewer in the environment. Point clouds are a good format candidate for distributing virtual reality (VR) worlds.

[0007] In many applications, it is very important to be able to distribute dynamic point clouds to end users (or store them in a server) while consuming only a reasonable amount of bits (or the storage space of the application), while maintaining an acceptable (or preferably very good) quality of experience. To make the distribution chain of many immersive worlds practical, efficiently compressing these dynamic point clouds is a key point.

[0008] In view of the foregoing, at least one embodiment has been devised. SUMMARY OF THE INVENTION

[0009] A simplified overview of at least one current embodiment is presented below in order to provide a basic understanding of some aspects of the present disclosure. This summary of the invention is not an extensive overview of the embodiments. It is not intended to identify key or critical elements of the embodiments. The following overview merely presents some aspects of at least one current embodiment in a simplified form as a prelude to a more detailed description provided elsewhere in this document.

[0010] According to a general aspect of at least one embodiment, a method is provided that includes signaling in a bitstream a first syntax element representing a Point Local Reconstruction mode, the Point Local Reconstruction mode representing at least one parameter defining a mode for reconstructing at least one point of a point cloud frame.

[0011] According to one embodiment, the Point Local Reconstruction mode is an index value of an entry in a lookup table, and the entry in the lookup table defines a relationship between the index value and the at least one parameter.

[0012] According to one embodiment, the first syntax element is signaled for each block or each patch, where a patch is a set of at least one block of 2D samples that represent an orthogonal projection of at least one 3D sample of a point cloud frame onto a projection plane.

[0013] According to one embodiment, a patch is a set of at least one block of 2D samples that represent an orthogonal projection of at least one 3D sample of a point cloud frame onto a projection plane, and the method further includes signaling a second syntax element for each patch, the second syntax element indicating whether a single first syntax element is signaled for all blocks of the patch at once, or whether the first syntax element is signaled for each block of the patch.

[0014] According to one embodiment, the second syntax element is signaled only if the patch contains more than a given number of blocks.

[0015] According to one embodiment, the method further includes signaling a third syntax element representing a default Point Local Reconstruction mode.

[0016] According to one embodiment, the third syntax element is signaled for each point cloud frame and / or each tile, where a tile is a set of at least one block of 2D samples that represent the orthogonal projection of at least one 3D sample of a point cloud frame onto a projection plane.

[0017] According to one embodiment, a tile is a set of at least one block of 2D samples that represent the orthogonal projection of at least one 3D sample of a point cloud frame onto a projection plane, and the method further includes signaling a fourth syntax element that is used to signal the granularity of the first syntax element for each block or each tile, and the use of the third syntax element for each tile.

[0018] According to one embodiment, the first syntax element is coded and decoded differently from at least one previously signaled first syntax element.

[0019] One or more of at least one current embodiment also provide an apparatus, a computer program, a signal, and a non-transitory computer-readable storage medium.

[0020] From the following description of examples in conjunction with the drawings, the specific nature of at least one current embodiment and other objects, advantages, features, and uses of the at least one current embodiment will become apparent. Description of the Drawings

[0021] In the drawings, examples of several embodiments are shown. The drawings show:

[0022] Figure 1 A schematic block diagram showing an example of a two-layer based point cloud coding structure according to at least one current embodiment;

[0023] Figure 2 A schematic block diagram showing an example of a two-layer based point cloud decoding structure according to at least one current embodiment;

[0024] Figure 3 A schematic block diagram showing an example of an image-based point cloud encoder according to at least one current embodiment;

[0025] Figure 3a An example of a canvas showing two tiles and their 2D bounding boxes;

[0026] Figure 4 A schematic block diagram showing an example of an image-based point cloud decoder according to at least one current embodiment;

[0027] Figure 5 A schematic syntax example of a bitstream representing a base layer BL according to at least one current embodiment;

[0028] Figure 6A schematic block diagram showing an example of a system in which various aspects and embodiments are implemented;

[0029] Figure 7 A schematic block diagram showing an example of a method for signaling a local reconstruction mode of a signaling point according to at least one current embodiment;

[0030] Figures 8a - 8b An example of a lookup table LUT according to an embodiment of step 710 is shown;

[0031] Figure 8c An example of the syntax for signaling a lookup table LUT according to an embodiment of step 710 is shown;

[0032] Figures 8d - 8e An example of an embodiment of step 710 is shown;

[0033] Figure 9 Shows Figure 7 An example of an embodiment of the method of;

[0034] Figure 10 Shows Figure 7 An example of a variant of an embodiment of the method of;

[0035] Figure 11 Shows Figure 8c An example of a variant of;

[0036] Figures 12a - 12b Shows Figure 7 An example of an embodiment of the method of;

[0037] Figure 13 Shows Figure 12b An example of a variant of the method of;

[0038] Figure 14 Shows Figure 7 An example of a variant of the method of; and

[0039] Figure 15 Shows Figure 7 An example of an embodiment of the method of. Detailed Description

[0040] Hereinafter, at least one current embodiment will be described more fully with reference to the accompanying drawings, in which examples of at least one current embodiment are shown. However, the embodiments can be embodied in many alternative forms and should not be construed as limited to the examples set forth herein. Thus, it should be understood that the embodiments are not intended to be limited to the particular forms disclosed. On the contrary, the present disclosure is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present application.

[0041] When a figure is presented as a flowchart, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / process.

[0042] Similar or identical elements in the drawings are referenced by the same reference numerals.

[0043] Some figures represent syntax tables widely used in V-PCC, which are used to define the structure of the bitstream conforming to V-PCC. In those syntax tables, the term "…" represents the unchanged part of the syntax regarding the original definition given in V-PCC, and these parts are removed from the figures for ease of reading. The bold terms in the figures indicate that the value of the term is obtained by parsing the bitstream. The right column of the syntax table indicates the number of bits used to encode the data of the syntax element. For example, u(4) indicates that 4 bits are used to encode the data, u(8) indicates 8 bits, and ae(v) represents a syntax element for context-adaptive arithmetic entropy coding and decoding.

[0044] The aspects described and considered below can be implemented in many different forms. The following Figures 1 - 15 provides some embodiments, but other embodiments are also considered, and Figures 1 - 15 the discussion does not limit the breadth of the implementation.

[0045] At least one aspect generally relates to point cloud encoding and decoding, and at least one other aspect generally relates to transmitting the generated or encoded bitstream.

[0046] More precisely, the various methods and other aspects described herein can be used to modify modules, for example, modules related to metadata encoding (such as the encoding occurring in Figure 3 the fragment information encoder 3700) and modules related to metadata decoding (such as the module occurring in Figure 4 the fragment information decoder 4400, or the decoding occurring in the reconstruction process in Figure 4 the geometry generation module 4500).

[0047] Furthermore, this aspect is not limited to the MPEG standard (such as MPEG-1 Part 5 related to point cloud compression), and can be applied to, for example, other standards and recommendations (whether existing or future developed), and extensions of any such standards and recommendations (including MPEG-1 Part 5). Unless otherwise stated, or technically excluded, the aspects described in this application can be used alone or in combination.

[0048] In the following text, image data refers to data, such as one or several arrays of 2D samples in a specific image / video format. The specific image / video format may specify information related to the pixel values of the image (or video). The specific image / video format may also specify information that can be used, for example, by a display and / or any other device to visualize and / or decode the image (or video). An image typically includes a first component in the shape of a first 2D sample array, which typically represents the luminance (or brightness) of the image. The image may also include second and third components in the form of other 2D sample arrays, which typically represent the chrominance (or color concentration) of the image. Some embodiments use a set of 2D color sample arrays (such as the traditional three-color RGB representation) to represent the same information.

[0049] In one or more embodiments, the pixel value is represented by a vector of C values, where C is the number of components. Each value of the vector is typically represented by a plurality of bits, which may define the dynamic range of the pixel value.

[0050] An image block refers to a set of pixels belonging to an image. The pixel values of an image block (or image block data) refer to the pixel values belonging to that image block. An image block can have any shape, although a rectangle is common.

[0051] A point cloud can be represented by a 3D sample data set within a 3D volume space, which has unique coordinates and may also have one or more attributes.

[0052] The 3D samples of the data set can be defined by its spatial position (X, Y, and Z coordinates in 3D space) and may be defined by one or more associated attributes, such as color, transparency, reflectivity, a two-component normal vector, or any feature representing the characteristics of the sample, as represented in the RGB or YUV color space. For example, a 3D sample can be defined by 6 components (X, Y, Z, R, G, B) or equivalently (X, Y, Z, y, U, V), where (X, Y, Z) defines the coordinates of the point in 3D space, and (R, G, B) or (y, U, V) defines the color of the 3D sample. The same type of attribute may occur multiple times. For example, multiple color attributes can provide color information from different perspectives.

[0053] A point cloud can be static or dynamic, depending on whether the cloud changes over time. Instances of static or dynamic point clouds are typically represented as point cloud frames. It should be noted that in the case of a dynamic point cloud, the number of points is usually not constant, but rather typically changes over time. More generally, if anything changes over time, such as the number of points, the position of one or more points, or any attribute of any point, the point cloud can be considered dynamic.

[0054] For example, a 2D sample point can be defined by six components (u, v, Z, R, G, B) or equivalently (u, v, Z, y, U, V). (u, v) defines the coordinates of the 2D sample point in the 2D space of the projection plane. Z is the depth value of the 3D sample point projected onto this projection plane. (R, G, B) or (y, U, V) defines the color of the 3D sample point.

[0055] Figure 1 FIG. shows a schematic block diagram of an example of a two - layer based point cloud coding structure 1000 according to at least one current embodiment.

[0056] The two - layer based point cloud coding structure 1000 can provide a bitstream B representing an input point cloud frame IPCF (input point cloud frame). Possibly, the input point cloud frame IPCF represents a frame of a dynamic point cloud. Then, the frame of the dynamic point cloud can be encoded by the two - layer based point cloud coding structure 1000 independently of another frame.

[0057] Basically, the two - layer based point cloud coding structure 1000 can provide the ability to structure the bitstream B into a base layer BL and an enhancement layer EL. The base layer BL can provide a lossy representation of the input point cloud frame IPCF, while the enhancement layer EL can provide a higher quality (possibly lossless) representation by encoding the isolated points not represented by the base layer BL.

[0058] The base layer BL can be provided by an image - based encoder 3000, as Figure 3 shown. The image - based encoder 3000 can provide a geometric / texture image representing the geometry / attributes of the 3D sample points of the input point cloud frame IPCF. It can allow discarding isolated 3D sample points. The base layer BL can be decoded by Figure 4 the image - based decoder 4000 shown, which can provide an intermediate reconstructed point cloud frame IRPCF (intermediate reconstructed point cloud frame).

[0059] Then, returning to Figure 1 the two - layer based point cloud coding 1000 in, a comparator COMP can compare the 3D sample points of the input point cloud frame IPCF with the 3D sample points of the intermediate reconstructed point cloud frame IRPCF in order to detect / localize the lost / isolated 3D sample points. Next, an encoder ENC can encode the lost 3D sample points and can provide the enhancement layer EL. Finally, the base layer BL and the enhancement layer EL can be multiplexed together by a multiplexer MUX in order to generate the bitstream B.

[0060] According to one embodiment, the encoder ENC may include a detector that may detect the 3D reference sample point R of the intermediate reconstructed point cloud frame IRPCF and associate it with the missing 3D sample point M.

[0061] For example, according to a given metric, the 3D reference sample point R associated with the missing 3D sample point M may be the nearest neighbor of M.

[0062] According to one embodiment, then, the encoder ENC may encode the spatial position and its attributes of the missing 3D sample point M as a difference determined according to the spatial position and attributes of the 3D reference sample point R.

[0063] In a variant, these differences may be encoded separately.

[0064] For example, for a missing 3D sample point M with spatial coordinates x(M), y(M), and z(M), the x-coordinate position difference Dx(M), y-coordinate position difference Dy(M), z-coordinate position difference Dz(M), R attribute component difference Dr(M), G attribute component difference Dg(M), and B attribute component difference Db(M) may be calculated as follows:

[0065] Dx(M) = x(M) - x(R),

[0066] where x(M) and x(R) are respectively Figure 3 the x-coordinates of the 3D sample points M and R in the provided geometric image,

[0067] Dy(M) = y(M) - y(R)

[0068] where y(M) and x(R) are respectively Figure 3 the y-coordinates of the 3D sample points M and R in the provided geometric image,

[0069] Dz(M) = z(M) - z(R)

[0070] where z(M) and x(R) are respectively Figure 3 the z-coordinates of the 3D sample points M and R in the provided geometric image,

[0071] Dr(M) = R(M) - R(R).

[0072] where R(M) and R(R) are respectively the r-color components of the color attributes of the 3D sample points M and R,

[0073] Dg(M) = G(M) - G(R).

[0074] where G(M) and G(R) are respectively the g-color components of the color attributes of the 3D sample points M and R,

[0075] Db(M) = B(M) - B(R).

[0076] Where B(M) and B(R) are the b-color components of the color attributes of the 3D sample points M and R, respectively.

[0077] Figure 2 FIG. shows a schematic block diagram of an example of a two-layer based point cloud decoding structure 2000 according to at least one current embodiment.

[0078] The behavior of the two-layer based point cloud decoding structure 2000 depends on its capabilities.

[0079] As Figure 4 shown, a two-layer based point cloud decoding structure 2000 with limited capabilities can access only the base layer BL from the bitstream B by using a demultiplexer DMUX, and then can provide a faithful (but lossy) version IRPCF of the input point cloud frame IPCF by decoding the base layer BL by the point cloud decoder 4000.

[0080] A two-layer based point cloud decoding structure 2000 with full capabilities can access the base layer BL and the enhancement layer EL from the bitstream B by using a demultiplexer DMUX. As Figure 4 shown, the point cloud decoder 4000 can determine an intermediate reconstructed point cloud frame IRPCF from the base layer BL. The decoder DEC can determine a complementary point cloud frame CPCF (complementary point cloud frame) from the enhancement layer EL. Then, the combiner COM can combine the intermediate reconstructed point cloud frame IRPCF and the complementary point cloud frame CPCF together, thereby providing a higher quality (possibly lossless) representation (reconstruction) CRPCF of the input point cloud frame IPCF.

[0081] Figure 3 FIG. shows a schematic block diagram of an example of an image-based point cloud encoder 3000 according to at least one current embodiment.

[0082] The image-based point cloud encoder 3000 utilizes existing video codecs to compress the geometric and texture (attribute) information of dynamic point clouds. This is mainly achieved by converting the point cloud data into a set of different video sequences.

[0083] In a particular embodiment, two videos can be generated and compressed using existing video codecs, one for capturing the geometric information of the point cloud data and the other for capturing the texture information. An example of an existing video codec is the HEVC main profile encoder / decoder (ITU-T H.265, Telecommunication Standardization Sector of ITU (02 / 2018), Series H: Audiovisual and Multimedia Systems, Audiovisual service infrastructure - Moving picture coding, High Efficiency Video Coding, Recommendation ITU-T H.265).

[0084] Additional metadata for explaining two videos is typically also generated and compressed separately. Such additional metadata includes, for example, an occupancy map (OM) and / or patch information (PI).

[0085] Then, the generated video bitstream and metadata can be multiplexed together to generate a combined bitstream.

[0086] It should be noted that metadata typically represents a small amount of overall information. Most of the information is in the video bitstream.

[0087] The test model category 2 algorithm (also denoted as V-PCC) gives an example of such a point cloud encoding / decoding process, which implements the MPEG draft standard defined in ISO / IEC JTC 1 / SC29 / WG11 MPEG 2019 / w 18180 (January 2019, Marrakech).

[0088] In step 3100, the module PGM can generate at least one patch by decomposing the 3D samples of the dataset representing the input point cloud frame (IPCF) into 2D samples on a projection plane using a strategy that provides optimal compression.

[0089] A patch can be defined as a set of 2D samples.

[0090] For example, in V-PCC, as described by Hoppe et al. (Hugues Hoppe, Tony DeRose, Tom Duchamp, John McDonald, Werner Stuetzle, Surface Reconstruction from Unorganized Points, Proceedings of the ACM SIGGRAPH 1992 Conference, pp. 71-78), the normal vector at each 3D sample point is first estimated. Next, an initial clustering of the input point cloud frame IPCF is obtained by associating each 3D sample point with one of the six oriented planes of the 3D bounding box that encompasses the 3D sample points of the input point cloud frame IPCF. More precisely, each 3D sample point is clustered and associated with the oriented plane having the closest normal vector (i.e., maximizing the dot product of the point normal vector and the plane normal vector). The 3D sample points are then projected onto their associated planes. A set of 3D sample points that form a connected region within their plane is called a connected component. A connected component is a set of at least one 3D sample point having similar normal vectors and the same associated oriented plane. Then, the initial clustering is refined by iteratively updating the clustering associated with each 3D sample point based on the normal vector of each 3D sample point and the clustering of its nearest neighbor sample points. The last step includes generating a patch from each connected component, which is done by projecting the 3D sample points of each connected component onto the oriented plane associated with the connected component. The patch is associated with auxiliary patch information PI, which represents the auxiliary patch information defined for each patch to interpret the projected 2D sample points corresponding to geometric and / or attribute information.

[0091] For example, in V-PCC, the auxiliary patch information PI includes 1) information indicating one of the six oriented planes of the 3D bounding box that encompasses the 3D sample points of the connected component; 2) information with respect to the plane normal vector; 3) information determining the 3D position of the connected component with respect to the patch represented by depth, tangential offset, and bi-tangential offset; and 4) information such as the coordinates (u0, v0, u1, v1) in the projection plane that defines the 2D bounding box encompassing the patch.

[0092] In step 3200, the patch packing module PPM (patch packing module) can map (place) at least one generated patch onto a 2D grid (also called a canvas) in a way that generally minimizes the unused space without any overlap, and can ensure that each T×T (e.g., 16×16) block of the 2D grid is associated with a unique patch. The given minimum block size T×T of the 2D grid can specify the minimum distance between different patches placed on the 2D grid. The 2D grid resolution can depend on the input point cloud size and its width W and height H, and the block size T can be sent as metadata to the decoder.

[0093] The auxiliary patch information PI may also include information regarding the association between the blocks and patches relative to the 2D grid.

[0094] In V-PCC, the auxiliary information PI may include block-to-patch index information (BlockToPatch), which determines the association between the blocks of the 2D grid and the patch indices.

[0095] Figure 3a An example of a canvas C is shown, which includes two patches P1 and P2 and their associated 2D bounding boxes B1 and B2. Note that, as Figure 3a shown, the two bounding boxes may overlap within the canvas C. The 2D grid (the partitioning of the canvas) is represented only within the bounding boxes, but the partitioning of the canvas also occurs outside these bounding boxes. The bounding boxes associated with the patches may be divided into T×T blocks, typically T = 16.

[0096] The T×T blocks containing the 2D sample points belonging to a patch may be considered as the occupied blocks in the corresponding occupancy map OM. Then, the blocks of the occupancy map OM may indicate whether the block is occupied, i.e., whether it contains 2D sample points belonging to a patch.

[0097] In Figure 3a , the occupied blocks are represented by white blocks, while the light gray blocks represent the unoccupied blocks. The image generation process (steps 3300 and 3400) stores the geometry and texture of the input point cloud frame IPCF as an image by using at least one generated mapping of the patch to the 2D grid calculated during step 3200.

[0098] In step 3300, the geometry image generator GIG (Geometry image generator) may generate at least one geometry image GI from the input point cloud frame IPCF, the occupancy map OM, and the auxiliary patch information PI. The geometry image generator GIG may utilize the occupancy map information to detect (locate) the occupied blocks, and thus detect (locate) the non-blank pixels in the geometry image GI.

[0099] The geometry image GI may represent the geometry of the input point cloud frame IPCF and may be, for example, a monochrome image of W×H pixels represented in the YV420 - 8 bit format.

[0100] To better handle the case where multiple 3D sample points are projected (mapped) to the same 2D sample point (along the same projection direction (line)) on the projection plane, multiple images called layers may be generated. Thus, different depth values D1, …, Dn may be associated with the 2D sample points of the patch, and then multiple geometry images may be generated.

[0101] In V-PCC, the 2D samples of a fragment are projected onto two layers. The first layer, also known as the near layer, can store depth values D0 associated with 2D samples having a smaller depth, for example. The second layer, called the far layer, can store depth values D1 associated with 2D samples having a greater depth, for example. Alternatively, the second layer can store the difference between depth values D1 and D0. For example, the information stored by the second depth image can be within the interval [0, Δ] corresponding to depth values in the range [D0, D0 + Δ], where Δ is a user-defined parameter describing the surface thickness.

[0102] In this way, the second layer can contain significant contour-like high-frequency features. Thus, it is clear that the second depth image may be difficult to encode and decode using legacy video codecs, and thus the depth values may be poorly reconstructed from the decoded second depth image, which results in poor geometric quality of the reconstructed point cloud frame.

[0103] According to one embodiment, the geometric image generation module GIG can encode and decode (derive) depth values associated with the 2D samples of the first and second layers by using auxiliary fragment information PI.

[0104] In V-PCC, the position of a 3D sample in a fragment with a corresponding connected component can be represented by depth δ(u, v), tangential displacement s(u, v), and bi-tangential displacement r(u, v), as follows:

[0105] δ(u, v) = δ0 + g(u,v)

[0106] s(u, v) = s0 - u0 + u

[0107] r(u, v) = r0 - v0 + v

[0108] where g(u, v) is the luminance component of the geometric image, (u, v) is the pixel associated with the 3D sample on the projection plane, (δ0, s0, r0) is the 3D position of the corresponding fragment to which the 3D sample belongs, and (u0, v0, u1, v1) are the coordinates in the projection plane defining a 2D bounding box that encompasses the projection of the fragment associated with the connected component.

[0109] Thus, the geometric image generation module GIG can encode and decode (derive) the depth values associated with the 2D samples of the layer (the first layer or the second layer or both) as the luminance component g(u, v), which is given by: g(u, v) = δ(u, v) - δ0. Note that this relationship can be used to reconstruct the 3D sample position (δ0, s0, r0) from the reconstructed geometric image g(u, v) with accompanying auxiliary fragment information PI.

[0110] According to one embodiment, a projection mode can be used to indicate whether a first geometric image GI0 can store depth values of 2D sample points of a first layer or a second layer, and whether a second geometric image GI1 can store depth values associated with 2D sample points of the second layer or the first layer.

[0111] For example, when the projection mode is equal to 0, the first geometric image GI0 can store depth values of 2D sample points of the first layer, and the second geometric image GI1 can store depth values associated with 2D sample points of the second layer. Conversely, when the projection mode is equal to 1, the first geometric image GI0 can store depth values of 2D sample points of the second layer, and the second geometric image GI1 can store depth values associated with 2D sample points of the first layer.

[0112] According to one embodiment, a frame projection mode can be used to indicate whether a fixed projection mode is used for all fragments, or whether a variable projection mode is used where a different projection mode can be used for each fragment.

[0113] The projection mode and / or the frame projection mode can be sent as metadata.

[0114] For example, in Section 2.2.1.3.1 of V-PPC, a frame projection mode decision algorithm can be provided.

[0115] According to one embodiment, when the frame projection indicates that a variable projection mode can be used, a fragment projection mode can be used to indicate an appropriate mode for (de)projecting the fragment.

[0116] The fragment projection mode can be sent as metadata and may be information included in the auxiliary fragment information PI.

[0117] For example, in Section 2.2.1.3.2 of V-PCC, a fragment projection mode decision algorithm is provided.

[0118] According to an embodiment of step 3300, the pixel value in the first geometric image (e.g., GI0) corresponding to the 2D sample point (u, v) of the fragment can represent a depth value associated with at least one in-between 3D sample point defined along the projection line corresponding to the 2D sample point (u, v). The in-between 3D sample point exists along the projection line and shares the same coordinates of the 2D sample point (u, v), and the depth value D1 of the 2D sample point (u, v) is encoded and decoded in the second geometric image (e.g., GI1). In addition, the in-between 3D sample point can have a depth value between the depth value D0 and the depth value D1. A specified bit can be associated with each of the in-between 3D sample points, and if the in-between 3D sample point exists, the specified bit is set to 1, otherwise it is set to 0.

[0119] Then, all the specified bits along the projection line can be concatenated to form a codeword, hereinafter referred to as an enhanced-delta-depth (EDD) code. Finally, all the EDD codes can be packed in an image, e.g., in the first geometry image GI1 or the occupancy map OM.

[0120] In step 3400, a texture image generator TIG (texture image generator) can generate at least one texture image TI by deriving the geometry of a reconstructed point cloud frame from at least one of an input point cloud frame IPCF, an occupancy map OM, auxiliary fragment information PI, and a decoded geometry image DGI (decoded geometry image) output from the video decoder VDEC ( Figure 4 from step 4200 therein).

[0121] The texture image TI can represent the texture of the input point cloud frame IPCF and can be, for example, an image of W×H pixels represented in the YV420-8 bit format.

[0122] The texture image generator TG can utilize the occupancy map information to detect (locate) occupied blocks, thereby detecting (locating) non-blank pixels in the texture image.

[0123] The texture image generator TIG can be adapted to generate the texture image TI and associate it with each geometry image / layer DGI.

[0124] According to one embodiment, the texture image generator TIG can encode / decode (store) the texture (attribute) value T0 associated with the 2D samples of the first layer as the pixel values of the first texture image TI0, and encode / decode (store) the texture value T1 associated with the 2D samples of the second layer as the pixel values of the second texture image TI1.

[0125] Alternatively, the texture image generation module TIG can encode / decode (store) the texture value T1 associated with the 2D samples of the second layer as the pixel values of the first texture image TI0, and encode / decode (store) the texture value D0 associated with the 2D samples of the first layer as the pixel values of the second geometry image GI1.

[0126] For example, the color of 3D samples can be obtained as described in Sections 2.2.3, 2.2.4, 2.2.5, 2.2.8, or 2.5 of V-PCC.

[0127] According to one embodiment, a filling process can be applied to the geometry and / or texture image. The filling process can be used to fill the blank spaces between fragments to generate a piecewise smooth image suitable for video compression.

[0128] Sections 2.2.6 and 2.2.7 of V-PCC provide an example of image filling.

[0129] At step 3500, the video codec VENC may encode the generated images / layers TI and GI.

[0130] For example, as detailed in Section 2.2.2, at step 3600, the encoder OMENC may encode the occupancy map as an image. Lossy or lossless encoding may be used.

[0131] According to one embodiment, the video encoder ENC and / or OMENC may be an HEVC-based encoder.

[0132] At step 3700, the encoder PIENC may encode the auxiliary fragment information PI and possible additional metadata, such as the block size T, width W, and height H of the geometry / texture image.

[0133] According to one embodiment, the auxiliary fragment information may be differentially encoded (e.g., as defined in Section 2.4.1 of V-PCC).

[0134] At step 3800, a multiplexer may be applied to the generated outputs of steps 3500, 3600, and 3700, and as a result, these outputs may be multiplexed together to generate a bitstream representing the base layer BL. It should be noted that the metadata information represents a small part of the entire bitstream. A large amount of information is compressed using the video codec.

[0135] Figure 4 A schematic block diagram showing an example of an image-based point cloud decoder 4000 according to at least one current embodiment is shown.

[0136] At step 4100, a demultiplexer DMUX may be applied to demultiplex the encoded information representing the bitstream of the base layer BL.

[0137] At step 4200, the video decoder VDEC may decode the encoded information to derive at least one decoded geometry image DGI and at least one decoded texture image DTI (decoded texture image).

[0138] At step 4300, the decoder OMDEC may decode the encoded information to derive a decoded occupancy map DOM (decoded occupancy map).

[0139] According to one embodiment, the video decoder VDEC and / or OMDEC may be an HEVC-based decoder.

[0140] In step 4400, the decoder PIDEC may decode the encoded information to derive decoded auxiliary patch information DPI.

[0141] Possibly, metadata may also be derived from the bitstream BL.

[0142] In step 4500, the geometry generating module GGM may derive the geometry RG of the reconstructed point cloud frame IRPCF from at least one decoded geometry image DGI, the decoded occupancy map DOM, the decoded auxiliary patch information DPI, and possibly additional metadata.

[0143] The geometry generating module GGM may utilize the decoded occupancy map information DOM to locate non - blank pixels in at least one decoded geometry image DGI. Then, the 3D coordinates of the reconstructed 3D sample points associated with the non - blank pixels may be derived from the coordinates of the non - blank pixels and the values of the reconstructed 2D samples.

[0144] According to one embodiment, the geometry generating module GGM may derive the 3D coordinates of the reconstructed 3D sample points from the coordinates of the non - blank pixels.

[0145] According to one embodiment, the geometry generating module GGM may derive the 3D coordinates of the reconstructed 3D sample points from the coordinates of the non - blank pixels, the values of the non - blank pixels of at least one of the decoded geometry images DGI, the decoded auxiliary patch information, and possibly from additional metadata.

[0146] The use of non - blank pixels is based on the relationship between 2D pixels and 3D sample points. For example, using the projection in V - PCC, the 3D coordinates of the reconstructed 3D sample points can be expressed in terms of depth δ(u, v), tangential displacement s(u, v), and bi - tangential displacement r(u, v) as follows:

[0147] δ(u, v)=δ0 + g(u, v)

[0148] s(u, v)=s0 - u0+u

[0149] r(u, v)=r0 - v0+v

[0150] where g(u, v) is the luminance component of the decoded geometry image DGI, (u, v) is the pixel associated with the reconstructed 3D sample point, (δ0, s0, r0) is the 3D position of the connected component to which the reconstructed 3D sample point belongs, and (u0, v0, u1, v1) are the coordinates in the projection plane defining the 2D bounding box that encompasses the projection of the patch associated with the connected component.

[0151] In step 4600, the texture generation module TGM may derive the texture of the reconstructed point cloud frame IRPCF from the geometry RG and at least one decoded texture image DTI.

[0152] Figure 5 Schematically shows an example syntax of a bitstream representing a base layer BL according to at least one current embodiment.

[0153] The bitstream includes a bitstream header SH (Bitstream Header) and at least one group of frame stream GOFS (Group Of Frame Stream).

[0154] The group of frame stream GOFS includes a header HS, at least one syntax element OMS representing an occupancy map OM, at least one syntax element GVS representing at least one geometry image (or video), at least one syntax element TVS representing at least one texture image (or video), and at least one syntax element PIS representing auxiliary fragment information and other additional metadata.

[0155] In a variant, the group of frame stream GOFS includes at least one frame stream.

[0156] Figure 6 Shows a schematic block diagram illustrating an example of a system in which various aspects and embodiments are implemented.

[0157] System 6000 may be implemented as one or more devices including the various components described below, and is configured to perform one or more aspects described in this document. Examples of devices that may form all or part of system 6000 include personal computers, laptop computers, smartphones, tablets, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, connected vehicles and their associated processing systems, head-mounted display devices (HMDs, see-through glasses), projectors, "caves" (systems including multiple displays), servers, video encoders, video decoders, post-processors that process the output from video decoders, pre-processors that provide input to video encoders, web servers, set-top boxes, and any other devices or other communication devices for processing point clouds, videos, or images. The elements of system 6000 may be implemented singly or in combination in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 6000 may be distributed across multiple ICs and / or discrete components. In various embodiments, system 6000 may be communicatively coupled to other similar systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 6000 may be configured to implement one or more aspects described in this document.

[0158] System 6000 may include at least one processor 6010, which is configured to execute instructions loaded therein for implementing various aspects described in this document, for example. The processor 6010 may include embedded memory, input / output interfaces, and various other circuits known in the art. System 6000 may include at least one memory 6020 (e.g., volatile memory devices and / or non-volatile memory devices). System 6000 may include a storage device 6040, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, the storage device 6040 may include internal storage devices, attached storage devices, and / or network-accessible storage devices.

[0159] System 6000 may include an encoder / decoder module 6030, which is configured to process data to provide encoded data or decoded data, for example, and the encoder / decoder module 6030 may include its own processor and memory. The encoder / decoder module 6030 may represent the (multiple) modules that may be included in a device to perform encoding and / or decoding functions. As is well known, a device may include one or both of an encoding module and a decoding module. Additionally, the encoder / decoder module 6030 may be implemented as a stand-alone element of System 6000 or may be incorporated within the processor 6010 as a combination of hardware and software known to those skilled in the art.

[0160] The program code to be loaded onto the processor 6010 or encoder / decoder 6030 to execute the various aspects described in this document may be stored in the storage device 6040 and subsequently loaded onto the memory 6020 for execution by the processor 6010. According to various embodiments, one or more of the processor 6010, memory 6020, storage device 6040, and encoder / decoder module 6030 may store one or more of various items during the execution of the processes described in this document. Such stored items may include but are not limited to point cloud frames, encoded / decoded geometric / texture video / images, or portions of encoded / decoded geometric / texture video / images, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operation logic.

[0161] In several embodiments, the memory internal to the processor 6010 and / or encoder / decoder module 6030 may be used to store instructions and provide a working memory for the processing that may be performed during encoding or decoding.

[0162] However, in other embodiments, a memory external to the processing device (e.g., the processing device can be the processor 6010 or the encoder / decoder module 6030) can be used for one or more of these functions. The external memory can be the memory 6020 and / or the storage device 6040, such as dynamic volatile memory and / or non-volatile flash memory. In several embodiments, the external non-volatile flash memory can be used to store the operating system of the television. In at least one embodiment, a fast external dynamic volatile memory such as RAM can be used as a working memory for video encoding and decoding operations, such as for MPEG-2 Part 2 (also known as ITU-T Recommendation H.262 and ISO / IEC 13818-2, also known as MPEG-2 video), HEVC (High Efficiency Video Coding), or VVC (Versatile Video Coding).

[0163] As shown in block 6130, the inputs of the elements of the system 6000 can be provided by various input devices. Such input devices include, but are not limited to, (i) an RF section that can receive an RF signal transmitted over the air, for example, by a broadcaster, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.

[0164] In various embodiments, the input devices of block 6130 can have corresponding input processing elements known in the art. For example, the RF section can be associated with the elements required for (i) selecting a desired frequency (also known as the selected signal, or band-limiting the signal frequency band), (ii) down-converting the selected signal, (iii) band-limiting again to a narrower frequency band to select a signal frequency band that can be called a channel in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF section of various embodiments can include one or more elements that perform these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF section can include a tuner that performs various of these functions, including, for example, down-converting the received signal to a lower frequency (e.g., an intermediate frequency or a near-baseband frequency) or to baseband.

[0165] In one set-top box embodiment, the RF section and its associated input processing elements can receive an RF signal transmitted through a wired (e.g., cable) medium. Then, the RF section can perform frequency selection by filtering, down-converting, and filtering again to a desired frequency band.

[0166] Various embodiments rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements performing similar or different functions.

[0167] Adding elements may include inserting elements between existing elements, such as inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF portion may include an antenna.

[0168] In addition, the USB and / or HDMI terminals may include corresponding interface processors for connecting the system 6000 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented in, for example, a separate input processing IC or processor 6010 as needed. Similarly, various aspects of USB or HDMI interface processing may be implemented in a separate interface IC or processor 6010 as needed. The streams of demodulation, error correction, and demultiplexing may be provided to various processing elements (including, for example, processor 6010 and encoder / decoder 6030) operating in conjunction with memory and storage elements to process the data stream required for presentation on the output device.

[0169] The various elements of the system 6000 may be provided within an integrated housing. Within the integrated housing, appropriate connection arrangements 6140 (such as internal buses known in the art, including I2C buses, wiring, and printed circuit boards) may be used to interconnect the various elements and send data between them.

[0170] The system 6000 may include a communication interface 6050 that is capable of enabling communication with other devices via a communication channel 6060. The communication interface 6050 may include, but is not limited to, a transceiver configured to send and receive data via the communication channel 6060. The communication interface 6050 may include, but is not limited to, a modem or a network card, and the communication channel 6060 may be implemented, for example, in a wired and / or wireless medium.

[0171] In various embodiments, data may be streamed to the system 6000 using Wi-Fi such as IEEE 802.11. The Wi-Fi signals of these embodiments may be received via a communication channel 6060 and a communication interface 6050 suitable for Wi-Fi communication. The communication channel 6060 of these embodiments may generally be connected to an access point or a router that provides access to an external network including the Internet to allow streaming applications and other over-the-top communications.

[0172] Other embodiments may provide streamed data to the system 6000 using a set-top box that delivers data via an HDMI connection to the input block 6130.

[0173] Other embodiments may use the RF connection of input block 6130 to provide streaming data to system 6000.

[0174] It should be understood that signaling can be accomplished in various ways. For example, in various embodiments, one or more syntax elements, flags, etc. can be used to signal information to the corresponding decoder.

[0175] System 6000 can provide output signals to various output devices, including display 6100, speaker 6110, and other peripheral devices 6120. In various examples of the embodiments, other peripheral devices 6120 may include one or more of a standalone DVR, disk player, stereo system, lighting system, and other devices that provide functions based on the output of system 3000.

[0176] In various embodiments, signaling such as AV.Link (Audio / Video Link), CEC (Consumer Electronics Control), or other communication protocols that support device-to-device control with or without user intervention can be used to transmit control signals between system 6000 and display 6100, speaker 6110, or other peripheral devices 6120.

[0177] The output devices can be communicatively coupled to system 6000 via dedicated connections through corresponding interfaces 6070, 6080, and 6090.

[0178] Alternatively, the output devices can be connected to system 6000 using communication channel 6060 via communication interface 6050. Display 6100 and speaker 6110 can be integrated with other components of system 6000 in an electronic device (such as a television) in a single unit.

[0179] In various embodiments, display interface 6070 can include a display driver, such as a timing controller (T Con) chip.

[0180] For example, if the RF portion of input 6130 is part of a separate set-top box, display 6100 and speaker 6110 can optionally be separated from one or more other components. In various embodiments where display 6100 and speaker 6110 can be external components, the output signal can be provided via a dedicated output connection, including for example an HDMI port, a USB port, or a COMP output.

[0181] PLR (Point Local Reconstruction) represents point local reconstruction. PLR is a reconstruction method that can be used to generate additional 3D sample points for a point cloud frame. PLR is typically applied immediately after 2D to 3D projection and before any other processing (such as geometric or texture smoothing).

[0182] The PLR takes a 3D sample layer as input and applies a set of filters driven by PLR metadata to generate additional 3D samples along with their geometry and texture.

[0183] The PLR can be local because the PLR metadata can vary for each T×T block of the canvas, and the set of filters can use a small neighborhood to generate additional 3D samples. Note that the same PLR metadata must be used in both the encoder and decoder.

[0184] In general, using a single-layer depth and texture image with PLR provides better BD-rate performance than a two-layer depth and texture image. By projecting onto the depth and texture images, the encoding and decoding of PLR metadata takes fewer bits than traditional 3D sample encoding and decoding, thus reducing the overall bitrate. At the same time, the additional 3D samples compensate for the quality loss due to using fewer layers.

[0185] V-PCC includes an implementation of PLR, where the PLR is determined by RDO (Rate-Distortion Optimization) on the encoder side. The described implementation of PLR defines multiple modes represented as PLRM (Point Local Reconstruction Mode) for reconstructing (generating) at least one 3D sample of a point cloud frame. Each PLRM is determined by a specific value of PLRM metadata, which defines how to use the filters.

[0186] For example, in Section "9.4.4" of V-PCC, multiple PLRMs are determined by four parameters described in Section "7.4.35 Point Local Reconstruction Semantics" of V-PCC. These four parameters are sent as PLRM metadata in the bitstream:

[0187] · point_local_reconstruction_mode_interpolate_flag:

[0188] This parameter equal to 1 indicates that point interpolation is used during the reconstruction method; this parameter equal to 0 indicates that point interpolation is not used during the reconstruction method;

[0189] · point_local_reconstruction_mode_filling_flag

[0190] This parameter equal to 1 indicates that the filling mode is used during the reconstruction method; this parameter equal to 0 indicates that the filling mode is not used during the reconstruction process.

[0191] · point_local_reconstruction_mode_minimum_depth_minus1:

[0192] This parameter specifies the minimum depth value minus 1 used during the reconstruction method.

[0193] ·point_local_reconstruction_mode_neighbour_minus1:

[0194] This parameter specifies the size of the 2D neighbours minus 1 used during the reconstruction method.

[0195] A syntax for encoding and decoding PLR is described in detail in the section "7.3.35 Point Local Reconstruction Syntax" of V-PCC. This syntax describes the PLRM metadata sent for a T×T block (occupying the packed block size) via the "blockToPatch" information. The blockToPatch structure indicates for each T×T pixel block the patch to which it belongs. A size of 16×16 pixels for the blockToPatch block is a typical value used in the V-PCC test model software. The specification (and at least one current embodiment) can support other block sizes for the blockToPatch index and PLRM metadata.

[0196] The PLR metadata can be retrieved from the bitstream as follows: Iterate cyclically in scan order (e.g., raster scan order) over all T×T blocks of the canvas. For each block of the canvas, if the occupancy map OM indicates that the block is not occupied, proceed to the next block. Otherwise, i.e., the block is occupied, retrieve the blockToPatch information from the bitstream to obtain the patch ID of the block, and retrieve the PLR parameters of the occupied block from the bitstream.

[0197] In the current approach to point cloud encoding and decoding (V-PCC), decoding the PLR metadata from the bitstream requires intensive computational resources (CPU, GPU, memory) because four parameters must be decoded for each T×T block.

[0198] According to a general aspect of at least one embodiment, a method is provided for signaling a syntax element representing a point local reconstruction mode, the point local reconstruction mode representing at least one parameter defining a mode for reconstructing at least one point of a point cloud frame.

[0199] Signaling the syntax element instead of four parameters per block can reduce the computational resources on the decoding side as defined in V-PCC because the syntax element can be signaled per patch or per point cloud frame. This also provides flexibility for signaling the PLRM metadata.

[0200] Figure 7 A schematic block diagram showing an example of a method for signaling a point local reconstruction mode according to at least one current embodiment is shown.

[0201] In step 710, the module may add at least one syntax element SE1 representing a point local reconstruction mode (denoted as PLRM) to the bitstream. The PLRM represents at least one parameter defining a mode for reconstructing at least one 3D sample point of a point cloud frame.

[0202] In step 740, the bitstream is sent.

[0203] In step 770, the module may retrieve (read) at least one first syntax element SE1 from the bitstream (the received bitstream), representing at least one PLRM for reconstructing at least one 3D sample point of the point cloud.

[0204] At least one parameter is retrieved from the at least one syntax element SE1, and then at least one 3D sample point of the point cloud is reconstructed using the at least one PLRM.

[0205] According to an embodiment of step 710, the PLRM may be an index value of an entry in a look-up table LUT (look-up-table). Each entry of the LUT defines a relationship between the index value and at least one parameter, and the at least one parameter defines a specific mode for reconstructing at least one point of the point cloud frame.

[0206] Figures 8a - 8b An example of a look-up table LUT according to an embodiment of step 710 is shown.

[0207] The left column indicates different PLRM values, and each PLRM defines a combination of 4 parameters I, F, D1min, and N, which uses a combination of the following tools explained earlier to determine the reconstruction method:

[0208] point_local_reconstruction_mode_interpolate_flag;

[0209] point_local_reconstruction_mode_filling_flag;

[0210] point_local_reconstruction_mode_minimum_depth_minus1; and

[0211] point_local_reconstruction_mode_neighbour_minus1

[0212] Figure 8a A PLRM of 0 indicates the use of "no PLRM metadata", i.e., no PLR metadata is sent.

[0213] FromFigure 8b This PLRM mode is removed from the look-up table LUT. In this case, the "no PLRM metadata" mode is signaled elsewhere.

[0214] Figure 8c A syntax example of the look-up table LUT for signaling an embodiment according to step 710 is shown.

[0215] According to an embodiment of step 710, as Figure 8d shown, a first syntax element SE1 can be signaled for each block, where point_local_reconstruction_mode[p][i] refers to the syntax element SE1 associated with block i of slice p.

[0216] According to an embodiment of step 710, as Figure 8e shown, a first syntax element SE1 can be signaled for each slice, where point_local_reconstruction_mode[p] refers to the first syntax element SE1 associated with slice p.

[0217] In step 720 (optional), the module can add a second syntax element SE2 in the bitstream for each slice, which indicates whether a single first syntax element SE1 is signaled once for all blocks of the slice, or whether a first syntax element SE1 is signaled for each block of the slice.

[0218] When the second syntax element SE2 indicates that a single first syntax element is signaled once for all blocks of the slice, the single first syntax element is used for all blocks of the slice and is thus signaled at the slice level.

[0219] As Figure 9 shown, point_local_reconstruction_patch_level[p] refers to the second syntax element SE2: point_local_reconstruction_patch_level[p]=0 means that a first syntax element point_local_reconstruction_mode[p][i] is signaled for each block of the slice, while point_local_reconstruction_patch_level[p]=1 means that a first syntax element point_local_reconstruction_mode[p] is signaled once for all blocks of slice p.

[0220] According to a variant of step 720, the second syntax element SE2 can be signaled only if the slice contains more than a given number of blocks.

[0221] As Figure 10 shown, BlockCountThreshold refers to a given number of blocks. When the number of blocks in a fragment is greater than or equal to BlockCountThreshold, the second syntax element point_local_reconstruction_patch_level[p] is set to 0 (signaling the first syntax element point_local_reconstruction_mode[p] for each block). Otherwise it is set to 1 (signaling the first syntax element point_local_reconstruction_mode[p] for each fragment).

[0222] In step 730 (optional), the module may add at least one third syntax element SE3 to the bitstream, which represents the default point local reconstruction mode and is also referred to as the preferred point local reconstruction mode.

[0223] As Figures 12a - 12b shown, point_local_reconstruction_preferred_mode refers to the third syntax element SE3.

[0224] Figure 11 An example of a lookup table LUT (first syntax element SE1) and point_local_reconstruction_preferred_mode for signaling according to the embodiment of step 710 is shown.

[0225] According to the embodiment of step 730, the third syntax element SE3 may be signaled for each point cloud frame and / or each fragment.

[0226] As Figure 13 shown, point_local_reconstruction_preferred_mode_patch[p] refers to the third syntax element SE3 signaled for each fragment, and point_local_reconstruction_preferred_mode refers to the third syntax element SE3 signaled for each point cloud frame.

[0227] As Figure 15 shown, point_local_reconstruction_preferred_mode refers to the third syntax element SE3 signaled for each point cloud frame.

[0228] According to an embodiment of step 740, the module may add to the bitstream the use of a fourth syntax element SE4 and a third syntax element SE3 (default point local reconstruction mode signaled for each fragment), where the fourth syntax element SE4 is used to signal the granularity of a first syntax element SE1 (PLRM for each block or each fragment).

[0229] As Figures 12a - 12b shown, point_local_reconstruction_patch_level[p] refers to the fourth syntax element. When point_local_reconstruction_patch_level[p] = 0, then for each block i the first syntax element SE1 (point_local_reconstruction_mode[p][i]) is signaled; when point_local_reconstruction_patch_level[p] = 1, then for each fragment p the first syntax element SE1 (point_local_reconstruction_mode[p]) is signaled, otherwise, for all blocks of fragment p the third syntax element SE3 (point_local_reconstruction_preferred_mode) is signaled.

[0230] In Figure 12a it, point_local_reconstruction_preferred_mode is signaled for each point cloud frame.

[0231] In Figure 12b it, point_local_reconstruction_preferred_mode_patch[p] is signaled for each fragment.

[0232] Figure 13 Shows an example of a Figure 12b variant when the second syntax element SE2 is signaled only when a fragment contains more than a given number of blocks BlockCountThreshold.

[0233] Figure 14 Shows an example of the variant when the first syntax element point_local_reconstruction_mode[p][i]( Figure 8d ) is signaled for each block: point_local_reconstruction_mode_delta[p][i] instead of point_local_reconstruction_mode[p][i].[[]END]

[0234] The variations can also be applied to other examples for signaling a first syntax element SE1, e.g., as Figure 8e , Figure 9 , Figure 10 , Figures 12a - 12b and Figure 13 shown.

[0235] The variations can also be applied to differentially encoding a second syntax element SE2, e.g., as Figure 9 , Figure 10 , Figures 12a - 12b and Figure 13 shown.

[0236] In Figures 1 - 15 , various methods are described herein, and each method includes one or more steps or actions for implementing the described method. Unless a particular order of steps or actions is required for the correct operation of the method, the order and / or use of specific steps and / or actions can be modified or combined.

[0237] Some examples are described with reference to block diagrams and operational flowcharts. Each block represents a circuit element, a module, or a code portion, and the code portion includes one or more executable instructions for implementing a particular logical function(s). It should also be noted that in other implementations, the function(s) mentioned in the blocks may not occur in the order indicated. For example, depending on the functions involved, two consecutively shown blocks may actually be executed substantially simultaneously, or the blocks may sometimes be executed in the reverse order.

[0238] The implementations and aspects described herein can be implemented, for example, as a method or process, an apparatus, a computer program, a data stream, a bit stream, or a signal. Even if discussed only in the context of a single form of implementation (e.g., only as a method), the implementation of the discussed features can be implemented in other forms (e.g., an apparatus or a computer program).

[0239] These methods can be implemented, for example, in a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device.

[0240] Additionally, these methods can be implemented by instructions executed by a processor, and such instructions (and / or data values generated by the implementation) can be stored on a computer-readable storage medium. The computer-readable storage medium can take the form of a computer-readable program product embodied in one or more computer-readable media, and having thereon computer-readable program code executable by a computer. Given the inherent ability to store information therein and the inherent ability to provide retrieval of information therefrom, the computer-readable storage medium used herein can be considered a non-transitory storage medium. The computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. It should be understood that although more specific examples of computer-readable storage media to which this embodiment can be applied are provided below, they are merely illustrative and not an exhaustive list readily understandable by those of ordinary skill in the art: portable computer floppy disks; hard disks; read-only memory (ROM); erasable programmable read-only memory (EPROM or flash memory); portable compact disc read-only memory (CD-ROM); optical storage devices; magnetic storage devices; or any suitable combination of the foregoing.

[0241] The instructions can form an application program tangibly embodied on a processor-readable medium.

[0242] The instructions can be, for example, in hardware, firmware, software, or any combination thereof. The instructions can be found, for example, in an operating system, a separate application program, or a combination of both. Thus, a processor can be characterized as, for example, a device configured to execute a process and a device including a processor-readable medium (e.g., a storage device) having instructions for executing the process. Additionally, in addition to or instead of instructions, the processor-readable medium can store data values generated by the implementation.

[0243] The apparatus can be implemented using, for example, appropriate hardware, software, and firmware. Examples of such apparatus include personal computers, laptop computers, smartphones, tablets, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, head-mounted display devices (HMDs, see-through glasses), projectors, "caves" (systems including multiple displays), servers, video encoders, video decoders, post-processors that process the output from a video decoder, pre-processors that provide input to a video encoder, web servers, set-top boxes, and any other device or other communication device for processing point clouds, videos, or images. It should be clear that the device can be mobile and can even be installed in a moving vehicle.

[0244] Computer software can be implemented by the processor 6010 or by hardware or by a combination of hardware and software. As a non-limiting example, the embodiments can also be implemented by one or more integrated circuits. As a non-limiting example, the memory 6020 can be of any type suitable for the technical environment and can be implemented using any appropriate data storage technology (e.g., optical storage devices, magnetic storage devices, semiconductor-based storage devices, fixed memory, and removable memory). As a non-limiting example, the processor 6010 can be of any type suitable for the technical environment and can encompass one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture).

[0245] It will be apparent to those of ordinary skill in the art that the implementations can generate various signals that are formatted to carry information such as can be stored or transmitted. The information can include, for example, instructions for performing a method or data generated by one of the described implementations. For example, the signal can be formatted to carry the bit stream of the described embodiments. Such a signal can be formatted as, for example, an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or a baseband signal. The formatting can include, for example, encoding the data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. As is well known, signals can be transmitted over a variety of different wired or wireless links. Signals can be stored on a processor-readable medium.

[0246] The terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting. As used herein, unless the context clearly dictates otherwise, the singular forms “a,” “an,” and “the” may also include the plural forms. It will be further understood that when the terms “comprises / comprising” and / or “includes / including” are used in this specification, they specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof. Further, when an element is referred to as being “responsive” or “connected” to another element, it can be directly responsive or connected to the other element, or intervening elements may be present. In contrast, when an element is referred to as being “directly responsive” or “directly connected” to other elements, no intervening elements are present.

[0247] It should be understood that the use of any symbol / term such as “ / ”, “and / or” and “at least one”, e.g., in the cases of “A / B”, “A and / or B” and “at least one of A and B”, may be intended to cover only the first-listed option (A), or only the second-listed option (B), or both options (A and B). As another example, in the cases of “A, B and / or C” and “at least one of A, B and C”, such wording is intended to cover only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or only the first and second-listed options (A and B), or only the first and third-listed options (A and C), or only the second and third-listed options (B and C), or all three options (A and B and C). As will be clear to those of ordinary skill in the art and related fields, this can be extended to any number of items listed.

[0248] A variety of numerical values may be used in this application. The specific values may be for illustrative purposes, and the aspects described are not limited to these specific values.

[0249] It should be understood that although terms such as first, second, etc. may be used to describe various elements, these elements are not limited by these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the teachings of this application. There is no implied order between the first element and the second element.

[0250] The mention of “an embodiment” or “embodiments” or “an implementation” or “implementations” and other variants thereof are often used to convey that a particular feature, structure, characteristic, etc. (described in connection with the embodiment / implementation) is included in at least one embodiment / implementation. Thus, the appearance of the phrase “in an embodiment” or “in embodiments” or “in an implementation” or “in implementations” and any other variants that appear throughout this application do not necessarily all refer to the same embodiment.

[0251] Similarly, the mention herein of “according to an embodiment / example / implementation” or “in an embodiment / example / implementation” and other variants thereof are often used to convey that a particular feature, structure, or characteristic (described in connection with the embodiment / example / implementation) may be included in at least one embodiment / example / implementation. Thus, the expressions “according to an embodiment / example / implementation” or “in an embodiment / example / implementation” that appear in different places in the specification do not necessarily all refer to the same embodiment / example / implementation, nor are they necessarily separate or alternative embodiments / examples / implementations that are mutually exclusive of other embodiments / examples / implementations.

[0252] Reference numerals appearing in the claims are for illustration purposes only and have no limiting effect on the scope of the claims.The present embodiments / examples and variations may be used in any combination or sub-combination even if not explicitly described.

[0253] When the figure is presented as a flow chart, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when the figure is presented as a block diagram, it should be understood that it also provides a flow chart of the corresponding method / process.

[0254] Although some diagrams include arrows on communication paths to illustrate a primary direction of communication, it should be understood that communication can occur in the opposite direction of the illustrated arrows.

[0255] Various implementations involve decoding. "Decoding" as used in this application may encompass, for example, all or part of a process performed on a received point cloud frame (which may include a received bitstream encoding one or more point cloud frames) to produce a final output suitable for display or further processing in a reconstructed point cloud domain. In various embodiments, such a process includes one or more processes typically performed by an image-based decoder.

[0256] As a further example, in one embodiment, "decoding" may refer only to entropy decoding, in another embodiment, "decoding" may refer only to differential decoding, and in another embodiment, "decoding" may refer to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" refers specifically to a subset of operations or generally to a broader decoding process will be clear based on the context of the specific description and is considered to be well understood by those skilled in the art.

[0257] Various implementations involve encoding. In a manner similar to the discussion above about "decoding", "encoding" as used in this application can encompass, for example, all or part of the processes performed on an input point cloud frame in order to produce an encoded bitstream. In various embodiments, such processes include one or more processes typically performed by an image-based encoder.

[0258] As a further example, in one embodiment, "encoding" may refer only to entropy encoding, in another embodiment, "encoding" may refer only to differential encoding, and in another embodiment, "encoding" may refer to a combination of differential encoding and entropy encoding. Whether the phrase "encoding process" is intended to refer specifically to a subset of operations or generally to a broader encoding process will be clear based on the context of the particular description and is considered to be well understood by those skilled in the art.

[0259] Note that the syntax elements used in this document, e.g., the flag point_local_reconstruction_mode_present_flag, are descriptive terms. Thus, they do not preclude the use of other syntax element names.

[0260] Various embodiments relate to rate-distortion optimization. In particular, during the encoding process, given the limitations of computational complexity, a balance or trade-off between rate and distortion is typically considered. Rate-distortion optimization can generally be formulated as minimizing a rate-distortion function, which is a weighted sum of the rate and the distortion. There are different ways to solve the rate-distortion optimization problem. For example, these methods can be based on an extensive test of all encoding options, including all considered modes or codec parameter values, and a complete evaluation of their encoding and decoding costs and the associated distortion of the encoded and decoded and reconstructed signals. Faster methods can also be used to save encoding complexity, especially by calculating an approximate distortion based on the predicted or prediction residual signal rather than the reconstructed signal. A hybrid of these two methods can also be used, e.g., using approximate distortion only for some of the possible encoding options and complete distortion for other encoding options. Other methods only evaluate a subset of the possible encoding options. More generally, many methods employ any of a variety of techniques to perform the optimization, but the optimization does not necessarily involve a complete evaluation of the encoding and decoding costs and the associated distortion.

[0261] In addition, this application may refer to "determining" various information. Determining information can include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from a memory.

[0262] In addition, this application may refer to "accessing" various information. Accessing information can include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0263] In addition, this application may refer to "receiving" various information. Like "accessing", receiving is a broad term. Receiving information can include, for example, one or more of accessing information or retrieving information (e.g., from a memory). Further, during operations such as storing information, processing information, sending information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information, "receiving" is typically involved in one way or another.

[0264] Furthermore, as used herein, the word "signaling" particularly refers to something directed to a corresponding decoder. For example, in some embodiments, an encoder signals a particular syntax element SE and / or PLR metadata. Thus, in an embodiment, the same parameter (PLR metadata) can be used on the encoder side and the decoder side. Therefore, for example, an encoder can send (explicit signaling) a particular parameter to a decoder such that the decoder can use the same particular parameter. Conversely, if a decoder already has a particular parameter and other parameters, signaling (implicit signaling) can be used to simply allow the decoder to know and select the particular parameter without sending it. By avoiding the sending of any actual functionality, bit savings are achieved in various embodiments. It should be understood that signaling can be done in various ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder. Although the foregoing relates to the verb form of the word "signaling", the word "signaling" can also be used as a noun herein.

[0265] Numerous implementations have been described. However, it should be understood that various modifications can be made. For example, elements of different implementations can be combined, supplemented, modified, or removed to yield other implementations. In addition, those of ordinary skill in the art will understand that other structures and processes can be substituted for those disclosed, and the resulting implementations will perform at least substantially the same (multiple) functions in at least substantially the same (multiple) ways to achieve at least substantially the same (multiple) results as the disclosed implementations. Accordingly, this application contemplates these and other embodiments.

Claims

1. A method, comprising: generating at least one additional 3D sample for a reconstructed point cloud frame from a layer storing depth values of 3D samples of the reconstructed point cloud frame, using a point-local reconstruction mode indicated by a first syntax element, the point-local reconstruction mode including at least one parameter defining the point-local reconstruction mode; signaling, in a bitstream, a second syntax element for each tile of the point cloud frame, the second syntax element indicating whether the first syntax element is signaled once per tile and the same first syntax element is used for all blocks of the tile, or whether the first syntax element is signaled for each block of the tile; signaling the first syntax element in the bitstream depending on the second syntax element.

2. The method according to claim 1, wherein, The first syntax element is an index value of an entry of a lookup table, the entries of the lookup table defining an association between the index value and a specific value of at least one parameter of the point-local reconstruction mode.

3. The method according to claim 1, wherein, A tile is a collection of at least one block of at least one 2D sample, the 2D samples representing orthogonal projections of at least one 3D sample of the point cloud frame onto a projection plane.

4. The method according to claim 1, wherein Signaling the second syntax element in response to determining that a tile contains more than a given number of blocks.

5. A method, comprising: generating at least one additional 3D sample for a reconstructed point cloud frame from a layer storing depth values of 3D samples of the reconstructed point cloud frame, using a point-local reconstruction mode indicated by a first syntax element, the point-local reconstruction mode including at least one parameter defining the point-local reconstruction mode, decoding, from a bitstream, a second syntax element for each tile of the point cloud frame, the second syntax element indicating whether the first syntax element is decoded once per tile and the same first syntax element is used for all blocks of the tile, or whether the first syntax element is decoded for each block of the tile; decoding the first syntax element from the bitstream depending on the second syntax element.

6. The method according to claim 5, wherein, Decoding the second syntax element in response to determining that a tile contains more than a given number of blocks.

7. The method according to claim 5, wherein, The first syntax element is an index value of an entry of a lookup table, the entries of the lookup table defining an association between the index value and a specific value of at least one parameter of the point-local reconstruction mode.

8. The method according to claim 5, wherein the point-local reconstruction mode includes four parameters, the values of the four parameters defining a mode for reconstructing at least one point of the point cloud frame, wherein the four parameters include a first flag indicating whether point interpolation is used to reconstruct the at least one point, a second flag indicating whether a padding mode is used to reconstruct the at least one point, a first value indicating a minimum depth value minus 1 used to reconstruct the at least one point, and a second value indicating a size of 2D neighbors minus 1 used to reconstruct the at least one point.

9. The method according to claim 7, wherein the lookup table is signaled in the bitstream.

10. An apparatus comprising one or more processors configured to: Generate at least one additional 3D sample for a reconstructed point cloud frame from a layer storing depth values of 3D samples of the reconstructed point cloud frame, using a point-local reconstruction mode indicated by a first syntax element, the point-local reconstruction mode including at least one parameter defining the point-local reconstruction mode. Signaling in a bitstream a second syntax element for each tile of the point cloud frame, the second syntax element indicating whether the first syntax element is signaled once per tile and the same first syntax element is used for all blocks of the tile, or whether the first syntax element is signaled for each block of the tile. Signaling the first syntax element in the bitstream depending on the second syntax element.

11. The apparatus according to claim 10, wherein, A tile is a collection of at least one block of at least one 2D sample, the 2D sample representing an orthogonal projection of at least one 3D sample of the point cloud frame onto a projection plane.

12. The apparatus according to claim 10, wherein, Signaling the second syntax element in response to determining that a tile contains more than a given number of blocks.

13. An apparatus comprising one or more processors configured to: Generate at least one additional 3D sample for a reconstructed point cloud frame from a layer storing depth values of 3D samples of the reconstructed point cloud frame, using a point-local reconstruction mode indicated by a first syntax element, the point-local reconstruction mode including at least one parameter defining the point-local reconstruction mode, decoding a second syntax element for each tile of the point cloud frame from the bitstream, the second syntax element indicating whether the first syntax element is decoded once per tile and the same first syntax element is used for all blocks of the tile, or whether the first syntax element is decoded for each block of the tile. Decoding the first syntax element from the bitstream depending on the second syntax element.

14. The apparatus according to claim 13, wherein, Decoding the second syntax element in response to determining that a tile contains more than a given number of blocks.

15. The device according to claim 13, wherein, The first syntax element is an index value of an entry in a lookup table, the entries of the lookup table defining an association between the index value and a specific value of at least one parameter of the point-local reconstruction mode.

16. The apparatus according to claim 13, wherein the point-local reconstruction mode includes four parameters, the values of the four parameters defining a mode for reconstructing at least one point of the point cloud frame, wherein the four parameters include a first flag indicating whether point interpolation is used to reconstruct the at least one point, a second flag indicating whether a padding mode is used to reconstruct the at least one point, a first value indicating a minimum depth value minus 1 used to reconstruct the at least one point, and a second value indicating a size of 2D neighbors minus 1 used to reconstruct the at least one point.

17. The apparatus according to claim 15, wherein the lookup table is signaled in the bitstream.

18. A non-transitory computer-readable medium comprising a bitstream, the bitstream including: Encoding point cloud frames data, including a layer storing depth values of 3D samples of the point cloud frames for reconstructing the point cloud frames, wherein a first syntax element indicates a point local reconstruction mode for generating at least one additional 3D sample for the reconstruction of the point cloud frames, and the point local reconstruction mode includes at least one parameter defining the point local reconstruction mode. - A second syntax element for each tile of the point cloud frame, the second syntax element indicating whether the first syntax element is signaled once per tile and the same first syntax element is used for all blocks of the tile, or whether the first syntax element is signaled for each block of the tile. Depending on the second syntax element signaling the first syntax element.

19. A non-transitory computer-readable medium comprising instructions for causing one or more processors to perform the method according to claim 1 or 5.