Method and apparatus for encoding point cloud

By projecting time-domain continuous pictures and traditional encoding processing of point clouds representing 3D objects, the problems of high resource consumption and poor compression performance in the prior art are solved, and more efficient point cloud encoding and decoding are achieved.

CN111095362BActive Publication Date: 2025-05-27INTERDIGITAL VC HOLDINGS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201880057536.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-07-13
Filing Date
2018-07-09
Publication Date
2025-05-27
Estimated Expiration
2038-07-09

AI Technical Summary

Technical Problem

The prior art faces problems such as occlusion, redundancy and time domain stability when encoding and decoding point clouds representing 3D objects, resulting in high resource consumption and poor compression performance.

Method used

By obtaining a time-domain continuous collection of pictures and projecting them onto multiple cube surfaces, video information of texture and depth is generated, and then encoded using a traditional encoder.

Benefits of technology

Improve the encoding efficiency of point clouds, reduce the size of the bitstream, while maintaining a good quality of experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111095362B_ABST
    Figure CN111095362B_ABST
Patent Text Reader

Abstract

Methods and apparatus for encoding a point cloud representing a three-dimensional (3D) object. A set or sets of temporally consecutive pictures are obtained. Each picture in the set or sets includes a first set of images that are spatially arranged in the same way in each picture in the set or sets. A second set of projections is associated with the set or sets, and a unique projection is associated with each image in such a way that the same projection is associated with only a single one of the images and all projections are associated with an image. First information representing the projections is encoded. The point cloud may then be encoded based on the obtained pictures. Corresponding methods and apparatus for decoding a bitstream comprising data representing the encoded pictures are also described.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of encoding and decoding point clouds representing the geometric structure and texture of 3D objects. Specifically, but not exclusively, the technical field of the present disclosure relates to the encoding / decoding of 3D image data using texture and depth projection schemes. Background Art

[0002] This section is intended to introduce the reader to aspects of the technology that may be related to various aspects of the present disclosure described and / or claimed below. It is believed that this discussion will help to provide background information to the reader to facilitate a better understanding of the various aspects of the invention. Accordingly, these statements should be understood in this light and should not be admitted as prior art.

[0003] A point cloud is a collection of points, typically intended to represent the outer surface of a 3D object and also more complex geometric structures (such as hair, fur) that may not be effectively represented in other data formats (such as meshes). Each point of a point cloud is typically defined by a 3D spatial position (X, Y, and Z coordinates in 3D space) and can be defined by other relevant attributes such as color (e.g., represented in the RGB or YUV color space), transparency, reflectivity, two-component normal vector, etc.

[0004] One can consider a colored point cloud, i.e., a collection of 6-component points (X, Y, Z, R, G, B) or equivalently (X, Y, Z, Y, U, V), where (X, Y, Z) defines the spatial position of the point in 3D space and (R, G, B) or (Y, U, V) defines the color of the point.

[0005] A point cloud can be static or dynamic, depending on whether the cloud evolves over time. It should be noted that in the case of a dynamic point cloud, the number of points is not constant; rather, the number of points evolves over time. Thus, a dynamic point cloud is a chronological list of collections of points.

[0006] In fact, point clouds can be used for various purposes such as cultural heritage / buildings, where objects such as statues or buildings are 3D scanned in order to share the spatial structure of the object without sending or accessing it. At the same time, it is also a way to ensure the preservation of knowledge of the object in case it may be destroyed (e.g., a temple is destroyed by an earthquake). Such colored point clouds are typically static and huge.

[0007] Another use case is in topography and cartography, where by using 3D representations, maps are not limited to a plane and can also include relief.

[0008] The automotive industry and autonomous vehicles are also areas where point clouds can be used. Autonomous vehicles should be able to "detect" their environment to make safe driving decisions based on the actual situation in their immediate vicinity. Typical sensors generate dynamic point clouds that are used by the decision-making engine. These point clouds are not intended for human viewing and are usually small, not necessarily colored, dynamic, and have a high capture frequency. They may have other properties, such as reflectivity, which is useful information about the material of the sensed object and may contribute to decision-making.

[0009] Virtual reality and immersive worlds have become a hot topic recently and are foreseen by many as the future of 2D flat video. The basic idea is to immerse the viewer in the environment around him, in stark contrast to a standard TV where he can only see the virtual world in front of him. There are some levels of immersion, which depend on the degree of freedom of the viewer in the environment. Colored point clouds are a good candidate format for distributing virtual reality (or VR) worlds. They can be static or dynamic and are usually of an average size (that is, no more than a few million points at a time).

[0010] Point cloud compression will only be successful in the storage / transmission of 3D objects for immersive worlds when the size of the bitstream is small enough to allow for practical storage / transmission to the end user.

[0011] It is also crucial to be able to distribute dynamic point clouds to end users with a reasonable bandwidth consumption while maintaining an acceptable (or preferably very good) quality of experience.

[0012] A well-known method is to project a colored point cloud representing the outer surface of a 3D object onto the surface of a cube enclosing the 3D object to obtain a video about texture and depth, and use a traditional encoder such as 3D-HEVC (an extension of HEVC, whose specification can be found in Annex G and I of ITU website T Recommendation H series h265, http: / / www.itu.int / rec / T-rec-H.265-201612-I / en) to encode the texture and depth videos. Some projections may be required to handle occlusion. To obtain high compression efficiency, temporal inter-frame prediction of texture (or color) and / or depth from other already encoded pictures can be implemented.

[0013] The compression performance is close to that of video compression for each projected point, but when considering dynamic point clouds, some content may be more complex due to occlusion, redundancy, and temporal stability.

[0014] Regarding occlusion, it is almost impossible to obtain the complete geometry of a complex topology without using many projections. Therefore, the resources (computing power, storage memory) required for encoding / decoding all these projections are usually too high.

[0015] In terms of redundancy, if a point is seen twice on two different projections, its coding efficiency is divided by two, and this can easily get worse if a large number of projections are used. Non-overlapping patches can be used before projection, but this makes the partition boundaries of the projections non-smooth and thus difficult to code, and this seriously affects the coding performance.

[0016] In terms of temporal stability, non-overlapping patches before projection can be optimized for an object at a given time. However, when the object moves, the patch boundaries also move, and the temporal stability of the regions that are difficult to code (equal to the boundaries) is lost. In fact, the compression performance is not much better than full intra-frame coding because, in this case, temporal inter-frame prediction is inefficient.

[0017] Therefore, there is a trade-off between seeing a point at most once but having a poorly compressed projection image (bad boundaries) and having a well-compressed projection image but seeing some points multiple times, resulting in encoding more points in the projection image than actually belong to the model. SUMMARY OF THE INVENTION

[0018] References in the specification to "one embodiment", "an embodiment", "example embodiment", "specific embodiment" indicate that the described embodiment may include a particular feature, structure, or characteristic, but each embodiment does not necessarily include the particular feature, structure, or characteristic. Moreover, these phrases do not necessarily refer to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is understood that such feature, structure, or characteristic is within the knowledge of those skilled in the art in connection with other embodiments (whether or not explicitly described).

[0019] The present disclosure relates to a method of encoding a point cloud representing a three-dimensional object, the method comprising:

[0020] - obtaining at least one set of temporally consecutive pictures of the point cloud,

[0021] each picture in the at least one set including a first set of images, the images in the first set having the same position in each picture in the at least one set,

[0022] a second set of projections associated with the at least one set, different projections of the second set being associated with each image of the first set;

[0023] - encoding first information representing the projections in the second set; and

[0024] - encoding the point cloud based on the obtained pictures.

[0025] The present disclosure also relates to an apparatus adapted to encode a point cloud representing a three-dimensional object, the apparatus including a memory associated with a processor, the processor being configured to:

[0026] - obtain at least one set of temporally consecutive pictures of the point cloud,

[0027] each picture in the at least one set including a first set of images, the images in the first set having the same position in each picture in the at least one set,

[0028] a second set of projections associated with the at least one set, different projections of the second set being associated with each image of the first set;

[0029] - encode first information representing the projections in the second set; and

[0030] - encode the point cloud according to the obtained pictures.

[0031] The present disclosure also relates to an apparatus adapted to encode a point cloud representing a three-dimensional object, the apparatus including:

[0032] - means for obtaining at least one set of temporally consecutive pictures of the point cloud,

[0033] each picture in the at least one set including a first set of images, the images in the first set having the same position in each picture in the at least one set,

[0034] a second set of projections associated with the at least one set, different projections of the second set being associated with each image of the first set;

[0035] - means for encoding first information representing the projections in the second set; and

[0036] - means for encoding the point cloud according to the obtained pictures.

[0037] The present disclosure relates to a method for decoding a point cloud representing a three-dimensional object from at least one bitstream, the method including:

[0038] - obtaining the at least one bitstream, the at least one bitstream including encoded data of at least one set of temporally consecutive pictures of the point cloud,

[0039] each picture in the at least one set including a first set of images, the images in the first set having the same position in each picture in the at least one set,

[0040] - decoding the at least one set of temporally consecutive pictures from the at least one bitstream;

[0041] - Decode first information representing a second set of projections, the second set of projections being associated with the at least one set of temporally consecutive pictures, different projections in the second set being associated with each image in the first set; and

[0042] - Decode the point cloud based on the decoded pictures.

[0043] The present disclosure also relates to an apparatus adapted to decode a point cloud representing a three-dimensional object, the apparatus including a memory associated with a processor, the processor being configured to:

[0044] - Obtain the at least one bitstream, the at least one bitstream including encoded data of at least one set of temporally consecutive pictures of the point cloud,

[0045] each picture in the at least one set including a first set of images, the images in the first set having the same position in each picture in the at least one set,

[0046] - Decode the at least one set of temporally consecutive pictures from the at least one bitstream;

[0047] - Decode first information representing a second set of projections, the second set of projections being associated with the at least one set of temporally consecutive pictures, different projections in the second set being associated with each image in the first set; and

[0048] - Decode the point cloud based on the decoded pictures.

[0049] The present disclosure also relates to an apparatus adapted to decode a point cloud representing a three-dimensional object, the apparatus including:

[0050] - Means for obtaining the at least one bitstream, the at least one bitstream including encoded data of at least one set of temporally consecutive pictures of the point cloud,

[0051] each picture in the at least one set including a first set of images, the images in the first set having the same position in each picture in the at least one set,

[0052] - Means for decoding the at least one set of temporally consecutive pictures from the at least one bitstream;

[0053] - Means for decoding first information representing a second set of projections, the second set of projections being associated with the at least one set of temporally consecutive pictures, different projections in the second set being associated with each image in the first set; and

[0054] - Means for decoding the point cloud based on the decoded pictures.

[0055] The present disclosure also relates to a signal carrying data representing at least one picture of at least one set of temporally consecutive pictures of a point cloud, each picture in the at least one set including a set of images having the same position in each picture of the at least one set, the signal further carrying first information representing a set of projections associated with the at least one set of temporally consecutive pictures, different projections being associated with each image in the set of images.

[0056] The present disclosure also relates to a computer program product including instructions of program code which, when the program is executed on a computer, are executed by at least one processor to perform the above encoding and / or decoding method.

[0057] The present disclosure also relates to a (non-transitory) processor-readable medium storing instructions for causing a processor to perform at least the above encoding and / or decoding method. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] The present disclosure will be better understood and other specific features and advantages will become apparent after reading the following description with reference to the drawings, in which:

[0059] Figure 1 shows an example of an encoding / decoding scheme for a point cloud according to an embodiment;

[0060] Figure 2 shows a first example of a process for encoding a point cloud according to a non-limiting embodiment for a Figure 1 scheme;

[0061] Figure 3 shows a second example of a process for encoding a point cloud according to a non-limiting embodiment for a Figure 1 scheme;

[0062] Figure 4A and 4B each show an example of a picture representing a point cloud according to a non-limiting embodiment of Figure 1 ;

[0063] Figure 5 shows an example of a projection for obtaining a picture of Figure 4A according to a non-limiting embodiment;

[0064] Figure 6 shows an example of a projection for obtaining a picture of Figure 4B according to a non-limiting embodiment;

[0065] Figure 7 shows a set of pictures representing a point cloud of Figure 1 according to a non-limiting embodiment;

[0066] Figure 8 illustrates an example of a process for obtaining an image such as Figure 4A or 4B;

[0067] Figure 9 illustrates an example of a process for generating a reference image associated with an image such as Figure 4A ;

[0068] Figure 10 illustrates a third example of a process for encoding a point cloud of a Figure 1 scheme;

[0069] Figure 11 illustrates an example of a process implemented in the Figure 10 encoding;

[0070] Figure 12 illustrates a first example of a process for decoding a bitstream to obtain a Figure 1 decoded point cloud;

[0071] Figure 13 illustrates a second example of a process for decoding a bitstream to obtain a Figure 1 decoded point cloud;

[0072] Figure 14 illustrates a third example of a process for decoding a bitstream to obtain a Figure 1 decoded point cloud;

[0073] Figure 15 illustrates an example of a process implemented in the Figure 14 decoding;

[0074] Figure 16 illustrates an example of the architecture of an apparatus for implementing at least a portion of a Figure 1 coding / decoding scheme;

[0075] Figure 17 illustrates two remote devices (such as Figure 16 devices) communicating via a communication network;

[0076] Figure 18 illustrates an example of the syntax of a signal for transmitting a bitstream obtained by a Figure 1 scheme; and

[0077] Figure 19 illustrates an example of a process for representing Figure 1An example of an octree of at least a portion of a point cloud. Detailed Description

[0078] The subject matter is now described with reference to the accompanying drawings, in which like reference numerals are used throughout to refer to like elements. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the subject matter. It will be evident, however, that the subject matter embodiments may be practiced without these specific details.

[0079] According to non-limiting embodiments, methods and apparatuses for encoding a point cloud in a bitstream are disclosed. Methods and apparatuses for decoding an encoded point cloud from the bitstream are also disclosed. A syntax of a signal including the bitstream is also disclosed.

[0080] Hereinafter, an image includes an array of one or some samples (pixel values) of a specific image / video format, which array specifies all information related to the pixel values of the image (or video), and all information used by a display and / or any other device to visualize and / or decode, for example, the image (or video). An image includes at least one component, which is shaped as a first sample array, typically a luma (or luminance) component, and (possibly) at least one other component, which is shaped as at least one other sample array, typically a color component. Alternatively, equivalently, the same information may also be represented by a set of color sample arrays, such as a conventional three-color RGB representation.

[0081] Hereinafter, a picture may be regarded as an image (i.e., an array of samples), or as a set of images.

[0082] A pixel value is represented by a vector of nv values, where nv is the number of components. Each value of the vector is represented by the number of bits defining the maximum dynamic range of the pixel value.

[0083] Embodiments of a method (and a configured apparatus) for encoding a point cloud representing a three-dimensional (3D) object are described. One or more sets of temporally consecutive pictures are obtained (e.g., generated or received), one or more sets forming, for example, an intra period. Each picture in the one or more sets includes a first set of images that are arranged in the same way spatially (i.e., have the same positions) in each picture of the one or more sets. A second set of projections is associated with (or appended to) the one or more sets, a unique projection being associated with each image in such a way that the same projection is associated with only a single image, and all projections are associated with a plurality of images. First information representing the projections is encoded, the first information including, for example, information about the mapping between the images and the projections, information about the parameters of the projections, and / or information about the positions of the images (and thus the projections) in the pictures. Then, the point cloud may be encoded based on the obtained pictures.

[0084] One or more specific embodiments of a corresponding method (and a configured device) for decoding a point cloud from a bitstream comprising encoded data representing a picture of the point cloud will also be described.

[0085] Using pictures with the same image arrangement enables providing temporal consistency between pictures, which can improve temporal prediction between pictures and thus improve coding efficiency.

[0086] Although described with reference to a single picture of a point cloud, this embodiment also applies in the same way to a sequence of pictures, and for each picture of at least a part of the picture sequence, a reference picture can be obtained.

[0087] Figure 1 A diagram schematically showing an encoding / decoding scheme of a point cloud 111 according to a specific and non-limiting embodiment.

[0088] Through an encoding process 11 implemented in a module M11, the point cloud 111 is encoded into encoded data in the form of a bitstream 112. This bitstream is sent to a module M12, and the module M12 implements a decoding process 12 to decode the encoded data to obtain a decoded point cloud 113. The modules M11 and M12 can be hardware, software, or a combination of hardware and software.

[0089] The point cloud 111 corresponds to a set of a large number of points representing an object (such as the outer surface or external shape of an object). The point cloud can be regarded as a vector-based structure, where each point has its coordinates (for example, three-dimensional coordinates XYZ, or depth / distance from a given viewing point) and one or more components. Examples of components are color components that can be represented in different color spaces, such as RGB (red, green, and blue) or YUV (Y is the luminance component, and UV are two chrominance components). The point cloud can be a representation of an object seen from one or more viewpoints. The point cloud can be obtained in different ways, for example:

[0090] - Capture of a real object taken by one or more cameras, optionally supplemented by a depth active sensing device;

[0091] - Capture of a virtual / synthetic object taken by one or more virtual cameras in a modeling tool;

[0092] - A mixture of both real and virtual objects.

[0093] The point cloud 111 can be a dynamic point cloud that evolves over time, that is, the number of points can change over time and / or the position of one or more points (for example, at least one of the coordinates X, Y, and Z) can change over time. The evolution of the point cloud can correspond to the movement of the object represented by the point cloud and / or correspond to any shape change of the object or a part of the object.

[0094] The point cloud 111 may be represented in a picture, or in a set or sets of temporally consecutive pictures, each picture including a representation of the point cloud at a determined time "t". The set or sets of temporally consecutive pictures may form a video representing at least a part of the point cloud 111.

[0095] The encoding process 11 may implement, for example, intra-picture encoding and / or inter-picture encoding. Intra-picture encoding is based on intra-picture prediction that exploits spatial redundancy (i.e., the correlation between pixels within a picture), and calculates a predicted value by extrapolating from already encoded pixels for efficient differential encoding. Inter-picture encoding is based on inter-picture prediction that exploits temporal redundancy. So-called internal pictures "I" that are independently encoded in the time domain use only intra-picture encoding. Predicted pictures "P" (or "B") that are encoded in the time domain may use intra-picture and inter-picture prediction.

[0096] The decoding process 12 may correspond, for example, to the inverse operation of the encoding process 11 to decode the data encoded by the encoding process.

[0097] Figure 2 Operations for encoding the point cloud 111 according to specific and non-limiting embodiments are shown. These operations may be part of the encoding process 11 and may be implemented by Figure 16 the apparatus 16.

[0098] In operation 20, the data of the picture 201 of the point cloud is encoded by the encoder ENC1. The picture 201 is, for example, part of a group of pictures (GOP) and includes data representing the point cloud at a determined time "t". The picture 201 includes a set of images, at least one image of which includes a first attribute corresponding to at least a part of the data of the picture 201. Each image including the first attribute is referred to as a first image. The first attribute may be obtained by projecting a part of the point cloud onto each first image according to a first projection, and the first attribute corresponds to the attributes of the points of the part of the point cloud projected onto each of the first images. The attribute (and thus the obtained first attribute) may correspond to texture (or color) information and / or depth (or distance to the viewpoint) information. The set of images of the picture 201 may further include one or more second images that do not include any attributes generated by the projection of the points of the point cloud. For example, the data associated with each second image may correspond to default data, such as a determined gray level for texture information or a determined depth value (e.g., 0) for depth information. Examples of the picture 201 are provided in Figure 4A and 4B and will be described with Figure 4A and 4BTo provide a detailed description of Picture 201. The encoder ENC1 is compatible with traditional encoders, such as:

[0099] · JPEG, Specification ISO / CEI 10918-1 UIT-T Recommendation T.81, https: / / www.itu.int / rec / T-rec-T.81 / en,

[0100] · AVC, also known as MPEG-4 AVC or h264. Specified in UIT-T H.264 and ISO / CEI MPEG-4 Part 10 (ISO / CEI 14496-10), http: / / www.itu.int / rec / T-rec-H.264 / en,

[0101] · HEVC (whose specification can be found on the ITU website, Recommendation T Series H.265, http: / / www.ITU.int / rec / T-rec-H.265-201612-I / en), or

[0102] · 3D-HEVC (an extension of HEVC, whose specification can be found on the ITU website, Recommendation T Series H.265, http: / / www.ITU.int / rec / T-rec-H.265-201612-I / en, Annexes G and I).

[0103] The encoded data of Picture 201 can be stored in Bitstream 112 and / or sent in Bitstream 112.

[0104] In Operation 21, the encoded data of Picture 201 is decoded by the decoder DEC1. The decoder DEC1 is compatible with the encoder ENC1, and is, for example, compatible with traditional decoders, such as:

[0105] · JPEG,

[0106] · AVC, also known as MPEG-4 AVC or h264,

[0107] · HEVC, or

[0108] · 3D-HEVC (an extension of HEVC).

[0109] In Operation 21, the first attribute encoded in Operation 20 is decoded and retrieved, and stored in the buffer memory, for example, for use in generating the reference picture 202 associated with Picture 201.

[0110] In operation 22 implemented by module M22, each first image is deprojected (i.e., the inverse operation of the first projection is performed, e.g., based on metadata containing the parameters of the first projection) according to the first projection associated with each of the first images. The deprojection of the first attribute of the decoded first image enables obtaining a three-dimensional (3D) representation of the point cloud or a portion thereof represented in picture 201. The 3D representation of the point cloud takes, for example, the form of a reconstructed point cloud corresponding to the point cloud of picture 201, with possible differences, for example, due to encoding operation 20 and / or decoding operation 21. The differences may also be due to occlusion during the projection process of generating the first image from the point cloud.

[0111] In operation 23 implemented by module M23, a second attribute is obtained from the 3D encoded representation of the point cloud obtained at operation 22. The second attribute may be obtained, for example, for a second image of picture 201, i.e., for an image in the set of images of picture 201 that includes default data and does not include points from the point cloud. For each second image of the set of images of picture 201, a set of second attributes can be obtained by projecting the points or a portion of the points of the 3D representation according to the projection parameters of the second projection associated with each second image. A specific second projection may be associated with each second image, and each specific second projection is different from the first projection associated with the first images of picture 201.

[0112] A reference picture 202 can be obtained from picture 201 by fusing the decoded first attribute obtained from operation 21 with the second attribute obtained from operation 23. The reference picture may include the same structure as picture 201, i.e., the same spatial arrangement of the set of images, but with different data, i.e., with the decoded first attribute and the obtained second attribute. The following Figure 9 description provides a detailed description of an example of the process of obtaining the reference picture.

[0113] Reference picture 202 can be used in operation 24 implemented by module M24. Operation 24 includes, for example, generating a predictor for inter-frame prediction for encoding one or more pictures of the point cloud that are different from picture 201 (e.g., pictures of the point cloud at a determined time different from the time "t" of picture 201). Then, the point cloud 111 or a picture representing the point cloud can be encoded by referring to reference picture 201. According to a variant, module M24 is part of encoder ENC1.

[0114] The use of the reference picture that has been completed with data (i.e., the second attribute obtained from the data (i.e., the first attribute of the associated picture)) enables increasing the possibility of selecting more inter-frame coding modes of pictures of the point cloud by referring to the obtained reference picture, thereby improving the compression efficiency.

[0115] Of course, multiple reference images can be obtained in the same way as the reference image 202, and each reference image of the multiple reference images can be obtained from a specific image representing the point cloud. The encoding of the point cloud 111 refers to one or more reference images.

[0116] Figure 3 An operation for encoding the point cloud 111 according to another specific and non - restrictive embodiment is shown. Figure 3 The embodiment of... can be regarded as Figure 2 an alternative embodiment of the embodiment of...

[0117] Operations 20, 21, 22, 23 and 24 are the same as those of the embodiment of Figure 2 ...

[0118] An operation 31 implemented by the module M31 is added. Operation 31 includes: fusing the 3D representation of the point cloud obtained through operation 22 with the supplementary point cloud 301 or the supplementary part of the point cloud. The supplementary part of the point cloud corresponds, for example, to the occluded part of the point cloud not represented in the image 201. The supplementary part 301 of the point cloud is represented, for example, by an octree.

[0119] The octree O includes a root node, at least one leaf node and possibly intermediate nodes. A leaf node is a node of the octree O that has no children. All other nodes have children. Each node of the octree is associated with a cube. Thus, the octree O includes a set {Cj} of at least one cube Cj associated with the nodes. A leaf cube is a cube associated with a leaf node of the octree. In Figure 19 the example shown, the cube associated with the root node (depth 0) is divided into 8 sub - cubes (depth 1), and then two sub - cubes of depth 1 are divided into 8 sub - cubes (final depth = maximum depth = 2). Cubes of the same depth generally have the same size, but this embodiment is not limited to this example. When a cube is divided, specific processing can also determine different numbers of sub - cubes for each depth and / or multiple sizes of cubes of the same depth according to its depth.

[0120] The leaf cube associated with the leaf node of the octree O may or may not include points representing at least a part (preferably located at the center of the leaf cube) of the point cloud 111. A color value can be attached to each point included in the leaf cube of the octree O. The color value can be represented in an image OI associated with the octree O such that the color of the pixel of the image OI is associated with the color of the point included in the leaf cube of the octree O. The association can be obtained through a predetermined scan order of the leaf nodes of the octree O and a predetermined scan order of the pixels of the image OI. For example, the predetermined scan order of the octree O can be a 3D recursive Hilbert path, and the predetermined scan order of the image OI can be a raster scan order.

[0121] Thus, the representation of an octree representing a portion of a colored point cloud may include the following information:

[0122] - A first set of flags indicating whether to divide (or not divide) a node into (eight) child nodes,

[0123] - A second set of flags indicating the presence (or absence) of a point at the center of a leaf cube associated with a leaf node,

[0124] - A color image OI, and / or

[0125] - A scan order of leaf cubes associated with leaf nodes and pixels of the image OI.

[0126] Once completed, the 3D representation is projected according to one or more second projections to obtain a second attribute at operation 23.

[0127] Figure 4A And 4B Each shows a picture of the point cloud 111 according to a specific and non - limiting embodiment, such as picture 201.

[0128] Figure 4A Illustrates a first example of picture 40 of a point cloud, e.g., a picture of a GOP of a point cloud, such as picture 201. Picture 40 consists of a set of n images 401, 402, 403,... 40n, where n is an integer greater than or equal to 2. Each of the images 401 to 40n corresponds to a pixel array, and the size and / or clarity of the pixel array may vary from one image to another. For example, the clarity of images 401 and 40n is the same, while the clarity of picture 402 is different from each other and different from the clarity of images 401 and 40n. In Figure 4AIn the example, images 401 to 40n are spatially arranged to cover the entire picture 40 without overlap between the images. According to a variant, images 401 to 40n do not cover the entire picture, and there is space between images 401 to 402, or at least some of them, i.e., the edges of two adjacent images may not touch. Data such as texture information and / or depth information can be associated with each pixel of images 401 to 40n. For example, texture information can be stored in the form of a gray level associated with each channel of a color space (e.g., RGB color space or YUV color space), e.g., representing the gray level of each channel with a first determined number of bits (e.g., 8, 10, or 12 bits). Depth information can be stored, for example, in an alpha channel having a second determined number of bits (e.g., 8, 10, or 12 bits) in the form of a value. Thus, for example, four components RGBα or YUVα (e.g., four 10-bit channels) can be associated with each pixel in picture 40 to represent a point cloud at a determined time "t". According to a variant, a first picture 40 is used to store texture information (e.g., 3 components RGB or YUV), and a second picture having the same image arrangement is used to store depth information, and both pictures represent a point cloud at time "t".

[0129] The set of images forming picture 40 can, for example, include one or more first images and possibly one or more second images. For example, as Figure 5 shown, a first image can be obtained by projecting points of a point cloud according to a first projection (e.g., a different first projection for each first image).

[0130] Figure 5 FIG. illustrates a cube 51 enclosing at least a portion of a point cloud according to a particular and non-limiting embodiment.

[0131] For example, cube 51 is subdivided into 8 sub-cubes at the first subdivision level (only one sub-cube 52 out of the 8 sub-cubes is shown for clarity). At the second subdivision level, sub-cube 52 is also subdivided into 8 sub-cubes (only one sub-cube 53 out of the 8 sub-cubes is shown for clarity). At each subdivision level, a part of the points of the point cloud can be projected (e.g., according to an orthogonal projection) onto one or more faces of the cube (e.g., the faces filled with gray). For example, the points of the point cloud are projected onto face 501 of cube 51, face 502 of cube 52, and face 503 of cube 53. For example, these faces are discretized to form a pixel array, the clarity / size of which depends on the subdivision level of the cube. For example, for the pixels of the face of the cube, the points of the point cloud projected onto the pixels correspond to the points of the point cloud that are closest to the pixels when a ray is emitted from the pixel and orthogonal to the face including the pixel. The attribute associated with the pixel corresponds to the attribute (texture and / or depth) of the point projected onto the pixel.

[0132] Face 501 is used, for example, to form image 401, face 502 is used to form image 402, and face 503 is used to form image 403.

[0133] One or more images of the image set 401 to 40n may not be generated from the first projection and thus may not receive data (attributes) from the points of the point cloud. Such an image or these images are called second images. For such second images, default texture information and / or default depth information is assigned to each pixel of the second image. The default texture / depth information may be a system or user-defined value.

[0134] Figure 4B A second example of picture 41 of the points of the point cloud is illustrated, e.g., a picture of the GOP of the point cloud, e.g., picture 201. Picture 41 consists of m images 411, 412, 413, 414, 41m, where m is an integer greater than or equal to 2. The arrangement of images 411 to 41m may be different from the arrangement in image 40, e.g., there may be free space between images 411 to 41m. Images 411 to 41m may have varying sizes and / or clarities. Each picture may receive attributes from the points of the point cloud, and the attributes are associated with at least some pixels of each of images 411 to 41m. For example, the part of each image that receives attributes from the point cloud is shown as a gray area, while the part of the image that does not receive attributes from the point cloud is shown as a white area, and the white area may be filled with a default value, such as the free space between images. Just like Figure 4AIn the picture 40, the data associated with the pixels of the images 411 to 41n can correspond to texture information and / or depth information. In a variant, a first picture 41 is used to store texture information (e.g., 3-component RGB or YUV), and a second picture 41 having the same arrangement as the images 411 to 41n is used to store depth information. Both pictures represent the point cloud at time "t".

[0135] For example, the set of images forming the picture 41 can include one or more first images and possibly one or more second images. For example, according to a first projection (as Figure 6 shown, different first projections of each first image), the first images (at least the gray areas of each first image) can be obtained by projecting the points of the point cloud.

[0136] Figure 6 Illustrated is the obtaining of the first images of the set of images forming the picture 41 according to a non-limiting embodiment. The point cloud representing the 3D object 6 is segmented into a plurality of 3D parts, such as 50, 100, 1000 or more 3D parts, Figure 6 showing 3 of them (i.e., 3D parts 62, 63, and 64). The 3D part 64 includes the points of the point cloud representing the head of a person, the 3D part 62 includes the points of the point cloud representing the armpit of a person, and the 3D part 63 includes the points of the point cloud representing the hand of a person. One or more first images of each 3D part or a part of the 3D part are generated to represent each 3D part in two dimensions, i.e., according to 2D parameterization. For example, 2D parameterization 601 is obtained for the 3D part 64, 2D parameterization 602 is obtained for the 3D part 62, and 2 different 2D parameterizations 603 and 604 are obtained for the 3D part 63. The 2D parameterizations may differ from each other depending on the 3D part. For example, the 2D parameterization 601 associated with the 3D part 64 is a linear perspective projection, while the 2D parameterization 602 associated with the 3D part 62 is LLE, and the 2D parameterizations 603 and 604 associated with the 3D part 63 are both orthogonal projections according to different viewpoints. According to a variant, all the 2D parameterizations associated with all the 3D parts are of the same type, such as linear perspective projection or orthogonal projection. According to a variant, different 2D parameterizations can be used for the same 3D part.

[0137] The 2D parameterization associated with a given 3D part of the point cloud corresponds to a 2D view of the given 3D part of the point cloud, thus allowing sampling of the given 3D part, i.e., a 2D representation of the content (i.e., points (multiple points)) of the given 3D part including a plurality of samples (possibly corresponding to the pixels of the first image). The number of samples depends on the sampling step applied. The 2D parameterization can be obtained in various ways, such as by implementing any of the following methods:

[0138] - linearly perspective-project the points of the 3D portion of the point cloud onto a plane associated with the view point, where the parameters representing the linear perspective projection include the position of the virtual camera, the spatial sampling step, and the 2D field of view;

[0139] - orthogonally project the points of the 3D portion of the point cloud onto a surface, where the parameters representing the orthogonal projection include the geometric structure (shape, size, and orientation) of the projection surface and the spatial sampling step;

[0140] - LLE (Locally-Linear Embedding) corresponding to the mathematical operation of dimensionality reduction, which is here applied to the 3D-to-2D conversion / transformation, where the parameters representing LLE include the transformation coefficients.

[0141] Each first image (and second image) advantageously has a rectangular shape to simplify the packing process on picture 41.

[0142] Figure 7 A group of pictures (GOP) 7 according to a non-limiting embodiment is shown. Each picture 701, 702, 703, 704, 705, 706, 707, 708, and 709 of the GOP may correspond, for example, to picture 201 or picture 40 or picture 41. The GOP 7 may include, for example, different types of pictures, such as, an I picture 701 (i.e., an intra-coded picture), a P picture 709 (i.e., a predictive-coded picture), and "B" pictures 702 to 708 (i.e., bi-predictive-coded pictures). The arrow diagram illustrates the coding relationship between the pictures. For example, the P picture 709 is coded by referring to the I picture 701, the B picture 705 is coded by using references to pictures 701 and 709, the B picture 703 is coded by using references to pictures 701 and 705, and the B picture 702 is coded by using references to pictures 701 and 703. The GOP may be part of an intra period, i.e., a sequence of pictures included between two I pictures, where the first I picture belongs to the intra period and indicates the start of the intra period, while the second (temporarily) I picture does not belong to the intra period but belongs to a subsequent intra period. The intra period starts from the I picture 701 and may include multiple P pictures, some B pictures included between the I picture 701 and the first (temporarily) P picture 709, some B pictures included between the first P picture and the second (temporarily) P picture 709, etc.

[0143] An I picture is a picture that is coded independently of all other pictures. Each intra period starts with this type of picture (in decoding order).

[0144] A P picture includes motion compensation difference information relative to a previously decoded picture. In compression standards such as MPEG-1, H.262 / MPEG-2, each P picture can only refer to one picture, and this picture must be before the P picture in both the display order and the decoding order, and must be an I or P picture. These restrictions do not apply to more recent standards such as H.264 / MPEG-4 AVC and HEVC.

[0145] A B picture includes motion compensation difference information relative to a previously decoded picture. In standards such as MPEG-1 and H.262 / MPEG-2, each B picture can only refer to two pictures, one of which is before the B picture in the display order and the other follows immediately, and all the pictures referred to must be I or P pictures. These restrictions do not apply to more recent standards such as H.264 / MPEG-4 AVC and HEVC.

[0146] Figure 8 An example of a process for obtaining a picture such as picture 201, 40, and / or 41, or at least a structure of such a picture, according to a non-limiting embodiment is shown.

[0147] Pictures 80 and 81 respectively correspond to two pictures of a point cloud at different times t1 and t2, for example, t2 is greater than t1. Pictures 80 and 81 are, for example, part of a GOP, such as GOP 7. Other pictures of the GOP are not shown in Figure 8 . Picture 80 includes a set of images 801 to 809, and image 81 includes a set of images 811 to 819 arranged in the same manner as images 801 to 809 in picture 80.

[0148] In each of pictures 80 and 81, the images filled with gray shading (i.e., images 801, 802, 803, 808, and 809 for picture 80 and images 812, 815, 817, and 818 for picture 81) correspond to the first image, i.e., an image having data associated with its pixels obtained, for example, by projection or by 2D parameterization from the point cloud.

[0149] In each of pictures 80 and 81, the images filled with diagonally upward stripes (i.e., images 805, 807 of picture 80 and images 811, 813, 819 of picture 81) correspond to the second image, i.e., an image without data obtained from the point cloud, and the second image has a corresponding first image with data obtained from the point cloud in another picture of the GOP. For example, image 815 of picture 81 corresponds to image 805 of picture 80, i.e., images 805 and 815 have the same position and the same size in each of pictures 80 and 81.

[0150] In each of the pictures 80 and 81, the unfilled images (i.e., white images, e.g., images 804, 806 for picture 80, and images 814, 816 for picture 81) correspond to the second images, i.e., images that do not have data obtained from the point cloud, and the second images do not have any corresponding images with data obtained from the point cloud in another picture of the GOP. For example, image 804 corresponds to image 814, and image 806 corresponds to image 816, and neither image 804, 814 (or 806, 816 respectively) has data obtained from the point cloud.

[0151] The letters A, B, C, D, E, F, and G indicate the types of the first projections for obtaining data from the point cloud. Pictures 80 and 81 are compatible in the sense that the type of the first projection used in one image does not conflict with the other type of the first projection in the corresponding picture in the other picture. For example, the first projection D (or B respectively) exists in both pictures, but is associated with the corresponding images 808, 818 (or 802, 812 respectively) located at the same position and having the same size. The first projection A is only used in picture 80 and is associated with the first image 801, but the corresponding image 811 in picture 81 does not have the associated first projection, and there is no conflict.

[0152] Now regarding Figure 8 the right - hand part, an example of the arrangement of images in the pictures of the GOP is explained.

[0153] A set S of compatible pictures is given (e.g., the set including pictures 80 and 81), and another picture Pi of the GOP (not belonging to the set S) is selected. The aim is to arrange the images in picture Pi without conflicting with the images of the pictures belonging to the set S.

[0154] In the first operation 82, an empty rectangle R with the same size as each picture of the GOP is generated.

[0155] In the second operation 83, the projections of picture Pi that also exist in the set S are located at the same positions in picture Pi. For example, if it is considered that the projections A, C, H, and I are used for picture Pi, the projections A and C that exist in at least one picture (i.e., picture 80) of the set S are arranged at the positions of the first images associated with the projections A and C in the set S. The remaining projections H and I will still be located in the rectangle R. To achieve this goal, the remaining space R' (shown with diagonally downward black stripes) in the rectangle R is determined.

[0156] In the third operation 84, the projections H and I are arranged in the remaining space R' of the rectangle R.

[0157] According to a variation of the first example of the arrangement of images in the pictures of the GOP obtained, instead of using the set S of compatible pictures, a determined picture with an associated projection is selected from among the plurality of pictures and used as the starting picture for defining the position of the projection and the remaining space. The determined picture selected is, for example, the first picture of the GOP or the picture with the largest number of projections compared to the other pictures of the GOP. Then, as explained in operations 82, 83, and 84, the other pictures of the GOP are selected one by one to determine the placement positions of the remaining projections.

[0158] According to another example of the arrangement of images in the pictures of the GOP obtained, a list of projections used in all the pictures of the GOP is determined. Then, the projections are arranged in the empty rectangle R, for example, using a method of optimizing the space. Once each projection has been arranged in the area of the empty rectangle, the pattern of the images is applied to each picture of the GOP to project the data of the point cloud onto the right region of the picture according to the associated first projection.

[0159] Obtaining the same arrangement of images enables providing temporal consistency between pictures of a GOP or multiple GOPs or intra-cycles, which can improve temporal prediction between pictures of the GOP(s) / intra-cycle and thus improve the coding efficiency.

[0160] Figure 9 An example of the process of generating a reference picture (e.g., Figure 2 reference picture 202) according to a non-limiting embodiment is shown.

[0161] In Figure 9 , picture 81 represents a picture to be encoded, for example, by implementing the Figure 2 encoding process. Picture 81 includes a set of images 811 to 819, where images 812, 815, 817, and 818 are first images because their data is obtained by projecting the points of the point cloud according to the first projections B, F, G, and D, respectively. To benefit from the efficiency of inter-frame coding, for example, one or more reference pictures and motion information at the block level are required.

[0162] Pictures 90 and 91 each represent a reference picture associated with the pictures of the GOP (or intra-cycle) including the picture 81 to be encoded. Picture 90 is, for example, a reference picture associated with (and obtained from) Figure 8 picture 80. The data of the first image 812 obtained by projection B can be encoded by referring to the corresponding first images 902 and 912 obtained from the same projection B (i.e., from the same part of the point cloud with the same projection parameters at different times), because these first images 902 and 912 have been directly obtained from the associated pictures of the GOP.

[0163] To encode the data of the first image 815, the corresponding data in reference pictures 90 and / or 91 needs to be referred to. As Figure 8 shown, there is no data corresponding to image 905 in associated picture 80. Image 805 (which now corresponds to image 905 in picture 90) is a second image, i.e., an image of a picture with default attributes and not having attributes obtained from a point cloud. To obtain the data associated with the pixels of image 905 of reference picture 90, the data (i.e., attributes) of images 901, 902, 903, 908, and 909 are deprojected (using the metadata of associated projections A, B, C, D, E respectively) to obtain a 3D representation of the point cloud that has been encoded in picture 80 used to obtain reference picture 90. The deprojection performs an inverse projection (relative to the projection used to obtain the considered first image) using the metadata of the projection associated with the depth information (the depth information being associated with the pixels of the considered first image) to obtain a three-dimensional representation of the reconstructed point cloud. The texture information associated with the pixels is assigned to the reconstructed points of the point cloud obtained by deprojection. To obtain the data (i.e., depth and / or texture information) of image 905, the obtained 3D representation is projected according to projection F associated with image 905, and the projection is associated with each image during the processing described above Figure 8 . The same operation is repeated to obtain the data (attributes) of images 907, 915, and 918 (shown with a hollow diamond grid fill pattern), which can be used to encode the data of images 815, 817, and 818 of picture 81. Through such deprojection / reprojection processing, the parts of the reference pictures that do not have data directly obtained from the first image of the associated pictures of the GOP can be made complete.

[0164] Therefore, effective inter-picture prediction between reference pictures 90 and 91 and picture 81 to be encoded is now possible. For example, once the deprojection / reprojection processing has been performed, image 815 can now be better predicted by images 905 and 915.

[0165] Figure 10 Operations for encoding point cloud 111 according to a specific and non-limiting embodiment are shown. The operations can be part of encoding process 11 and can be implemented by Figure 16 device 16.

[0166] In a first operation 101, a set or multiple sets of temporally consecutive pictures are obtained, each picture including data representing a point cloud (or at least a part thereof) at different times t. Each picture of the set or multiple sets has the same structure, i.e., consists of a first set of images that are spatially arranged in the same way in each picture. A projection is associated with each image of the first set, the projection being specific to the image with which it is associated and being different from image to image, and the projections form a second set of projections.

[0167] Each picture includes, for example, one or more first images and possibly one or more second images. The data of one or more first images is obtained by projecting at least one point of the point cloud according to a first projection associated with the one or more first images under consideration and retrieving the attributes of the projected points as the data of the one or more first images under consideration. The data of one or more second images corresponds to the data of a default setting, i.e., to determined data representing default values. The default data indicates, for example, the absence of data directly obtained from the point cloud.

[0168] One or more sets of temporally consecutive pictures are obtained, for example, from a storage device (e.g., the storage device of device 16 or a remote storage device such as a server). According to another example, the one or more sets are obtained by performing the processing described above with respect to Figure 8 associating different projections with each image and arranging all the projections / images spatially within the picture.

[0169] In a second operation 102, first information representing a second set of projections is encoded. The encoded first information can be stored in the bitstream 1001 and / or the encoded first information is sent in the bitstream 1001. The first information includes, for example, a set of metadata, a subset of the metadata describing each projection with a parameter list, and information representing the position of each image within the picture (e.g., the indices of the picture column and picture row of a reference pixel of each image, where the reference pixel is, for example, the top - left pixel or the bottom - right pixel of each image).

[0170] The metadata describing the projection can be based on an octree - based structure of the projection. The octree - based structure of the projection is an octree in which each parent node can include at most eight child nodes and in which a cube is associated with each node. The root node (depth 0) is the only node that does not have any parent node, and each child node (depth greater than 0) has a separate parent node.

[0171] Attach the cube C j to each node of the octree - based structure. The index j refers to the index of the cube of the octree - based structure of the projection. The face F j of the cube C i,j of the octree - based structure of the projection is selected according to the orthogonal projection of the point cloud 111 on these faces. The index i refers to the index of the face (1 to 6) attached to the cube.

[0172] The octree - based structure of the projection can be obtained by recursively dividing an initial cube associated with the root node and enclosing the point cloud 111. Thus, the octree - based structure of the projection includes at least one cube C associated with the node(s).j The set {C j}. The stop condition for the partitioning process can be checked when the maximum octree depth is reached, or when the size of the cube associated with the node is less than the threshold, or when the number of points of the point cloud 111 contained in the cube does not exceed the minimum number.

[0173] In Figure 19 the example shown, the cube associated with the root node (depth 0) is divided into 8 sub-cubes (depth 1), and then two sub-cubes of depth 1 are divided into 8 sub-cubes (final depth = maximum depth = 2), as shown in the left part regarding Figure 5 .

[0174] The projection metadata can thus include:

[0175] - Information data (e.g., a flag) for each cube of the octree-based structure of the projection, which indicates whether the cube associated with the node is divided,

[0176] - And / or face information data (e.g., 6 flags for each cube), which indicates which face or faces of the cube are used for projection.

[0177] According to Figure 2 the embodiment shown, the node information data is a binary flag: equal to 1 to indicate that the cube associated with the node is divided; otherwise 0. And, the face information data is 6-bit data, each bit equal to 1 to indicate that the face is used for projection, otherwise equal to 0.

[0178] In a third operation 103, the point cloud is encoded by using the picture obtained in operation 101. For example, the point cloud is encoded by an encoder compatible with a conventional encoder such as JPEG, AVC, HEVC, or 3D-HEVC. According to a variant, the point cloud is encoded with Figure 2 or Figure 3 processing, and the picture obtained at operation 101 corresponds to the picture 201 at the input of operation 20.

[0179] According to a particular embodiment, attributes are obtained for each picture of at least one group by projecting the points of the point cloud in at least one image of each picture according to a second set of projections associated with at least one image.

[0180] According to another particular embodiment, the first information includes:

[0181] - Projection parameters associated with the projection; and

[0182] - Information representing the position of the image associated with the projection within the picture.

[0183] According to another specific embodiment, the first information is appended to the intra-frame.

[0184] According to another specific embodiment, for each picture in at least one group, second information is appended to each picture, the second information identifying at least one projection for obtaining a second set of attributes from the point cloud, the attributes being associated with at least one image of the first set included in each picture.

[0185] Figure 11 An example of the processing of pictures for generating one or more groups of pictures according to a non-limiting embodiment is shown. Figure 11 The embodiments of can correspond, for example, to the Figure 10 operation 101.

[0186] In a first operation 1101, a set of projections is selected. The selection includes: determining which projections will be used to represent the point cloud in a picture sequence, each picture of the sequence representing the point cloud at a different time t, and the point cloud 111 being dynamic in terms of its evolution over time. The projections can correspond, for example, to the projections described with respect to Figure 1 and Figure 6 . The selection is performed for each picture of one or more GOPs (e.g., for each picture of the intra-period). At the end of operation 1101, a list of projections is obtained, which includes, for example, the set of projections used to represent the point cloud during the entire intra-period.

[0187] The selection of the projections can be performed according to a measure Q(F j ) of the ability to represent the projection (texture and / or depth) image associated with the face F i,j of the cube C i,j in order to effectively compress the projections of the points of the point cloud 111 included in the cube C j onto the face F i,j .

[0188] The measure Q(F i,j ) can respond to the ratio of the total number of pixels N_total(i, j) and the number of newly seen points N_new(i, j), where N_total(i, j) corresponds to the projection of a portion of the point cloud 111 included in the cube C j . A point is considered "newly seen" if it has not been projected onto a previously selected face. If no new point projections are seen onto the face F i,j through the projection of a portion of the input colored point cloud, the ratio becomes infinite. Conversely, if all points are new, the ratio is equal to 1.

[0189] When the measure Q(F i,j ) is less than or equal to a threshold Q_acceptable, the face F i,j is selected:

[0190] Q(F i,j ) ≤ Q_acceptable

[0191] Then, either no face is selected for the cube or at least one face is selected. The threshold Q_acceptable can be a given coding parameter.

[0192] In a second operation 1102, each projection of the list obtained at operation 1101 is associated with an image, and the set of images thus obtained forms a picture (a picture with texture information or a picture with depth information or a picture including both texture and depth information). Determine the arrangement of the images in the set, for example, as described with respect to Figure 8 to optimize the space within the picture. Within each image of a GOP or intra period, the arrangement of the images is the same, and only the data associated with the pixels of the image (i.e., texture and / or depth information) varies from one picture to another. During operation 1102, first information representing each projection in the set of projections is generated. The first information may include:

[0193] - Metadata about the parameters of each projection;

[0194] - Information representing the mapping between the set of projections and the set of images, assigning one projection of the set to one image of the set; and / or

[0195] - Information representing the position of each image within the picture (e.g., the coordinates of the reference pixel of each image (e.g., column index and row index in the picture)).

[0196] The first information can be stored in the bitstream 1001 and / or the first information can be sent in the bitstream 1001.

[0197] In a third operation 1103, the point cloud 111 is projected into the images of each picture of a GOP or intra period according to the projections associated with the images to obtain the attributes to be assigned to the image pixels. Each picture represents the point cloud at a different time t, and the projections used to represent the point cloud may vary from one picture to another. Second information can be further generated during operation 1103 to signal which projections are used to represent the point cloud in a given picture of a GOP or intra period for each picture. The second information can be generated for each picture and the second information may include, for example, a list of the projections performed for each picture (e.g., identified by an ID) to obtain the data / attributes of the so-called first image of the picture. The second information can be stored in the bitstream 1001 together with the first information and / or the second information can be sent in the bitstream 1001 together with the first information.

[0198] In a fourth operation 1104, a picture of a GOP or an intra period is generated by packing / collecting pictures within a corresponding picture (having attributes obtained for each picture at operation 1103), according to first information providing information about the position of an image within a picture and the mapping between the image and its associated projection.

[0199] The pictures thus obtained can then be encoded, e.g., with respect to Figure 2 , 3 or 10, as described for picture 201.

[0200] Figure 12 Operations for decoding an encoded version of a point cloud 111 from a bitstream 112 are shown according to a particular and non - limiting embodiment. The operations can be part of a decoding process 12 and can be implemented by a device 16 of Figure 16 .

[0201] In operation 121, an encoder DEC2 decodes encoded data of one or more pictures of the point cloud (e.g., one or more pictures of a GOP or an intra period) from the received bitstream 112. The bitstream 112 includes the encoded data of one or more pictures. Each picture includes a set of images, at least one of the images in the set including a first attribute corresponding to at least a part of the data of the encoded picture. Each image including the first attribute is referred to as a first image. The first attribute can be obtained by projecting a part of the point cloud onto each first image according to a first projection, and the first attribute corresponds to the attributes of the points of the part of the point cloud projected onto each of the first images. The attribute and thus the first attribute obtained can correspond to texture (or color) information and / or depth (or distance to the viewpoint) information. The set of images of a picture can also include one or more second images that do not include any attribute resulting from the projection of points of the point cloud. The data associated with each second image can correspond, e.g., to default data, e.g., a determined gray level for texture information or a determined depth value for depth information. The decoder DEC2 can correspond to the decoder DEC1 of Figure 2 and, e.g., be compatible with the following conventional decoders, such as:

[0202] ·JPEG,

[0203] ·AVC (also known as MPEG - 4 AVC or H264),

[0204] ·HEVC or

[0205] ·3D - HEVC (extended HEVC).

[0206] Retrieve the first attribute decoded at operation 121, e.g., store it in a buffer memory for use in the generation of one or more reference pictures 1201, each reference picture being associated with a picture. Hereinafter, for the purposes of clarity and conciseness, only one reference picture associated with one picture will be considered.

[0207] In operation 122 implemented by module M122 (which may be the same as module M22 of Figure 2 ), each of the first images is deprojected according to a first projection associated with each first image, i.e., for example, the inverse operation of the first projection is performed based on metadata including the parameters of the first projection. The deprojection of the first attribute decoded for the first image enables obtaining a three-dimensional (3D) representation of the point cloud or a part thereof represented in the picture.

[0208] In operation 123 implemented by module M123 (which may be the same as module M23 of Figure 2 ), a second attribute is obtained from the 3D representation of the point cloud obtained in operation 122. For example, a second attribute may be obtained for a second image of the picture, i.e., an image of the set of images of the picture that includes default data and does not include points from the point cloud. For each second image of the set of images of the picture, a set of second attributes may be obtained by projecting the points or a part of the points of the 3D representation according to the projection parameters of a second projection associated with each second image. A specific second projection may be associated with each second image, and each specific second projection is different from the first projection associated with the first image of the picture.

[0209] Reference picture 1201 (which may be the same as reference picture 2 of Figure 2 ) may be obtained from the image by fusing the decoded first attribute obtained from operation 121 with the second attribute obtained from operation 123. The reference picture may include the same structure as the picture, i.e., the same spatial arrangement as the set of images, but with different data, i.e., with the decoded first attribute and the obtained second attribute. The description provided above in conjunction with Figure 9 provides a detailed description of an example of the process for obtaining a reference picture.

[0210] Reference picture 1201 may be used in operation 124 implemented by module M124. Operation 124 includes, for example, generating a predictor for inter-frame prediction from the decoding of the encoded data included in the bitstream. The data associated with the generation of the predictor may include:

[0211] - A prediction type, e.g., a flag indicating whether the prediction mode is intra-frame or inter-frame;

[0212] - A motion vector, and / or

[0213] - An index for indicating a reference picture from a list of reference pictures.

[0214] Naturally, multiple reference pictures can be obtained in the same manner as reference picture 1201, where each of the multiple reference pictures is obtained from decoded data of a specific picture representing a point cloud. Decoding of the data of bitstream 112 can be based on one or some reference pictures to obtain decoded point cloud 113.

[0215] According to a specific embodiment, attributes are obtained for each picture in at least one group by projecting points of the point cloud in at least one image of each picture according to a second set of projections associated with at least one image.

[0216] According to another specific embodiment, the first information includes:

[0217] - Projection parameters associated with the projection; and

[0218] - Information indicating the position of the image associated with the projection within the picture.

[0219] According to another specific embodiment, the first information is appended to an intra-frame.

[0220] According to another specific embodiment, for each picture in at least one group, second information is appended to each picture, the second information identifying at least one projection of a second set used to obtain attributes from the point cloud, the attributes being associated with at least one image of the first set included in each picture.

[0221] Figure 13 An operation for decoding an encoded version of point cloud 111 from bitstream 112 is shown according to another specific and non - restrictive embodiment. Figure 13 The embodiment of... can be considered as Figure 12 an alternative embodiment of the embodiment of...

[0222] Operations 121, 122, 123, and 124 are the same as those of the Figure 12 embodiment of...

[0223] An operation 131 implemented by module M131 (which can be the same as module M31 of Figure 3 ...) is added. Operation 131 includes: fusing the 3D representation of the point cloud obtained during operation 122 with a supplementary point cloud 1310 or with a supplementary part of the point cloud. The supplementary part of the point cloud corresponds, for example, to the occluded part of the point cloud not represented in the decoded picture. The supplementary part 1310 of the point cloud is represented, for example, by an octree, and the data representing these supplementary parts is included, for example, in bitstream 112.

[0224] The octree O includes a root node, at least one leaf node, and possibly intermediate nodes. A leaf node is a node of the octree O that has no children. All other nodes have children. Each node of the octree is associated with a cube. Thus, the octree O includes a set {C j} of at least one cube C j associated with the nodes. A leaf cube is a cube associated with a leaf node of the octree.

[0225] In Figure 19 the illustrated embodiment, the cube associated with the root node (depth 0) is divided into 8 sub-cubes (depth 1), and then two of the sub-cubes at depth 1 are divided into 8 sub-cubes (final depth = maximum depth = 2). Cubes at the same depth typically have the same size, but this embodiment is not limited to this example. When a cube is divided, specific processing can also determine a different number of sub-cubes at each depth, and / or multiple sizes of cubes at the same depth, or according to their depth.

[0226] The leaf cube associated with a leaf node of the octree O may then include or not include points representing at least a portion (preferably located at the center of the leaf cube) of the point cloud 111.

[0227] A color value can be attached to each point included in a leaf cube of the octree O. The color value can be represented in an image OI associated with the octree O such that the color of a pixel of the image OI is associated with the color of the point included in the leaf cube of the octree O. The association can be obtained through a predetermined scan order of the leaf nodes of the octree O and a predetermined scan order of the pixels of the image OI. For example, the predetermined scan order of the octree O can be a 3D recursive Hilbert path, and the predetermined scan order of the image OI can be a raster scan order.

[0228] Thus, a representation of an octree representing a portion of a colored point cloud can contain the following information:

[0229] - A first set of flags indicating whether a node is divided (or not divided) into (eight) children,

[0230] - A second set of flags indicating the presence (or absence) of a point at the center of the leaf cube associated with a leaf node,

[0231] - A color image OI, and / or

[0232] - A scan order of the leaf cubes associated with the leaf nodes and the pixels of the image OI.

[0233] Once completed, the 3D representation is projected according to one or more second projections to obtain a second property at operation 123.

[0234] Figure 14 illustrates operations for decoding an encoded version of point cloud 111 from bitstream 1001 according to another specific and non - limiting embodiment. The operations may be part of decoding process 12 and may be implemented by Figure 16 device 16.

[0235] In a first operation 141, one or more temporally - consecutive pictures are decoded from one or more received bitstreams 1001. At least one received bitstream 1001 includes encoded data of pictures of one or more GOPs, the pictures representing a point cloud (or at least a part thereof) at different times. Each picture has the same structure, namely consisting of a first set of images spatially arranged in the same way in each picture. A projection is associated with each image of the first set, the projection being specific to the image with which it is associated and different from image to image, and the projections form a second set of projections.

[0236] Each picture includes, for example, one or more first images and possibly one or more second images. Data of one or more first images can be obtained by projecting at least one point of the point cloud according to a first projection associated with the one or more first images under consideration and retrieving the attributes of the projected points as data of the one or more first images under consideration. Data of one or more second images corresponds to data of a default setting, i.e., corresponds to determined data representing default values. The default data indicates, for example, the absence of data directly obtained from the point cloud.

[0237] In a second operation 142, first information representing the projections of the second set is decoded. The decoded first information can be stored in the memory of device 16. The first information includes, for example, a set of metadata, a subset of the metadata describing each projection with a parameter list, and information representing the position of each image within the picture (e.g., indices of the picture column and picture row of a reference pixel of each image, the reference pixel being, for example, the upper - left pixel or the lower - right pixel of each image).

[0238] The metadata describing the projections can be based on an octree - based structure of the projections. The octree - based structure of the projections is an octree in which each parent node can include at most eight child nodes, and in which a cube is associated with each node. The root node (depth 0) is the only node that does not have any parent node, and each child node (depth greater than 0) has a separate parent node.

[0239] Attach cube C j to each node of the octree - based structure of the projections. Index j refers to the index of the cube of the octree - based structure of the projections. Cube C of the octree - based structure of the projections is selected according to the orthogonal projection of the point cloud 111 on these faces.j face F i,j The index i refers to the index attached to the faces (1 to 6) of the cube.

[0240] The projected octree-based structure can be obtained by recursively dividing an initial cube associated with a root node and enclosing the point cloud 111. Thus, the projected octree-based structure includes at least one cube C associated with a node (or nodes). j set {C j}. When the maximum octree depth is reached, or when the size of the cube associated with a node is less than a threshold, or when the number of points of the point cloud 111 contained in the cube does not exceed a minimum number, the stop condition for the division process can be checked.

[0241] In Figure 19 the example shown, the cube associated with the root node (depth 0) is divided into 8 sub-cubes (depth 1), and then two sub-cubes of depth 1 are divided into 8 sub-cubes (final depth = maximum depth = 2).

[0242] The projected metadata can thus include:

[0243] - Information data (e.g., a flag) for each cube of the projected octree-based structure, which indicates whether the cube associated with a node is divided.

[0244] - And / or face information data (e.g., 6 flags for each cube), which indicates which face or faces of the cube are used for projection.

[0245] According to Figure 2 the embodiment shown, the node information data is a binary flag: equal to 1 to indicate that the cube associated with a node is divided; otherwise 0. And the face information data is 6-bit data, each bit equal to 1 to indicate that the face is used for projection, otherwise equal to 0.

[0246] In a third operation 143, the point cloud is decoded by using the decoded data of the picture obtained in operation 141. For example, the point cloud is decoded by a decoder compatible with a conventional decoder such as JPEG, AVC, HEVC, or 3D-HEVC.

[0247] Figure 15 shows an example of decoding an encoded point cloud from a bitstream 1001 according to another non-limiting embodiment.

[0248] In a first operation 151, the encoded data included in the bitstream 1001 is decoded to obtain a sequence of temporally consecutive pictures forming one or more GOPs or intra-cycles. The encoded data is decoded, for example, by a decoder compatible with conventional decoders such as JPEG, AVC, HEVC, or 3D-HEVC. According to a variant, a first sequence and a second sequence of temporally consecutive pictures can be obtained. For example, the first sequence corresponds to pictures including texture information, and the second sequence corresponds to pictures including depth information.

[0249] In a second operation 152, each of the pictures obtained at operation 151 is decoded to obtain a set of decoded images, each picture including a set of images. The decoding of operation 152 is based on, for example, first information received in the signal including the bitstream 1001. The first information includes: a list of projections associated with one or more GOPs or with an intra-cycle, information indicating the position of the image in the picture, and information mapping each projection to the picture. For example, the first information is signaled together with each I picture (i.e., at the start of the intra-cycle).

[0250] In a third operation 153, each of the decoded images obtained at operation 152 is de-projected to obtain a decoded dynamic point cloud, i.e., a 3D representation of the point cloud at consecutive times. The de-projection is based on, for example, second information received in the signal including the bitstream 1001 and the first information. The second information includes a list of projections that each picture has used, a specific list of projections being associated with each picture and forming part of the list of projections of the intra-cycle included in the first information received for each received first picture. The de-projection also uses the projection parameters / metadata associated with each projection and included in the first information. For example, the second information is signaled together with each picture of the intra-cycle. According to a variant, the second information is signaled by the GOP, for example, together with each I and P picture of the intra-cycle.

[0251] Figure 16 An example of the architecture of a device 16 for implementing at least a part of the Figure 1 encoding / decoding scheme according to a non-limiting embodiment is shown. The device 16 can be configured, for example, to implement at least a part of the operations described with respect to Figures 2 to 15 the above.

[0252] The device 16 includes the following elements linked together by a data and address bus 161:

[0253] - A microprocessor 162 (or CPU), for example, a DSP (or digital signal processor);

[0254] - A ROM (or read-only memory) 163;

[0255] - A RAM (or random access memory) 164;

[0256] - Storage interface 165;

[0257] - I / O interface 166 for receiving data to be sent from an application; and

[0258] - A power source, e.g., a battery.

[0259] According to the example, the power source is external to the device. In each of the mentioned memories, the term "register" used in the specification may correspond to a small-capacity area (certain bits), or may correspond to a very large area (e.g., the entire program or a large amount of received or decoded data). The ROM 163 includes at least programs and parameters. The ROM 163 may store algorithms and instructions to execute the technology according to the present embodiment. When powered on, the CPU 162 uploads the program to the RAM and runs the corresponding instructions.

[0260] The RAM 164 includes in registers the program to be run by the CPU 162 and uploaded after the device 16 is turned on, includes input data in registers, includes intermediate data in different states of the method in registers, and includes other variables for running the method in registers.

[0261] The implementations described herein may be implemented, for example, as a method or process, a device, a computer program product, a data stream, or a signal. Even if discussed only in the context of a single implementation form (e.g., discussed only as a method or a device), the implementation of the discussed functions may be implemented in other forms (e.g., a program). The device may be implemented, for example, in appropriate hardware, software, and firmware. The method may be implemented, for example, in a device such as a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. For example, the processor also includes communication devices such as a computer, a mobile phone, a portable / personal digital assistant (“PDA”), and other devices that facilitate information communication between end users.

[0262] According to an example of an encoder or an encoding, point clouds and associated data (e.g., the depth and texture of the points of the point cloud) are obtained from a source. For example, the source belongs to a set including the following:

[0263] - Local memory (163 or 164), e.g., video memory or RAM (or random access memory), flash memory, ROM (or read-only memory), hard disk;

[0264] - Storage interface (165), e.g., an interface with a mass storage, RAM, flash memory, ROM, optical disc, or magnetic support;

[0265] - A communication interface (166), e.g., a wired interface (e.g., a bus interface, a wide area network interface, a local area network interface) or a wireless interface (such as an IEEE 802.11 interface or a Bluetooth interface); and

[0266] - A user interface, such as a graphical user interface that enables a user to input data.

[0267] According to an example of a decoder or a decoding unit, send the decoded point cloud or the reconstructed 3D representation of the point cloud to a destination. Specifically, the destination belongs to a set including the following:

[0268] - A local memory (163 or 164), e.g., a video memory or RAM, a flash memory, a hard disk;

[0269] - A storage interface (165), e.g., an interface with a mass storage, RAM, flash memory, ROM, an optical disc or a magnetic support; and

[0270] - A communication interface (166), e.g., a wired interface (e.g., a bus interface (e.g., USB (or Universal Serial Bus)), a wide area network interface, a local area network interface, an HDMI (High-Definition Multimedia Interface) interface) or a wireless interface (such as an IEEE 802.11 interface, or a Bluetooth interface).

[0271] According to an example of an encoder or an encoding unit, send the bitstream 112 and / or 1001 to a destination. As an example, the bitstream is stored in a local or remote memory, e.g., a video memory (164) or RAM (164), a hard disk (163). In a variant, the bitstream is sent to a storage interface (165), e.g., an interface with a mass storage, flash memory, ROM, an optical disc, or a magnetic carrier, and / or sent through a communication interface (166), the communication interface being, for example, an interface with a point-to-point link, a communication bus, a point-to-multipoint link or a broadcast network.

[0272] According to an example of a decoder or a decoding unit or a renderer, obtain a bitstream from a source. Exemplarily, read the bitstream from a local memory (e.g., a video memory (164), RAM (164), ROM (163), flash memory (163) or hard disk (163)). In a variant, receive the bitstream from a storage interface (165) (e.g., an interface with a mass storage, RAM, ROM, flash memory, an optical disc or a magnetic support), and / or receive the bitstream from a communication interface (165) (e.g., an interface with a point-to-point link, a bus, a point-to-multipoint link or a broadcast network).

[0273] According to an example, the device 16 is configured to implement the method described in conjunction with Figures 2 to 11 and belongs to a set including the following:

[0274] - Mobile device;

[0275] - Communication device;

[0276] - Gaming device;

[0277] - Tablet (or tablet computer);

[0278] - Laptop computer;

[0279] - Still image camera;

[0280] - Video camera;

[0281] - Encoding chip;

[0282] - Server (e.g., broadcast server, video - on - demand server, or web server).

[0283] According to the example, device 16 is configured to implement the decoding method described in conjunction with Figures 12 to 15 and belongs to the set including the following:

[0284] - Mobile device;

[0285] - Communication device;

[0286] - Gaming device;

[0287] - Set - top box;

[0288] - Television set;

[0289] - Tablet (or tablet computer);

[0290] - Laptop computer; and

[0291] - Display (e.g., such as an HMD).

[0292] According to Figure 17 the embodiment shown in, in the transmission context between two remote devices 171 and 172 (of type device 16) on the communication network NET 170, device 171 includes components configured to implement the method for encoding data as described with respect to Figures 2 to 11 and device 172 includes components configured to implement the decoding method as described with respect to Figures 12 to 15 described.

[0293] According to the example, network 170 is a LAN or WLAN network, adapted to broadcast a still image or video picture with associated audio information from device 171 to a decoding / rendering device including device 172.

[0294] According to another example, the network is a broadcast network suitable for broadcasting the encoded point cloud from device 171 to a decoding device including device 172.

[0295] The bitstreams 112, 1001 that may carry the first and / or second information can be carried by signals intended to be sent by device 171.

[0296] Figure 18 An embodiment showing the syntax of such a signal when transmitting data through a packet-based transport protocol, for example, between two remote devices 171 and 172, is shown. Each transmitted packet P includes a header H and payload data PAYLOAD.

[0297] According to an embodiment, the payload PAYLOAD may include at least one of the following elements:

[0298] - Bits representing at least one picture representing the point cloud at a determined time t. For example, the bits may represent texture information and / or depth information associated with the pixels of at least one picture;

[0299] - Bits representing the position of the image in at least one picture;

[0300] - Bits representing projection information data and the mapping between the projection of at least one picture and the image;

[0301] - Bits enabling the identification of the attributes (i.e., texture information and / or depth information) of the pixels of one or more images for at least one picture using which projections in at least one picture.

[0302] According to a particular embodiment, the signal further carries second information associated with at least one picture, the second information identifying at least one projection in a set of projections used to obtain an attribute from the point cloud, the attribute being associated with at least one image included in at least one picture.

[0303] Naturally, the present disclosure is not limited to the embodiments described previously.

[0304] The present disclosure is not limited to methods for encoding and / or decoding point clouds, but also extends to methods and devices for sending bitstreams obtained by encoding point clouds, and / or methods and devices for receiving bitstreams obtained by encoding point clouds. The present disclosure also extends to methods and devices for rendering and / or displaying decoded point clouds (i.e., images of 3D objects represented by the decoded point clouds, with a viewpoint associated with each image).

[0305] The implementations described herein can be implemented, for example, as a method or process, apparatus, software program, data stream, or signal. Even if discussed only in the context of a single form of implementation (e.g., only as a method or device), the implementation of the functions discussed can be implemented in other forms (e.g., a program). The apparatus can be implemented, for example, in appropriate hardware, software, and firmware. The method can be implemented, for example, in an apparatus such as a processor, which generally refers to a processing device, including, for example, a smartphone, tablet, computer, microprocessor, integrated circuit, or programmable logic device. For example, a processor also includes communication devices such as a computer, mobile phone, portable / personal digital assistant (“PDA”), and other devices that facilitate information communication between end users.

[0306] The implementations of the various processes and functions described herein can be implemented in a variety of different devices or applications, specifically, for example, devices or applications associated with data encoding, data decoding, view generation, texture processing, and other processing of images and associated texture information and / or depth information. Examples of such devices include encoders, decoders, post-processors that process decoder outputs, pre-processors that provide inputs to encoders, video encoders, video decoders, video codecs, web servers, set-top boxes, laptop computers, personal computers, mobile phones, PDAs, HMDS (head-mounted displays), smart glasses, and other communication devices. It should be clear that the devices can be mobile and can even be installed in a mobile vehicle.

[0307] In addition, the method can be implemented by instructions executed by a processor, and such instructions (and / or data values generated by the implementation) can be stored in a processor-readable medium or other storage device, a processor-readable medium such as an integrated circuit, software carrier, and other storage devices such as a hard disk, compact disc (CD), optical disc (e.g., DVD, commonly referred to as a digital versatile disc or digital video disc), random access memory (“RAM”), or read-only memory (“ROM”). The instructions can form an application program tangibly implemented on the processor-readable medium. The instructions can be in, for example, hardware, firmware, software, or a combination thereof. The instructions can be found, for example, in an operating system, a separate application program, or a combination of both. Thus, the characteristics of a processor can be described as, for example, a device configured to perform a process and a device including a processor-readable medium (e.g., a storage device) having instructions for performing the process. In addition, the processor-readable medium can store data values generated by the implementation in addition to or instead of the instructions.

[0308] It will be apparent to those skilled in the art that the implementation can generate various signals that are formatted to carry information that can be stored or transmitted, for example. The information can include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal can be formatted to carry, as data, rules for writing or reading the syntax of the described implementations, or actual syntax values written by the described implementations. Such signals can be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or baseband signals. Formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. As is well known, signals can be transmitted over a variety of different wired or wireless links. Signals can be stored on a processor-readable medium.

[0309] Numerous implementations have been described. However, it should be understood that various modifications can be made. For example, elements of different implementations can be combined, supplemented, modified, or removed to produce other implementations. Additionally, one of ordinary skill in the art should understand that other structures and processes can replace the disclosed structures and processes, and the resulting implementations will perform at least substantially the same functions in at least substantially the same manner to achieve at least substantially the same results as the implementations of the present disclosure. Accordingly, the present application contemplates these and other embodiments.

Claims

1. A method for encoding a point cloud representing a three-dimensional object, the method comprises: - obtaining at least one set of temporally consecutive pictures of the point cloud, each picture in the at least one set of temporally consecutive pictures comprising a set of images, the images in the set of images having the same position in each picture in the at least one set of temporally consecutive pictures, a set of projections associated with the at least one set of temporally consecutive pictures, different projections in the set of projections being associated with each image in the set of images; - encoding first information representing the projections in the set of projections; and - encoding the point cloud based on the obtained pictures.

2. An apparatus adapted to encode a point cloud representing a three-dimensional object, the apparatus comprising a memory associated with a processor, the processor being configured to: - obtain at least one set of temporally consecutive pictures of the point cloud, each picture in the at least one set of temporally consecutive pictures comprising a set of images, the images in the set of images having the same position in each picture in the at least one set of temporally consecutive pictures, a set of projections associated with the at least one set of temporally consecutive pictures, different projections in the set of projections being associated with each image in the set of images; - encode first information representing the projections in the set of projections; and - encode the point cloud based on the obtained pictures.

3. The method according to claim 1 or the apparatus according to claim 2, wherein, attributes are obtained for each picture in the at least one set of temporally consecutive pictures by projecting points of the point cloud into at least one image of each picture according to the projections in the set of projections associated with at least one image.

4. The method according to claim 1 or the apparatus according to claim 2, wherein the first information comprises: - projection parameters associated with the projections; and - information representing the position of the image associated with the projection within the picture.

5. The method according to claim 1 or the apparatus according to claim 2, wherein the first information is appended to an intra-frame.

6. The method according to claim 1 or the apparatus according to claim 2, wherein, for each picture in the at least one set of temporally consecutive pictures, second information is appended to each picture, the second information identifying at least one projection in the set of projections used to obtain an attribute from the point cloud, the attribute being associated with at least one image in the set of images included in each picture.

7. A method for decoding from at least one bitstream a point cloud representing a three-dimensional object, the method comprises: - obtaining the at least one bitstream, the at least one bitstream comprising encoded data of at least one set of temporally consecutive pictures of the point cloud, each picture in the at least one set of temporally consecutive pictures comprising a set of images, the images in the set of images having the same position in each picture in the at least one set of temporally consecutive pictures, - decoding the at least one set of temporally consecutive pictures from the at least one bitstream; - Decode first information representing a set of projections, the set of projections being associated with the at least one set of temporally consecutive pictures, and different projections in the set of projections being associated with each image in the set of images; And - Decode the point cloud according to the decoded pictures.

8. An apparatus adapted to decode a point cloud representing a three-dimensional object, the apparatus comprising a memory associated with a processor, the processor being configured to: - Obtain at least one bitstream, the at least one bitstream including encoded data of at least one set of temporally consecutive pictures of the point cloud, Each picture in the at least one set of temporally consecutive pictures includes a set of images, and the images in the set of images have the same position in each picture in the at least one set of temporally consecutive pictures, - Decode the at least one set of temporally consecutive pictures from the at least one bitstream; - Decode first information representing a set of projections, the set of projections being associated with the at least one set of temporally consecutive pictures, and different projections in the set of projections being associated with each image in the set of images; And - Decode the point cloud according to the decoded pictures.

9. The method according to claim 7 or the apparatus according to claim 8, Wherein, For each picture in the at least one set of temporally consecutive pictures, obtain an attribute by projecting points of the point cloud into at least one image of each picture according to the projections in the set of projections associated with at least one image.

10. The method according to claim 7 or the apparatus according to claim 8, wherein the first information Comprises: - Projection parameters associated with the projections; And - Information indicating the position of the image associated with the projection within the picture.

11. The method according to claim 7 or the apparatus according to claim 8, wherein the first information is appended to an intra-frame.

12. The method according to claim 7 or the apparatus according to claim 8, Wherein, For each picture in the at least one set of temporally consecutive pictures, second information is appended to each picture, the second information identifying at least one projection in the set of projections used to obtain an attribute from the point cloud, the attribute being associated with at least one image in the set of images included in each picture.

13. A computer-readable medium storing data representing at least one picture of at least one set of temporally consecutive pictures of a point cloud, each picture in the at least one set of temporally consecutive pictures includes a set of images, the images having the same position in each picture in the at least one set of temporally consecutive pictures, the data further carrying first information representing a set of projections, the projections being associated with the at least one set of temporally consecutive pictures, and different projections being associated with each image in the set of images.

14. The computer-readable medium according to claim 13, further carrying second information associated with the at least one picture, the second information identifying at least one of the set of projections for obtaining an attribute from the point cloud, the attribute being associated with at least one image included in the at least one picture.

15. A non-transitory processor-readable medium storing instructions for causing a processor to perform the method according to claim 1 or 9.

Citation Information

Patent Citations

  • Plan-view projections of depth image data for object tracking

    US7003136B1