Information processing device and method
The proposed information processing device and method enable scalable decoding of point cloud data by decoding geometry and attribute data independently, addressing structural mismatches and simplifying processing, thus enhancing decoding efficiency.
Patent Information
- Application Number
- JP2025040224
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-06-22
- Filing Date
- 2025-03-13
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2041-06-08
AI Technical Summary
Existing methods for scalable decoding of point cloud data face challenges due to mismatches between the reference structure of attribute data and the tree structure of geometry data, requiring complex processing when geometry data is scaled.
An information processing device and method that decodes encoded geometry data without applying scalable encoding, based on scalable encoding information, and decodes attribute data at an intermediate resolution of the geometry data's tree structure, independent of geometry scaling.
Facilitates easier realization of scalable decoding of point cloud data by preventing mismatches between reference and tree structures, reducing the need for complex processing and potential cost increases.
Smart Images

Figure 0007786631000002 
Figure 0007786631000003 
Figure 0007786631000004
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an information processing device and method, and more particularly to an information processing device and method that enable scalable decoding of point cloud data to be more easily realized. [Background technology]
[0002] Conventionally, methods for encoding 3D data representing a three-dimensional structure such as a point cloud have been considered (see, for example, Non-Patent Document 1). Also, an encoding method has been proposed that enables scalable decoding of encoded data of this point cloud (see, for example, Non-Patent Document 2). In the method described in Non-Patent Document 2, scalable decoding is achieved by making the reference structure of attribute data similar to the tree structure of geometry data.
[0003] Incidentally, a method has been proposed for scaling geometry data and thinning out points when encoding such a point cloud (see, for example, Non-Patent Document 3). [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] R. Mekuria, Student Member IEEE, K. Blom, P. Cesar., Member, IEEE, "Design, Implementation and Evaluation of a Point Cloud Codec for Tele-Immersive Video",tcsvt_paper_submitted_february.pdf [Non-patent document 2] Ohji Nakagami, Satoru Kuma, "[G-PCC] Spatial scalability support for G-PCC", ISO / IEC JTC1 / SC29 / WG11 MPEG2019 / m47352, March 2019, Geneva, CH [Non-patent document 3] Xiang Zhang, Wen Gao, Sehoon Yea, Shan Liu, "[G-PCC][New proposal] Signaling delta QPs for adaptive geometry quantization in point cloud coding", ISO / IEC JTC1 / SC29 / WG11 MPEG2019 / m49232, July 2019 Gothenburg, Sweden Summary of the Invention [Problem to be solved by the invention]
[0005] However, scaling geometry data and thinning points as described in Non-Patent Document 3 changes the tree structure of the geometry data. As a result, there is a risk that a mismatch occurs between the reference structure of the attribute data and the tree structure of the geometry data, making scalable decoding impossible. In other words, in order to achieve scalable decoding when scaling geometry data as described in Non-Patent Document 3, it is necessary to form the reference structure of the attribute data in accordance with the scaling of the geometry data. In other words, in order to encode point cloud data so as to enable scalable decoding, the attribute data reference structure must be changed depending on whether or not the geometry data is scaled, which may require complicated processing.
[0006] The present disclosure has been made in light of the above circumstances, and aims to make it easier to realize scalable decoding of point cloud data. [Means for solving the problem]
[0007] An information processing device according to one aspect of the present technology is an information processing device that includes: a geometry data decoding unit that decodes encoded data of geometry data that has been encoded by applying geometry scaling, which is an encoding method that involves changing the tree structure of the geometry data, in encoding a point cloud that represents a three-dimensional object as a collection of points; and an attribute data decoding unit that decodes the encoded data of the attribute data without applying scalable encoding, based on scalable encoding information corresponding to the fact that scalable encoding that enables attribute data to be decoded at an intermediate resolution of the tree structure of the geometry data is not applied depending on the application of the geometry scaling.
[0008] An information processing method according to one aspect of the present technology is an information processing method including: decoding, by an information processing device, encoded data of geometry data that has been encoded using geometry scaling, an encoding method that involves changing the tree structure of the geometry data, in encoding a point cloud that represents a three-dimensional object as a collection of points; and decoding, based on scalable encoding information corresponding to the fact that scalable encoding that enables attribute data to be decoded at an intermediate resolution of the tree structure of the geometry data is not applied depending on the application of the geometry scaling, the encoded data of the attribute data without applying the scalable encoding.
[0009] In an information processing device and method according to one aspect of the present technology, in encoding a point cloud that represents a three-dimensional object as a collection of points, the encoded data of the geometry data that has been encoded using geometry scaling, which is an encoding method that involves changing the tree structure of the geometry data, is decoded, and the encoded data of the attribute data is decoded without applying scalable encoding based on scalable encoding information corresponding to the fact that scalable encoding that enables attribute data to be decoded at an intermediate resolution of the tree structure of the geometry data is not applied depending on the application of geometry scaling. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 10 is a diagram illustrating an example of hierarchical structure of geometry data. [Figure 2] FIG. 10 is a diagram illustrating an example of lifting. [Figure 3] FIG. 10 is a diagram illustrating an example of lifting. [Figure 4] FIG. 10 is a diagram illustrating an example of lifting. [Figure 5] FIG. 10 is a diagram illustrating an example of quantization. [Figure 6] FIG. 10 is a diagram illustrating an example of hierarchical structure of attribute data. [Figure 7] FIG. 10 is a diagram illustrating an example of inverse hierarchical organization of attribute data. [Figure 8] FIG. 10 is a diagram illustrating an example of geometry scaling. [Figure 9] FIG. 10 is a diagram illustrating an example of geometry scaling. [Figure 10] FIG. 10 is a diagram illustrating an example of hierarchical structure of attribute data. [Figure 11] FIG. 10 is a diagram illustrating an example of encoding control. [Figure 12] FIG. 10 is a diagram illustrating an example of semantics. [Figure 13] FIG. 10 is a diagram illustrating an example of a profile. [Figure 14] FIG. 10 is a diagram illustrating an example of syntax. [Figure 15] FIG. 1 is a block diagram illustrating an example of the main configuration of an encoding device. [Figure 16] FIG. 2 is a block diagram illustrating an example of the main configuration of a geometry data encoding unit. [Figure 17] 10 is a block diagram showing an example of the main configuration of an attribute data encoding unit. FIG. [Figure 18] 10 is a flowchart illustrating an example of the flow of an encoding control process. [Figure 19] 10 is a flowchart illustrating an example of the flow of an encoding process. [Figure 20] 10 is a flowchart illustrating an example of the flow of a geometry data encoding process. [Figure 21] 10 is a flowchart illustrating an example of the flow of an attribute data encoding process. [Figure 22] FIG. 2 is a block diagram illustrating an example of the main configuration of a decoding device. [Figure 23] 10 is a flowchart illustrating an example of the flow of a decoding process. [Figure 24] FIG. 1 is a block diagram illustrating an example of the main configuration of a computer. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, modes for carrying out the present disclosure (hereinafter referred to as embodiments) will be described in the following order. 1. Encoding Control 2. First embodiment (encoding device) 3. Second embodiment (decoding device) 4. Notes
[0012] <1. Encoding control> <References supporting technical content and technical terminology> The scope of disclosure of the present technology includes not only the contents described in the embodiments but also the contents described in the following non-patent documents that were publicly known at the time of filing.
[0013] Non-patent document 1: (mentioned above) Non-patent document 2: (mentioned above) Non-patent document 3: (mentioned above) Non-patent document 4: Khaled Mammou, Alexis Tourapis, Jungsun Kim, Fabrice Robinet, Valery Valentin, Yeping Su, "Lifting Scheme for Lossy Attribute Encoding in TMC1", ISO / IEC JTC1 / SC29 / WG11 MPEG2018 / m42640, April 2018, San Diego, US
[0014] In other words, the contents of the above-mentioned non-patent documents and the contents of other documents referenced in the above-mentioned non-patent documents are also used as the basis for determining the support requirements.
[0015] <Point Cloud> Previously, 3D data existed, such as point clouds, which represent three-dimensional structures using point location information and attribute information, and meshes, which are composed of vertices, edges, and faces and define three-dimensional shapes using polygonal representations.
[0016] For example, in the case of a point cloud, a three-dimensional structure (a three-dimensional object) is represented by a large number of points. Point cloud data (also referred to as point cloud data) is composed of geometry data (also referred to as position information) and attribute data (also referred to as attribute information) for each point. The attribute data can include any information. For example, the attribute data may include color information, reflectance information, normal information, etc. for each point. In this way, point cloud data has a relatively simple data structure, and by using a sufficient number of points, it is possible to represent any three-dimensional structure with sufficient accuracy.
[0017] <Quantization of position information using voxels> Since such point cloud data has a relatively large amount of data, an encoding method using voxels was devised to compress the data volume through encoding, etc. A voxel is a three-dimensional region for quantizing geometry data (position information).
[0018] That is, the three-dimensional region (also called a bounding box) containing the point cloud is divided into small three-dimensional regions called voxels, and each voxel indicates whether it contains a point. In this way, the position of each point is quantized in voxel units. Therefore, by converting point cloud data into such voxel data (also called voxel data), it is possible to suppress an increase in the amount of information (typically, to reduce the amount of information).
[0019] For example, as shown in FIG. 1A, a bounding box 10 is divided into a plurality of voxels 11-1, each represented by a small rectangle. For simplicity, the three-dimensional space is described as a two-dimensional plane. In reality, the bounding box 10 is a three-dimensional spatial region, and the voxels 11-1 are small rectangular parallelepiped (including cube) regions. The geometry data (position information) of each point 12-1 in the point cloud data, represented by a black circle in FIG. 1A, is corrected so that it is positioned at each voxel 11-1. In other words, the geometry data is quantized using the voxel as a unit. All the squares within the bounding box 10 in FIG. 1A are voxels 11-1, and all the black circles in FIG. 1A are points 12-1.
[0020] <Tree structure (Octree)> Furthermore, a method has been devised to enable scalable decoding of geometry data by structuring the geometry data in a tree. In other words, by making it possible to decode nodes from the top layer of the tree structure to any layer, it becomes possible to restore not only geometry data at the highest resolution (lowest layer), but also at lower resolutions (intermediate layers). In other words, it is possible to decode at any resolution without decoding information at unnecessary layers (resolutions).
[0021] This tree structure may be of any type. For example, it may be a KD tree or an octree. An octree is an octet tree, and is suitable for dividing a three-dimensional space region (dividing it into two in each of the x, y, and z directions). In other words, as described above, it is suitable for a structure that divides the bounding box 10 into multiple voxels.
[0022] For example, one voxel is divided into two in each of the x, y, and z directions (i.e., divided into eight) to form voxels at the next lower level (also called LoD). In other words, two voxels aligned in each of the x, y, and z directions (i.e., eight voxels) are combined to form a voxel at the next higher level (LoD). By recursively repeating this structure, an octree can be constructed using voxels.
[0023] The voxel data indicates whether each voxel contains a point. In other words, the voxel data expresses the position of a point at a resolution of the voxel size. Therefore, by constructing an octree using voxel data, scalability of the resolution of the geometry data can be achieved. In other words, constructing an octree of geometry data by structuring voxel data into a tree is easier than constructing a tree of points scattered at arbitrary positions.
[0024] For example, in the case of A in FIG. 1, since it is two-dimensional, four voxels 11-1 arranged vertically and horizontally as shown in B in FIG. 1 are integrated to form voxel 11-2, which is one level higher and is indicated by a bold line. Geometry data is then quantized using this voxel 11-2. That is, if point 12-1 (A in FIG. 1) exists within voxel 11-2, the point 12-1 is converted to point 12-2 corresponding to voxel 11-2 by correcting its position. Note that although only one voxel 11-1 is indicated by a symbol in B in FIG. 1, all squares indicated by dotted lines within the bounding box 10 in B in FIG. 1 are voxels 11-1. Similarly, although only one voxel 11-2 is indicated by a symbol in B in FIG. 1, all squares indicated by bold lines within the bounding box 10 in B in FIG. 1 are voxels 11-2. Similarly, although only one point 12-2 is labeled in FIG. 1B, all of the black circles shown in FIG. 1B are points 12-2.
[0025] Similarly, as shown in FIG. 1C, four voxels 11-2 arranged vertically and horizontally are integrated to form the next higher voxel 11-3 indicated by a bold line. Geometry data is then quantized using this voxel 11-3. That is, if point 12-2 (FIG. 1B) exists within voxel 11-3, the point 12-2 is converted to point 12-3 corresponding to voxel 11-3 by correcting its position. Note that although only one voxel 11-2 is labeled in FIG. 1C, all squares indicated by dotted lines within the bounding box 10 of FIG. 1C are voxels 11-2. Similarly, although only one voxel 11-3 is labeled in FIG. 1C, all squares indicated by bold lines within the bounding box 10 of FIG. 1C are voxels 11-3. Similarly, although only one point 12-3 is labeled in FIG. 1C, all of the black circles shown in FIG. 1C are points 12-3.
[0026] Similarly, as shown in FIG. 1D, four voxels 11-3 arranged vertically and horizontally are integrated to form the next higher voxel 11-4, indicated by a bold line. Geometry data is then quantized using this voxel 11-4. That is, if point 12-3 (FIG. 1C) exists within voxel 11-4, the point 12-3 is converted into point 12-4 corresponding to voxel 11-4 by correcting its position. Although only one voxel 11-3 is labeled in FIG. 1D, all of the squares indicated by dotted lines within the bounding box 10 of FIG. 1D are voxels 11-3.
[0027] By doing this, the geometry data is organized into a tree structure (octree).
[0028] <Lifting> In contrast, when encoding attribute data, the geometry data, including any degradation due to encoding, is assumed to be known, and encoding is performed using the positional relationships between points. As a method for encoding such attribute data, methods using a transformation called RAHT (Region Adaptive Hierarchical Transform) or lifting as described in Non-Patent Document 4 have been considered. By applying these technologies, it is possible to hierarchize (structure the reference structure) of attribute data (referencing relationships) like the octree of geometry data.
[0029] For example, in the case of lifting, the attribute data of each point is coded as a difference value from a predicted value derived using the attribute data of other points. Then, points for deriving the difference value (i.e., deriving the predicted value) are selected hierarchically.
[0030] For example, in the hierarchy shown in A of Fig. 2, of the points (P0 to P9) indicated by circles, points P7, P8, and P9 indicated by white circles are selected as prediction points, which are points from which predicted values are derived, and the other points P0 to P6 are set to be selected as reference points, which are points whose attribute data are referenced when deriving the predicted values. In other words, in this hierarchy, for each of the prediction points P7 to P9, a difference value between the attribute data and its predicted value is derived.
[0031] For the sake of simplicity, the three-dimensional space is illustrated as a two-dimensional plane in Fig. 2. In other words, the points P0 to P6 are actually arranged in the three-dimensional space.
[0032] Each arrow in A in Figure 2 indicates a reference relationship when deriving a predicted value. For example, the predicted value of prediction point P7 is derived with reference to attribute data of reference points P0 and P1. The predicted value of prediction point P8 is derived with reference to attribute data of reference points P2 and P3. The predicted value of prediction point P9 is derived with reference to attribute data of reference points P4 to P6. Then, for each of prediction points P7 to P9, a difference value between the predicted value calculated as described above and the attribute data is derived.
[0033] In the next higher layer, as shown in B of Figure 2, the points (P0 to P6) selected as reference points in the layer A of Figure 2 (the layer one level lower) are classified (sorted) into prediction points and reference points in the same way as in the layer A of Figure 2.
[0034] For example, in B of Fig. 2, points P1, P3, and P6 indicated by gray circles are selected as prediction points, and points P0, P2, P4, and P5 indicated by black circles are selected as reference points. That is, in this hierarchy, for each of the prediction points P1, P3, and P6, a difference value between the attribute data and its prediction value is derived.
[0035] Each arrow in FIG. 2B indicates a reference relationship when deriving a predicted value. For example, the predicted value of prediction point P1 is derived with reference to the attribute data of reference points P0 and P2. The predicted value of prediction point P3 is derived with reference to the attribute data of reference points P2 and P4. The predicted value of prediction point P6 is derived with reference to the attribute data of reference points P4 and P5. Then, for each of prediction points P1, P3, and P6, a difference value between the predicted value calculated as described above and the attribute data is derived.
[0036] At the next higher level, as shown in C of Figure 2, the points (P0, P2, P4, P5) selected as reference points at level B of Figure 2 (the next lower level) are classified (sorted), predicted values for each prediction point are derived, and the difference value between the predicted value and the attribute data is derived.
[0037] By repeating this classification recursively for the reference points at the next lower level, the reference structure of the attribute data is hierarchized.
[0038] <Point classification> The procedure for classifying (sorting) points in such lifting will now be described in more detail. In lifting, points are classified in order from the lower layer to the upper layer, as described above. In each layer, first, the points are sorted in Morton code order. Next, the first point in the sequence of points sorted in Morton code order is selected as the reference point. Next, points located near the reference point (neighboring points) are searched for, and the searched points (neighboring points) are set as prediction points (also called index points).
[0039] For example, as shown in Fig. 3, points are searched for within a circle 22 of radius R centered on a reference point 21 to be processed. This radius R is set in advance for each layer. In the example of Fig. 3, points 23-1 to 23-4 are detected and set as predicted points.
[0040] For simplicity of explanation, the three-dimensional space is illustrated as a two-dimensional plane in Fig. 3. In other words, in reality, each point is arranged in three-dimensional space, and the search for the point is performed within a spherical region of radius R.
[0041] Next, the remaining points are classified in the same way. That is, among the points that have not yet been selected as either a reference point or a prediction point, the first point in the Morton code order is selected as the reference point, and points near that reference point are searched for and set as prediction points.
[0042] Once the above process has been repeated until all points have been classified, processing for that layer ends and processing moves to the next higher layer. The above procedure is then repeated for that layer. In other words, each point selected as a reference point in the next lower layer is sorted in Morton code order and classified into reference points and prediction points as described above. By repeating the above process, the reference structure of the attribute data is hierarchically organized.
[0043] <Derivation of predicted values> As described above, in the case of lifting, the predicted value of the attribute data of a prediction point is derived using the attribute data of reference points around the prediction point. For example, as shown in Fig. 4, the predicted value of a prediction point Q(i,j) is derived by referring to the attribute data of reference points P1 to P3.
[0044] For the sake of simplicity, the three-dimensional space is illustrated as a two-dimensional plane in Fig. 4. In other words, each point is actually arranged in three-dimensional space.
[0045] In this case, the attribute data of each reference point is weighted and integrated by a weight value (α(P, Q(i, j))) corresponding to the inverse of the distance (actually the distance in three-dimensional space) between the prediction point and the reference point, as shown in the following equation (1), where A(P) represents the attribute data of point P.
[0046]
number
[0047] In the method described in Non-Patent Document 4, the highest resolution (ie, lowest layer) position information is used to derive the distance between the predicted point and its reference point. <Quantization> After the attribute data is layered as described above, it is quantized and encoded. During the quantization, the attribute data (difference value) of each point is weighted according to the layer structure, as shown in the example of FIG. 5. This weight value (Quantization Weight) W is derived for each point using the weight value of the lower layer, as shown in FIG. 5. Note that this weight value can also be used in lifting (layering of attribute data) to improve compression efficiency.
[0048] <Tree structure inconsistency> In the case of lifting described in Non-Patent Document 4, as described above, the method of hierarchizing the reference structure of attribute data is different from the method of tree-structuring (e.g., octree-structuring) geometry data. Therefore, it is not guaranteed that the reference structure of the attribute data matches the tree structure of the geometry data. Therefore, in order to decode the attribute data, it is necessary to decode the geometry data down to the lowest layer, regardless of its hierarchy. In other words, it is difficult to achieve scalable decoding of point cloud data without decoding unnecessary information.
[0049] <Achieving scalable decoding of point cloud data> Therefore, as described in Non-Patent Document 2, a method has been proposed in which the reference structure of attribute data is made similar to the tree structure of geometry data. More specifically, when constructing the reference structure of attribute data, a prediction point is selected so that a point also exists in a voxel one layer higher than the voxel in which a point in the current layer exists. In this way, it is possible to decode point cloud data at a desired resolution without decoding unnecessary information. In other words, it is possible to achieve scalable decoding of point cloud data.
[0050] For example, as shown in A of Fig. 6, in a bounding box 100, which is a predetermined three-dimensional spatial region, points 102-1 to 102-9 are arranged for each voxel 101-1 at a predetermined level. Note that in Fig. 6, for simplicity of explanation, the three-dimensional space is described as a two-dimensional plane. In other words, in reality, the bounding box is a three-dimensional spatial region, and the voxels are small regions of a rectangular parallelepiped (including a cube). The points are arranged in the three-dimensional space.
[0051] When it is not necessary to distinguish between the voxels 101-1 to 101-3, they will be referred to as voxels 101. When it is not necessary to distinguish between the points 102-1 to 102-9, they will be referred to as points 102.
[0052] In this hierarchy, points 102-1 to 102-9 are classified into prediction points and reference points so that points also exist in voxel 101-2, which is one hierarchy higher than voxel 101-1 in which points 102-1 to 102-9 exist, as shown in B of Fig. 6. In the example of B of Fig. 6, points 102-3, 102-5, and 102-8, indicated by white circles, are set as prediction points, and the other points are set as reference points.
[0053] Similarly, at the next higher level, points 102-1, 102-2, 102-4, 102-6, 102-7, and 102-9 exist in voxel 101-2, and voxel 101-3 at the next higher level also has points 102. These points 102 are classified into prediction points and reference points (C in FIG. 6). In the example of C in FIG. 6, points 102-1, 102-4, and 102-7, indicated by gray circles, are set as prediction points, and the other points are set as reference points.
[0054] By doing this, as shown in D of Fig. 6, the hierarchy is created so that voxel 101-3, which has point 102 in the lower layer, now has one point 102. This process is performed for each layer. In other words, by performing this process when constructing the reference structure of the attribute data (when classifying prediction points and reference points in each layer), the reference structure of the attribute data can be made similar to the tree structure (octree) of geometry data.
[0055] Decoding is performed in the reverse order of FIG. 6, as shown in FIG. 7, for example. For example, as shown in A of FIG. 7, in a given bounding box 100, points 102-2, 102-6, and 102-9 are arranged for each voxel 101-3 of a given hierarchy (similar to the state shown in D of FIG. 6). Note that, also in FIG. 7, for simplicity of explanation, the three-dimensional space is explained as a two-dimensional plane. That is, in reality, the bounding box is a three-dimensional spatial region, and the voxel is a small region of a rectangular parallelepiped (including a cube). The points are arranged in the three-dimensional space.
[0056] At the next lower level, as shown in B of FIG. 6, attribute data for points 102-2, 102-6, and 102-9 for each voxel 101-3 are used to derive predicted values for points 102-1, 102-4, and 102-7, which are then added to the differential values to restore the attribute data for point 102 for each voxel 101-2 (similar to C of FIG. 6).
[0057] Furthermore, at the next lower level, as shown in C of FIG. 7, the attribute data of points 102-1, 102-2, 102-4, 102-6, 102-7, and 102-9 for each voxel 101-2 are used to derive predicted values for points 102-3, 102-5, and 102-8, which are then added to the differential values to restore the attribute data (similar to B of FIG. 6).
[0058] By doing this, the attribute data of the point 102 for each voxel 101-1 is restored as shown in D of Fig. 7 (the same state as A of Fig. 6). In other words, as in the case of an octree, the attribute data of each layer can be restored using the attribute data of the upper layer.
[0059] By doing so, the reference structure (hierarchical structure) of the attribute data can be associated with the tree structure (hierarchical structure) of the geometry data. Therefore, since geometry data corresponding to each attribute data can be obtained even at intermediate resolutions, the geometry data and attribute data can be correctly decoded at the intermediate resolutions. In other words, scalable decoding of point cloud data can be realized.
[0060] <Geometry scaling> Meanwhile, as described in Non-Patent Document 3, a method called geometry scaling has been proposed in which geometry data is scaled and points are thinned out when encoding a point cloud. Geometry scaling quantizes the geometry data of nodes during encoding. This process allows points to be thinned out depending on the characteristics of the area to be encoded.
[0061] For example, as shown in A of Fig. 8, six points (point1 to point6) with X coordinates of 100, 101, 102, 103, 104, and 105 are to be processed. Note that, for simplicity of explanation, only the X coordinate will be explained here. That is, since points are actually arranged in a three-dimensional space, the same processing as for the X coordinate, which will be explained below, is also performed on the Y coordinate and Z coordinate.
[0062] When geometry scaling is applied to these six points, the X coordinate of each point is scaled according to the value of the quantization parameter baseQP, as shown in the table in FIG. 8B. When baseQP = 0, no scaling is performed, and the X coordinate of each point remains as shown in FIG. 8A. For example, when baseQP = 4, the X coordinate of point2 is scaled from 101 to 102, the X coordinate of point4 is scaled from 103 to 104, and the X coordinate of point6 is scaled from 105 to 106.
[0063] If the X coordinates of multiple points overlap due to this scaling, they can be merged. For example, if mergeDuplicatePoint=1, such duplicate points are merged into one point. mergeDuplicatePoint is flag information indicating whether or not to merge such duplicate points. In other words, such merging thins out points (reduces the number of points). For example, in each row of the table in B of FIG. 8, such merging can thin out points shown in gray. For example, if baseQP=4, the number of points for the four points (point2 to point5) shown in bold lines is reduced by half.
[0064] In order to achieve scalable decoding of point cloud data, when attribute data is layered as described in Non-Patent Document 2, the tree structure of the geometry data is estimated based on the decoding results of the coded data of the geometry data. In other words, the reference structure of the attribute data is formed so that it is similar to the tree structure of the estimated geometry data.
[0065] <Tree structure inconsistency due to geometry scaling> However, when geometry scaling is applied as described above, points may be thinned out in the decoded result of the encoded data of the geometry data. If points are thinned out, it may be impossible to estimate the tree structure of the actual geometry data. In other words, there is a risk that the estimated tree structure may not match the actual tree structure (the tree structure corresponding to the points before thinning out). In such a case, there is a risk that scalable decoding of point cloud data may not be achieved.
[0066] An example will be described for the case where baseQP = 4 in B of Fig. 8. Assume that the tree structure corresponding to each point (point1 to point6) in the table shown in A of Fig. 8 is the tree structure shown in Fig. 9. Nodes 121-1 to 121-6 in this tree structure correspond to each point (point1 to point6) in the table shown in A of Fig. 8. In other words, assume that the X coordinates of nodes 121-1 to 121-6 in the lowest layer (third layer) of this tree structure are 100, 101, 102, 103, 104, and 105, respectively.
[0067] The following two rules are applied to the tree structure in Figure 9. The first rule is that in the layer one level above the third layer (the second layer), the nodes with X coordinates of 100 and 101, the nodes with X coordinates of 102 and 103, and the nodes with X coordinates of 104 and 105 in the third layer are grouped together. The second rule is that in the top layer (the first layer), all nodes in the second layer are grouped together.
[0068] That is, nodes 121-1 and 121-2 in the third layer belong to node 122-1 in the second layer. Nodes 121-3 and 121-4 in the third layer belong to node 122-2 in the second layer. Nodes 121-5 and 121-6 in the third layer belong to node 122-3 in the second layer. Nodes 122-1 to 122-3 in the second layer belong to node 123 in the first layer.
[0069] When geometry scaling is performed, the X coordinate of node 121-2 is scaled from 101 to 102, the X coordinate of node 121-4 is scaled from 103 to 104, and the X coordinate of node 121-6 is scaled from 105 to 106. As a result, the X coordinates of node 121-2 and node 121-3 overlap, so they are merged and node 121-3 is thinned out. Similarly, node 121-4 and node 121-5 are merged and node 121-5 is thinned out. As a result, nodes 121-1, 121-2, 121-4, and 121-6 are encoded in the third layer.
[0070] In encoding attribute data, the highest resolution geometry data is obtained as a result of decoding the encoded data of geometry data. That is, in the example of Figure 9, nodes 121-1, 121-2, 121-4, and 121-6 are obtained. Then, the tree structure of the geometry data is estimated from these four points (four nodes).
[0071] In this case, the same rules as in the example of Fig. 9 are applied to form a similar tree structure. That is, in the layer one level above the third layer (second layer), the nodes with X coordinates of 100 and 101, the nodes with X coordinates of 102 and 103, and the nodes with X coordinates of 104 and 105 in the third layer are grouped together. In addition, in the top layer (first layer), all nodes in the second layer are grouped together.
[0072] As a result, a tree structure such as that shown in Fig. 10 is estimated. That is, node 121-1 in the third layer belongs to node 124-1 in the second layer. Node 121-2 in the third layer belongs to node 124-2 in the second layer. Node 121-4 in the third layer belongs to node 124-3 in the second layer. Node 121-6 in the third layer belongs to node 124-4 in the second layer. Furthermore, nodes 124-1 to 124-4 in the second layer belong to node 125 in the first layer.
[0073] As is clear from a comparison of Figures 9 and 10, these tree structures do not match. For example, suppose the value of the parameter skipOctreeLayer, which indicates the layer to be decoded, is "1" (skipOctreeLayer = 1). That is, when decoding the second layer, three nodes are obtained for the geometry data, as shown in Figure 9. In contrast, four nodes are obtained for the attribute data, as shown in Figure 10. As such, the tree structures do not match, making it difficult to obtain the decoded results for the desired layer without decoding unnecessary information. In other words, it has been difficult to achieve scalable decoding of point cloud data.
[0074] In other words, to achieve scalable decoding of point cloud data, the tree structure must be estimated taking into account such geometry scaling. In other words, complicated processing is required, such as changing the tree structure estimation method depending on whether geometry scaling is performed or not. Furthermore, preparing multiple estimation methods may increase costs.
[0075] Furthermore, in geometry scaling, when thinning out overlapping points, the thinning method (which points to thin out among multiple overlapping points) is not specified and depends on the design. In other words, a tree structure estimation method corresponding to geometry scaling must be newly designed according to the design of the geometry scaling, which could increase costs.
[0076] <Encoding method restrictions> Therefore, as shown in the top row of the table in FIG. 11, applicable encoding methods are restricted, and the combined use of scalable decoding encoding and processing for updating the tree structure of geometry data is prohibited.
[0077] In other words, when encoding point clouds, which represent three-dimensional objects as a collection of points, the system is controlled to prohibit the combined use of scalable encoding, which is an encoding method that generates encoded data that can be decoded in a scalable manner, and scaling encoding, which is an encoding method that involves changing the tree structure of geometry data.
[0078] For example, in an information processing device, when encoding a point cloud that represents a three-dimensional object as a collection of points, an encoding control unit is provided that controls to prohibit the combined use of scalable encoding, which is an encoding method that generates encoded data that can be decoded in a scalable manner, and scaling encoding, which is an encoding method that involves changing the tree structure of geometry data.
[0079] In other words, when scalable coding is applied, the application of scaling coding is prohibited, and when scaling coding is applied, the application of scalable coding is prohibited. By doing so, when scalable coding is performed, it is possible to suppress the occurrence of a mismatch between the reference structure of attribute data and the tree structure of geometry data. Therefore, it is possible to more easily realize scalable decoding of point cloud data.
[0080] Note that scalable coding may be any coding method that generates coded data that can be decoded in a scalable manner. For example, it may be lifting scalability, as described in Non-Patent Document 2, in which attribute data is coded by lifting using a reference structure similar to the tree structure of geometry data. In other words, coding may be controlled so as to prohibit the combined use of lifting scalability and scaling coding.
[0081] Furthermore, scaling encoding may be any encoding method that involves changing the tree structure of geometry data. For example, it may be geometry scaling, which scales and encodes geometry data, as described in Non-Patent Document 3. In other words, control may be exercised to prohibit the combined use of scalable encoding and geometry scaling.
[0082] Of course, as shown in the second row from the top of the table in FIG. 11, it is also possible to prohibit the combined use of Lifting Scalability and Geometry Scaling (Method 1).
[0083] In this case, the coding control unit may perform control so as to prohibit the application of scaling coding when scalable coding is applied. For example, as shown in the third row from the top of the table in Fig. 11, the application of geometry scaling may be prohibited when lifting scalability is applied (method 1-1).
[0084] Furthermore, the coding control unit may perform control so as to prohibit the application of scalable coding when scaling coding is applied. For example, as shown in the fourth row from the top of the table in Fig. 11, the application of lifting scalability may be prohibited when geometry scaling is applied (method 1-2).
[0085] In order to perform such control, flag information (permission flag or prohibition flag) indicating whether or not to permit (or prohibit) the application of scalable coding and scaling coding may be set, for example, as shown in the fifth row from the top of the table in Figure 11 (Method 2).
[0086] For example, the coding control unit may control the signaling of a scalable coding enable flag, which is flag information regarding the application of scalable coding, and a scaling coding enable flag, which is flag information regarding the application of scaling coding.
[0087] For example, as shown in the sixth row from the top of the table in Fig. 11, such a restriction may be specified in the semantics (Method 2-1). For example, when the semantics signal a scalable coding enable flag with a value indicating the application of scalable coding, it may be specified that a scaling coding enable flag with a value indicating non-application of scaling coding is to be signaled, and the coding control unit may perform signaling in accordance with the semantics. Also, when the semantics signal a scaling coding enable flag with a value indicating the application of scaling coding, it may be specified that a scalable coding enable flag with a value indicating non-application of scalable coding is to be signaled, and the coding control unit may perform signaling in accordance with the semantics.
[0088] FIG. 12 shows an example of semantics in this case. For example, semantics 161 shown in FIG. 12 specifies that if the value of geom_scaling_enabled_flag is greater than 0, the value of lifting_scalability_enabled_flag must be set to 0. Here, geom_scaling_enabled_flag is flag information indicating whether or not geometry scaling is applied. If geom_scaling_enabled_flag = 1, geometry scaling is applied. Also, if geom_scaling_enabled_flag = 0, geometry scaling is not applied. lifting_scalability_enabled_flag is flag information indicating whether or not lifting scalability is applied. If lifting_scalability_enabled_flag = 1, lifting scalability is applied. Also, if lifting_scalability_enabled_flag = 0, lifting scalability is not applied.
[0089] In other words, in this semantics 161, when geometry scaling is applied, the application of lifting scalability is prohibited. Conversely, the semantics may also specify that when lifting scalability is applied, the application of geometry scaling is prohibited. In other words, the semantics may specify that when the value of lifting_scalability_enabled_flag is greater than 0, the value of geom_scaling_enabled_flag must be 0.
[0090] Furthermore, for example, as shown in the seventh row from the top of the table in Fig. 11, signaling based on such restrictions may be performed in the profile (method 2-2). For example, when scalable coding is applied, the coding control unit may perform control to signal, in the profile, a scalable coding enable flag with a value indicating the application of scalable coding and a scaling coding enable flag with a value indicating the non-application of scaling coding. Furthermore, when scaling coding is applied, the coding control unit may perform control to signal, in the profile, a scaling coding enable flag with a value indicating the application of scaling coding and a scalable coding enable flag with a value indicating the non-application of scalable coding.
[0091] An example of a profile when lifting scalability is applied is shown in Fig. 13. For example, profile 162 shown in Fig. 13 signals lifting_scalability_enabled_flag = 1 and geom_scaling_enabled_flag = 0. In other words, this profile 162 indicates that lifting scalability is applied and geometry scaling is not applied.
[0092] In addition, in a profile when geometry scaling is applied, geom_scaling_enabled_flag = 1 and lifting_scalability_enabled_flag = 0 may be signaled.
[0093] Furthermore, such a restriction may be specified in the syntax, for example, as described in the bottom row of the table shown in Fig. 11 (Method 2-3). For example, when a scalable coding enable flag with a value indicating the application of scalable coding is signaled, the coding control unit may perform control to omit signaling of the scaling coding enable flag according to the above-mentioned syntax. Also, when a scaling coding enable flag with a value indicating the application of scaling coding is signaled, the coding control unit may perform control to omit signaling of the scalable coding enable flag according to the above-mentioned syntax.
[0094] An example of syntax in this case is shown in Fig. 14. For example, in syntax 163 shown in Fig. 14, lifting_scalability_enabled_flag is signaled only when geom_scaling_enabled_flag is not signaled (i.e., the value is set to "0" and geometry scaling is not applied). In other words, in this case, lifting scalability can be applied. In other words, when geometry scaling is applicable (i.e., when geom_scaling_enabled_flag is signaled), lifting_scalability_enabled_flag is not signaled (i.e., the value is set to "0" and lifting scalability is not applied).
[0095] Conversely, geom_scaling_enabled_flag may be signaled only when lifting_scalability_enabled_flag is not signaled (i.e., the value is set to "0" and lifting scalability is not applied). In other words, in this case, geometry scaling can be applied. In other words, when lifting scalability is applicable (i.e., when lifting_scalability_enabled_flag is signaled), geom_scaling_enabled_flag may not be signaled (i.e., the value is set to "0" and geometry scaling is not applied).
[0096] 2. First Embodiment <Encoding device> Next, a device to which the present technology described above in <1. Encoding Control> is applied will be described. Fig. 15 is a block diagram showing an example of the configuration of an encoding device, which is one aspect of an information processing device to which the present technology is applied. The encoding device 200 shown in Fig. 15 is a device that encodes a point cloud (3D data). The encoding device 200 encodes the point cloud by applying the present technology described above in <1. Encoding Control>.
[0097] Note that Fig. 15 shows the main processing units, data flows, etc., and does not necessarily show everything. In other words, in encoding device 200, there may be processing units that are not shown as blocks in Fig. 15, and there may be processing and data flows that are not shown as arrows, etc. in Fig. 15.
[0098] As shown in FIG. 15, the encoding device 200 includes an encoding control unit 201, a geometry data encoding unit 211, a geometry data decoding unit 212, a point cloud generation unit 213, an attribute data encoding unit 214, and a bitstream generation unit 215.
[0099] The encoding control unit 201 performs processing related to control of encoding of point cloud data. For example, the encoding control unit 201 controls the geometry data encoding unit 211. The encoding control unit 201 also controls the attribute data encoding unit 214. For example, as described above in <1. Encoding Control>, the encoding control unit 201 controls these processing units so as to prohibit the combined use of scalable encoding, which is an encoding method that generates scalably decodable encoded data, and scaling encoding, which is an encoding method that involves changing the tree structure of geometry data. The encoding control unit 201 also controls the bitstream generation unit 215 to control signaling between a scalable encoding enable flag (e.g., Lifting_scalability_enabled_flag) that is flag information regarding the application of scalable encoding and a scaling encoding enable flag (e.g., geom_scaling_enabled_flag) that is flag information regarding the application of scaling encoding.
[0100] The geometry data encoding unit 211 encodes geometry data (position information) of the point cloud (3D data) input to the encoding device 200 and generates the encoded data. Any encoding method may be used. For example, processing such as filtering or quantization for noise suppression (denoising) may be performed. However, the geometry data encoding unit 211 performs this encoding under the control of the encoding control unit 201. In other words, the geometry data encoding unit 211 applies geometry scaling during this encoding under the control of the encoding control unit 201. The geometry data encoding unit 211 supplies the generated encoded data of the geometry data to the geometry data decoding unit 212 and the bitstream generation unit 215.
[0101] The geometry data decoding unit 212 acquires the coded data of the geometry data supplied from the geometry data encoding unit 211 and decodes the coded data. Any decoding method may be used as long as it is compatible with the coding performed by the geometry data encoding unit 211. For example, processing such as filtering or inverse quantization for denoising may be performed. The geometry data decoding unit 212 supplies the generated geometry data (decoded result) to the point cloud generation unit 213.
[0102] The point cloud generation unit 213 acquires attribute data (attribute information) of the point cloud input to the encoding device 200 and geometry data (decoded result) supplied from the geometry data decoding unit 212. The point cloud generation unit 213 performs processing (recolor processing) to associate the attribute data with the geometry data (decoded result). The point cloud generation unit 213 supplies the attribute data associated with the geometry data (decoded result) to the attribute data encoding unit 214.
[0103] The attribute data encoding unit 214 acquires the geometry data (decoded result) and attribute data supplied from the point cloud generation unit 213. The attribute data encoding unit 214 encodes the attribute data using the geometry data (decoded result) to generate encoded data of the attribute data. However, the attribute data encoding unit 214 performs this encoding under the control of the encoding control unit 201. That is, the attribute data encoding unit 214 applies lifting scalability to this encoding under the control of the encoding control unit 201. The attribute data encoding unit 214 supplies the generated encoded data of the attribute data to the bitstream generation unit 215.
[0104] The bitstream generation unit 215 acquires coded data of geometry data supplied from the geometry data encoding unit 211. The bitstream generation unit 215 also acquires coded data of attribute data supplied from the attribute data encoding unit 214. The bitstream generation unit 215 generates a bitstream including these coded data. Furthermore, under the control of the encoding control unit 201, the bitstream generation unit 215 signals control information such as a scalable coding enable flag (e.g., Lifting_scalability_enabled_flag) that is flag information regarding the application of scalable coding and a scaling coding enable flag (e.g., geom_scaling_enabled_flag) that is flag information regarding the application of scaling coding (includes the control information in the bitstream). The bitstream generation unit 215 outputs the generated bitstream to the outside of the encoding device 200 (e.g., the decoding side).
[0105] By adopting such a configuration, the encoding device 200 can prohibit the combined use of scalable encoding, which is an encoding method that generates encoded data that can be decoded in a scalable manner, and scaling encoding, which is an encoding method that involves changing the tree structure of geometry data, making it easier to achieve scalable decoding of point cloud data.
[0106] These processing units (the encoding control unit 201, the geometry data encoding unit 211, and the bitstream generation unit 215) may have any configuration. For example, each processing unit may be configured with a logic circuit that realizes the above-described processing. Furthermore, each processing unit may have, for example, a central processing unit (CPU), read-only memory (ROM), random access memory (RAM), etc., and may execute a program using these to realize the above-described processing. Of course, each processing unit may have both of these configurations, and may realize part of the above-described processing using a logic circuit and other parts by executing a program. The configurations of the processing units may be independent of each other. For example, some processing units may realize part of the above-described processing using a logic circuit, other processing units may execute a program to realize the above-described processing, and still other processing units may realize the above-described processing using both a logic circuit and by executing a program.
[0107] <Geometry data encoding part> Fig. 16 is a block diagram showing an example of the main configuration of the geometry data encoding unit 211. Note that Fig. 16 shows the main processing units, data flows, etc., and does not necessarily show everything. In other words, the geometry data encoding unit 211 may have processing units that are not shown as blocks in Fig. 16, or may have processing or data flows that are not shown as arrows, etc. in Fig. 16.
[0108] As shown in FIG. 16, the geometry data encoding unit 211 includes a voxel generation unit 231 , a tree structure generation unit 232 , a selection unit 233 , a geometry scaling unit 234 , and an encoding unit 235 .
[0109] The voxel generation unit 231 performs processing related to the generation of voxel data. For example, the voxel generation unit 231 sets a bounding box for the input point cloud data and sets voxels to divide the bounding box. The voxel generation unit 231 then quantizes the geometry data of each point in units of voxels to generate voxel data. The voxel generation unit 231 supplies the generated voxel data to the tree structure generation unit 232.
[0110] The tree structure generation unit 232 performs processing related to the generation of a tree structure. For example, the tree structure generation unit 232 acquires voxel data supplied from the voxel generation unit 231. The tree structure generation unit 232 also converts the voxel data into a tree structure. For example, the tree structure generation unit 232 generates an octree using the voxel data. The tree structure generation unit 232 supplies data of the generated octree to the selection unit 233.
[0111] The selection unit 233 performs processing related to control of whether or not to apply geometry scaling. For example, the selection unit 233 acquires octree data supplied from the tree structure generation unit 232. Furthermore, the selection unit 233 selects a supply destination of the octree data in accordance with the control of the encoding control unit 201. That is, the selection unit 233 selects whether to supply the octree data to the geometry scaling unit 234 or the encoding unit 235 in accordance with the control of the encoding control unit 201, and supplies the octree data to the selected supply destination.
[0112] For example, when the encoding control unit 201 instructs the application of geometry scaling, the selection unit 233 supplies the octree data to the geometry scaling unit 234. On the other hand, when the encoding control unit 201 instructs the non-application of geometry scaling, the selection unit 233 supplies the octree data to the encoding unit 235.
[0113] The geometry scaling unit 234 performs processing related to geometry scaling. For example, the geometry scaling unit 234 acquires octree data supplied from the selection unit 233. The geometry scaling unit 234 also performs geometry scaling on the octree data, scaling the geometry data, and merging points. The geometry scaling unit 234 supplies the octree data that has been subjected to geometry scaling to the encoding unit 235.
[0114] The encoding unit 235 performs processing related to encoding of octree data (octree-arranged voxel data (i.e., geometry data)). For example, the encoding unit 235 acquires octree data supplied from the selection unit 233 or the geometry scaling unit 234. For example, when the encoding control unit 201 instructs the encoding unit 235 to apply geometry scaling, the encoding unit 235 acquires octree data that has been subjected to geometry scaling and that has been supplied from the geometry scaling unit. Furthermore, when the encoding control unit 201 instructs the encoding unit 235 not to apply geometry scaling, the encoding unit 235 acquires octree data that has not been subjected to geometry scaling and that has been supplied from the selection unit 233.
[0115] The encoding unit 235 encodes the acquired octree data to generate encoded data of the geometry data. Any encoding method may be used. The encoding unit 235 supplies the generated encoded data of the geometry data to the geometry data decoding unit 212 and the bitstream generation unit 215 (both shown in FIG. 15).
[0116] <Attribute data encoding part> Fig. 17 is a block diagram showing an example of the main configuration of the attribute data encoding unit 214. Note that Fig. 17 shows the main processing units, data flows, etc., and is not limited to what is shown in Fig. 17. In other words, the attribute data encoding unit 214 may include processing units that are not shown as blocks in Fig. 17, or processing and data flows that are not shown as arrows, etc. in Fig. 17.
[0117] As shown in FIG. 17, the attribute data encoding unit 214 includes a selection unit 251 , a scalable layering processing unit 252 , a layering processing unit 253 , a quantization unit 254 , and an encoding unit 255 .
[0118] The selection unit 251 performs processing related to control of whether lifting scalability is applied or not. For example, the selection unit 251 selects a supply destination of attribute data, geometry data (decoded results), etc. generated by the point cloud generation unit 213 ( FIG. 15 ) under the control of the encoding control unit 201. That is, under the control of the encoding control unit 201, the selection unit 251 selects whether to supply the data to the scalable layering processing unit 252 or the layering processing unit 253, and supplies the data to the selected supply destination.
[0119] For example, when the encoding control unit 201 instructs the application of lifting scalability, the selection unit 251 supplies attribute data, geometry data (decoded results), etc. to the scalable layering processing unit 252. On the other hand, when the encoding control unit 201 instructs the non-application of lifting scalability, the selection unit 251 supplies attribute data, geometry data (decoded results), etc. to the layering processing unit 253.
[0120] The scalable layering processor 252 performs processing related to lifting of attribute data (forming a reference structure). For example, the scalable layering processor 252 acquires attribute data and geometry data (decoded results) supplied from the selector 251. The scalable layering processor 252 uses the geometry data to layer the attribute data (i.e., form a reference structure). In this case, the scalable layering processor 252 performs layering using the method described in Non-Patent Document 2. That is, the scalable layering processor 252 estimates a tree structure (octree) based on the geometry data, forms a reference structure for the attribute data so as to correspond to the estimated tree structure, derives a predicted value according to the reference structure, and derives a difference value between the predicted value and the attribute data. The scalable layering processor 252 supplies the attribute data (difference value) generated in this manner to the quantizer 254.
[0121] The layering processing unit 253 performs processing related to lifting of attribute data (forming a reference structure). For example, the layering processing unit 253 acquires attribute data and geometry data (decoded results) supplied from the selection unit 251. The layering processing unit 253 uses the geometry data to layer the attribute data (i.e., form a reference structure). At this time, the layering processing unit 253 performs layering by applying the method described in Non-Patent Document 4. In other words, the layering processing unit 253 forms a reference structure for the attribute data independently of the tree structure (octree) of the geometry data, derives a predicted value according to the reference structure, and derives a difference value between the predicted value and the attribute data. In other words, the layering processing unit 253 does not estimate the tree structure of the geometry data. The layering processing unit 253 supplies the attribute data (difference value) generated in this manner to the quantization unit 254.
[0122] The quantization unit 254 acquires attribute data (difference values) supplied from the scalable layering processing unit 252 or the layering processing unit 253. The quantization unit 254 quantizes the attribute data (difference values). The quantization unit 254 supplies the quantized attribute data (difference values) to the encoding unit 255.
[0123] The encoding unit 255 acquires the quantized attribute data (difference value) supplied from the quantization unit 254. The encoding unit 255 encodes the quantized attribute data (difference value) to generate coded data of the attribute data. This coding method is arbitrary. The encoding unit 255 supplies the generated coded data of the attribute data to the bitstream generation unit 215 (FIG. 15).
[0124] With the above-described configuration, the encoding device 200 can prohibit the combined use of scalable encoding, which is an encoding method that generates scalably decodable encoded data, and scaling encoding, which is an encoding method that involves changing the tree structure of geometry data, when encoding a point cloud that represents a three-dimensional object as a collection of points. Therefore, for example, when encoding attribute data, scalable decoding of point cloud data can be achieved without the need for complex processing, such as preparing multiple methods for estimating the tree structure of geometry data and selecting one from among them. Furthermore, since there is no need to newly design a method for estimating the tree structure, costs can be reduced. In other words, scalable decoding of point cloud data can be more easily achieved.
[0125] <Encoding control process flow> Next, we will explain the processing executed by the encoding device 200. The encoding control unit 201 of the encoding device 200 controls the encoding of point cloud data by executing encoding control processing. An example of the flow of this encoding control processing will be explained with reference to the flowchart in Fig. 18.
[0126] When the encoding control process starts, the encoding control unit 201 determines whether or not to apply lifting scalability in step S101. If it is determined that lifting scalability is to be applied, the process proceeds to step S102.
[0127] In step S102, the encoding control unit 201 prohibits geometry scaling when encoding the geometry data. In step S103, the encoding control unit 201 applies lifting scalability when encoding the attribute data. As described above in <1. Encoding Control>, the encoding control unit 201 signals control information such as Lifting_scalability_enabled_flag and geom_scaling_enabled_flag to correspond to these controls.
[0128] When the process of step S103 ends, the encoding control process ends.
[0129] Also, if it is determined in step S101 that lifting scalability is not to be applied, the process proceeds to step S104.
[0130] In step S104, the encoding control unit 201 determines whether or not to apply geometry scaling. If it is determined that geometry scaling is to be applied, the process proceeds to step S105.
[0131] In step S105, the encoding control unit 201 applies geometry scaling when encoding the geometry data. In step S106, the encoding control unit 201 disables lifting scalability when encoding the attribute data. As described above in <1. Encoding Control>, the encoding control unit 201 also signals control information such as Lifting_scalability_enabled_flag and geom_scaling_enabled_flag to correspond to these controls.
[0132] When the process of step S106 ends, the encoding control process ends.
[0133] If it is determined in step S104 that geometry scaling is not to be applied, the process proceeds to step S107.
[0134] In step S107, the encoding control unit 201 disables geometry scaling when encoding the geometry data. In step S108, the encoding control unit 201 disables lifting scalability when encoding the attribute data. As described above in <1. Encoding Control>, the encoding control unit 201 also signals control information such as Lifting_scalability_enabled_flag and geom_scaling_enabled_flag to correspond to these controls.
[0135] When the process of step S108 ends, the encoding control process ends.
[0136] By executing the encoding control process as described above, the encoding control unit 201 can more easily realize scalable decoding of point cloud data.
[0137] <Encoding process flow> The encoding device 200 encodes the point cloud data by executing an encoding process. An example of the flow of this encoding process will be described with reference to the flowchart in FIG.
[0138] When the encoding process starts, in step S201, the geometry data encoding unit 211 of the encoding device 200 executes the geometry data encoding process to encode the geometry data of the input point cloud and generate encoded data of the geometry data.
[0139] In step S202, the geometry data decoding unit 212 decodes the coded data of the geometry data generated in step S201 to generate geometry data (decoded result).
[0140] In step S203, the point cloud generation unit 213 performs recolor processing using the attribute data of the input point cloud and the geometry data (decoded result) generated in step S202, and associates the attribute data with the geometry data.
[0141] In step S204, the attribute data encoding unit 214 performs attribute data encoding processing to encode the attribute data that has been recolored in step S203, and generate encoded data of the attribute data.
[0142] In step S205, the bitstream generating unit 215 generates and outputs a bitstream including the coded data of the geometry data generated in step S201 and the coded data of the attribute data generated in step S204.
[0143] When the process of step S205 is completed, the encoding process ends.
[0144] <Geometry data encoding process flow> Next, an example of the flow of the geometry data encoding process executed in step S201 of FIG. 19 will be described with reference to the flowchart of FIG.
[0145] When the geometry data encoding process starts, the voxel generation unit 231 of the geometry data encoding unit 211 generates voxel data in step S221.
[0146] In step S222, the tree structure generating unit 232 generates a tree structure (octree) of geometry data using the voxel data generated in step S221.
[0147] In step S223, the selection unit 233 determines whether or not to perform geometry scaling under the control of the encoding control unit 201. If it is determined that geometry scaling is to be performed, the process proceeds to step S22.
[0148] In step S224, the geometry scaling unit 234 performs geometry scaling on the tree-structured geometry data (i.e., octree data) generated in step S222. When the processing of step S224 ends, the process proceeds to step S225. Also, if it is determined in step S223 that geometry scaling is not to be performed, the process of step S224 is skipped and the process proceeds to step S225.
[0149] In step S225, the encoding unit 235 encodes the tree-structured geometry data (i.e., octree data) generated in step S222, or the geometry data (i.e., octree data) that has been subjected to geometry scaling in step S224, to generate encoded data of the geometry data.
[0150] When the process of step S225 ends, the geometry data encoding process ends, and the process returns to FIG.
[0151] <Attribute data encoding process flow> Next, an example of the flow of the attribute data encoding process executed in step S204 of FIG. 19 will be described with reference to the flowchart of FIG.
[0152] When the attribute data encoding process starts, in step S241, the selection unit 251 of the attribute data encoding unit 214 determines whether or not to apply lifting scalability under the control of the encoding control unit 201. If it is determined that lifting scalability is to be applied, the process proceeds to step S242.
[0153] In step S242, the scalable hierarchical processing unit 252 performs lifting using the method described in Non-Patent Document 2. That is, the scalable hierarchical processing unit 252 estimates the tree structure of the geometry data and hierarchizes the attribute data (forms a reference structure) according to the estimated tree structure. The scalable hierarchical processing unit 252 then derives a predicted value according to the reference structure and derives a difference value between the attribute data and the predicted value.
[0154] When the process of step S242 ends, the process proceeds to step S244. Also, if it is determined in step S241 that lifting scalability is not to be applied, the process proceeds to step S243.
[0155] In step S243, the layering processing unit 253 performs lifting using the method described in Non-Patent Document 4. That is, the layering processing unit 253 layers the attribute data (forms a reference structure) independently of the tree structure of the geometry data. The layering processing unit 253 then derives a predicted value according to the reference structure, and derives a difference value between the attribute data and the predicted value. When the processing of step S243 ends, the process proceeds to step S244.
[0156] In step S244, the quantization unit 254 performs a quantization process to quantize each of the difference values derived in step S242 or step S243.
[0157] In step S245, the encoding unit 255 encodes the difference value quantized in step S244 to generate encoded data of the attribute data. When the process of step S245 ends, the attribute data encoding process ends, and the process returns to FIG. 19.
[0158] By performing each process as described above, the encoding device 200 can more easily achieve scalable decoding of point cloud data.
[0159] 3. Second Embodiment <Decryption device> Fig. 22 is a block diagram showing an example of the configuration of a decoding device, which is one aspect of an information processing device to which the present technology is applied. The decoding device 300 shown in Fig. 22 is a device that decodes coded data of a point cloud (3D data). The decoding device 300 decodes the coded data of the point cloud generated by the coding device 200, for example.
[0160] Note that Fig. 22 shows the main processing units, data flows, etc., and does not necessarily show everything. That is, in the decoding device 300, there may be processing units that are not shown as blocks in Fig. 22, and there may be processing or data flows that are not shown as arrows, etc. in Fig. 22.
[0161] As shown in FIG. 22, the decoding device 300 includes an encoded data extraction unit 311, a geometry data decoding unit 312, an attribute data decoding unit 313, and a point cloud generation unit 314.
[0162] The coded data extraction unit 311 acquires and stores the bitstream input to the decoding device 300. The coded data extraction unit 311 extracts coded data of geometry data and attribute data from the top layer to a desired layer from the stored bitstream. If the coded data supports scalable decoding, the coded data extraction unit 311 can extract coded data up to intermediate layers. If the coded data does not support scalable decoding, the coded data extraction unit 311 extracts coded data of all layers.
[0163] The coded data extraction unit 311 supplies the coded data of the extracted geometry data to the geometry data decoding unit 312. The coded data extraction unit 311 supplies the coded data of the extracted attribute data to the attribute data decoding unit 313.
[0164] The geometry data decoding unit 312 acquires the coded data of position information supplied from the coded data extraction unit 311. The geometry data decoding unit 312 decodes the coded data of the geometry data by performing the inverse process of the geometry data coding process performed by the geometry data coding unit 211 of the coding device 200, and generates geometry data (decoded result). The geometry data decoding unit 312 supplies the generated geometry data (decoded result) to the attribute data decoding unit 313 and the point cloud generation unit 314.
[0165] The attribute data decoding unit 313 acquires the coded data of the attribute data supplied from the coded data extraction unit 311. The attribute data decoding unit 313 acquires the geometry data (decoded result) supplied from the geometry data decoding unit 312. The attribute data decoding unit 313 performs the inverse process of the attribute data coding process performed by the attribute data coding unit 214 of the coding device 200, thereby decodes the coded data of the attribute data using the geometry data (decoded result) and generates attribute data (decoded result). The attribute data decoding unit 313 supplies the generated attribute data (decoded result) to the point cloud generation unit 314.
[0166] The point cloud generation unit 314 acquires the geometry data (decoded result) supplied from the geometry data decoding unit 312. The point cloud generation unit 314 acquires the attribute data (decoded result) supplied from the attribute data decoding unit 313. The point cloud generation unit 314 generates a point cloud (decoded result) using the geometry data (decoded result) and the attribute data (decoded result). The point cloud generation unit 314 outputs the generated point cloud (decoded result) data to the outside of the decoding device 300.
[0167] With the above-described configuration, the decoding device 300 can correctly decode the coded data of the point cloud data generated by the coding device 200. In other words, scalable decoding of the point cloud data can be more easily realized.
[0168] These processing units (the encoded data extraction unit 311 to the point cloud generation unit 314) may have any configuration. For example, each processing unit may be configured with a logic circuit that realizes the above-described processing. Furthermore, each processing unit may have, for example, a CPU, ROM, RAM, etc., and may execute a program using these to realize the above-described processing. Of course, each processing unit may have both of these configurations, and may realize part of the above-described processing using a logic circuit and the other part by executing a program. The configurations of the processing units may be independent of each other. For example, some processing units may realize part of the above-described processing using a logic circuit, other processing units may execute a program to realize the above-described processing, and still other processing units may realize the above-described processing using both a logic circuit and by executing a program.
[0169] <Decryption process flow> Next, a description will be given of the processing executed by the decoding device 300. The decoding device 300 decodes the encoded data of the point cloud by executing a decoding process. An example of the flow of this decoding process will be described with reference to the flowchart in Fig. 23.
[0170] When the decoding process starts, in step S301, the coded data extraction unit 311 of the decoding device 300 acquires and holds the bitstream, and extracts coded data of geometry data and attribute data up to the LoD depth to be decoded.
[0171] In step S302, the geometry data decoding unit 312 decodes the coded data of the geometry data extracted in step S301 to generate geometry data (decoded result).
[0172] In step S303, the attribute data decoding unit 313 decodes the coded data of the attribute data extracted in step S301 to generate attribute data (decoded result).
[0173] In step S304, the point cloud generation unit 314 generates and outputs a point cloud (decoding result) using the geometry data (decoding result) generated in step S302 and the attribute data (decoding result) generated in step S303.
[0174] When the process of step S304 ends, the decoding process ends.
[0175] By performing the processing of each step in this manner, the decoding device 300 can correctly decode the coded data of the point cloud data generated by the coding device 200. In other words, scalable decoding of point cloud data can be more easily realized.
[0176] <4. Notes> <Hierarchization / reverse hierarchy method> In the above, lifting has been described as an example of a method for layering and delayering attribute data, but the method for layering and delayering attribute data may be other than lifting, such as RAHT.
[0177] <Control information> In the above embodiments, an enable flag has been described as an example of control information related to the present technology, but any other control information may be signaled.
[0178] <Surroundings / neighborhood> In this specification, the positional relationship such as "nearby" or "surrounding" may include not only a spatial positional relationship but also a temporal positional relationship.
[0179] <Computer> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs constituting the software are installed on a computer. Here, the term "computer" includes computers built into dedicated hardware, and general-purpose personal computers, etc., that can execute various functions by installing various programs.
[0180] FIG. 24 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program.
[0181] In a computer 900 shown in FIG. 24, a CPU (Central Processing Unit) 901, a ROM (Read Only Memory) 902, and a RAM (Random Access Memory) 903 are interconnected via a bus 904.
[0182] An input / output interface 910 is also connected to the bus 904. To the input / output interface 910, an input unit 911, an output unit 912, a storage unit 913, a communication unit 914, and a drive 915 are connected.
[0183] The input unit 911 includes, for example, a keyboard, a mouse, a microphone, a touch panel, an input terminal, etc. The output unit 912 includes, for example, a display, a speaker, an output terminal, etc. The storage unit 913 includes, for example, a hard disk, a RAM disk, a non-volatile memory, etc. The communication unit 914 includes, for example, a network interface. The drive 915 drives removable media 921 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0184] In a computer configured as above, the CPU 901 performs the above-described series of processes by, for example, loading a program stored in the storage unit 913 into the RAM 903 via the input / output interface 910 and the bus 904 and executing the program. The RAM 903 also stores data necessary for the CPU 901 to execute various processes as appropriate.
[0185] The program executed by the computer can be applied by recording it on removable media 921 such as package media, for example. In this case, the program can be installed in storage unit 913 via input / output interface 910 by inserting removable media 921 into drive 915.
[0186] This program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, digital satellite broadcasting, etc. In this case, the program can be received by the communication unit 914 and installed in the storage unit 913.
[0187] Alternatively, this program can be installed in advance in the ROM 902 or the storage unit 913 .
[0188] <Applicable targets of this technology> While the above describes the application of this technology to the encoding and decoding of point cloud data, this technology is not limited to these examples and can be applied to the encoding and decoding of 3D data of any standard. For example, when encoding and decoding mesh data, the mesh data may be converted into point cloud data, and the encoding and decoding may be performed using this technology. In other words, as long as it does not conflict with the above-described technology, various processes such as encoding and decoding methods, and specifications of various data such as 3D data and metadata, are arbitrary. Furthermore, as long as it does not conflict with the above-described technology, some of the above-described processes and specifications may be omitted.
[0189] Furthermore, although the encoding device 200 and the decoding device 300 have been described above as application examples of the present technology, the present technology can be applied to any configuration.
[0190] For example, this technology can be applied to various electronic devices, such as transmitters and receivers (e.g., television sets and mobile phones) used in satellite broadcasting, cable TV and other wired broadcasting, distribution over the Internet, and distribution to terminals via cellular communications, or devices (e.g., hard disk recorders and cameras) that record images on media such as optical disks, magnetic disks, and flash memories, or play images from these storage media.
[0191] Furthermore, for example, the present technology can also be implemented as a part of an apparatus, such as a processor (e.g., a video processor) as a system LSI (Large Scale Integration), a module (e.g., a video module) using multiple processors, a unit (e.g., a video unit) using multiple modules, or a set in which other functions are added to a unit (e.g., a video set).
[0192] Furthermore, for example, the present technology can also be applied to a network system configured with multiple devices. For example, the present technology may be implemented as cloud computing in which multiple devices share and collaborate on processing via a network. For example, the present technology may be implemented in a cloud service that provides image (video)-related services to any terminal, such as a computer, AV (Audio Visual) equipment, a portable information processing terminal, or an IoT (Internet of Things) device.
[0193] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all the components are contained in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.
[0194] <Fields and applications where this technology can be applied> Systems, devices, processing units, etc. to which the present technology is applied can be used in any field, such as transportation, medical care, crime prevention, agriculture, livestock farming, mining, beauty, factories, home appliances, weather, and nature monitoring. In addition, the applications thereof are also arbitrary.
[0195] <Other> In this specification, a "flag" refers to information for identifying multiple states, and includes not only information used to identify two states, true (1) or false (0), but also information capable of identifying three or more states. Therefore, the value that this "flag" can take may be, for example, two values, 1 / 0, or three or more values. In other words, the number of bits constituting this "flag" is arbitrary, and may be one bit or multiple bits. Furthermore, identification information (including flags) can be assumed not only to include the identification information in the bit stream, but also to include difference information of the identification information relative to certain reference information in the bit stream. Therefore, in this specification, "flag" and "identification information" include not only the information itself, but also difference information relative to the reference information.
[0196] Furthermore, various types of information (metadata, etc.) related to the coded data (bitstream) may be transmitted or recorded in any form as long as they are associated with the coded data. Here, the term "associate" means, for example, that one piece of data can be used (linked) when processing the other piece of data. In other words, data associated with each other may be combined into one piece of data or may be individual pieces of data. For example, information associated with coded data (image) may be transmitted over a transmission path separate from that of the coded data (image). Furthermore, for example, information associated with coded data (image) may be recorded on a recording medium separate from that of the coded data (image) (or on a different recording area of the same recording medium). Note that this "association" may refer not to the entire data, but to only part of the data. For example, an image and information corresponding to that image may be associated with each other in any unit, such as multiple frames, one frame, or a portion of a frame.
[0197] In this specification, terms such as "composite," "multiplex," "add," "integrate," "include," "store," "embed," "insert," and the like refer to combining multiple items into one, such as combining encoded data and metadata into one piece of data, and refer to one method of "associating" as described above.
[0198] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present technology.
[0199] For example, a configuration described as one device (or processing unit) may be divided and configured as multiple devices (or processing units). Conversely, configurations described above as multiple devices (or processing units) may be combined and configured as one device (or processing unit). Of course, configurations other than those described above may be added to the configuration of each device (or each processing unit). Furthermore, as long as the configuration and operation of the entire system are substantially the same, part of the configuration of one device (or processing unit) may be included in the configuration of another device (or other processing unit).
[0200] Furthermore, for example, the above-described program may be executed in any device, as long as the device has the necessary functions (functional blocks, etc.) and can obtain the necessary information.
[0201] Also, for example, each step of a single flowchart may be executed by one device, or may be shared and executed by multiple devices. Furthermore, when one step includes multiple processes, the multiple processes may be executed by one device, or may be shared and executed by multiple devices. In other words, multiple processes included in one step can be executed as multiple step processes. Conversely, processes described as multiple steps can be executed collectively as one step.
[0202] For example, the steps of a program executed by a computer may be executed in chronological order in the order described herein, or may be executed in parallel or individually at the required timing, such as when a call is made. In other words, as long as no contradiction occurs, the steps may be executed in an order different from the order described above. Furthermore, the steps of this program may be executed in parallel with the processing of another program, or may be executed in combination with the processing of another program.
[0203] Furthermore, for example, multiple technologies related to the present technology can be implemented independently and independently, as long as no contradiction occurs. Of course, any multiple technologies can also be implemented in combination. For example, part or all of the present technology described in any embodiment can be implemented in combination with part or all of the present technology described in another embodiment. Furthermore, part or all of any of the above-described present technologies can be implemented in combination with other technologies not described above.
[0204] The present technology can also be configured as follows. (1) A coding control unit that controls the use of scalable coding, which is a coding method that generates scalably decodable coded data, in the coding of point clouds that represent three-dimensional objects as a set of points, in combination with scaling coding, which is a coding method that involves changing the tree structure of geometry data. An information processing device comprising: (2) The scalable coding is lifting scalability, which encodes attribute data by lifting it using a reference structure similar to the tree structure of the geometry data. An information processing device according to (1). (3) The scaling encoding is geometry scaling, which scales and encodes the geometry data. An information processing device according to (1) or (2). (4) When the scalable coding is applied, the coding control unit controls to prohibit application of the scaling coding. An information processing device according to any one of (1) to (3). (5) The encoding control unit controls signaling of a scalable encoding enable flag, which is flag information regarding application of the scalable encoding, and a scaling encoding enable flag, which is flag information regarding application of the scaling encoding. (4) An information processing device according to the present invention. (6) When the encoding control unit signals the scalable encoding enable flag with a value indicating application of the scalable encoding, the encoding control unit controls to signal the scaling encoding enable flag with a value indicating non-application of the scaling encoding. (5) An information processing device according to (5). (7) When the scalable coding is applied, the coding control unit controls to signal, in a profile, the scalable coding enable flag having a value indicating application of the scalable coding and the scaling coding enable flag having a value indicating non-application of the scaling coding. An information processing device according to (5) or (6). (8) When the encoding control unit signals the scalable encoding enable flag having a value indicating application of the scalable encoding, the encoding control unit performs control so as to omit signaling of the scaling encoding enable flag. An information processing device according to any one of (5) to (7). (9) When the scaling coding is applied, the coding control unit controls to prohibit application of the scalable coding. An information processing device according to any one of (1) to (3). (10) The encoding control unit controls signaling of a scaling encoding enable flag, which is flag information regarding application of the scaling encoding, and a scalable encoding enable flag, which is flag information regarding application of the scalable encoding. (9) An information processing device according to (9). (11) When the encoding control unit signals the scaling encoding enable flag with a value indicating application of the scaling encoding, the encoding control unit controls to signal the scalable encoding enable flag with a value indicating non-application of the scalable encoding. (10) An information processing device according to (10). (12) When the scaling coding is applied, the coding control unit controls to signal, in a profile, the scaling coding enable flag having a value indicating application of the scaling coding and the scalable coding enable flag having a value indicating non-application of the scalable coding. The information processing device according to (10) or (11). (13) When the encoding control unit signals the scaling encoding enable flag having a value indicating application of the scaling encoding, the encoding control unit performs control so as to omit signaling of the scalable encoding enable flag. An information processing device according to any one of (10) to (12). (14) a geometry data encoding unit that encodes the geometry data of the point cloud under control of the encoding control unit and generates encoded data of the geometry data; an attribute data encoding unit that encodes attribute data of the point cloud under control of the encoding control unit and generates encoded data of the attribute data; The information processing device according to any one of (1) to (13), further comprising: (15) The geometry data encoding unit a selection unit that selects whether to apply scaling to the geometry data under control of the encoding control unit; a geometry scaling unit that scales and merges the geometry data when the selection unit selects application of scaling to the geometry data; an encoding unit that encodes the geometry data that has been scaled and merged by the geometry scaling unit when the selection unit has selected to apply scaling to the geometry data, and encodes the geometry data that has not been scaled and merged when the selection unit has selected not to apply scaling to the geometry data; The information processing device according to (14) is provided with: (16) The geometry data encoding unit a tree structure generating unit for generating a tree structure of the geometry data; Furthermore, The geometry scaling unit performs the scaling and the merging to update the tree structure generated by the tree structure generation unit. (15) An information processing device according to (15). (17) The attribute data encoding unit a selection unit that selects whether to apply scalable layering to the attribute data, which lifts the attribute data using a reference structure similar to the tree structure of the geometry data, under the control of the encoding control unit; a scalable layering unit that performs the scalable layering on the attribute data when the selection unit selects application of the scalable layering; an encoding unit that encodes the attribute data that has been scalably layered by the scalable layering unit when the selection unit selects application of the scalable layering, and encodes the attribute data that has not been scalably layered when the selection unit selects not application of the scalable layering; An information processing device according to any one of (14) to (16). (18) a geometry data decoding unit that decodes the coded data of the geometry data generated by the geometry data coding unit and generates the geometry data; a recolor processing unit that performs a recolor process on the attribute data using the geometry data generated by the geometry data decoding unit; Furthermore, The attribute data encoding unit encodes the attribute data that has been recolored by the recolor processing unit. An information processing device according to any one of (14) to (17). (19) A bitstream generating unit that generates a bitstream including the coded data of the geometry data generated by the geometry data coding unit and the coded data of the attribute data generated by the attribute data coding unit. The information processing device according to any one of (14) to (18), further comprising: (20) In encoding of point clouds, which represent three-dimensional objects as a set of points, the present invention provides control to prohibit the combined use of scalable encoding, which is an encoding method that generates encoded data that can be decoded in a scalable manner, and scaling encoding, which is an encoding method that involves changing the tree structure of geometry data. Information processing methods. [Explanation of symbols]
[0205] 200 encoding device, 201 encoding control unit, 211 geometry data encoding unit, 212 geometry data decoding unit, 213 point cloud generation unit, 214 attribute data encoding unit, 215 bitstream generation unit, 231 voxel generation unit, 232 tree structure generation unit, 233 selection unit, 234 geometry scaling unit, 235 encoding unit, 251 selection unit, 252 scalable layering processing unit, 253 layering processing unit, 254 quantization unit, 255 encoding unit, 300 decoding device, 311 encoded data extraction unit, 312 geometry data decoding unit, 313 attribute data decoding unit, 314 point cloud generation unit
Claims
1. a geometry data decoding unit that decodes coded data of geometry data that has been coded using geometry scaling, which is a coding method that involves changing the tree structure of the geometry data, in coding of a point cloud that represents a three-dimensional object as a set of points; an attribute data decoding unit that decodes the coded data of the attribute data without applying scalable coding, based on scalable coding information corresponding to the fact that scalable coding that enables attribute data to be decoded at an intermediate resolution of the tree structure of the geometry data is not applied in accordance with the application of the geometry scaling; and An information processing device comprising:
2. The scalable coding is lifting scalability, which encodes the attribute data by lifting it using a reference structure similar to the tree structure of the geometry data. The information processing device according to claim 1 .
3. The encoded geometry data is data in which points have been thinned out by the geometry scaling. The information processing device according to claim 1 .
4. the geometry data decoding unit decodes the coded data of the geometry data based on the signaled scalable coding enable flag, which is flag information regarding application of the scalable coding, and the signaled scaling coding enable flag, which is flag information regarding application of the geometry scaling; The attribute data decoding unit decodes the coded data of the attribute data based on the signaled scalable coding enable flag and the scaling coding enable flag. The information processing device according to claim 1 .
5. When the scalable coding enable flag is signaled with a value indicating application of the scalable coding, the scaling coding enable flag is signaled with a value indicating non-application of the geometry scaling. The information processing device according to claim 4 .
6. When the scalable coding is applied, the scalable coding enable flag is signaled in a profile with a value indicating the application of the scalable coding and a value indicating the non-application of the geometry scaling. The information processing device according to claim 4 .
7. If the scalable coding enable flag is signaled with a value indicating application of the scalable coding, signaling of the scaling coding enable flag is omitted. The information processing device according to claim 4 .
8. the geometry data decoding unit decodes encoded data of the geometry data that has been encoded without applying the geometry scaling; The attribute data decoding unit decodes the coded data of the attribute data by applying the scalable coding based on the scalable coding information corresponding to the fact that the scalable coding is applied in accordance with non-application of the geometry scaling. The information processing device according to claim 1 .
9. the geometry data decoding unit decodes the coded data of the geometry data based on the signaled scalable coding enable flag, which is flag information regarding application of the scalable coding, and the signaled scaling coding enable flag, which is flag information regarding application of the geometry scaling; The attribute data decoding unit decodes the coded data of the attribute data based on the signaled scalable coding enable flag and the scaling coding enable flag. The information processing device according to claim 8 .
10. When the scaling encoding enable flag is signaled with a value indicating application of the geometry scaling, the scalable encoding enable flag is signaled with a value indicating non-application of the scalable encoding. The information processing device according to claim 9 .
11. If the geometry scaling is applied, the scaling coding enable flag is signaled in a profile with a value indicating the application of the geometry scaling and a value indicating the non-application of the scalable coding. The information processing device according to claim 9 .
12. If the scaling encoding enable flag is signaled with a value indicating application of the geometry scaling, signaling of the scalable encoding enable flag is omitted. The information processing device according to claim 9 .
13. The geometry data decoding unit decodes the encoded data of the geometry data, and if the geometry scaling was applied during encoding, performs inverse scaling and inverse merging of the geometry data. The information processing device according to claim 1 .
14. The geometry data decoding unit generates a tree structure of the geometry data, and updates the tree structure by inverse scaling and inverse merging the geometry data. The information processing device according to claim 13.
15. The attribute data decoding unit decodes the encoded data of the attribute data, and when scalable layering is applied to the attribute data during encoding, the attribute data is lifted using a reference structure similar to the tree structure of the geometry data, and performs inverse scalable layering on the attribute data. The information processing device according to claim 1 .
16. The attribute data decoding unit decodes the coded data of the attribute data that has been recolored during coding. The information processing device according to claim 1 .
17. further comprising an encoded data extraction unit that extracts encoded data of the geometry data and encoded data of the attribute data from the bitstream, from the highest hierarchy to a desired hierarchy; the geometry data decoding unit decodes the coded data of the geometry data extracted from the bitstream; The attribute data decoding unit decodes the coded data of the attribute data extracted from the bitstream. The information processing device according to claim 1 .
18. The apparatus further includes a point cloud generation unit that generates a point cloud using the geometry data obtained by decoding the geometry data decoding unit and the attribute data obtained by decoding the attribute data decoding unit. The information processing device according to claim 1 .
19. The information processing device In encoding a point cloud that represents a three-dimensional object as a set of points, decoding encoded data of the geometry data that has been encoded using geometry scaling, which is an encoding method that involves changing the tree structure of the geometry data; decoding the coded data of the attribute data without applying scalable coding, based on scalable coding information corresponding to non-application of scalable coding that enables the attribute data to be decoded at an intermediate resolution of the tree structure of the geometry data, in accordance with application of the geometry scaling; An information processing method including:
Citation Information
Patent Citations
Scalable point cloud compression with transform, and corresponding decompression
US20170347122A1
US2019/80483A1
Method and apparatus for video coding
US20200107048A1
Image processing device and method
WO2020071115A1