Information processing device and method
The information processing apparatus and method address the challenge of scalable decoding of point cloud data by decoding geometry data with geometry scaling and attribute data without scalable encoding, facilitating easier scalable decoding.
Patent Information
- Application Number
- JP2025040224
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-06-22
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-06-08
AI Technical Summary
Existing methods for scalable decoding of point cloud data face challenges when geometry data is scaled and points are decimated, leading to mismatches between the reference structure of attribute data and the tree structure of geometry data.
An information processing apparatus and method that decodes geometry data encoded with geometry scaling and decodes attribute data without applying scalable encoding, based on scalable encoding information that corresponds to the non-application of scalable encoding in response to geometry scaling.
This approach enables easier realization of scalable decoding of point cloud data by avoiding the need for complex processing to align the reference structure of attribute data with the scaled geometry data tree structure.
Smart Images

Figure 2025090771000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an information processing apparatus and method, and more particularly to an information processing apparatus and method capable of more easily realizing scalable decoding of point cloud data.
Background Art
[0002] Conventionally, for example, an encoding method for 3D data representing a three-dimensional structure such as a point cloud has been considered (see, for example, Non-Patent Document 1). In addition, an encoding method has been proposed that enables scalable decoding of the encoded data of this point cloud (see, for example, Non-Patent Document 2). In the case of the method described in Non-Patent Document 2, scalable decoding is realized by making the reference structure of the attribute data the same as the tree structure of the geometric data.
[0003] By the way, when encoding such a point cloud, a method has been proposed in which geometric data is scaled and points are decimated (see, for example, Non-Patent Document 3).
Prior Art Documents
Non-Patent Documents
[0004]
Non-Patent Document 1
Non-Patent Document 2
[0005] However, as described in Non-Patent Document 3, when the geometry data is scaled and points are decimated, the tree structure of the geometry data changes. Therefore, a mismatch may occur between the reference structure of the attribute data and the tree structure of the geometry data, which may prevent scalable decoding. In other words, when scaling the geometry data as described in Non-Patent Document 3, in order to achieve scalable decoding, it is necessary to form the reference structure of the attribute data in correspondence with the scaling of the geometry data. That is, in order to encode the point cloud data so that scalable decoding is possible, the reference structure of the attribute data has to be changed depending on whether or not the geometry data is scaled, which may require complicated processing.
[0006] The present disclosure has been made in view of such a situation, and enables easier realization of scalable decoding of point cloud data. [Means for Solving the Problems]
[0007] An information processing apparatus according to an aspect of the present technology includes a geometry data decoding unit that decodes encoded data of the geometry data encoded by applying geometry scaling, which is an encoding method involving a change in the tree structure of geometry data, in encoding a point cloud representing a three-dimensional object as a set of points; and an attribute data decoding unit that decodes the encoded data of the attribute data without applying the scalable encoding based on scalable encoding information corresponding to the fact that scalable encoding that enables decoding the attribute data at an intermediate resolution of the tree structure of the geometry data is not applied in response to the application of the geometry scaling.
[0008] An information processing method according to an aspect of the present technology includes: the information processing apparatus decoding the encoded data of the geometry data encoded by applying geometry scaling, which is an encoding method involving a change in the tree structure of geometry data, in encoding a point cloud representing a three-dimensional object as a set of points; and decoding the encoded data of the attribute data without applying the scalable encoding based on scalable encoding information corresponding to the fact that scalable encoding that enables decoding the attribute data at an intermediate resolution of the tree structure of the geometry data is not applied in response to the application of the geometry scaling.
[0009] In an information processing apparatus and method according to an aspect of the present technology, in the encoding of a point cloud representing a three-dimensional object as a set of points, geometric scaling, which is an encoding method involving a change in the tree structure of geometric data, is applied and the encoded data of the geometric data is decoded. Based on scalable encoding information corresponding to the fact that scalable encoding, which enables the decoding of attribute data at an intermediate resolution of the tree structure of the geometric data, is not applied in response to the application of geometric scaling, the encoded data of the attribute data is decoded without applying the scalable encoding.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Mode for Carrying Out the Invention
[0011] Hereinafter, a mode for carrying out the present disclosure (hereinafter referred to as the embodiment) will be described. The description will be made in the following order. 1. Symbolization control 2. First Embodiment (Symbolization Device) 3. Second Embodiment (Decoding Device) 4. Supplementary Note
[0012] <1. Symbolization control> <Literature etc. Supporting Technical Content and Technical Terms> The scope disclosed in this technology includes not only the content described in the embodiment but also the content described in the following non-patent documents that were known at the time of filing.
[0013] Non-Patent Document 1: (described above) Non-Patent Document 2: (above-mentioned) Non-Patent Document 3: (above-mentioned) Non-Patent Document 4: Khaled Mammou, Alexis Tourapis, Jungsun Kim, Fabrice Robinet, Valery Valentin, Yeping Su, "Lifting Scheme for Lossy Attribute Encoding in TMC1", ISO / IEC JTC1 / SC29 / WG11 MPEG2018 / m42640, April 2018, San Diego, US
[0014] That is, the content described in the above-mentioned non-patent documents, the content of other documents referred to in the above-mentioned non-patent documents, etc. also serve as the basis for judging the support requirements.
[0015] <Point cloud> Conventionally, there have been 3D data such as a point cloud (Point cloud) that represents a three-dimensional structure by point position information, attribute information, etc., and a mesh (Mesh) that is composed of vertices, edges, and faces and defines a three-dimensional shape using a polygonal representation.
[0016] For example, in the case of a point cloud, a three-dimensional structure (a three-dimensional object) is represented by a large number of points. The data of the point cloud (also referred to as point cloud data) is composed of the geometry data (also referred to as position information) of each point and the attribute data (also referred to as attribute information). The attribute data can include any information. For example, color information, reflectivity information, normal information, etc. of each point may be included in the attribute data. In this way, the point cloud data has a relatively simple data structure and can represent any three-dimensional structure with sufficient accuracy by using a sufficiently large number of points.
[0017] <Quantization of position information using voxels> Since such point cloud data has a relatively large data volume, in order to compress the data volume by encoding or the like, an encoding method using voxels has been considered. A voxel is a three-dimensional region for quantizing geometric data (position information).
[0018] That is, a three-dimensional region (also referred to as a bounding box) that encloses the point cloud is divided into small three-dimensional regions called voxels, and for each voxel, it is made to indicate whether it encloses a point or not. By doing so, the position of each point is quantized in voxel units. Therefore, by converting point cloud data into such voxel data (also referred to as voxel data), an increase in the amount of information can be suppressed (typically, the amount of information is reduced).
[0019] For example, as shown in A of FIG. 1, assume that the bounding box 10 is divided into a plurality of voxels 11-1 shown as small squares. Here, for the sake of simplifying the explanation, the three-dimensional space is described as a two-dimensional plane. That is, actually, the bounding box 10 is a three-dimensional space region, and the voxel 11-1 is a small region of a rectangular parallelepiped (including a cube). For each voxel 11-1, the geometric data (position information) of each point 12-1 of the point cloud data shown as a black circle in A of FIG. 1 is corrected so as to be arranged. That is, the geometric data is quantized in units of this voxel. Note that all the squares within the bounding box 10 in A of FIG. 1 are voxels 11-1, and all the black circles shown in A of FIG. 1 are points 12-1.
[0020] <Octree structure> Furthermore, a method was conceived to enable scalable decoding of geometric data by structuring the geometric data into a tree structure. That is, by enabling decoding of nodes from the topmost layer to any layer of the tree structure, it becomes possible not only to restore geometric data at the highest resolution (lowermost layer), but also to restore geometric data at a lower resolution (intermediate layer). That is, it is possible to decode at any resolution without decoding information of unnecessary layers (resolutions).
[0021] This tree structure can be of any kind. For example, there are KD trees (KD Trees), octrees, etc. An octree is an octal tree and is suitable for dividing a three-dimensional space region (two-way division in each of the x, y, and z directions). That is, as described above, it is suitable for a structure that divides the bounding box 10 into a plurality of voxels.
[0022] For example, one voxel is divided into two in each of the x, y, and z directions (that is, divided into eight), and voxels of one lower layer (also referred to as LoD) are formed. In other words, two voxels arranged in each of the x, y, and z directions (that is, eight voxels) are integrated to form a voxel of one upper layer (LoD). By recursively repeating such a structure, an octree can be constructed using voxels.
[0023] And in voxel data, it is indicated whether each voxel contains a point. In other words, in voxel data, the position of a point is represented at the resolution of the voxel size. Therefore, by constructing an octree using voxel data, scalability of the resolution of geometric data can be realized. That is, it is possible to more easily construct an octree of geometric data by structuring voxel data than by structuring points scattered at arbitrary positions.
[0024] For example, in the case of A in FIG. 1, since it is two-dimensional, four voxels 11-1 arranged vertically, horizontally, and diagonally as in B of FIG. 1 are integrated to form a voxel 11-2 at the first layer indicated by the thick line. Then, geometric data is quantized using this voxel 11-2. That is, when a point 12-1 (in A of FIG. 1) exists within the voxel 11-2, by correcting its position, the point 12-1 is converted into a point 12-2 corresponding to the voxel 11-2. In FIG. 1B, only one voxel 11-1 is labeled, but all the squares indicated by the dotted lines within the bounding box 10 of FIG. 1B are voxels 11-1. Similarly, in FIG. 1B, only one voxel 11-2 is labeled, but all the squares indicated by the thick lines within the bounding box 10 of FIG. 1B are voxels 11-2. Similarly, in FIG. 1B, only one point 12-2 is labeled, but all the black circles shown in FIG. 1B are points 12-2.
[0025] Similarly, as in C of FIG. 1, four voxels 11-2 arranged vertically, horizontally, and diagonally are integrated to form a voxel 11-3 at the next higher level indicated by the thick line. Then, geometric data is quantized using this voxel 11-3. That is, when a point 12-2 (in B of FIG. 1) exists within the voxel 11-3, by correcting its position, the point 12-2 is converted into a point 12-3 corresponding to the voxel 11-3. In FIG. 1C, only one voxel 11-2 is labeled, but all the squares indicated by the dotted lines within the bounding box 10 of FIG. 1C are voxels 11-2. Similarly, in FIG. 1C, only one voxel 11-3 is labeled, but all the squares indicated by the thick lines within the bounding box 10 of FIG. 1C are voxels 11-3. Similarly, in FIG. 1C, only one point 12-3 is labeled, but all the black circles shown in FIG. 1C are points 12-3.
[0026] Similarly, as shown by D in FIG. 1, four voxels 11-3 arranged vertically, horizontally, and diagonally are integrated to form one higher-order voxel 11-4 indicated by a thick line. Then, the geometric data is quantized using this voxel 11-4. That is, when a point 12-3 (C in FIG. 1) exists within the voxel 11-4, by correcting its position, the point 12-3 is converted into a point 12-4 corresponding to the voxel 11-4. In FIG. 1D, only one voxel 11-3 is labeled, but all the squares indicated by the dotted lines within the bounding box 10 in FIG. 1D are voxels 11-3.
[0027] By doing so, the geometric data is structured like a tree (octree).
[0028] <Lifting> On the other hand, when encoding the attribute data, assuming the geometric data is known including the degradation due to encoding, the encoding is performed using the positional relationship between the points. As such a method for encoding the attribute data, methods using RAHT (Region Adaptive Hierarchical Transform) or a transform called lifting as described in Non-Patent Document 4 are considered. By applying these techniques, it is also possible to hierarchically structure (tree-structure) the reference structure (reference relationship) of the attribute data like the octree of the geometric data.
[0029] For example, in the case of lifting, the attribute data of each point is encoded as a difference value from a predicted value derived using the attribute data of other points. Then, the points for deriving the difference value (that is, deriving the predicted value) are hierarchically selected.
[0030] For example, in the hierarchy shown at A in FIG. 2, among each point (P0 to P9) indicated by a circle, points P7, P8, and P9 indicated by white circles are selected as prediction points for which predicted values are derived, and the other points P0 to P6 are set to be selected as reference points that are points to which attribute data is referred when deriving their predicted values. That is, in this hierarchy, for each of prediction points P7 to P9, a difference value between the attribute data and its predicted value is derived.
[0031] Note that in FIG. 2, for simplicity of explanation, the three-dimensional space is described as a two-dimensional plane. That is, actually, each of points P0 to P6 is arranged in a three-dimensional space.
[0032] Each arrow in A of FIG. 2 indicates a reference relationship when deriving a predicted value. For example, the predicted value of prediction point P7 is derived by referring to the attribute data of reference points P0 and P1. Also, the predicted value of prediction point P8 is derived by referring to the attribute data of reference points P2 and P3. Further, the predicted value of prediction point P9 is derived by referring to the attribute data of reference points P4 to P6. And for each of prediction points P7 to P9, a difference value between the predicted value calculated as described above and the attribute data is derived.
[0033] In the one-level higher hierarchy, as shown in B of FIG. 2, for the points (P0 to P6) selected as reference points in the hierarchy of A in FIG. 2 (one-level lower hierarchy), the same classification (sorting) of prediction points and reference points as in the case of the hierarchy of A in FIG. 2 is performed.
[0034] For example, in B of FIG. 2, points P1, P3, and P6 indicated by gray circles are selected as prediction points, and points P0, P2, P4, and P5 indicated by black circles are selected as reference points. That is, in this hierarchy, for each of prediction points P1, P3, and P6, a difference value between the attribute data and its predicted value is derived.
[0035] Each arrow at B in FIG. 2 indicates a reference relationship when deriving a predicted value. For example, the predicted value of prediction point P1 is derived by referring to the attribute data of reference points P0 and P2. Also, the predicted value of prediction point P3 is derived by referring to the attribute data of reference points P2 and P4. Furthermore, the predicted value of prediction point P6 is derived by referring to the attribute data of reference points P4 and P5. And for each of prediction points P1, P3, and P6, a difference value between the predicted value calculated as described above and the attribute data is derived.
[0036] In the one-higher-level hierarchy, as shown in C of FIG. 2, classification (sorting) of the points (P0, P2, P4, P5) selected as reference points in the hierarchy (one-lower-level hierarchy) of B in FIG. 2, derivation of the predicted value of each prediction point, and derivation of the difference value between the predicted value and the attribute data are performed.
[0037] By recursively repeating such classification for the reference points in the one-lower-level hierarchy, the reference structure of the attribute data is hierarchized.
[0038] <Classification of Points> The procedure for classifying (sorting) points in such lifting will be described more specifically. In lifting, the classification of points is performed in the order from the lower layer to the upper layer as described above. In each layer, first, each point is sorted in Morton code order. Next, the point at the head of the column of points arranged in that Morton code order is selected as the reference point. Next, points (neighboring points) located in the vicinity of the reference point are searched for, and the searched points (neighboring points) are set as prediction points (also referred to as index points).
[0039] For example, as shown in FIG. 3, points are searched within a circle 22 with a radius R centered on a reference point 21 to be processed. This radius R is preset for each layer. In the example of FIG. 3, points 23-1 to 23-4 are detected and set as prediction points.
[0040] Note that in FIG. 3, for simplicity of explanation, the three-dimensional space is described as a two-dimensional plane. That is, actually, each point is arranged in a three-dimensional space, and the search for points is performed within a spherical region with a radius R.
[0041] Next, the same classification is performed for the remaining points. That is, among the points that have not been selected as reference points or prediction points at the current time, the point at the head in Morton code order is selected as the reference point, the points near the reference point are searched, and set as prediction points.
[0042] When the above processing is repeated until all points are classified, the processing for that layer ends, and the processing target moves to the next higher layer. Then, for that layer, the above procedure is repeated. That is, each point selected as a reference point in the next lower layer is sorted in Morton code order and classified into reference points and prediction points as described above. By repeating the above processing, the reference structure of the attribute data is hierarchically organized.
[0043] <Derivation of Prediction Value> Also, as described above in the case of lifting, the predicted value of the attribute data of the prediction point is derived using the attribute data of the reference points around the prediction point. For example, as shown in FIG. 4, it is assumed that the predicted value of the prediction point Q(i,j) is derived with reference to the attribute data of the reference points P1 to P3.
[0044] Note that in FIG. 4, for simplicity of explanation, the three-dimensional space is described as a two-dimensional plane. That is, actually, each point is arranged in a three-dimensional space.
[0045] In this case, as shown in the following formula (1), the attribute data of each reference point is weighted and integrated by a weight value (α(P, Q(i, j))) corresponding to the reciprocal of the distance between the prediction point and the reference point (actually the distance in three-dimensional space), and then derived. Here, A(P) represents the attribute data of point P.
[0046]
Equation
[0047] In the method described in Non-Patent Document 4, the position information of the highest resolution (i.e., the lowest layer) was used to derive the distance between the prediction point and the reference point. <Quantization> Also, after the attribute data is hierarchically structured as described above, it is quantized and encoded. During this quantization, the attribute data (difference value) of each point is weighted as in the example of FIG. 5 according to the hierarchical structure. This weight value (Quantization Weight) W is derived for each point using the weight value of the lower layer as shown in FIG. 5. Note that this weight value can also be used in lifting (hierarchical structuring of attribute data) to improve the compression efficiency.
[0048] <Discrepancy in tree structure> In the case of lifting described in Non-Patent Document 4, as described above, the method of hierarchical structuring of the reference structure of attribute data is different from the case of tree structuring (e.g., octree structuring) of geometric data. Therefore, it is not guaranteed that the reference structure of attribute data matches the tree structure of geometric data. For this reason, in order to decode the attribute data, it was necessary to decode the geometric data to the lowest layer regardless of the layer. That is, it was difficult to achieve scalable decoding of point cloud data without decoding unnecessary information.
[0049] <Realization of scalable decoding of point cloud data> Therefore, as described in Non-Patent Document 2, a method has been proposed to make the reference structure of attribute data the same as the tree structure of geometric data. More specifically, when constructing the reference structure of attribute data, predicted points are selected such that there are also points in the voxel of the upper layer to which the voxel where the point of the current layer exists belongs. By doing so, it is possible to decode the point cloud data at a desired resolution without decoding unnecessary information. That is, scalable decoding of point cloud data can be realized.
[0050] For example, as shown in A of FIG. 6, in a bounding box 100 which is a predetermined three-dimensional space region, it is assumed that points 102-1 to 102-9 are arranged for each voxel 101-1 of a predetermined layer. In FIG. 6, for simplicity of explanation, the three-dimensional space is described as a two-dimensional plane. That is, actually, the bounding box is a three-dimensional space region, and the voxel is a small region of a rectangular parallelepiped (including a cube). The points are arranged in three-dimensional space.
[0051] When there is no need to distinguish and explain voxels 101-1 to 101-3 from each other, they are referred to as voxel 101. Also, when there is no need to distinguish and explain points 102-1 to 102-9 from each other, they are referred to as point 102.
[0052] In this layer, as shown in B of FIG. 6, points 102-1 to 102-9 are classified into predicted points and reference points such that there are also points in voxel 101-2 of the upper layer of voxel 101-1 where points 102-1 to 102-9 exist. In the example of B of FIG. 6, points 102-3, 102-5, and 102-8 shown as white circles are set as predicted points, and the other points are set as reference points.
[0053] Similarly, in one higher-level hierarchy, points are classified into prediction points and reference points such that there are points in the voxel 101-3 at one higher level of the voxel 101-2 where points 102-1, 102-2, 102-4, 102-6, 102-7, and 102-9 exist (C in FIG. 6). In the example of C in FIG. 6, points 102-1, 102-4, and 102-7 indicated by gray circles are set as prediction points, and the other points are set as reference points.
[0054] By doing so, as shown in D of FIG. 6, the voxels 101-3 where points 102 exist in the lower layer are hierarchically arranged such that there is one point 102. Such processing is performed for each layer. That is, by performing such processing when constructing the reference structure of the attribute data (when classifying the prediction points and reference points in each layer), the reference structure of the attribute data can be made the same as the geometric data tree structure (octree).
[0055] Decoding is performed, for example, in the reverse order of FIG. 6 as shown in FIG. 7. For example, as shown in A of FIG. 7, in a predetermined bounding box 100, it is assumed that points 102-2, 102-6, and 102-9 are arranged for each voxel 101-3 at a predetermined level (the same state as D in FIG. 6). Also in FIG. 7, for simplicity of explanation, the three-dimensional space is described as a two-dimensional plane. That is, actually, the bounding box is a three-dimensional space region, the voxel is a small region of a rectangular parallelepiped (including a cube), and the points are arranged in three-dimensional space.
[0056] In the next lower layer, as shown in B of FIG. 6, using the attribute data of points 102-2, 102-6, and 102-9 for each voxel 101-3, the predicted values of points 102-1, 102-4, and 102-7 are derived and added to the difference value to restore the attribute data of point 102 for each voxel 101-2 (the same state as C in FIG. 6).
[0057] Furthermore, similarly in the next lower layer, as shown in C of FIG. 7, using the attribute data of points 102-1, 102-2, 102-4, 102-6, 102-7, and 102-9 for each voxel 101-2, the predicted values of points 102-3, 102-5, and 102-8 are derived and added to the difference value to restore the attribute data (the same state as B in FIG. 6).
[0058] By doing so, as shown in D of FIG. 7, the attribute data of point 102 for each voxel 101-1 is restored (the same state as A in FIG. 6). That is, similar to the case of the octree, the attribute data of each layer can be restored using the attribute data of the upper layer.
[0059] By doing so, the reference structure (hierarchical structure) of the attribute data can be associated with the tree structure (hierarchical structure) of the geometry data. Therefore, since the geometry data corresponding to each attribute data can be obtained even at the intermediate resolution, the geometry data and the attribute data can be correctly decoded at that intermediate resolution. That is, scalable decoding of the point cloud data can be realized.
[0060] <Geometry Scaling> By the way, as described in Non-Patent Document 3, a method called Geometry Scaling has been proposed in which geometric data is scaled and points are decimated during the encoding of point clouds. In Geometry Scaling, the geometric data of nodes is quantized during encoding. By this process, points can be decimated according to the characteristics of the region to be encoded.
[0061] For example, as shown in A of FIG. 8, six points (point1 to point6) with X coordinates of 100, 101, 102, 103, 104, and 105 respectively are to be processed. Here, for simplicity of explanation, only the X coordinate will be described. That is, actually, since the points are arranged in a three-dimensional space, the same processing as in the case of the X coordinate described below will be performed for the Y coordinate and the Z coordinate as well.
[0062] When Geometry Scaling is applied to such six points, the X coordinates of each point are scaled as shown in the table of B of FIG. 8 according to the value of the quantization parameter baseQP. When baseQP = 0, no scaling is performed, so the X coordinates of each point remain the same as the coordinates shown in A of FIG. 8. For example, when baseQP = 4, the X coordinate of point2 is scaled from 101 to 102, the X coordinate of point4 is scaled from 103 to 104, and the X coordinate of point6 is scaled from 105 to 106.
[0063] When the X coordinates of multiple points overlap due to this scaling, they can be merged. For example, when mergeDuplicatePoint = 1, such overlapping points are merged and become one point. mergeDuplicatePoint is flag information indicating whether to merge such overlapping points. That is, due to such merging, points are thinned out (the number of points is reduced). For example, in each row of the table in B of FIG. 8, such merging can thin out the points shown in gray. For example, when baseQP = 4, for the four points (point2 to point5) indicated by the thick line, the number of points is reduced to half.
[0064] In order to realize scalable decoding of point cloud data, when hierarchizing attribute data as described in Non-Patent Document 2, an estimation of the geometric data tree structure is performed based on the decoding result of the encoded data of the geometric data. That is, a reference structure of the attribute data is formed so as to be the same as the estimated geometric data tree structure.
[0065] <Discrepancy in the tree structure due to geometric scaling> However, when geometric scaling is applied as described above, in the decoding result of the encoded data of the geometric data, there is a possibility that points are thinned out. If points are thinned out, it may not be possible to estimate the actual geometric data tree structure. That is, there is a possibility of a discrepancy between the estimated tree structure and the actual tree structure (the tree structure corresponding to the points before thinning). In such a case, there was a possibility that scalable decoding of the point cloud data could not be realized.
[0066] Taking the case where baseQP = 4 in B of FIG. 8 as an example. Assume that the tree structure corresponding to each point (point1 to point6) in the table shown in A of FIG. 8 is the tree structure shown in FIG. 9. Nodes 121-1 to 121-6 of this tree structure correspond to each point (point1 to point6) in the table shown in A of FIG. 8. That is, assume that the X coordinates of nodes 121-1 to 121-6 in the lowest layer (the third layer) of this tree structure are 100, 101, 102, 103, 104, and 105 respectively.
[0067] The tree structure in FIG. 9 applies the following two rules. The first rule is that in the layer one level above the third layer (the second layer), the nodes with X coordinates of 100 and 101 in the third layer, the nodes with X coordinates of 102 and 103 in the third layer, and the nodes with X coordinates of 104 and 105 in the third layer are grouped together respectively. The second rule is that in the topmost layer (the first layer), all nodes in the second layer are grouped together.
[0068] That is, nodes 121-1 and 121-2 in the third layer belong to node 122-1 in the second layer. Nodes 121-3 and 121-4 in the third layer belong to node 122-2 in the second layer. Nodes 121-5 and 121-6 in the third layer belong to node 122-3 in the second layer. Also, nodes 122-1 to 122-3 in the second layer belong to node 123 in the first layer.
[0069] When geometric scaling is performed, the X coordinate of node 121-2 is scaled from 101 to 102, the X coordinate of node 121-4 is scaled from 103 to 104, and the X coordinate of node 121-6 is scaled from 105 to 106. As a result, the X coordinates of node 121-2 and node 121-3 overlap and are merged, and node 121-3 is thinned out. Similarly, node 121-4 and node 121-5 are merged and node 121-5 is thinned out. As a result, in the third layer, nodes 121-1, 121-2, 121-4, and 121-6 are encoded.
[0070] In the encoding of attribute data, as the decoded result of the encoded data of geometric data, the highest-resolution geometric data is obtained. That is, in the case of the example in FIG. 9, nodes 121-1, 121-2, 121-4, and 121-6 are obtained. Then, the tree structure of the geometric data is estimated from these four points (four nodes).
[0071] In that case, in order to form a similar tree structure, the same rules as in the example of FIG. 9 are applied. That is, in the layer one level above the third layer (the second layer), the nodes with X coordinates of 100 and 101 in the third layer, the nodes with X coordinates of 102 and 103 in the third layer, and the nodes with X coordinates of 104 and 105 in the third layer are grouped together respectively. Also, in the topmost layer (the first layer), all the nodes in the second layer are grouped together.
[0072] As a result, a tree structure as shown in FIG. 10 is estimated. That is, node 121-1 in the third layer belongs to node 124-1 in the second layer. Node 121-2 in the third layer belongs to node 124-2 in the second layer. Node 121-4 in the third layer belongs to node 124-3 in the second layer. Node 121-6 in the third layer belongs to node 124-4 in the second layer. Also, nodes 124-1 to 124-4 in the second layer belong to node 125 in the first layer.
[0073] As is clear from comparing FIG. 9 and FIG. 10, these tree structures do not match. For example, assume that the value of the parameter skipOctreeLayer indicating the layer to be decoded is "1" (skipOctreeLayer = 1). That is, when decoding the second layer, as shown in FIG. 9, three nodes are obtained for the geometric data. In contrast, as shown in FIG. 10, four nodes are obtained for the attribute data. Thus, due to the mismatch of the tree structures, it was difficult to obtain the decoded result of the desired layer without decoding unnecessary information. That is, it was difficult to achieve scalable decoding of the point cloud data.
[0074] In other words, in order to achieve scalable decoding of point cloud data, it is necessary to estimate the tree structure in consideration of such geometric scaling. That is, complicated processing such as changing the tree structure estimation method depending on whether geometric scaling is performed or not was required. Also, there was a risk of increased cost by preparing multiple estimation methods.
[0075] Also, in geometric scaling, when thinning out duplicate points, the thinning method (which of the multiple overlapping points to thin out) is not defined and depends on the design. That is, the tree structure estimation method corresponding to geometric scaling has to be newly designed according to the design of the geometric scaling, and there was a risk of increased cost.
[0076] <Limitations of Encoding Methods> Therefore, as described in the top row of the table shown in FIG. 11, the applicable encoding methods are restricted, and the combined use of scalable decodable encoding and the process of updating the tree structure of geometric data is prohibited.
[0077] That is, in the encoding of a point cloud that represents a three-dimensional shape object as a set of points, control is performed so as to prohibit the combined use of scalable encoding, which is an encoding method for generating scalable decodable encoded data, and scaling encoding, which is an encoding method involving a change in the tree structure of geometric data.
[0078] For example, in an information processing apparatus, in the encoding of a point cloud that represents a three-dimensional shape object as a set of points, an encoding control unit is provided to control so as to prohibit the combined use of scalable encoding, which is an encoding method for generating scalable decodable encoded data, and scaling encoding, which is an encoding method involving a change in the tree structure of geometric data.
[0079] That is, when applying scalable encoding, the application of scaling encoding is prohibited, and when applying scaling encoding, the application of scalable encoding is prohibited. By doing so, when performing scalable encoding, it is possible to suppress the occurrence of a mismatch between the reference structure of the attribute data and the tree structure of the geometry data. Therefore, scalable decoding of the point cloud data can be more easily realized.
[0080] Note that the scalable encoding may be any method as long as it is an encoding method that generates encodable data that can be decoded in a scalable manner. For example, it may be Lifting Scalability that lifts and encodes attribute data with a reference structure similar to the tree structure of geometry data as described in Non-Patent Document 2. That is, the encoding may be controlled so as to prohibit the combined use of Lifting Scalability and scaling encoding.
[0081] Also, the scaling encoding may be any method as long as it is an encoding method that involves changing the tree structure of the geometry data. For example, it may be Geometry Scaling that scales and encodes geometry data as described in Non-Patent Document 3. That is, the control may be such that the combined use of scalable encoding and geometry scaling is prohibited.
[0082] Of course, as described in the second row from the top of the table shown in FIG. 11, the combined use of Lifting Scalability and Geometry Scaling may be prohibited (Method 1).
[0083] At that time, when the encoding control unit applies scalable encoding, it may control to prohibit the application of scaling encoding. For example, as described in the third row from the top of the table shown in FIG. 11, when applying lifting scalability, the application of geometric scaling may be prohibited (Method 1-1).
[0084] Also, when the encoding control unit applies scaling encoding, it may control to prohibit the application of scalable encoding. For example, as described in the fourth row from the top of the table shown in FIG. 11, when applying geometric scaling, the application of lifting scalability may be prohibited (Method 1-2).
[0085] Also, in order to perform such control, for example, as described in the fifth row from the top of the table shown in FIG. 11, flag information (permission flag or prohibition flag) indicating whether to permit (or prohibit) the application of scalable encoding and scaling encoding may be set (Method 2).
[0086] For example, the encoding control unit may control the signaling of a scalable encoding enable flag, which is flag information regarding the application of scalable encoding, and a scaling encoding enable flag, which is flag information regarding the application of scaling encoding.
[0087] For example, as described in the sixth row from the top of the table shown in FIG. 11, such a restriction may be defined in the semantics (Method 2-1). For example, in the semantics, when signaling the scalable coding enable flag of a value indicating the application of scalable coding, it may be defined to signal the scalable coding enable flag of a value indicating the non-application of scalable coding, and the coding control unit may perform signaling according to the semantics. Also, in the semantics, when signaling the scalable coding enable flag of a value indicating the application of scalable coding, it may be defined to signal the scalable coding enable flag of a value indicating the non-application of scalable coding, and the coding control unit may perform signaling according to the semantics.
[0088] FIG. 12 shows an example of the semantics in that case. For example, in the semantics 161 shown in FIG. 12, it is defined that when the value of geom_scaling_enabled_flag is greater than 0, the value of lifting_scalability_enabled_flag must be set to 0. Here, geom_scaling_enabled_flag is flag information indicating whether to apply geometry scaling. When geom_scaling_enabled_flag = 1, geometry scaling is applied. Also, when geom_scaling_enabled_flag = 0, geometry scaling is not applied. lifting_scalability_enabled_flag is flag information indicating whether to apply lifting scalability. When lifting_scalability_enabled_flag = 1, lifting scalability is applied. Also, when lifting_scalability_enabled_flag = 0, lifting scalability is not applied.
[0089] That is, in this semantics 161, when geometric scaling is applied, the application of lifting scalability is prohibited. Conversely, in the semantics, when lifting scalability is applied, the application of geometric scaling may be prohibited. That is, in the semantics, it may be stipulated that when the value of the lifting_scalability_enabled_flag is greater than 0, the value of the geom_scaling_enabled_flag must be set to 0.
[0090] Also, for example, as described in the seventh row from the top of the table shown in FIG. 11, such signaling based on such a restriction may be performed in the profile (Method 2-2). For example, when applying scalable coding, the coding control unit may control to signal, in the profile, a scalable coding enable flag with a value indicating the application of scalable coding and a scalable coding enable flag with a value indicating the non-application of scalable coding. Also, when applying scalable coding, the coding control unit may control to signal, in the profile, a scalable coding enable flag with a value indicating the application of scalable coding and a scalable coding enable flag with a value indicating the non-application of scalable coding.
[0091] FIG. 13 shows an example of a profile when applying lifting scalability. For example, in the profile 162 shown in FIG. 13, lifting_scalability_enabled_flag = 1 and geom_scaling_enabled_flag = 0 are signaled. That is, in this profile 162, it is shown that lifting scalability is applied and geometric scaling is not applied.
[0092] In addition, in the profile when applying geometric scaling, it may be signaled that geom_scaling_enabled_flag = 1 and lifting_scalability_enabled_flag = 0.
[0093] Furthermore, for example, as described in the bottom row of the table shown in FIG. 11, such a restriction may be defined in the syntax (Method 2-3). For example, when signaling the scalable coding enable flag of the value indicating the application of scalable coding, the coding control unit may control to omit signaling the scaling coding enable flag according to the syntax as described above. Also, when signaling the scalable coding enable flag of the value indicating the application of scalable coding, the coding control unit may control to omit signaling the scaling coding enable flag according to the syntax as described above.
[0094] An example of the syntax in that case is shown in FIG. 14. For example, in Syntax 163 shown in FIG. 14, the lifting_scalability_enabled_flag is signaled only when the geom_scaling_enabled_flag is not signaled (that is, when the value is "0" and geometric scaling is not applied). That is, in this case, lifting scalability can be applied. In other words, when geometric scaling is applicable (that is, when the geom_scaling_enabled_flag is signaled), the lifting_scalability_enabled_flag is not signaled (that is, the value is "0" and lifting scalability is not applied).
[0095] Conversely, geom_scaling_enabled_flag may be signaled only when lifting_scalability_enabled_flag is not signaled (i.e., when the value is "0" and lifting scalability is not applicable). That is, in this case, geometry scaling can be applied. In other words, when lifting scalability is applicable (i.e., when lifting_scalability_enabled_flag is signaled), geom_scaling_enabled_flag may not be signaled (i.e., the value is "0" and geometry scaling is not applicable).
[0096] <2. First Embodiment> <Encoding Device> Next, an apparatus that applies the present technology described above in <1. Encoding Control> will be described. FIG. 15 is a block diagram showing an example of the configuration of an encoding device, which is an aspect of an information processing apparatus to which the present technology is applied. The encoding device 200 shown in FIG. 15 is an apparatus that encodes point clouds (3D data). The encoding device 200 encodes the point cloud by applying the present technology described above in <1. Encoding Control>.
[0097] Note that FIG. 15 shows the main components such as the processing unit and the data flow, and what is shown in FIG. 15 is not necessarily all. That is, in the encoding device 200, there may be a processing unit that is not shown as a block in FIG. 15, or there may be a process or data flow that is not shown as an arrow or the like in FIG. 15.
[0098] As shown in FIG. 15, the encoding device 200 includes an encoding control unit 201, a geometry data encoding unit 211, a geometry data decoding unit 212, a point cloud generation unit 213, an attribute data encoding unit 214, and a bitstream generation unit 215.
[0099] The symbol control unit 201 performs processes related to the control of the encoding of point cloud data. For example, the symbol control unit 201 controls the geometry data encoding unit 211. Also, the symbol control unit 201 controls the attribute data encoding unit 214. For example, as described above in <1. Symbol Control>, the symbol control unit 201 controls these processing units so as to prohibit the combined use of scalable encoding, which is an encoding method for generating scalable decodable encoded data, and scaling encoding, which is an encoding method involving a change in the tree structure of geometry data. Further, the symbol control unit 201 controls the bitstream generation unit 215 and controls the signaling of a scalable encoding enable flag (for example, Lifting_scalability_enabled_flag), which is flag information related to the application of scalable encoding, and a scaling encoding enable flag (for example, geom_scaling_enabled_flag), which is flag information related to the application of scaling encoding.
[0100] The geometry data encoding unit 211 encodes the geometry data (position information) of the point cloud (3D data) input to the encoding device 200 and generates the encoded data thereof. This encoding method is arbitrary. For example, processes such as filtering for noise suppression (denoising) and quantization may be performed. However, the geometry data encoding unit 211 performs this encoding in accordance with the control of the symbol control unit 201. That is, in this encoding, the geometry data encoding unit 211 applies geometry scaling in accordance with the control of the symbol control unit 201. The geometry data encoding unit 211 supplies the encoded data of the generated geometry data to the geometry data decoding unit 212 and the bitstream generation unit 215.
[0101] The geometry data decoding unit 212 acquires the encoded data of the geometry data supplied from the geometry data encoding unit 211 and decodes the encoded data. This decoding method can be arbitrary as long as it corresponds to the encoding by the geometry data encoding unit 211. For example, processing such as filtering for noise and inverse quantization may be performed. The geometry data decoding unit 212 supplies the generated geometry data (decoding result) to the point cloud generation unit 213.
[0102] The point cloud generation unit 213 acquires the attribute data (attribute information) of the point cloud input to the encoding device 200 and the geometry data (decoding result) supplied from the geometry data decoding unit 212. The point cloud generation unit 213 performs a process (recalibration process) of associating the attribute data with the geometry data (decoding result). The point cloud generation unit 213 supplies the attribute data associated with the geometry data (decoding result) to the attribute data encoding unit 214.
[0103] The attribute data encoding unit 214 acquires the geometry data (decoding result) and the attribute data supplied from the point cloud generation unit 213. The attribute data encoding unit 214 encodes the attribute data using the geometry data (decoding result) and generates the encoded data of the attribute data. However, the attribute data encoding unit 214 performs this encoding according to the control of the encoding control unit 201. That is, in this encoding, the attribute data encoding unit 214 applies lifting scalability according to the control of the encoding control unit 201. The attribute data encoding unit 214 supplies the generated encoded data of the attribute data to the bitstream generation unit 215.
[0104] The bit stream generation unit 215 acquires the encoded data of the geometry data supplied from the geometry data encoding unit 211. Also, the bit stream generation unit 215 acquires the encoded data of the attribute data supplied from the attribute data encoding unit 214. The bit stream generation unit 215 generates a bit stream including these encoded data. Further, the bit stream generation unit 215 performs signaling of control information such as a scalable encoding enable flag (e.g., Lifting_scalability_enabled_flag), which is flag information regarding the application of scalable encoding, and a scaling encoding enable flag (e.g., geom_scaling_enabled_flag), which is flag information regarding the application of scaling encoding, in accordance with the control of the encoding control unit 201 (includes the control information in the bit stream). The bit stream generation unit 215 outputs the generated bit stream to the outside of the encoding device 200 (e.g., the decoding side).
[0105] With such a configuration, the encoding device 200 can prohibit the combined use of scalable encoding, which is an encoding method for generating encodable data that can be decoded in a scalable manner, and scaling encoding, which is an encoding method involving a change in the tree structure of geometry data, and can more easily achieve scalable decoding of point cloud data.
[0106] Incidentally, these processing units (the encoding control unit 201, the geometry data encoding units 211 to 215 up to the bitstream generation unit 215) may have any configuration. For example, each processing unit may be configured by a logic circuit that realizes the above-described processing. Further, each processing unit may have, for example, a CPU (Central Processing Unit), a ROM (Read Only Memory), a RAM (Random Access Memory), etc., and may realize the above-described processing by executing a program using them. Of course, each processing unit may have both configurations, and may realize a part of the above-described processing by a logic circuit and the other part by executing a program. The configurations of the respective processing units may be independent of each other. For example, some of the processing units may realize a part of the above-described processing by a logic circuit, some other processing units may realize the above-described processing by executing a program, and still other processing units may realize the above-described processing by both a logic circuit and program execution.
[0107] <Geometry data encoding unit> FIG. 16 is a block diagram showing a main configuration example of the geometry data encoding unit 211. Note that in FIG. 16, main components such as processing units and data flows are shown, and not all of the components shown in FIG. 16 are necessarily included. That is, in the geometry data encoding unit 211, there may be a processing unit that is not shown as a block in FIG. 16, or a processing or data flow that is not shown as an arrow or the like in FIG. 16.
[0108] As shown in FIG. 16, the geometry data encoding unit 211 includes a voxel generation unit 231, a tree structure generation unit 232, a selection unit 233, a geometry scaling unit 234, and an encoding unit 235.
[0109] The voxel generation unit 231 performs processing related to the generation of voxel data. For example, the voxel generation unit 231 sets a bounding box for the input point cloud data and sets voxels so as to divide the bounding box. Then, the voxel generation unit 231 quantizes the geometric data of each point in units of the voxels to generate voxel data. The voxel generation unit 231 supplies the generated voxel data to the tree structure generation unit 232.
[0110] The tree structure generation unit 232 performs processing related to the generation of a tree structure. For example, the tree structure generation unit 232 acquires the voxel data supplied from the voxel generation unit 231. Also, the tree structure generation unit 232 structures the voxel data into a tree structure. For example, the tree structure generation unit 232 generates an octree using the voxel data. The tree structure generation unit 232 supplies the generated octree data to the selection unit 233.
[0111] The selection unit 233 performs processing related to the control of applying / non-applying geometric scaling. For example, the selection unit 233 acquires the octree data supplied from the tree structure generation unit 232. Also, the selection unit 233 selects the destination for supplying the octree data according to the control of the encoding control unit 201. That is, the selection unit 233 selects whether to supply the octree data to the geometric scaling unit 234 or the encoding unit 235 according to the control of the encoding control unit 201, and supplies the octree data to the selected destination.
[0112] For example, when the encoding control unit 201 instructs the application of geometric scaling, the selection unit 233 supplies the octree data to the geometric scaling unit 234. Also, when the encoding control unit 201 instructs the non-application of geometric scaling, the selection unit 233 supplies the octree data to the encoding unit 235.
[0113] The geometry scaling unit 234 performs processing related to geometry scaling. For example, the geometry scaling unit 234 acquires the octree data supplied from the selection unit 233. Further, the geometry scaling unit 234 performs geometry scaling on the octree data, and performs scaling of geometry data and merging of points. The geometry scaling unit 234 supplies the octree data subjected to geometry scaling to the encoding unit 235.
[0114] The encoding unit 235 performs processing related to encoding of octree data (octree - formed voxel data (i.e., geometry data)). For example, the encoding unit 235 acquires the octree data supplied from the selection unit 233 or the geometry scaling unit 234. For example, when the encoding control unit 201 instructs the application of geometry scaling, the encoding unit 235 acquires the octree data subjected to geometry scaling supplied from the geometry scaling unit. Also, when the encoding control unit 201 instructs the non - application of geometry scaling, the encoding unit 235 acquires the octree data not subjected to geometry scaling supplied from the selection unit 233.
[0115] The encoding unit 235 encodes the acquired octree data and generates encoded data of the geometry data. This encoding method is arbitrary. The encoding unit 235 supplies the generated encoded data of the geometry data to the geometry data decoding unit 212 and the bit - stream generation unit 215 (both in FIG. 15).
[0116] <Attribute Data Encoding Unit> FIG. 17 is a block diagram showing a main configuration example of the attribute data encoding unit 214. Note that in FIG. 17, main components such as processing units and data flows are shown, and not all components shown in FIG. 17 are necessarily included. That is, in the attribute data encoding unit 214, there may be a processing unit not shown as a block in FIG. 17, or a process or data flow not shown as an arrow or the like in FIG. 17.
[0117] As shown in FIG. 17, the attribute data encoding unit 214 includes a selection unit 251, a scalable hierarchical processing unit 252, a hierarchical processing unit 253, a quantization unit 254, and an encoding unit 255.
[0118] The selection unit 251 performs processes related to control of application / non-application of lifting scalability. For example, the selection unit 251 selects a supply destination such as the attribute data or geometric data (decoding result) generated by the point cloud generation unit 213 (FIG. 15) according to the control of the encoding control unit 201. That is, the selection unit 251 selects whether to supply such data to the scalable hierarchical processing unit 252 or the hierarchical processing unit 253 according to the control of the encoding control unit 201, and supplies the data to the selected supply destination.
[0119] For example, when the encoding control unit 201 instructs the application of lifting scalability, the selection unit 251 supplies the attribute data or geometric data (decoding result) to the scalable hierarchical processing unit 252. Also, when the encoding control unit 201 instructs the non-application of lifting scalability, the selection unit 251 supplies the attribute data or geometric data (decoding result) to the hierarchical processing unit 253.
[0120] The scalable hierarchical processing unit 252 performs processing related to the lifting of attribute data (formation of a reference structure). For example, the scalable hierarchical processing unit 252 acquires the attribute data and geometry data (decoding result) supplied from the selection unit 251. The scalable hierarchical processing unit 252 hierarchizes the attribute data (that is, forms a reference structure) using the geometry data. At that time, the scalable hierarchical processing unit 252 performs the hierarchical processing by applying the method described in Non-Patent Document 2. That is, the scalable hierarchical processing unit 252 estimates the tree structure (octree) based on the geometry data, forms a reference structure of the attribute data so as to correspond to the estimated tree structure, derives a predicted value according to the reference structure, and derives a difference value between the predicted value and the attribute data. The scalable hierarchical processing unit 252 supplies the attribute data (difference value) generated in this way to the quantization unit 254.
[0121] The hierarchical processing unit 253 performs processing related to the lifting of attribute data (formation of a reference structure). For example, the hierarchical processing unit 253 acquires the attribute data and geometry data (decoding result) supplied from the selection unit 251. The hierarchical processing unit 253 hierarchizes the attribute data (that is, forms a reference structure) using the geometry data. At that time, the hierarchical processing unit 253 performs the hierarchical processing by applying the method described in Non-Patent Document 4. That is, the hierarchical processing unit 253 forms a reference structure of the attribute data independently of the tree structure (octree) of the geometry data, derives a predicted value according to the reference structure, and derives a difference value between the predicted value and the attribute data. That is, the hierarchical processing unit 253 does not estimate the tree structure of the geometry data. The hierarchical processing unit 253 supplies the attribute data (difference value) generated in this way to the quantization unit 254.
[0122] The quantization unit 254 acquires the attribute data (difference value) supplied from the scalable hierarchical processing unit 252 or the hierarchical processing unit 253. The quantization unit 254 quantizes the attribute data (difference value). The quantization unit 254 supplies the quantized attribute data (difference value) to the encoding unit 255.
[0123] The encoding unit 255 acquires the quantized attribute data (difference value) supplied from the quantization unit 254. The encoding unit 255 encodes the quantized attribute data (difference value) to generate encoded data of the attribute data. This encoding method is arbitrary. The encoding unit 255 supplies the generated encoded data of the attribute data to the bit stream generation unit 215 (FIG. 15).
[0124] With the above configuration, the encoding device 200 is an encoding method that generates scalable decodable encoded data in the encoding of a point cloud that represents a three-dimensional object as a set of points. It is possible to prohibit the combined use of scalable encoding and scaling encoding, which is an encoding method involving a change in the tree structure of geometric data. Therefore, for example, in the encoding of attribute data, it is not necessary to prepare a plurality of estimation methods for the tree structure of geometric data and select from among them, and scalable decoding of point cloud data can be realized. In addition, since it is not necessary to newly design the estimation method of the tree structure, an increase in cost can be suppressed. That is, scalable decoding of point cloud data can be more easily realized.
[0125] <Flow of Encoding Control Process> Next, the processing executed by this encoding device 200 will be described. The encoding control unit 201 of the encoding device 200 controls the encoding of point cloud data by executing encoding control processing. An example of the flow of this encoding control processing will be described with reference to the flowchart of FIG. 18.
[0126] When the symbolic control process starts, the encoding control unit 201 determines whether to apply lifting scalability in step S101. If it is determined to apply lifting scalability, the process proceeds to step S102.
[0127] In step S102, the encoding control unit 201 prohibits geometric scaling in the encoding of geometric data. Also, in step S103, the encoding control unit 201 applies lifting scalability in the encoding of attribute data. Further, as described above in <1. Encoding Control>, the encoding control unit 201 signals control information such as Lifting_scalability_enabled_flag and geom_scaling_enabled_flag to correspond to these controls.
[0128] When the process of step S103 ends, the encoding control process ends.
[0129] Also, in step S101, if it is determined not to apply lifting scalability, the process proceeds to step S104.
[0130] In step S104, the encoding control unit 201 determines whether to apply geometric scaling. If it is determined to apply geometric scaling, the process proceeds to step S105.
[0131] In step S105, the encoding control unit 201 applies geometric scaling in the encoding of geometric data. Also, in step S106, the encoding control unit 201 prohibits lifting scalability in the encoding of attribute data. Further, as described above in <1. Encoding Control>, the encoding control unit 201 signals control information such as Lifting_scalability_enabled_flag and geom_scaling_enabled_flag to correspond to these controls.
[0132] When the process of step S106 ends, the encoding control process ends.
[0133] Also, in step S104, if it is determined not to apply geometric scaling, the process proceeds to step S107.
[0134] In step S107, the encoding control unit 201 prohibits geometric scaling in the encoding of geometric data. Also, in step S108, the encoding control unit 201 prohibits lifting scalability in the encoding of attribute data. Also, as described above in <1. Encoding Control>, the encoding control unit 201 signals control information such as Lifting_scalability_enabled_flag and geom_scaling_enabled_flag to correspond to these controls.
[0135] When the process of step S108 ends, the encoding control process ends.
[0136] By executing the encoding control process as described above, the encoding control unit 201 can more easily realize scalable decoding of point cloud data.
[0137] <Flow of Encoding Process> The encoding device 200 encodes the point cloud data by executing an encoding process. An example of the flow of this encoding process will be described with reference to the flowchart of FIG. 19.
[0138] When the encoding process is started, the geometric data encoding unit 211 of the encoding device 200 encodes the geometric data of the input point cloud by executing the geometric data encoding process in step S201, and generates encoded data of the geometric data.
[0139] In step S202, the geometry data decoding unit 212 decodes the encoded data of the geometry data generated in step S201, and generates geometry data (decoding result).
[0140] In step S203, the point cloud generation unit 213 performs recoloring processing using the input attribute data of the point cloud and the geometry data (decoding result) generated in step S202, and associates the attribute data with the geometry data.
[0141] In step S204, the attribute data encoding unit 214 encodes the attribute data recolored in step S203 by executing attribute data encoding processing, and generates encoded data of the attribute data.
[0142] In step S205, the bitstream generation unit 215 generates and outputs a bitstream including the encoded data of the geometry data generated in step S201 and the encoded data of the attribute data generated in step S204.
[0143] When the processing of step S205 ends, the encoding processing ends.
[0144] <Flow of geometry data encoding process> Next, an example of the flow of the geometry data encoding process executed in step S201 of FIG. 19 will be described with reference to the flowchart of FIG. 20.
[0145] When the geometry data encoding process is started, the voxel generation unit 231 of the geometry data encoding unit 211 generates voxel data in step S221.
[0146] In step S222, the tree structure generation unit 232 generates a tree structure (octree) of the geometry data using the voxel data generated in step S221.
[0147] In step S223, the selection unit 233 determines whether to perform geometric scaling according to the control of the encoding control unit 201. If it is determined to perform geometric scaling, the process proceeds to step S22.
[0148] In step S224, the geometric scaling unit 234 performs geometric scaling on the geometric data of the tree structure (i.e., the data of the octree) generated in step S222. When the process of step S224 ends, the process proceeds to step S225. Also, in step S223, if it is determined not to perform geometric scaling, the process of step S224 is skipped and the process proceeds to step S225.
[0149] In step S225, the encoding unit 235 encodes the geometric data of the tree structure (i.e., the data of the octree) generated in step S222, or the geometric data (i.e., the data of the octree) subjected to geometric scaling in step S224, and generates encoded data of the geometric data.
[0150] When the process of step S225 ends, the geometric data encoding process ends and the process returns to FIG. 19.
[0151] <Flow of Attribute Data Encoding Process> Next, an example of the flow of the attribute data encoding process executed in step S204 of FIG. 19 will be described with reference to the flowchart of FIG. 21.
[0152] When the attribute data encoding process is started, the selection unit 251 of the attribute data encoding unit 214 determines whether to apply lifting scalability according to the control of the encoding control unit 201 in step S241. If it is determined to apply lifting scalability, the process proceeds to step S242.
[0153] In step S242, the scalable hierarchical processing unit 252 performs lifting by the method described in Non-Patent Document 2. That is, the scalable hierarchical processing unit 252 estimates the tree structure of the geometry data, and hierarchizes (forms a reference structure) the attribute data according to the estimated tree structure. Then, the scalable hierarchical processing unit 252 derives a predicted value according to the reference structure, and derives a difference value between the attribute data and the predicted value.
[0154] When the process of step S242 ends, the process proceeds to step S244. Also, in step S241, if it is determined not to apply lifting scalability, the process proceeds to step S243.
[0155] In step S243, the hierarchical processing unit 253 performs lifting by the method described in Non-Patent Document 4. That is, the hierarchical processing unit 253 hierarchizes (forms a reference structure) the attribute data independently of the tree structure of the geometry data. Then, the hierarchical processing unit 253 derives a predicted value according to the reference structure, and derives a difference value between the attribute data and the predicted value. When the process of step S243 ends, the process proceeds to step S244.
[0156] In step S244, the quantization unit 254 quantizes each difference value derived in step S242 or step S243 by executing quantization processing.
[0157] In step S245, the encoding unit 255 encodes the difference value quantized in step S244, and generates encoded data of the attribute data. When the process of step S245 ends, the attribute data encoding process ends, and the process returns to FIG. 19.
[0158] By performing each process as described above, the encoding device 200 can more easily realize scalable decoding of point cloud data.
[0159] <3. Second Embodiment> <Decoder> FIG. 22 is a block diagram showing an example of the configuration of a decoder which is an aspect of an information processing apparatus to which the present technology is applied. The decoder 300 shown in FIG. 22 is an apparatus that decodes encoded data of point cloud (3D data). The decoder 300 decodes, for example, the encoded data of the point cloud generated in the encoder 200.
[0160] Note that in FIG. 22, main components such as a processing unit and data flow are shown, and not all components shown in FIG. 22 are necessarily included. That is, in the decoder 300, there may be a processing unit not shown as a block in FIG. 22, or there may be a process or data flow not shown as an arrow or the like in FIG. 22.
[0161] As shown in FIG. 22, the decoder 300 includes an encoded data extraction unit 311, a geometry data decoding unit 312, an attribute data decoding unit 313, and a point cloud generation unit 314.
[0162] The encoded data extraction unit 311 acquires and holds the bit stream input to the decoder 300. The encoded data extraction unit 311 extracts the encoded data of the geometry data and the attribute data from the held bit stream from the highest bit to the desired layer. When the encoded data supports scalable decoding, the encoded data extraction unit 311 can extract the encoded data up to an intermediate layer. When the encoded data does not support scalable decoding, the encoded data extraction unit 311 extracts the encoded data of all layers.
[0163] The encoded data extraction unit 311 supplies the extracted encoded data of the geometry data to the geometry data decoding unit 312. The encoded data extraction unit 311 supplies the extracted encoded data of the attribute data to the attribute data decoding unit 313.
[0164] The geometry data decoding unit 312 acquires the encoded data of the position information supplied from the encoded data extraction unit 311. The geometry data decoding unit 312 performs the reverse process of the geometry data encoding process performed by the geometry data encoding unit 211 of the encoding device 200, thereby decoding the encoded data of the geometry data and generating geometry data (decoding result). The geometry data decoding unit 312 supplies the generated geometry data (decoding result) to the attribute data decoding unit 313 and the point cloud generation unit 314.
[0165] The attribute data decoding unit 313 acquires the encoded data of the attribute data supplied from the encoded data extraction unit 311. The attribute data decoding unit 313 acquires the geometry data (decoding result) supplied from the geometry data decoding unit 312. The attribute data decoding unit 313 performs the reverse process of the attribute data encoding process performed by the attribute data encoding unit 214 of the encoding device 200, thereby decoding the encoded data of the attribute data using the geometry data (decoding result) and generating attribute data (decoding result). The attribute data decoding unit 313 supplies the generated attribute data (decoding result) to the point cloud generation unit 314.
[0166] The point cloud generation unit 314 acquires the geometry data (decoding result) supplied from the geometry data decoding unit 312. The point cloud generation unit 314 acquires the attribute data (decoding result) supplied from the attribute data decoding unit 313. The point cloud generation unit 314 generates a point cloud (decoding result) using the geometry data (decoding result) and the attribute data (decoding result). The point cloud generation unit 314 outputs the data of the generated point cloud (decoding result) to the outside of the decoding device 300.
[0167] By having the configuration as described above, the decoding device 300 can correctly decode the encoded data of the point cloud data generated by the encoding device 200. That is, scalable decoding of the point cloud data can be more easily realized.
[0168] Note that these processing units (encoding data extraction unit 311 to point cloud generation unit 314) have an arbitrary configuration. For example, each processing unit may be configured by a logic circuit that realizes the above-described processing. Also, each processing unit may have, for example, a CPU, ROM, RAM, etc., and realize the above-described processing by executing a program using them. Of course, each processing unit may have both configurations, realizing a part of the above-described processing by a logic circuit and the other part by executing a program. The configurations of the respective processing units may be independent of each other. For example, some processing units may realize a part of the above-described processing by a logic circuit, some other processing units may realize the above-described processing by executing a program, and still other processing units may realize the above-described processing by both a logic circuit and program execution.
[0169] <Flow of Decoding Process> Next, the processing executed by this decoding device 300 will be described. The decoding device 300 decodes the encoded data of the point cloud by executing a decoding process. An example of the flow of this decoding process will be described with reference to the flowchart of FIG. 23.
[0170] When the decoding process is started, the encoding data extraction unit 311 of the decoding device 300 acquires and holds a bitstream in step S301, and extracts the encoded data of the geometry data and attribute data up to the LoD depth to be decoded.
[0171] In step S302, the geometry data decoding unit 312 decodes the encoded data of the geometry data extracted in step S301 and generates geometry data (decoding result).
[0172] In step S303, the attribute data decoding unit 313 decodes the encoded data of the attribute data extracted in step S301, and generates attribute data (decoding result).
[0173] In step S304, the point cloud generation unit 314 generates and outputs a point cloud (decoding result) using the geometry data (decoding result) generated in step S302 and the attribute data (decoding result) generated in step S303.
[0174] When the process of step S304 ends, the decoding process ends.
[0175] By performing the processes of each step in this way, the decoding device 300 can correctly decode the encoded data of the point cloud data generated by the encoding device 200. That is, scalable decoding of point cloud data can be more easily realized.
[0176] <4. Supplementary Note> <Hierarchical and Inverse Hierarchical Method> In the above, lifting has been described as an example of the hierarchical and inverse hierarchical method of attribute data. However, the hierarchical and inverse hierarchical method of attribute data may be other than lifting, such as RAHT, for example.
[0177] <Control Information> In each of the above embodiments, the enable flag has been described as an example of the control information related to the present technology. However, any other control information may also be signaled.
[0178] <Periphery·Proximity> Note that in this specification, the positional relationships such as "proximity" and "periphery" may include not only spatial positional relationships but also temporal positional relationships.
[0179] <Computer> The above-described series of processes can be executed by hardware or by software. When the series of processes is executed by software, the program constituting the software is installed in a computer. Here, the computer includes a computer incorporated in dedicated hardware, or a general-purpose personal computer or the like that can execute various functions by installing various programs.
[0180] FIG. 24 is a block diagram showing a configuration example of the hardware of a computer that executes the above-described series of processes by a program.
[0181] In the computer 900 shown in FIG. 24, a CPU (Central Processing Unit) 901, a ROM (Read Only Memory) 902, and a RAM (Random Access Memory) 903 are interconnected via a bus 904.
[0182] An input / output interface 910 is also connected to the bus 904. Connected to the input / output interface 910 are an input unit 911, an output unit 912, a storage unit 913, a communication unit 914, and a drive 915.
[0183] The input unit 911 includes, for example, a keyboard, a mouse, a microphone, a touch panel, an input terminal, etc. The output unit 912 includes, for example, a display, a speaker, an output terminal, etc. The storage unit 913 includes, for example, a hard disk, a RAM disk, a non-volatile memory, etc. The communication unit 914 includes, for example, a network interface. The drive 915 drives a removable medium 921 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0184] In the computer configured as described above, for example, the CPU 901 loads and executes a program stored in the storage unit 913 via the input / output interface 910 and the bus 904 into the RAM 903, thereby performing the series of processes described above. The RAM 903 also appropriately stores data and the like necessary for the CPU 901 to execute various processes.
[0185] The program executed by the computer can be recorded and applied, for example, on a removable medium 921 such as a package medium. In that case, the program can be installed in the storage unit 913 via the input / output interface 910 by mounting the removable medium 921 on the drive 915.
[0186] Also, this program can be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting. In that case, the program can be received by the communication unit 914 and installed in the storage unit 913.
[0187] In addition, this program can be installed in advance in the ROM 902 or the storage unit 913.
[0188] <Applicable Object of the Present Technology> In the above, the case of applying the present technology to the encoding and decoding of point cloud data has been described. However, the present technology is not limited to these examples and can be applied to the encoding and decoding of 3D data of any standard. For example, in the encoding and decoding of mesh data, the mesh data may be converted into point cloud data and the present technology may be applied for encoding and decoding. That is, as long as it does not conflict with the present technology described above, various processes such as the encoding and decoding method, and the specifications of various data such as 3D data and metadata are arbitrary. Also, as long as it does not conflict with the present technology, some of the processes and specifications described above may be omitted.
[0189] Also, in the above description, the encoding device 200 and the decoding device 300 have been described as application examples of the present technology. However, the present technology can be applied to any configuration.
[0190] For example, the present technology can be applied to various electronic devices such as transmitters and receivers (e.g., television receivers and mobile phones) in satellite broadcasting, cable broadcasting such as cable TV, distribution on the Internet, and distribution to terminals via cellular communication, or devices that record images on media such as optical disks, magnetic disks, and flash memories, and reproduce images from these storage media (e.g., hard disk recorders and cameras).
[0191] Also, for example, the present technology can be implemented as a part of the configuration of a device such as a processor (e.g., a video processor) as a system LSI (Large Scale Integration), a module (e.g., a video module) using a plurality of processors, etc., a unit (e.g., a video unit) using a plurality of modules, etc., or a set (e.g., a video set) obtained by adding other functions to the unit.
[0192] Also, for example, the present technology can also be applied to a network system composed of a plurality of devices. For example, the present technology may be implemented as cloud computing in which processing is shared and jointly performed by a plurality of devices via a network. For example, the present technology may be implemented in a cloud service that provides services related to images (moving images) to any terminal such as a computer, an AV (Audio Visual) device, a portable information processing terminal, and an IoT (Internet of Things) device.
[0193] In this specification, the term "system" refers to a collection of multiple components (devices, modules (parts), etc.), regardless of whether all the components are in the same housing. Therefore, a plurality of devices housed in separate enclosures and connected via a network, as well as a single device with multiple modules housed in one enclosure, are both systems.
[0194] <Fields and Applications Applicable to this Technology> Systems, devices, processing units, etc. to which this technology is applied can be used in any field, such as transportation, medical, security, agriculture, livestock, mining, beauty, factories, household appliances, meteorology, natural monitoring, etc. Also, their applications are arbitrary.
[0195] <Others> In this specification, the term "flag" refers to information for identifying multiple states, and includes not only information used for identifying two states of true (1) or false (0), but also information capable of identifying three or more states. Therefore, the values that this "flag" can take may be, for example, two values of 1 / 0, or three or more values. That is, the number of bits constituting this "flag" is arbitrary and can be 1 bit or multiple bits. Also, identification information (including flags) is assumed not only in the form of including the identification information in a bit stream, but also in the form of including the difference information of the identification information with respect to a certain reference information in the bit stream. Therefore, in this specification, "flags" and "identification information" include not only the information itself, but also the difference information with respect to the reference information.
[0196] Also, various types of information (such as metadata) related to the encoded data (bitstream) may be transmitted or recorded in any form as long as it is associated with the encoded data. Here, the term "associate" means, for example, making it possible to use (link) the other data when processing one data. That is, the data associated with each other may be grouped as one data or may be individual data. For example, the information associated with the encoded data (image) may be transmitted on a transmission path different from that of the encoded data (image). Also, for example, the information associated with the encoded data (image) may be recorded on a recording medium different from that of the encoded data (image) (or a different recording area of the same recording medium). Note that this "association" may be for a part of the data, not the entire data. For example, an image and the information corresponding to the image may be associated with each other in any unit such as a plurality of frames, one frame, or a part within a frame.
[0197] Note that in this specification, terms such as "synthesize", "multiplex", "add", "integrate", "include", "store", "insert", "plug in", "insert" mean, for example, grouping a plurality of things into one, such as grouping the encoded data and the metadata into one data, and mean one method of the above-mentioned "associate".
[0198] Also, the embodiments of the present technology are not limited to the above-described embodiments, and various changes can be made without departing from the gist of the present technology.
[0199] For example, the configuration described as one device (or processing unit) may be divided and configured as a plurality of devices (or processing units). Conversely, the configurations described as a plurality of devices (or processing units) above may be combined into one device (or processing unit). Of course, other configurations than those described above may be added to the configuration of each device (or each processing unit). Furthermore, if the configuration and operation of the entire system are substantially the same, a part of the configuration of one device (or processing unit) may be included in the configuration of another device (or another processing unit).
[0200] Also, for example, the program described above may be executed on any device. In that case, the device may have necessary functions (such as function blocks) and be able to obtain necessary information.
[0201] Also, for example, each step of one flowchart may be executed by one device, or may be executed by a plurality of devices sharing the work. Furthermore, when a plurality of processes are included in one step, the plurality of processes may be executed by one device, or may be executed by a plurality of devices sharing the work. In other words, the plurality of processes included in one step can also be executed as processes of a plurality of steps. Conversely, the processes described as a plurality of steps can also be executed as one step combined together.
[0202] Also, for example, the program executed by a computer may be such that the processing of the steps of writing the program is executed in time series along the order described in this specification, or may be executed in parallel, or individually at necessary timings such as when a call is made. That is, as long as there is no contradiction, the processing of each step may be executed in an order different from the order described above. Furthermore, the processing of the steps of writing this program may be executed in parallel with the processing of another program, or may be executed in combination with the processing of another program.
[0203] In addition, for example, a plurality of techniques related to the present technology can be independently implemented individually as long as there is no contradiction. Of course, any plurality of the present technologies can also be implemented in combination. For example, part or all of the present technology described in any one of the embodiments can be implemented in combination with part or all of the present technology described in other embodiments. Also, part or all of any of the above-described present technologies can be implemented in combination with other technologies not described above.
[0204] Note that the present technology can also have the following configuration. (1) An encoding control unit that controls to prohibit the combined use of scalable encoding, which is an encoding method for generating scalable decodable encoded data in the encoding of a point cloud representing a three-dimensional object as a set of points, and scaling encoding, which is an encoding method involving a change in the tree structure of geometric data An information processing apparatus including the same. (2) The scalable encoding is lifting scalability in which attribute data is lifted and encoded with a reference structure similar to the tree structure of the geometric data The information processing apparatus according to (1). (3) The scaling encoding is geometry scaling in which the geometric data is scaled and encoded The information processing apparatus according to (1) or (2). (4) When applying the scalable encoding, the encoding control unit controls to prohibit the application of the scaling encoding The information processing apparatus according to any one of (1) to (3). (5) The encoding control unit controls the signaling of a scalable encoding enable flag, which is flag information related to the application of the scalable encoding, and a scaling encoding enable flag, which is flag information related to the application of the scaling encoding The information processing apparatus according to (4). (6) When the encoding control unit signals the scalable encoding enable flag having a value indicating application of the scalable encoding, it controls to signal the scalable encoding enable flag having a value indicating non-application of the scaling encoding. The information processing apparatus according to (5). (7) When the encoding control unit applies the scalable encoding, in a profile, it controls to signal the scalable encoding enable flag having a value indicating application of the scalable encoding and the scalable encoding enable flag having a value indicating non-application of the scaling encoding. The information processing apparatus according to (5) or (6). (8) When the encoding control unit signals the scalable encoding enable flag having a value indicating application of the scalable encoding, it controls to omit signaling of the scalable encoding enable flag. The information processing apparatus according to any one of (5) to (7). (9) When the encoding control unit applies the scaling encoding, it controls to prohibit application of the scalable encoding. The information processing apparatus according to any one of (1) to (3). (10) The encoding control unit controls signaling of the scalable encoding enable flag which is flag information regarding application of the scaling encoding and the scalable encoding enable flag which is flag information regarding application of the scalable encoding. The information processing apparatus according to (9). (11) When the encoding control unit signals the scalable encoding enable flag having a value indicating application of the scaling encoding, it controls to signal the scalable encoding enable flag having a value indicating non-application of the scalable encoding. The information processing apparatus according to (10). (12) When applying the scaling encoding, the encoding control unit controls to signal, in the profile, the scaling encoding enable flag having a value indicating the application of the scaling encoding and the scalable encoding enable flag having a value indicating the non-application of the scalable encoding. The information processing apparatus according to (10) or (11). (13) When signaling the scaling encoding enable flag having a value indicating the application of the scaling encoding, the encoding control unit controls to omit the signaling of the scalable encoding enable flag. (10) The information processing apparatus according to any one of (10) to (12). (14) A geometry data encoding unit that encodes the geometry data of the point cloud according to the control of the encoding control unit and generates encoded data of the geometry data; An attribute data encoding unit that encodes the attribute data of the point cloud according to the control of the encoding control unit and generates encoded data of the attribute data The information processing apparatus according to any one of (1) to (13), further comprising. (15) The geometry data encoding unit A selection unit that selects whether to apply scaling to the geometry data according to the control of the encoding control unit; When the application of the scaling of the geometry data is selected by the selection unit, a geometry scaling unit that performs the scaling and merging of the geometry data; When the application of the scaling of the geometry data is selected by the selection unit, the encoded data of the geometry data after the scaling and the merging are performed by the geometry scaling unit, and when the non-application of the scaling of the geometry data is selected by the selection unit, an encoding unit that encodes the geometry data without the scaling and the merging being performed The information processing apparatus according to (14), comprising. (16) The geometry data encoding unit further includes a tree structure generation unit that generates a tree structure of the geometry data and the geometry scaling unit updates the tree structure generated by the tree structure generation unit by performing the scaling and the merge The information processing apparatus according to (15). (17) The attribute data encoding unit includes a selection unit that selects whether to apply scalable hierarchicalization for lifting the attribute data in a reference structure similar to the tree structure of the geometry data according to the control of the encoding control unit, a scalable hierarchicalization unit that performs scalable hierarchicalization on the attribute data when the application of scalable hierarchicalization is selected by the selection unit, and an encoding unit that encodes the attribute data on which the scalable hierarchicalization has been performed by the scalable hierarchicalization unit when the application of scalable hierarchicalization is selected by the selection unit, and encodes the attribute data on which the scalable hierarchicalization has not been performed when the non-application of scalable hierarchicalization is selected by the selection unit The information processing apparatus according to any one of (14) to (16). (18) A geometry data decoding unit that decodes the encoded data of the geometry data generated by the geometry data encoding unit and generates the geometry data, and a recoloring processing unit that performs recoloring processing of the attribute data using the geometry data generated by the geometry data decoding unit and the attribute data encoding unit encodes the attribute data on which the recoloring processing has been performed by the recoloring processing unit The information processing apparatus according to any one of (14) to (17). A bitstream generation unit that generates a bitstream including the encoded data of the geometry data generated by the geometry data encoding unit and the encoded data of the attribute data generated by the attribute data encoding unit The information processing apparatus according to any one of (14) to (18), further comprising the same (20) In the encoding of a point cloud that represents a three-dimensional object as a set of points, control is performed to prohibit the combined use of scalable encoding, which is an encoding method for generating scalable decodable encoded data, and scaling encoding, which is an encoding method involving a change in the tree structure of geometry data Information processing method
Explanation of Signs
[0205] 200 Encoding device, 201 Encoding control unit, 211 Geometry data encoding unit, 212 Geometry data decoding unit, 213 Point cloud generation unit, 214 Attribute data encoding unit, 215 Bitstream generation unit, 231 Voxel generation unit, 232 Tree structure generation unit, 233 Selection unit, 234 Geometry scaling unit, 235 Encoding unit, 251 Selection unit, 252 Scalable hierarchical processing unit, 253 Hierarchical processing unit, 254 Quantization unit, 255 Encoding unit, 300 Decoding device, 311 Encoded data extraction unit, 312 Geometry data decoding unit, 313 Attribute data decoding unit, 314 Point cloud generation unit
Claims
1. a geometry data decoding unit that decodes encoded data of geometry data that has been encoded by applying geometry scaling, which is an encoding method that involves changing a tree structure of the geometry data, in encoding a point cloud that represents a three-dimensional object as a set of points; an attribute data decoding unit that decodes the coded data of the attribute data without applying scalable coding based on scalable coding information corresponding to the fact that scalable coding that enables attribute data to be decoded at an intermediate resolution of the tree structure of the geometry data is not applied in accordance with the application of the geometry scaling; An information processing device comprising:
2. The scalable coding is lifting scalability, which encodes the attribute data by lifting the attribute data using a reference structure similar to the tree structure of the geometry data. The information processing device according to claim 1 .
3. The encoded geometry data is data in which points have been thinned out by the geometry scaling. The information processing device according to claim 1 .
4. the geometry data decoding unit decodes encoded data of the geometry data based on a scalable encoding enable flag that is flag information regarding application of the scalable encoding and a scaling encoding enable flag that is flag information regarding application of the geometry scaling that are signaled; The attribute data decoding unit decodes the encoded data of the attribute data based on the signaled scalable encoding enable flag and the scaling encoding enable flag. The information processing device according to claim 1 .
5. If the scalable coding enable flag is signaled with a value indicating application of the scalable coding, then the scaling coding enable flag is signaled with a value indicating non-application of the geometry scaling. The information processing device according to claim 4.
6. If the scalable coding is applied, the scalable coding enable flag is signaled in a profile with a value indicating the application of the scalable coding and with a value indicating the non-application of the geometry scaling. The information processing device according to claim 4.
7. If the scalable coding enable flag of a value indicating application of the scalable coding is signaled, signaling of the scaling coding enable flag is omitted. The information processing device according to claim 4.
8. the geometry data decoding unit decodes encoded data of the geometry data that has been encoded without applying the geometry scaling; The attribute data decoding unit decodes the coded data of the attribute data by applying the scalable coding based on the scalable coding information corresponding to the scalable coding being applied in accordance with non-application of the geometry scaling. The information processing device according to claim 1 .
9. the geometry data decoding unit decodes encoded data of the geometry data based on a scalable encoding enable flag that is flag information regarding application of the scalable encoding and a scaling encoding enable flag that is flag information regarding application of the geometry scaling that are signaled; The attribute data decoding unit decodes the encoded data of the attribute data based on the signaled scalable encoding enable flag and the scaling encoding enable flag. The information processing device according to claim 8.
10. If the scaling encoding enable flag is signaled with a value indicating application of the geometry scaling, then the scalable encoding enable flag is signaled with a value indicating non-application of the scalable encoding. The information processing device according to claim 9.
11. If the geometry scaling is applied, the scaling encoding enable flag is signaled in a profile with a value indicating the application of the geometry scaling and with a value indicating the non-application of the scalable encoding. The information processing device according to claim 9.
12. If the scaling encoding enable flag of a value indicating application of the geometry scaling is signaled, signaling of the scalable encoding enable flag is omitted. The information processing device according to claim 9.
13. The geometry data decoding unit decodes the encoded data of the geometry data, and, if the geometry scaling was applied during encoding, performs inverse scaling and inverse merging of the geometry data. The information processing device according to claim 1 .
14. The geometry data decoding unit generates a tree structure of the geometry data, and updates the tree structure by inverse scaling and inverse merging the geometry data. The information processing device according to claim 13.
15. The attribute data decoding unit decodes the encoded data of the attribute data, and when scalable hierarchical coding is applied to the attribute data during encoding, the attribute data is lifted using a reference structure similar to the tree structure of the geometry data, and performs inverse scalable hierarchical coding on the attribute data. The information processing device according to claim 1 .
16. The attribute data decoding unit decodes the coded data of the attribute data that has been recolored during coding. The information processing device according to claim 1 .
17. The apparatus further includes an encoded data extraction unit that extracts encoded data of the geometry data and encoded data of the attribute data from the bit stream, from the highest level to a desired level; the geometry data decoding unit decodes the encoded data of the geometry data extracted from the bit stream; The attribute data decoding unit decodes the encoded data of the attribute data extracted from the bit stream. The information processing device according to claim 1 .
18. The apparatus further includes a point cloud generating unit that generates a point cloud using the geometry data obtained by being decoded by the geometry data decoding unit and the attribute data obtained by being decoded by the attribute data decoding unit. The information processing device according to claim 1 .
19. An information processing device, In encoding a point cloud that represents a three-dimensional object as a set of points, decoding encoded data of the geometry data that has been encoded using geometry scaling, which is an encoding method that involves changing a tree structure of the geometry data; decoding the coded data of the attribute data without applying the scalable coding based on scalable coding information corresponding to the fact that the scalable coding, which enables the attribute data to be decoded at an intermediate resolution of the tree structure of the geometry data, is not applied in accordance with the application of the geometry scaling; An information processing method comprising:
Citation Information
Patent Citations
Scalable point cloud compression with transform, and corresponding decompression
US20170347122A1
Point Cloud Compression
US20190080483A1
Method and apparatus for video coding
US20200107048A1
Image processing device and method
WO2020071115A1