Decoding device and method

By setting reference points based on the center of gravity in a hierarchical structure, the method enhances coding efficiency in point cloud encoding by reducing prediction accuracy losses.

JP7803401B2Active Publication Date: 2026-01-21SONY GROUP CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024228783
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-01-07
Filing Date
2024-12-25
Publication Date
2026-01-21
Estimated Expiration
2040-12-24

AI Technical Summary

Technical Problem

Existing methods for encoding attribute data in point clouds, such as the one described in Non-Patent Document 4, can lead to a decrease in coding efficiency due to potential reductions in prediction accuracy when selecting reference points.

Method used

A method that sets reference points based on the center of gravity of a limited number of points in a hierarchical structure, deriving prediction values and difference values to enhance coding efficiency.

Benefits of technology

This approach suppresses decreases in prediction accuracy and coding efficiency by ensuring reference points are closer to the center of gravity, thereby maintaining high encoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007803401000001
    Figure 0007803401000001
  • Figure 0007803401000002
    Figure 0007803401000002
  • Figure 0007803401000003
    Figure 0007803401000003
Patent Text Reader

Abstract

To enable a reduction in decoding efficiency to be suppressed.SOLUTION: A reference point is set on the basis of the center of gravity of a point in the case of hierarchizing attribute information by recursively repeating classification between a prediction point for deriving a difference value between the attribute information and a prediction value of the attribute information, and the reference point to be used to derive the prediction value about the attribute information of each point of point cloud representing a three-dimensional shape object as a set of points. This disclosure is applicable to, for example, an information processing device, an image processing device, an encoding device, a decoding device, an electronic apparatus, an information processing method, or a program or the like.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure provides: Decryption device and a method, in particular, capable of suppressing a decrease in coding efficiency. Decryption device and methods. [Background technology]

[0002] Conventionally, methods have been considered for encoding 3D data representing a three-dimensional structure, such as a point cloud (see, for example, Non-Patent Document 1). Point cloud data consists of geometry data (also referred to as position information) and attribute data (also referred to as attribute information) for each point. Therefore, point cloud encoding is performed for both the geometry data and the attribute data. Various methods have been proposed for encoding attribute data. For example, a technique called "Lifting" has been proposed (see, for example, Non-Patent Document 2). A method has also been proposed that enables attribute data to be decoded in a scalable manner (see, for example, Non-Patent Document 3).

[0003] In such a lifting scheme, attribute data is layered by recursively repeating a process of setting a point as a reference point or a prediction point for the reference point. According to this layered structure, a predicted value of the attribute data of the prediction point is derived using the attribute data of the reference point, and the difference between the predicted value and the attribute data is encoded. In layering attribute data in this way, a method has been proposed in which the first point and the last point in Morton order among the reference point candidates are alternately selected for each layer as the reference point (see, for example, Non-Patent Document 4). [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] R. Mekuria, Student Member IEEE, K. Blom, P. Cesar., Member, IEEE, "Design, Implementation and Evaluation of a Point Cloud Codec for Tele-Immersive Video",tcsvt_paper_submitted_february.pdf [Non-patent document 2] Khaled Mammou, Alexis Tourapis, Jungsun Kim, Fabrice Robinet, Valery Valentin, Yeping Su, "Lifting Scheme for Lossy Attribute Encoding in TMC1", ISO / IEC JTC1 / SC29 / WG11 MPEG2018 / m42640, April 2018, San Diego, US [Non-patent document 3] Ohji Nakagami, Satoru Kuma, "[G-PCC] Spatial scalability support for G-PCC", ISO / IEC JTC1 / SC29 / WG11 MPEG2019 / m47352, March 2019, Geneva, CH [Non-patent document 4] Hyejung Hur, Sejin Oh, "[G-PCC][New Proposal] on improved spatial scalable lifting", ISO / IEC JTC1 / SC29 / WG11 MPEG2019 / M51408, October 2019, Geneva, CH Summary of the Invention [Problem to be solved by the invention]

[0005] However, the method described in Non-Patent Document 4 is not always optimal, and other methods have been sought.

[0006] The present disclosure has been made in light of such circumstances, and makes it possible to suppress a decrease in coding efficiency. [Means for solving the problem]

[0007] One aspect of this technology Decryption device is a point cloud that represents a 3D object as a set of points. The system includes an inverse layering unit that acquires geometry data, sets a hierarchical structure based on the geometry data, classifies the points at each layer of the hierarchical structure into either prediction points or reference points that are closer to the center of gravity of a limited number of points including the prediction points by recursive processing, derives prediction values ​​for the reference points, and derives difference values ​​between attribute data corresponding to the geometry data and the prediction values ​​based on the prediction points, thereby deriving the attribute data. It is an information processing device.

[0008] A decoding method according to one aspect of the present technology includes: The decoding device This is a decoding method that acquires geometry data of a point cloud that represents a three-dimensional object as a collection of points, sets a hierarchical structure based on the geometry data, classifies the points at each level of the hierarchical structure through recursive processing into either prediction points or reference points that are closer to the center of gravity of a limited number of points including the prediction points, derives prediction values ​​for the reference points, and derives the difference between attribute data corresponding to the geometry data and the prediction value based on the prediction points, thereby deriving the attribute data.

[0015] One aspect of this technology Decryption device In the method, a point cloud that represents a three-dimensional object as a set of points is used. Geometry data is acquired, a hierarchical structure is established based on the geometry data, and points at each level of the hierarchical structure are classified by recursive processing into either a predicted point or a reference point that is closer to the center of gravity of a limited number of points including the predicted point, a predicted value for the reference point is derived, and the attribute data is derived by deriving the difference between the attribute data corresponding to the geometry data and the predicted value based on the predicted point. [Brief explanation of the drawings]

[0019] [Figure 1] FIG. 10 is a diagram illustrating an example of a lifting state. [Figure 2] 10A and 10B are diagrams illustrating an example of a method for setting reference points based on Morton's order. [Figure 3] 10A and 10B are diagrams illustrating an example of a method for setting reference points based on Morton's order. [Figure 4] 10A and 10B are diagrams illustrating an example of a method for setting reference points based on Morton's order. [Figure 5]10A and 10B are diagrams illustrating an example of a method for setting reference points based on Morton's order. [Figure 6] FIG. 10 is a diagram illustrating an example of a reference point setting method. [Figure 7] 10A and 10B are diagrams illustrating an example of a method for setting a reference point based on a center of gravity. [Figure 8] FIG. 10 is a diagram illustrating an example of a method for deriving a center of gravity. [Figure 9] FIG. 10 is a diagram illustrating an example of a method for deriving a center of gravity. [Figure 10] FIG. 10 is a diagram illustrating an example of a region from which a center of gravity is derived. [Figure 11] FIG. 10 is a diagram illustrating an example of a method for selecting points. [Figure 12] FIG. 10 is a diagram illustrating an example of the same conditions. [Figure 13] FIG. 1 is a block diagram illustrating an example of the main configuration of an encoding device. [Figure 14] FIG. 2 is a block diagram illustrating an example of the main configuration of an attribute information encoding unit. [Figure 15] FIG. 2 is a block diagram illustrating an example of the main configuration of a layering processing unit. [Figure 16] 10 is a flowchart illustrating an example of the flow of an encoding process. [Figure 17] 10 is a flowchart illustrating an example of the flow of an attribute information encoding process. [Figure 18] 10 is a flowchart illustrating an example of the flow of a layering process. [Figure 19] 10 is a flowchart illustrating an example of the flow of a reference point setting process. [Figure 20] FIG. 2 is a block diagram illustrating an example of the main configuration of a decoding device. [Figure 21] FIG. 10 is a block diagram illustrating an example of the main configuration of an attribute information decoding unit. [Figure 22] 10 is a flowchart illustrating an example of the flow of a decoding process. [Figure 23] 10 is a flowchart illustrating an example of the flow of an attribute information decoding process. [Figure 24] 10 is a flowchart illustrating an example of the flow of a layer inversion process. [Figure 25] FIG. 10 is a diagram illustrating an example of a table. [Figure 26] FIG. 10 is a diagram illustrating an example of a table and signaling. [Figure 27] 10 is a flowchart illustrating an example of the flow of a reference point setting process. [Figure 28] FIG. 10 is a diagram illustrating an example of information to be signaled. [Figure 29] FIG. 10 is a diagram showing an example of a signaling target. [Figure 30] FIG. 10 is a diagram illustrating an example of signaled information. [Figure 31] FIG. 10 is a diagram illustrating an example of signaled information. [Figure 32] FIG. 10 is a diagram illustrating an example of syntax for signaling fixed-length bit information. [Figure 33] FIG. 10 is a diagram illustrating an example of syntax for signaling fixed-length bit information. [Figure 34] FIG. 10 is a diagram illustrating an example of variable-length bit information to be signaled. [Figure 35] FIG. 10 is a diagram illustrating an example of variable-length bit information to be signaled. [Figure 36] FIG. 10 is a diagram illustrating an example of syntax for signaling variable-bit-length information. [Figure 37] FIG. 10 is a diagram illustrating an example of syntax for signaling variable-bit-length information. [Figure 38] 10 is a flowchart illustrating an example of the flow of a reference point setting process. [Figure 39] FIG. 10 is a diagram illustrating an example of a search order. [Figure 40] 10 is a flowchart illustrating an example of the flow of a reference point setting process. [Figure 41] FIG. 1 is a block diagram illustrating an example of the main configuration of a computer. DETAILED DESCRIPTION OF THE INVENTION

[0020] Hereinafter, modes for carrying out the present disclosure (hereinafter referred to as embodiments) will be described in the following order. 1. Setting a reference point 2. First embodiment (method 1) 3. Second embodiment (method 2) 4. Third embodiment (method 3) 5. Fourth embodiment (method 4) 6. Notes

[0021] <1. Setting the reference point> <References supporting technical content and technical terminology> The scope of disclosure of the present technology includes not only the contents described in the embodiments but also the contents described in the following non-patent documents that were publicly known at the time of filing.

[0022] Non-patent document 1: (mentioned above) Non-patent document 2: (mentioned above) Non-patent document 3: (mentioned above) Non-patent document 4: (mentioned above)

[0023] In other words, the contents of the above-mentioned non-patent documents and the contents of other documents referenced in the above-mentioned non-patent documents are also used as the basis for determining the support requirements.

[0024] <Point Cloud> Previously, 3D data existed, such as point clouds, which represent three-dimensional structures using point location information and attribute information, and meshes, which are composed of vertices, edges, and faces and define three-dimensional shapes using polygonal representations.

[0025] For example, in the case of a point cloud, a three-dimensional structure (a three-dimensional object) is represented as a collection of many points. Point cloud data (also referred to as point cloud data) consists of position information (also referred to as geometry data) and attribute information (also referred to as attribute data) for each point. The attribute data can include any information. For example, the attribute data may include color information, reflectance information, normal information, etc. for each point. In this way, point cloud data has a relatively simple data structure, and by using a sufficient number of points, it is possible to represent any three-dimensional structure with sufficient accuracy.

[0026] <Quantization of position information using voxels> Since such point cloud data has a relatively large amount of data, an encoding method using voxels was devised to compress the data volume through encoding, etc. A voxel is a three-dimensional region for quantizing geometry data (position information).

[0027] That is, the three-dimensional region (also called a bounding box) containing the point cloud is divided into small three-dimensional regions called voxels, and each voxel indicates whether it contains a point. In this way, the position of each point is quantized in voxel units. Therefore, by converting point cloud data into such voxel data (also called voxel data), it is possible to suppress an increase in the amount of information (typically, to reduce the amount of information).

[0028] <octree> Furthermore, it was considered to construct an octree using such voxel data for geometry data. An octree is a tree structure of voxel data. The value of each bit in the lowest node of this octree indicates whether or not each voxel has a point. For example, a value of "1" indicates a voxel that contains a point, and a value of "0" indicates a voxel that does not contain a point. In an octree, one node corresponds to eight voxels. In other words, each node in the octree is made up of eight bits of data, and these eight bits indicate whether or not each of the eight voxels has a point.

[0029] The upper nodes of the Octree indicate whether or not there is a point in the area that combines the eight voxels corresponding to the lower nodes belonging to that node. In other words, the upper nodes are generated by combining the information of the voxels of the lower nodes. Note that if a node has a value of "0", that is, if none of the eight corresponding voxels contain a point, the node is deleted.

[0030] By doing this, a tree structure (octree) consisting of nodes whose values ​​are not "0" is constructed. In other words, the octree can indicate whether or not there are points at each voxel resolution. By encoding the position information as an octree, the point cloud data at that resolution can be restored by decoding it from the highest resolution (top layer) to the desired layer (resolution). In other words, it is possible to easily decode at any resolution without decoding information at unnecessary layers (resolutions). In other words, it is possible to achieve scalability of voxels (resolutions).

[0031] Furthermore, by omitting nodes with a value of "0" as described above, it is possible to lower the resolution of voxels in areas where no points exist, thereby further suppressing the increase in the amount of information (typically reducing the amount of information).

[0032] <lifting> In contrast, when encoding attribute data (attribute information), the geometry data (position information) is assumed to be known, including degradation due to encoding, and encoding is performed using the positional relationships between points. As a method for encoding such attribute data, methods using a transformation called RAHT (Region Adaptive Hierarchical Transform) or Lifting as described in Non-Patent Document 2 have been considered. By applying these technologies, attribute data can also be hierarchically organized, like the octree of geometry data.

[0033] For example, in the case of Lifting described in Non-Patent Document 2, attribute data is layered by recursively repeating the process of setting a point as a reference point or a prediction point for the reference point. Then, according to this layered structure, a predicted value of the attribute data of the prediction point is derived using the attribute data of the reference point, and the difference value between the predicted value and the attribute data is encoded.

[0034] For example, suppose point P5 is selected as the reference point in Fig. 1. In this case, prediction points are searched for within a circular region with point P5 as the center and radius R. In this case, point P9 is located within that region, and is therefore set as the prediction point whose predicted value is derived with reference to point P5 (prediction point with point P5 as the reference point).

[0035] By such processing, for example, the differential values ​​of points P7 to P9 indicated by white circles, the differential values ​​of points P1, P3, and P6 indicated by diagonal lines, and the differential values ​​of points P0, P2, P4, and P5 indicated by gray circles are derived as differential values ​​of different hierarchical levels.

[0036] In reality, the point cloud is arranged in a three-dimensional space and the above-described processing is performed in the three-dimensional space, but for the sake of convenience, the three-dimensional space is shown schematically using a two-dimensional plane in Fig. 1. In other words, the explanation given with reference to Fig. 1 can be similarly applied to processing, phenomena, etc. in the three-dimensional space.

[0037] In the following, explanations of three-dimensional space are appropriately given using a two-dimensional plane. Unless otherwise specified, the explanations can basically be applied to processes and phenomena in three-dimensional space as well.

[0038] The reference point in this hierarchy was selected, for example, according to Morton order. For example, as in the tree structure shown in Figure 2, when selecting one reference point from multiple nodes in a certain hierarchy and making it the node one hierarchy higher, the multiple nodes were searched in Morton order, and the first node that appeared was selected as the reference point. In Figure 2, each circle represents a node, and the black circle represents the node selected as the reference point (i.e., selected as the node one hierarchy higher). In Figure 2, each node is sorted in Morton order from left to right. That is, in the example in Figure 2, the leftmost node is always selected.

[0039] In response to this, Non-Patent Document 4 proposes a method for layering attribute data in which the first point and the last point in Morton's order among the candidates for the reference point are selected alternately for each layer. That is, as shown in the example of Figure 3, in the LoD N layer, the first node in Morton's order is selected as the reference point, and in the next layer (LoD N-1), the last node in Morton's order is selected as the reference point.

[0040] FIG. 4 illustrates an example of how a reference point is selected in three-dimensional space using a two-dimensional plane. Each square in FIG. 4A represents a voxel at a certain level. Each circle represents a candidate reference point to be processed. When selecting a reference point from the 2x2 points shown in FIG. 4A, for example, the first point in Morton's order (gray point) is selected as the reference point. At the next higher level, as shown in FIG. 4B, the last point in Morton's order (gray point) of the 2x2 points is selected as the reference point. Furthermore, at the next higher level, as shown in FIG. 4C, the first point in Morton's order (gray point) of the 2x2 points is selected as the reference point.

[0041] The arrows shown in each of Figures 4A to 4C indicate the movement of the reference point. In this case, the movement range of the reference point is limited to a narrow range as shown in the dotted line frame in Figure 4C, so that the decrease in prediction accuracy is suppressed.

[0042] However, when a reference point is selected at the position shown in Figure 5 in the same way as in Figure 4, the position of the reference point moves as shown in Figures 5A to 5C. Figure 4 explains another example of the selection of a reference point in three-dimensional space using a two-dimensional plane. In other words, the range of movement of the reference point becomes wider than in Figure 4, as shown by the dotted frame in Figure 5C, which could reduce prediction accuracy.

[0043] As described above, in the method described in Non-Patent Document 4, there is a risk that prediction accuracy may decrease depending on the position of the point, resulting in a decrease in coding efficiency.

[0044] <How to set a reference point> Therefore, for example, in the layering of this attribute data, the center of gravity of the points may be found and a reference point may be set based on the center of gravity, as in Method 1 shown in the top row of the table in Fig. 6. For example, a point close to the derived center of gravity may be selected as the reference point.

[0045] Furthermore, for example, as in Method 2 shown in the second row from the top of the table in FIG. 6, in layering this attribute data, reference points may be selected according to the distribution pattern (distribution mode) of the points.

[0046] Furthermore, in the layering of this attribute data, information regarding the setting of reference points may be transmitted from the encoding side to the decoding side, as in Method 3 shown in the third row from the top of the table in Figure 6.

[0047] Furthermore, in layering this attribute data, for example, as shown in Method 4 in the fourth row from the top of the table in Figure 6, points close to the center of the bounding box and points far from it may be selected alternately for each layer as reference points.

[0048] By applying any of these methods, it is possible to suppress a decrease in coding efficiency. Note that the above-mentioned methods can be applied in any combination. Furthermore, the above-mentioned methods can be applied to the coding and decoding of attribute data that supports scalable decoding, and can also be applied to the coding and decoding of attribute data that does not support scalable decoding.

[0049] 2. First Embodiment <Method 1> A case where the above-mentioned "Method 1" is applied will be described. In this "Method 1," the center of gravity of the points is derived, and a reference point is selected based on the center of gravity. Any point relative to the derived center of gravity may be set as the reference point. For example, a point close to the derived center of gravity (e.g., a point located closer to the center of gravity) may be selected as the reference point.

[0050] FIG. 7A shows an example of a target region for setting a reference point. In FIG. 7A, squares represent voxels and circles represent points. That is, FIG. 7A is a diagram that schematically shows an example of a voxel structure in three-dimensional space using a two-dimensional plane. For example, if points A to C arranged as shown in FIG. 7A are candidates for reference points, point B, which is close to the center of gravity of these candidates, may be selected as the reference point, as shown in FIG. 7B. FIG. 7B shows the hierarchical structure of attribute data, similar to FIG. 2 etc., and black circles represent reference points. That is, point B is selected as the reference point from points A to C.

[0051] By using a point close to the center of gravity as a reference point in this way, it is possible to set a point close to more other points as a reference point, which means that it is possible to set a reference point so as to suppress a decrease in prediction accuracy for more prediction points, thereby suppressing a decrease in coding efficiency.

[0052] <How to derive the center of gravity> The method for deriving the center of gravity is arbitrary. For example, the center of gravity for any point may be used as the center of gravity used to select the reference point. For example, the center of gravity of points located within a predetermined range may be derived, and that center of gravity may be used to select the reference point. In this way, it is possible to prevent an increase in the number of points used to derive the center of gravity, and therefore an increase in the load.

[0053] The range of points used to derive this center of gravity (also referred to as the center of gravity derivation target range) may be any range. For example, the center of gravity of a reference point candidate may be derived as in method (1) shown in the second row from the top of the "center of gravity derivation method" table in FIG. 8. That is, for example, as shown in FIG. 9A, a voxel region consisting of 2x2x2 voxels in which points that are candidates for reference points exist may be used as the center of gravity derivation target range. In FIG. 9A, a 2x2x2 voxel region in three-dimensional space is schematically shown on a two-dimensional plane (as a 2x2 square). In this case, the centers of gravity of the three points indicated by circles in FIG. 9A are derived, and these centers of gravity are used to set the reference point.

[0054] By doing so, it is only necessary to derive the center of gravity of the point to be processed (candidate for reference point), so that it is not necessary to search for other points, and the center of gravity can be easily derived.

[0055] The voxel region to be used as the target range for center of gravity derivation is arbitrary and is not limited to 2x2x2. For example, the center of gravity of a point located within a voxel region consisting of NxNxN (N>=2) voxels may be derived. In other words, this NxNxN voxel region may be used as the target range for center of gravity derivation.

[0056] For example, as shown in A of Fig. 10, the voxel region that is the target range for deriving the center of gravity (the voxel region indicated by the thick line in A of Fig. 10) and the target voxel region for deriving the reference point (the voxel region indicated by the gray in A of Fig. 10) may be located at the same position (the two ranges may be completely coincident). Note that Fig. 10 shows a voxel region that is actually configured in a three-dimensional space, but is shown schematically on a two-dimensional plane. Also, for the sake of convenience, Fig. 10A shows the target range for deriving the center of gravity and the target voxel region for deriving the reference point slightly shifted from each other, but actually shows that the two ranges are completely coincident.

[0057] Also, for example, as shown in B of Fig. 10, the voxel region that is the target range for deriving the center of gravity (the voxel region indicated by the thick line in B of Fig. 10) may be wider than the target voxel region that is the target for deriving the reference point (the voxel region indicated by the gray in B of Fig. 10). In the example of B of Fig. 10, a 4x4x4 voxel region is the target range for deriving the center of gravity.

[0058] 10C, the center of the voxel region that is the target range for center of gravity derivation (the voxel region indicated by the thick line in FIG. 10C) does not have to coincide with the center of the target voxel region for deriving the reference point (the voxel region indicated by the gray in FIG. 10C). In other words, the target range for center of gravity derivation may be biased in a predetermined direction relative to the target voxel region for deriving the reference point. For example, the extent of the target range for center of gravity derivation may be biased in this way to prevent the target range for center of gravity derivation from going outside the bounding box near the edge of the bounding box.

[0059] Alternatively, the center of gravity of N neighboring points may be determined, as in method (2) shown in the third row from the top of the "Method for Deriving a Center of Gravity" table in FIG. 8. That is, for example, as shown in FIG. 9B, N points may be searched for starting from the center coordinates of a voxel region consisting of 2x2x2 voxels in which a point that is a candidate for a reference point exists, and the center of gravity of these N points may be derived. In FIG. 9B, the distribution of points actually arranged in a three-dimensional space is schematically shown on a two-dimensional plane. Also, the black circle indicates the center coordinate of the voxel region from which the reference point is to be derived. That is, N points (white circles) are selected in order from the closest to the black circle, and their centers of gravity are derived.

[0060] By doing so, the number of points to be searched can be limited to N, and an increase in the load due to the search can be suppressed.

[0061] Furthermore, for example, as in method (3) shown in the fourth row from the top of the "Method for Deriving a Center of Gravity" table in Fig. 8, reference point candidates (points existing in a voxel region consisting of 2x2x2 voxels) may be excluded from the nearby N points derived by method (2). In other words, as shown in Fig. 9C, the 2x2x2 voxel region may be excluded from the range of interest for derivation of the center of gravity, and the center of gravity of points located outside this 2x2x2 voxel region may be derived. In Fig. 9C, the distribution of points actually arranged in a three-dimensional space is shown schematically on a two-dimensional plane.

[0062] Furthermore, for example, as shown in the fifth row of the "Method for Deriving a Center of Gravity" table in FIG. 8, the center of gravity of points within a region of radius r centered on the central coordinates of a voxel region consisting of 2x2x2 voxels where a point that is a candidate for a reference point exists may be derived, as shown in FIG. 9D. That is, in this case, the center of gravity of points located within a region of radius r centered on the central coordinates of a voxel region consisting of 2x2x2 voxels, as shown in a dotted line frame, is derived. Note that FIG. 9D shows a schematic representation of the distribution of points actually arranged in three-dimensional space on a two-dimensional plane. Furthermore, black circles indicate the central coordinates of the voxel region from which the reference point is derived, and white circles indicate points within the region of radius r centered on the central coordinates of the voxel region consisting of 2x2x2 voxels.

[0063] By doing so, this technology can also be applied to Lifting, which does not use a voxel structure and is not scalable.

[0064] <How to select a reference point> As described above, in the case of "Method 1," for example, a point close to the derived center of gravity can be set as the reference point. If there are multiple "points close to the center of gravity," one of them is selected as the reference point. This selection method is arbitrary. For example, the reference point may be set as shown in the table of "Method for selecting a reference point from multiple candidates" in Figure 11.

[0065] For example, as shown in A of FIG. 12, there may be multiple points that are equidistant from the center of gravity. Also, as shown in B of FIG. 12, in order to suppress an increase in the load due to calculations, all points that are located sufficiently close may be considered to be "points close to the center of gravity." In the case of B of FIG. 12, all points that are located within a radius Dth from the center of gravity are considered to be "points close to the center of gravity." In this case, there may be multiple "points close to the center of gravity."

[0066] In such a case, for example, the point to be processed first in a predetermined search order may be selected, as shown in method (1) in the second row from the top of the table of "How to select a reference point from multiple candidates" in Figure 11.

[0067] Also, for example, the point to be processed first or last in a predetermined search order may be selected, as in method (2) shown in the third row from the top of the table of "Method for selecting a reference point from multiple candidates" in Fig. 11. For example, it may be possible to switch between selecting the point to be processed first in the predetermined search order and selecting the point to be processed last in the predetermined search order for each hierarchical level.

[0068] Furthermore, for example, as shown in method (3) in the fourth row from the top of the table of "How to select a reference point from multiple candidates" in Figure 11, the point to be processed may be selected in the middle (number / 2) in a specified search order.

[0069] Also, for example, as in method (4) shown in the fifth row from the top of the table of "Method of selecting a reference point from multiple candidates" in FIG. 11, a point to be processed may be selected in a specified order in a predetermined search order. In other words, a point to be processed Nth in the predetermined search order may be selected. This specified order (N) may be predetermined or may be set by a user, an application, or the like. Furthermore, if the specified order is settable, information regarding this specified order (N) may be signaled (transmitted).

[0070] As in the above methods (1) to (4), when there are multiple candidates with approximately the same conditions for the center of gravity, a reference point may be set from among the multiple candidates based on a predetermined search order.

[0071] The search order may be any order. For example, it may be Morton order or an order other than Morton order. The search order may be predetermined by a standard or the like, or may be settable by a user, an application, or the like. When the search order is settable, information regarding the search order may be signaled (transmitted).

[0072] Furthermore, for example, as in method (5) shown in the sixth row from the top of the table of "Method for selecting a reference point from multiple candidates" in Fig. 11, the range for center of gravity derivation may be made wider, the centers of gravity of points within this wider range for center of gravity derivation may be derived, and the points may be selected using the newly derived centers of gravity. In other words, the conditions may be changed and the centers of gravity may be re-derived.

[0073] <Encoding device> Next, a device to which the present technology is applied will be described. Fig. 13 is a block diagram showing an example of the configuration of an encoding device, which is one aspect of an information processing device to which the present technology ("Method 1") is applied. The encoding device 100 shown in Fig. 13 is a device that encodes a point cloud (3D data). The encoding device 100 encodes the point cloud by applying the present technology described in this embodiment.

[0074] Note that Fig. 13 shows the main processing units, data flows, etc., and is not necessarily all that is shown in Fig. 13. In other words, in encoding device 100, there may be processing units that are not shown as blocks in Fig. 13, and there may be processing and data flows that are not shown as arrows, etc. in Fig. 13.

[0075] As shown in FIG. 13, the encoding device 100 includes a position information encoding unit 101, a position information decoding unit 102, a point cloud generating unit 103, an attribute information encoding unit 104, and a bitstream generating unit 105.

[0076] The position information encoding unit 101 encodes geometry data (position information) of a point cloud (3D data) input to the encoding device 100. Any encoding method may be used as long as it is compatible with scalable decoding. For example, the position information encoding unit 101 layers the geometry data to generate an octree and encodes the octree. Furthermore, processing such as filtering and quantization for noise suppression (denoising) may also be performed. The position information encoding unit 101 supplies the generated encoded data of the geometry data to the position information decoding unit 102 and the bitstream generation unit 105.

[0077] The position information decoding unit 102 acquires the coded data of the geometry data supplied from the position information encoding unit 101 and decodes the coded data. This decoding method may be any method compatible with the coding performed by the position information encoding unit 101. For example, processing such as filtering or inverse quantization for denoising may be performed. The position information decoding unit 102 supplies the generated geometry data (decoded result) to the point cloud generation unit 103.

[0078] The point cloud generation unit 103 acquires attribute data (attribute information) of the point cloud input to the encoding device 100 and geometry data (decoded result) supplied from the position information decoding unit 102. The point cloud generation unit 103 performs processing (recolor processing) to match the attribute data with the geometry data (decoded result). The point cloud generation unit 103 supplies the attribute data associated with the geometry data (decoded result) to the attribute information encoding unit 104.

[0079] The attribute information encoding unit 104 acquires the geometry data (decoded result) and attribute data supplied from the point cloud generation unit 103. The attribute information encoding unit 104 uses the geometry data (decoded result) to encode the attribute data and generate encoded data of the attribute data.

[0080] At this time, the attribute information encoding unit 104 applies the above-described present technology (method 1) to encode the attribute data. The attribute information encoding unit 104 supplies the generated encoded data of the attribute data to the bit stream generation unit 105.

[0081] The bitstream generation unit 105 obtains coded data of geometry data supplied from the position information encoding unit 101. The bitstream generation unit 105 also obtains coded data of attribute data supplied from the attribute information encoding unit 104. The bitstream generation unit 105 generates a bitstream including these coded data. The bitstream generation unit 105 outputs the generated bitstream to the outside of the encoding device 100.

[0082] With this configuration, the encoding device 100 can determine the center of gravity of points in the layering of attribute data and set a reference point based on the center of gravity. By setting a point close to the center of gravity as the reference point in this way, it is possible to set a point close to more other points as the reference point. Therefore, it is possible to set a reference point so as to suppress a decrease in prediction accuracy for more prediction points, and it is possible to suppress a decrease in encoding efficiency.

[0083] Each of these processing units (position information encoding unit 101 to bitstream generation unit 105) of the encoding device 100 may have any configuration. For example, each processing unit may be configured with a logic circuit that realizes the above-described processing. Furthermore, each processing unit may have, for example, a central processing unit (CPU), read-only memory (ROM), random access memory (RAM), etc., and may execute a program using these to realize the above-described processing. Of course, each processing unit may have both of these configurations, and may implement part of the above-described processing using a logic circuit and other parts by executing a program. The configurations of each processing unit may be independent of each other. For example, some processing units may implement part of the above-described processing using a logic circuit, other processing units may implement the above-described processing by executing a program, and still other processing units may implement the above-described processing by both a logic circuit and by executing a program.

[0084] <Attribute information encoding part> Fig. 14 is a block diagram showing an example of the main configuration of the attribute information encoding unit 104 (Fig. 13). Note that Fig. 14 shows the main processing units, data flows, etc., and is not limited to what is shown in Fig. 14. In other words, the attribute information encoding unit 104 may include processing units not shown as blocks in Fig. 14, and may include processing and data flows not shown as arrows, etc. in Fig. 14.

[0085] As shown in FIG. 14, the attribute information encoding unit 104 includes a layering processing unit 111, a quantization unit 112, and an encoding unit 113.

[0086] The layering processing unit 111 performs processing related to layering of attribute data. For example, the layering processing unit 111 acquires attribute data and geometry data (decoded results) supplied from the point cloud generation unit 103. The layering processing unit 111 layers the attribute data using the geometry data. At this time, the layering processing unit 111 performs layering by applying the present technology (method 1) described above. That is, the layering processing unit 111 derives the center of gravity of points in each layer and selects a reference point based on the center of gravity. The layering processing unit 111 then sets a reference relationship in each layer of the hierarchical structure, and based on the reference relationship, derives a predicted value of the attribute data of each prediction point using the attribute data of the reference point, and derives a difference value between the attribute data and the predicted value. The layering processing unit 111 supplies the layered attribute data (difference value) to the quantization unit 112.

[0087] At this time, the layering processing unit 111 may also generate control information related to layering. The layering processing unit 111 may also supply the generated control information to the quantization unit 112 together with attribute data (difference values).

[0088] The quantization unit 112 acquires the attribute data (difference values) and control information supplied from the layering processing unit 111. The quantization unit 112 quantizes the attribute data (difference values). Any quantization method may be used. The quantization unit 112 supplies the quantized attribute data (difference values) and control information to the encoding unit 113.

[0089] The encoding unit 113 acquires the quantized attribute data (difference values) and control information supplied from the quantization unit 112. The encoding unit 113 encodes the quantized attribute data (difference values) to generate coded data of the attribute data. Any coding method may be used. The encoding unit 113 also includes control information in the generated coded data. In other words, the encoding unit 113 generates coded data of the attribute data including the control information. The encoding unit 113 supplies the generated coded data to the bitstream generation unit 105.

[0090] By performing layering as described above, the attribute information encoding unit 104 can set points close to the center of gravity as reference points, and can therefore set points close to more other points as reference points. This means that reference points can be set so as to suppress a decrease in prediction accuracy for more prediction points, and a decrease in encoding efficiency can be suppressed.

[0091] These processing units (layering processing unit 111 to encoding unit 113) may have any configuration. For example, each processing unit may be configured with a logic circuit that realizes the above-mentioned processing. Furthermore, each processing unit may have, for example, a CPU, ROM, RAM, etc., and may execute a program using these to realize the above-mentioned processing. Of course, each processing unit may have both of these configurations, and may realize part of the above-mentioned processing using a logic circuit and the other part by executing a program. The configurations of each processing unit may be independent of each other. For example, some processing units may realize part of the above-mentioned processing using a logic circuit, other processing units may execute a program to realize the above-mentioned processing, and still other processing units may realize the above-mentioned processing using both a logic circuit and by executing a program.

[0092] <Hierarchical Processing Unit> Fig. 15 is a block diagram showing an example of the main configuration of the layering processing unit 111 (Fig. 14). Note that Fig. 15 shows the main processing units, data flows, etc., and does not necessarily show everything. In other words, the layering processing unit 111 may include processing units that are not shown as blocks in Fig. 15, and may include processing and data flows that are not shown as arrows, etc. in Fig. 15.

[0093] As shown in FIG. 15, the layering processing unit 111 has a reference point setting unit 121, a reference relationship setting unit 122, an inversion unit 123, and a weight value derivation unit .

[0094] The reference point setting unit 121 performs processing related to setting reference points. For example, based on the geometry data of each point, the reference point setting unit 121 classifies the points to be processed into reference points whose attribute data are referenced and prediction points from which predicted values ​​of the attribute data are derived. That is, the reference point setting unit 121 sets reference points and prediction points. The reference point setting unit 121 recursively repeats this processing for the reference points. That is, the reference point setting unit 121 sets reference points and prediction points for the layer to be processed, with the reference point set in the previous layer as the processing target. In this way, a hierarchical structure is constructed. That is, the attribute data is layered. The reference point setting unit 121 supplies information indicating the reference points and prediction points for each layer that has been set to the reference relationship setting unit 122.

[0095] The reference relationship setting unit 122 performs processing related to setting the reference relationship for each layer based on the information supplied from the reference point setting unit 121. That is, for each prediction point in each layer, the reference relationship setting unit 122 sets a reference point (i.e., a reference destination) to be referenced to derive the predicted value of the prediction point. Then, the reference relationship setting unit 122 derives a predicted value of the attribute data of each prediction point based on the reference relationship. That is, the reference relationship setting unit 122 derives a predicted value of the attribute data of the prediction point using the attribute data of the reference point set as the reference destination. Furthermore, the reference relationship setting unit 122 derives a difference value between the derived predicted value and the attribute data of the prediction point. The reference relationship setting unit 122 supplies the derived difference value (layered attribute data) to the inversion unit 123 for each layer.

[0096] The reference point setting unit 121 generates control information and the like relating to the layering of the attribute data as described above, and supplies it to the quantization unit 112, which can then transmit it to the decoding side.

[0097] The inversion unit 123 performs processing related to layer inversion. For example, the inversion unit 123 acquires layered attribute data supplied from the reference relationship setting unit 122. In this attribute data, information for each layer is layered in the order in which it was generated. The inversion unit 123 inverts the layer of the attribute data. For example, the inversion unit 123 assigns a layer number (a number for identifying a layer in which the top layer is 0, the value is incremented by 1 for each layer down, and the lowest layer has the maximum value) to each layer of the attribute data in the reverse order of the order in which it was generated, so that the order of generation is from the lowest layer to the top layer. The inversion unit 123 supplies the layer-inverted attribute data to the weight value derivation unit 124.

[0098] The weight value derivation unit 124 performs processing related to weighting. For example, the weight value derivation unit 124 acquires attribute data supplied from the inversion unit 123. The weight value derivation unit 124 derives weight values ​​for the acquired attribute data. Any method for deriving these weight values ​​may be used. The weight value derivation unit 124 supplies the attribute data (difference value) and the derived weight values ​​to the quantization unit 112 (FIG. 14). The weight value derivation unit 124 may also supply the derived weight values ​​to the quantization unit 112 as control information, which may be transmitted to the decoding side.

[0099] In the layering processing unit 111 as described above, the reference point setting unit 121 can apply the above-described present technology. That is, the reference relationship setting unit 122 can apply the above-described "Method 1" to derive the center of gravity of the points and set the reference points based on the center of gravity. In this way, it is possible to suppress a decrease in prediction accuracy and a decrease in coding efficiency.

[0100] The layering procedure is arbitrary. For example, the processing by the reference point setting unit 121 and the processing by the reference relationship setting unit 122 may be executed in parallel. For example, for each layer, the reference point setting unit 121 may set reference points and prediction points, and the reference relationship setting unit 122 may set the reference relationship.

[0101] These processing units (reference point setting unit 121 to weight value derivation unit 124) may have any configuration. For example, each processing unit may be configured with a logic circuit that realizes the above-mentioned processing. Also, each processing unit may have, for example, a CPU, ROM, RAM, etc., and may realize the above-mentioned processing by executing a program using these. Of course, each processing unit may have both of these configurations, and may realize part of the above-mentioned processing by a logic circuit and other parts by executing a program. The configurations of each processing unit may be independent of each other. For example, some processing units may realize part of the above-mentioned processing by a logic circuit, other processing units may realize the above-mentioned processing by executing a program, and still other processing units may realize the above-mentioned processing by both a logic circuit and by executing a program.

[0102] <Encoding process flow> Next, a description will be given of the processing executed by the encoding device 100. The encoding device 100 encodes point cloud data by executing an encoding process. An example of the flow of this encoding process will be described with reference to the flowchart in FIG.

[0103] When the encoding process starts, in step S101, the position information encoding unit 101 of the encoding device 100 encodes the geometry data (position information) of the input point cloud, and generates encoded data of the geometry data.

[0104] In step S102, the position information decoding unit 102 decodes the coded data of the geometry data generated in step S101 to generate position information.

[0105] In step S103, the point cloud generation unit 103 performs recolor processing using the attribute data (attribute information) of the input point cloud and the geometry data (decoding result) generated in step S102, and associates the attribute data with the geometry data.

[0106] In step S104, the attribute information encoding unit 104 performs attribute information encoding processing to encode the attribute data recolored in step S103 and generate encoded data of the attribute data. At this time, the attribute information encoding unit 104 performs processing by applying the present technology (method 1) described above. For example, the attribute information encoding unit 104 derives the center of gravity of points in layering the attribute data and sets a reference point based on the center of gravity. The attribute information encoding processing will be described in detail later.

[0107] In step S105, the bitstream generating unit 105 generates and outputs a bitstream including the coded data of the geometry data generated in step S101 and the coded data of the attribute data generated in step S104.

[0108] When the process of step S105 is completed, the encoding process ends.

[0109] By performing the processes at each step in this manner, the encoding device 100 can suppress a decrease in prediction accuracy and a decrease in encoding efficiency.

[0110] <Attribute information encoding process flow> Next, an example of the flow of the attribute information encoding process executed in step S104 in FIG. 16 will be described with reference to the flowchart in FIG.

[0111] When the attribute information encoding process starts, in step S111, the layering processing unit 111 of the attribute information encoding unit 104 performs layering processing to layer the attribute data. That is, a reference point and a prediction point for each layer are set, and further, a reference relationship between them is set. At this time, the layering processing unit 111 performs layering by applying the above-described present technology (method 1). For example, when layering the attribute data, the attribute information encoding unit 104 derives the center of gravity of the points and sets the reference point based on the center of gravity. The layering process will be described in detail later.

[0112] In step S112, the layering processing unit 111 derives a predicted value of the attribute data for each prediction point in each layer of the attribute data layered in step S111, and derives a difference value between the attribute data of the prediction point and its predicted value.

[0113] In step S113, the quantization unit 112 quantizes each of the difference values ​​derived in step S112.

[0114] In step S114, the encoding unit 113 encodes the difference value quantized in step S112 to generate encoded data of the attribute data.

[0115] When the process of step S114 ends, the attribute information encoding process ends, and the process returns to FIG.

[0116] By performing the processing of each step in this manner, the layering processing unit 111 can apply the above-mentioned "Method 1" to derive the center of gravity of the points in layering the attribute data and set the reference point based on the center of gravity. Therefore, the layering processing unit 111 can layer the attribute data in a way that suppresses a decrease in prediction accuracy, thereby suppressing a decrease in encoding efficiency.

[0117] <Layering process flow> Next, an example of the flow of the layering process executed in step S111 of FIG. 17 will be described with reference to the flowchart of FIG.

[0118] When the layering process is started, in step S121, the reference point setting unit 121 of the layering processing unit 111 sets the value of the variable LoD Index, which indicates the layer to be processed, to an initial value (for example, "0").

[0119] In step S122, the reference point setting unit 121 executes a reference point setting process to set reference points in the layer to be processed (that is, to set prediction points as well). Details of the reference point setting process will be described later.

[0120] In step S123, the reference relationship setting unit 122 sets the reference relationship of the layer to be processed (which reference point is to be referenced in deriving the predicted value of each prediction point).

[0121] In step S124, the reference point setting unit 121 increments the LoD Index and sets the processing target to the next layer.

[0122] In step S125, the reference point setting unit 121 determines whether all points have been processed. If it is determined that unprocessed points exist, that is, if it is determined that layering has not been completed, the process returns to step S122 and the subsequent processes are repeated. In this way, each process from step S122 to step S125 is executed for each layer, and if it is determined in step S125 that all points have been processed, the process proceeds to step S126.

[0123] In step S126, the inverting unit 123 inverts the hierarchy of the attribute data generated as described above, and assigns a hierarchy number to each hierarchy in the reverse order of the generation.

[0124] In step S127, the weight value derivation unit 124 derives a weight value for the attribute data of each hierarchy.

[0125] When the process of step S127 ends, the process returns to FIG.

[0126] By performing the processing of each step in this manner, the layering processing unit 111 can apply the above-mentioned "Method 1" to derive the center of gravity of the points in layering the attribute data and set the reference point based on the center of gravity. Therefore, the layering processing unit 111 can layer the attribute data in a way that suppresses a decrease in prediction accuracy, thereby suppressing a decrease in encoding efficiency.

[0127] <Flow of reference relationship setting process> Next, an example of the flow of the reference point setting process executed in step S122 of FIG. 18 will be described with reference to the flowchart of FIG.

[0128] When the reference point setting process is started, the reference point setting unit 121 identifies a group of points to be used for deriving the center of gravity in step S141, and derives the center of gravity of the group of points to be processed. As described above, any method for deriving the center of gravity may be used. For example, the center of gravity may be derived using any of the methods shown in the table of FIG.

[0129] In step S142, the reference point setting unit 121 selects a point close to the center of gravity derived in step S141 as a reference point. The method for selecting this reference point is arbitrary. For example, the center of gravity may be derived using one of the methods shown in the table of FIG.

[0130] When the process of step S142 ends, the reference point setting process ends, and the process returns to FIG.

[0131] By performing the processing of each step in this manner, the reference point setting unit 121 can apply the above-mentioned "Method 1" to derive the center of gravity of the points in layering the attribute data and set the reference point based on the center of gravity. Therefore, the layering processing unit 111 can layer the attribute data so as to suppress a decrease in prediction accuracy, thereby suppressing a decrease in coding efficiency.

[0132] <Decryption device> Next, another example of a device to which the present technology is applied will be described. Fig. 20 is a block diagram showing an example of the configuration of a decoding device, which is one aspect of an information processing device to which the present technology is applied. The decoding device 200 shown in Fig. 20 is a device that decodes encoded data of a point cloud (3D data). The decoding device 200 decodes the encoded data of the point cloud by applying the present technology (method 1) described in this embodiment.

[0133] Note that Fig. 20 shows the main processing units, data flows, etc., and does not necessarily show everything. That is, in the decoding device 200, there may be processing units that are not shown as blocks in Fig. 20, and there may be processing and data flows that are not shown as arrows, etc. in Fig. 20.

[0134] As shown in FIG. 20, the decoding device 200 includes an encoded data extraction unit 201, a position information decoding unit 202, an attribute information decoding unit 203, and a point cloud generation unit 204.

[0135] The coded data extraction unit 201 acquires and holds a bitstream input to the decoding device 200. The coded data extraction unit 201 extracts coded data of geometry data (position information) and attribute data (attribute information) from the bitstream it holds. The coded data extraction unit 201 supplies the coded data of the extracted geometry data to the position information decoding unit 202. The coded data extraction unit 201 supplies the coded data of the extracted attribute data to the attribute information decoding unit 203.

[0136] The position information decoding unit 202 acquires the coded data of the geometry data supplied from the coded data extraction unit 201. The position information decoding unit 202 decodes the coded data of the geometry data to generate geometry data (decoded result). This decoding method may be any method similar to that used by the position information decoding unit 102 of the encoding device 100. The position information decoding unit 202 supplies the generated geometry data (decoded result) to the attribute information decoding unit 203 and the point cloud generation unit 204.

[0137] The attribute information decoding unit 203 acquires the coded data of the attribute data supplied from the coded data extraction unit 201. The attribute information decoding unit 203 acquires the geometry data (decoded result) supplied from the position information decoding unit 202. The attribute information decoding unit 203 uses the position information (decoded result) to decode the coded data of the attribute data by a method to which the present technology (method 1) described above is applied, and generates attribute data (decoded result). The attribute information decoding unit 203 supplies the generated attribute data (decoded result) to the point cloud generation unit 204.

[0138] The point cloud generation unit 204 acquires geometry data (decoding result) supplied from the position information decoding unit 202. The point cloud generation unit 204 acquires attribute data (decoding result) supplied from the attribute information decoding unit 203. The point cloud generation unit 204 generates a point cloud (decoding result) using the geometry data (decoding result) and attribute data (decoding result). The point cloud generation unit 204 outputs the generated point cloud (decoding result) data to the outside of the decoding device 200.

[0139] With this configuration, the decoding device 200 can select a point close to the center of gravity of points as a reference point in inverse layering. Therefore, the decoding device 200 can correctly decode the coded data of attribute data coded by, for example, the above-mentioned coding device 100. Therefore, it is possible to suppress a decrease in prediction accuracy and a decrease in coding efficiency.

[0140] These processing units (the encoded data extraction unit 201 to the point cloud generation unit 204) may have any configuration. For example, each processing unit may be configured with a logic circuit that realizes the above-described processing. Furthermore, each processing unit may have, for example, a CPU, a ROM, a RAM, etc., and may execute a program using these to realize the above-described processing. Of course, each processing unit may have both of these configurations, and may realize part of the above-described processing using a logic circuit and the other part by executing a program. The configurations of the processing units may be independent of each other. For example, some processing units may realize part of the above-described processing using a logic circuit, other processing units may execute a program to realize the above-described processing, and still other processing units may realize the above-described processing using both a logic circuit and by executing a program.

[0141] <Attribute information decoding unit> Fig. 21 is a block diagram showing a main example configuration of the attribute information decoding unit 203 (Fig. 20). Note that Fig. 21 shows main processing units, data flows, etc., and does not necessarily show everything. In other words, the attribute information decoding unit 203 may include processing units not shown as blocks in Fig. 21, or processes and data flows not shown as arrows, etc. in Fig. 21.

[0142] As shown in FIG. 21, the attribute information decoding unit 203 includes a decoding unit 211, an inverse quantization unit 212, and an inverse layer processing unit 213.

[0143] The decoding unit 211 performs processing related to decoding of the coded data of the attribute data. For example, the decoding unit 211 obtains the coded data of the attribute data supplied to the attribute information decoding unit 203.

[0144] The decoding unit 211 decodes the coded data of the attribute data to generate attribute data (decoded result). This decoding method may be any method that corresponds to the coding method used by the coding unit 113 (FIG. 14) of the coding device 100. The generated attribute data (decoded result) corresponds to the attribute data before coding, is a difference value between the attribute data and its predicted value, and is quantized. The decoding unit 211 supplies the generated attribute data (decoded result) to the inverse quantization unit 212.

[0145] If the coded data of the attribute data includes control information related to weight values ​​and control information related to the layering of the attribute data, the decoding unit 211 also supplies this control information to the inverse quantization unit 212 .

[0146] The inverse quantization unit 212 performs processing related to inverse quantization of attribute data. For example, the inverse quantization unit 212 acquires attribute data (decoding results) and control information supplied from the decoding unit 211.

[0147] The inverse quantization unit 212 inverse quantizes the attribute data (decoded result). At this time, if control information related to weight values ​​is supplied from the decoding unit 211, the inverse quantization unit 212 also acquires the control information and inverse quantizes the attribute data (decoded result) based on the control information (using weight values ​​derived based on the control information).

[0148] Furthermore, when control information relating to the layering of attribute data is supplied from the decoding unit 211, the inverse quantization unit 212 also acquires the control information.

[0149] The inverse quantization unit 212 supplies the inverse quantized attribute data (decoded result) to the inverse layering processing unit 213. Furthermore, when control information related to layering of the attribute data is acquired from the decoding unit 211, the inverse quantization unit 212 also supplies the control information to the inverse layering processing unit 213.

[0150] The inverse layering processing unit 213 acquires the inversely quantized attribute data (decoded result) supplied from the inverse quantization unit 212. As described above, this attribute data is a differential value. The inverse layering processing unit 213 also acquires the geometry data (decoded result) supplied from the position information decoding unit 202. Using the geometry data, the inverse layering processing unit 213 performs inverse layering on the acquired attribute data (differential value), which is the inverse process of the layering performed by the layering processing unit 111 ( FIG. 14 ) of the encoding device 100.

[0151] Here, the inverse layering will be described. For example, the inverse layering processor 213 layers the attribute data using a method similar to that of the encoding device 100 (layering processor 111) based on the geometry data supplied from the position information decoder 202. That is, the inverse layering processor 213 sets reference points and prediction points for each layer based on the decoded geometry data, and sets a hierarchical structure for the attribute data. The inverse layering processor 213 further uses the reference points and prediction points to set the reference relationships for each layer of the hierarchical structure (the reference destinations for each prediction point).

[0152] The inverse-layering processor 213 then inversely layers the acquired attribute data (difference values) using the hierarchical structure and the reference relationships between each layer. That is, the inverse-layering processor 213 derives a predicted value of the prediction point from the reference point according to the reference relationship, and adds the predicted value to the difference value to restore the attribute data of each prediction point. The inverse-layering processor 213 performs this process for each layer, from the upper layer to the lower layer. That is, the inverse-layering processor 213 uses a prediction point from which attribute data has been restored in a layer higher than the layer being processed as a reference point, and restores the attribute data of the prediction point in the layer being processed as described above.

[0153] In the de-hierarchization performed in this manner, the de-hierarchization processor 213 applies the present technology (Method 1) described above to set reference points when layering attribute data based on decoded geometry data. That is, the de-hierarchization processor 213 derives the center of gravity of the points and selects points close to the center of gravity as reference points. The de-hierarchization processor 213 supplies the de-hierarchized attribute data to the point cloud generator 204 ( FIG. 20 ) as the decoding result.

[0154] By performing inverse layering as described above, the inverse layering processor 213 can set a point close to the center of gravity as the reference point, thereby enabling layering of attribute data in a manner that suppresses a decrease in prediction accuracy. That is, the attribute information decoding unit 203 can correctly decode coded data coded using a similar method. For example, the attribute information decoding unit 203 can correctly decode coded data of attribute data coded by the above-described attribute information coding unit 104. Therefore, a decrease in coding efficiency can be suppressed.

[0155] These processing units (the decoding unit 211 to the inverse layering processing unit 213) may have any configuration. For example, each processing unit may be configured with a logic circuit that realizes the above-described processing. Furthermore, each processing unit may have, for example, a CPU, a ROM, a RAM, etc., and may realize the above-described processing by executing a program using these. Of course, each processing unit may have both of these configurations, and may realize part of the above-described processing by a logic circuit and the other by executing a program. The configurations of the processing units may be independent of each other. For example, some processing units may realize part of the above-described processing by a logic circuit, other processing units may realize the above-described processing by executing a program, and still other processing units may realize the above-described processing by both a logic circuit and by executing a program.

[0156] <Decryption process flow> Next, a description will be given of the processing executed by the decoding device 200. The decoding device 200 decodes the encoded data of the point cloud by executing a decoding process. An example of the flow of this decoding process will be described with reference to the flowchart in FIG.

[0157] When the decoding process starts, in step S201, the coded data extraction unit 201 of the decoding device 200 acquires and stores a bitstream, and extracts coded data of geometry data and coded data of attribute data from the bitstream.

[0158] In step S202, the position information decoding unit 202 decodes the coded data of the extracted geometry data to generate geometry data (decoded result).

[0159] In step S203, the attribute information decoding unit 203 executes attribute information decoding processing, decodes the coded data of the attribute data extracted in step S201, and generates attribute data (decoded result). At this time, the attribute information decoding unit 203 performs processing by applying the above-described present technology (method 1). For example, in layering the attribute data, the attribute information decoding unit 203 derives the center of gravity of the points and sets a point close to the center of gravity as a reference point. Details of the attribute information decoding processing will be described later.

[0160] In step S204, the point cloud generation unit 204 generates and outputs a point cloud (decoding result) using the geometry data (decoding result) generated in step S202 and the attribute data (decoding result) generated in step S203.

[0161] When the process of step S204 ends, the decoding process ends.

[0162] By performing the processing of each step in this manner, the decoding device 200 can correctly decode coded data of attribute data coded by a similar technique. For example, the decoding device 200 can correctly decode coded data of attribute data coded by the above-described coding device 100. Therefore, it is possible to suppress a decrease in prediction accuracy and a decrease in coding efficiency.

[0163] <Flow of attribute information decoding process> Next, an example of the flow of the attribute information decoding process executed in step S203 of FIG. 22 will be described with reference to the flowchart of FIG.

[0164] When the attribute information decoding process starts, in step S211, the decoding unit 211 of the attribute information decoding unit 203 decodes the coded data of the attribute data to generate attribute data (decoded result). This attribute data (decoded result) has been quantized as described above.

[0165] In step S212, the inverse quantization unit 212 performs inverse quantization processing to inverse quantize the attribute data (decoding result) generated in step S211.

[0166] In step S213, the inverse-layering processor 213 performs inverse-layering processing to inversely layer the attribute data (difference values) inversely quantized in step S212 and derive attribute data for each point. At this time, the inverse-layering processor 213 performs inverse layering by applying the present technology (method 1) described above. For example, in layering the attribute data, the inverse-layering processor 213 derives the center of gravity of the points and sets a point close to the center of gravity as a reference point. Details of the inverse-layering processing will be described later.

[0167] When the process of step S213 ends, the attribute information decoding process ends, and the process returns to FIG.

[0168] By performing the processing of each step in this manner, the attribute information decoding unit 203 can apply the above-described "Method 1" and set a point close to the center of gravity of the points as a reference point when layering the attribute data. Therefore, the inverse layering processing unit 213 can layer the attribute data so as to suppress a decrease in prediction accuracy. In other words, the attribute information decoding unit 203 can correctly decode coded data coded using a similar method. For example, the attribute information decoding unit 203 can correctly decode coded data of attribute data coded by the above-described attribute information coding unit 104. Therefore, a decrease in coding efficiency can be suppressed.

[0169] <Flow of reverse layering process> Next, an example of the flow of the layer inversion process executed in step S213 of FIG. 23 will be described with reference to the flowchart of FIG.

[0170] When the de-layering process starts, in step S221 the de-layering processor 213 performs layering process on the attribute data (decoded result) using the geometry data (decoded result), restores the reference points and prediction points of each layer set on the encoding side, and also restores the reference relationship between each layer. In other words, the de-layering processor 213 performs the same process as the layering process performed by the layering processor 111, sets the reference points and prediction points of each layer, and also sets the reference relationship between each layer.

[0171] For example, the inverse layering processor 213, like the layering processor 111, applies the above-mentioned "method 1" to derive the center of gravity of the points and set the points closest to this center of gravity as reference points.

[0172] In step S222, the inverse-layering processor 213 inversely layers the attribute data (decoded result) using this hierarchical structure and reference relationship, and restores the attribute data of each point. That is, the inverse-layering processor 213 derives a predicted value of the attribute data of the prediction point from the attribute data of the reference point based on the reference relationship, and adds the predicted value to the difference value of the attribute data (decoded result) to restore the attribute data.

[0173] When the process of step S222 is completed, the layer inverse process is completed, and the process returns to FIG.

[0174] By performing the processing of each step in this manner, the inverse layering processor 213 can achieve layering similar to that achieved during encoding. That is, the attribute information decoder 203 can correctly decode coded data that has been coded using a similar method. For example, the attribute information decoder 203 can correctly decode coded data of attribute data coded by the attribute information encoder 104 described above. This makes it possible to suppress a decrease in coding efficiency.

[0175] 3. Second Embodiment <Method 2> Next, a case where the above-mentioned "Method 2" is applied will be described with reference to Fig. 6. In the case of this "Method 2," in the layering of attribute data, reference points are selected according to the distribution pattern (distribution mode) of points.

[0176] For example, in the table (table information) shown in Fig. 25, the distribution pattern (distribution mode) of points in the processing target region where the reference point is set is associated with information (index) indicating the point to be selected in that case. For example, the second row from the top of this table shows that when the distribution pattern of points in the 2x2x2 voxel region where the reference point is set is "10100001," the point with index "2," i.e., the point that appears second, is selected. Each bit value of the distribution pattern "10100001" indicates the presence or absence of a point in each 2x2x2 voxel, with a value of "1" indicating that a point exists in the voxel to which that bit is assigned, and a value of "0" indicating that a point does not exist in the voxel to which that bit is assigned.

[0177] Similarly, this table shows the index of the point to be selected for each distribution pattern. In other words, when layering attribute data, this table is referenced and the point with the index corresponding to the distribution pattern of points in the processing target area for which the reference point is to be set is selected as the reference point.

[0178] In this way, the reference point can be selected more easily.

[0179] This table information may be any information that links the distribution pattern of points with information indicating the points to be selected. For example, it may be table information that selects points close to the center of gravity for each point distribution pattern, as in method (1) shown in the second row from the top of the "table" shown in A of Fig. 26. In other words, each distribution pattern may be linked to an index of a point close to the center of gravity for that distribution pattern.

[0180] Also, for example, as in method (2) shown in the third row from the top of the "table" shown in A of Fig. 26, it may be table information in which arbitrary points are selected for each point distribution pattern. Furthermore, for example, as in method (3) shown in the fourth row from the top of this "table," it may be possible to select a table to use from among multiple tables. For example, it may be possible to switch the table to be used depending on the hierarchy (depth of LoD).

[0181] This table information may be prepared in advance. For example, predetermined table information may be defined by a standard. In this case, signaling of the table information (transmission from the encoding side to the decoding side) is not required.

[0182] Alternatively, the position information decoding unit 102 may derive this table information from the geometry data. Alternatively, the position information decoding unit 202 may also derive this table information from the geometry data (decoded result). In this case, signaling of the table information (transmission from the encoding side to the decoding side) is not required.

[0183] Of course, this table information may be generated (or updated) by a user, an application, or the like. In this case, the generated (or updated) table information may be signaled. That is, for example, the encoding unit 113 may encode information related to this table information and signal it by including the encoded data in a bitstream.

[0184] As described above, this table information may be switched depending on the hierarchy (depth of LoD). In this case, the switching method may be specified in advance in a standard or the like, and information indicating the switching method may not be signaled, as in method (1) shown in the second row from the top of the "Table" shown in B of Fig. 26.

[0185] Alternatively, an index (identification information) indicating the selected table may be signaled, as in method (2) shown in the third row from the top of the "Table" table in B of Fig. 26. For example, this index may be signaled in an attribute parameter set (AttributeParameterSet).

[0186] Furthermore, the selected table information itself may be signaled, for example, as in method (3) shown in the fourth row from the top of the "Table" table shown in B of Fig. 26. For example, this table information may be signaled in an attribute brick header.

[0187] Alternatively, a portion of the selected table information may be signaled, as shown in method (4) in the fifth row from the top of the "Table" table in Figure 26B. In other words, it may be possible to partially update the table information. For example, this table information may be signaled in an AttributeBrickHeader.

[0188] When this method 2 is applied, the configurations of the encoding device 100 and the decoding device 200 are basically the same as when the above-described method 1 is applied. Therefore, the encoding device 100 can execute each process, such as the encoding process, the attribute information encoding process, and the layering process, in the same manner as in the first embodiment.

[0189] <Reference point setting process flow> An example of the flow of the reference point setting process in this case will be described with reference to the flowchart in Fig. 27. When the reference point setting process starts, in step S301, the reference point setting unit 121 refers to the table information and selects reference points according to the point distribution pattern.

[0190] In step S302, the reference point setting unit 121 determines whether or not to signal information about the used table. If it is determined that information should be signaled, the process proceeds to step S303.

[0191] In step S303, the reference point setting unit 121 signals information about the table used. When the process of step S303 ends, the reference point setting process ends, and the process returns to FIG.

[0192] By transmitting the table information from the encoding device 100 in this way, the decoding device 200 can perform decoding using the table information.

[0193] The decoding device 200 can execute the decoding process, the attribute information decoding process, the inverse layering process, and other processes in the same manner as in the first embodiment.

[0194] <4. Third Embodiment> <Method 3> Next, a case where the above-mentioned "Method 3" is applied will be described with reference to Fig. 6. In the case of this "Method 3", a reference point set in the hierarchical structure of attribute data may be signaled.

[0195] For example, as shown in method (1) in the second row from the top of the "Signaling Target" table in Fig. 28, information indicating whether to reference all nodes (all points), i.e., whether to treat them as reference points or prediction points, may be signaled. For example, as shown in A in Fig. 29, all nodes in all hierarchical levels may be sorted in Morton order, and an index (index 0 to index K) may be assigned to each node. In other words, each node (and the information assigned to each node) can be identified by index 0 to index K.

[0196] Also, for example, as in method (2) shown in the third row from the top of the "Signaling Target" table in Fig. 28, it is possible to signal information indicating which node (point) to use as a reference point for selection for some layers. For example, as shown in B in Fig. 29, it is possible to specify the layer (LoD) to be the target of signaling, sort all nodes in that layer in Morton order, and assign an index to each node.

[0197] For example, suppose the attribute data has a hierarchical structure as shown in A of Fig. 30. That is, one point each from point #0 of LoD2, point #1 of LoD2, and point #2 of LoD2 is selected as a reference point to form point #0 of LoD1. In this case, if indexes are assigned to the points of LoD2 in the search order shown in B of Fig. 30, each point of LoD1 #0 will be represented by the LoD2 index as shown in C of Fig. 30. In other words, the distribution pattern of points of LoD1 #0 can be represented by specifying "LoD2 0,1,0".

[0198] In this way, the index of LoD N-1 can specify a 2x2x2 voxel region of LoD N. In other words, one 2x2x2 voxel can be specified by the LoD (layer specification) and order (mth). By specifying by layer and index in this way, signaling can be performed for only some layers as needed, which makes it possible to suppress an increase in the amount of code compared to method (1) and to suppress a decrease in coding efficiency.

[0199] Furthermore, for example, as in method (3) shown in the fourth row from the top of the "Signaling Target" table in FIG. 28, the points to be signaled may be limited by the number of points in the lower NxNxN voxel region. In other words, information regarding the setting of reference points for points that satisfy a predetermined condition may be signaled. By doing so, it is possible to further reduce the number of nodes to be signaled, as shown in C of FIG. 29. This makes it possible to suppress a decrease in coding efficiency.

[0200] For example, suppose attribute data has a hierarchical structure as shown in A of FIG. 31. In this case, suppose the signaling target is limited to a 2x2x2 voxel region including three or more points. If indices are assigned to points of LoD2 in the search order as shown in B of FIG. 31, the voxel region shown to the right of LoD2 is excluded from the signaling target. Therefore, no index is assigned to this voxel region. Therefore, as shown in C of FIG. 31, the amount of signaling data can be reduced compared to C of FIG. 30, and an increase in the amount of coding can be suppressed.

[0201] It is also possible to apply a combination of methods (2) and (3), for example, as in method (4) shown in the fifth row from the top of the "Signaling Target" table in Figure 28.

[0202] <Fixed-length signaling> The above signaling may be performed using fixed-length data. For example, signaling may be performed using the syntax shown in A of Figure 32. In the syntax of A of Figure 32, num_Lod is a parameter indicating the number of LoDs to be signaled. lodNo[i] is a parameter indicating the LoD number. voxelType[i] is a parameter indicating the type of voxel to be signaled. By specifying this parameter, it is possible to limit the transmission target within the Lod. num_node is a parameter indicating the actual number of signals. This num_node can be derived from the geometry data. node[k] indicates the information to be signaled for each 2x2x2 voxel region. k indicates the node number in Morton order.

[0203] If parsing is required before geometry data is obtained due to parallel processing or the like, signaling may be performed using the syntax shown in B of Fig. 32. If parsing is required in this way, num_node can be signaled.

[0204] Furthermore, when signaling is performed using fixed-length data, the syntax shown in A of FIG. 33 may be applied. In this case, a flag Flag[k] that controls the signal is signaled. In this case, if parsing is required before geometry data is obtained due to parallel processing or the like, signaling may be performed using the syntax shown in B of FIG. 33. When parsing is required in this way, num_node can be signaled.

[0205] <Variable length signaling> Furthermore, the above signaling may be performed using variable-length data. For example, as shown in the example of Fig. 34, the position of a node within a 2x2x2 voxel region may be signaled. In this case, the bit length of the signaling may be set according to the number of nodes within the 2x2x2 voxel region, for example, based on the table information shown in A of Fig. 34.

[0206] On the decoding side, the number of nodes contained in a 2x2x2 voxel region can be found from the geometry data. Therefore, even if a bit string such as "10111010..." is input as shown in the example of B in Figure 34, the number of nodes contained in the 2x2x2 voxel region can be determined from the geometry data, so that division at an appropriate bit length can be performed and information for each voxel region can be correctly obtained.

[0207] Also, an index of the table information to be used may be signaled as in the example of Fig. 35. In this case, the bit length of the signaling may be made variable depending on the number of nodes in a 2x2x2 voxel region, for example, as shown in A of Fig. 35.

[0208] For example, based on the table information shown in A of FIG. 35, if the number of nodes is 5 to 8, 2 bits are assigned and the table information of B of FIG. 35 is selected. In this case, the bit string "00" indicates that the first node in a predetermined search order (e.g., Morton order) is selected. Also, the bit string "01" indicates that the first node in the reverse order (reverse) of the predetermined search order (e.g., Morton order) is selected. Furthermore, the bit string "10" indicates that the second node in the predetermined search order (no reverce) is selected. Also, the bit string "11" indicates that the second node in the reverse order (reverse) of the predetermined search order is selected.

[0209] Also, for example, based on the table information shown in A of FIG. 35, if the number of nodes is 3 or 4, 1 bit is assigned and the table information of C of FIG. 35 is selected. In this case, the bit string "0" indicates that the first node in a predetermined search order (e.g., Morton order) is selected. Also, the bit string "1" indicates that the first node in the reverse order of the predetermined search order (e.g., Morton order) is selected.

[0210] On the decoding side, the number of nodes contained in a 2x2x2 voxel region can be found from the geometry data. Therefore, even if a bit string such as "10111010..." is input as shown in the example of D in Figure 35, the number of nodes contained in the 2x2x2 voxel region can be determined from the geometry data, so that division at an appropriate bit length can be performed and information for each voxel region can be correctly obtained.

[0211] An example of syntax for variable length is shown in Fig. 36. In the syntax in Fig. 36, bitLength is a parameter indicating the bit length, and signalType[i] is a parameter indicating the variable length coding method for each LoD.

[0212] In addition, even in this variable length case, if parsing is required before the geometry data is obtained due to parallel processing, etc., it is possible to signal num_node or flag[j] as in the syntax shown in Figure 37.

[0213] <Reference point setting process flow> An example of the flow of the reference point setting process in this case will be described with reference to the flowchart in Fig. 38. When the reference point setting process starts, the reference point setting unit 121 selects a reference point in step S321.

[0214] In step S322, reference point setting unit 121 signals information about the reference point set in step S321. When the processing of step S322 ends, the reference point setting processing ends, and the processing returns to Fig. 18. By transmitting the table information from encoding device 100 in this way, decoding device 200 can perform decoding using the table information.

[0215] 5. Fourth Embodiment <Method 4> Next, a case where the above-mentioned "Method 4" is applied will be described with reference to Fig. 6. In this "Method 4", the reference point may be selected alternately for each layer between a point closest to the center of the bounding box and a point farthest from the center of the bounding box among the reference point candidates.

[0216] For example, when selecting reference points at the positions shown in Fig. 39, points closer to the center of the bounding box and points farther from the center of the bounding box are selected alternately for each layer, as shown in A to C of Fig. 39. By doing so, the movement range of the reference points is limited to a narrow range as shown in the dotted line frame, as shown in C of Fig. 39, thereby suppressing a decrease in prediction accuracy.

[0217] By using the center of the bounding box as a reference for the direction of point selection in this way, it is possible to suppress a decrease in prediction accuracy regardless of the position within the bounding box.

[0218] Note that such point selection may be achieved by changing the search order of points depending on their position within the bounding box. For example, the search order may be based on the distance from the center of the bounding box. In this way, the search order can be changed depending on their position within the bounding box. Also, for example, the search order may be changed for each of the eight divided regions obtained by dividing the bounding box into eight (two in each of the x, y, and z directions).

[0219] <Reference point setting process flow> An example of the flow of the reference point setting process in this case will be described with reference to the flowchart in Fig. 40. When the reference point setting process starts, in step S341, the reference point setting unit 121 determines whether or not a point closer to the center of the bounding box has been selected as the reference point for the previous layer.

[0220] If it is determined that the closer point has been selected, the process proceeds to step S342.

[0221] In step S342, the reference point setting unit 121 selects, from among the reference point candidates, the point farthest from the center of the bounding box as the reference point. When the processing of step S342 ends, the reference point setting processing ends, and the processing returns to FIG. 18.

[0222] If it is determined in step S341 that the closer point has not been selected, the process proceeds to step S343.

[0223] In step S343, the reference point setting unit 121 selects, from among the reference point candidates, the point closest to the center of the bounding box as the reference point. When the processing of step S343 ends, the reference point setting processing ends, and the processing returns to FIG. 18.

[0224] By selecting reference points in this way, the encoding apparatus 100 can suppress a decrease in prediction accuracy of the reference points, thereby suppressing a decrease in encoding efficiency.

[0225] <6. Notes> <Hierarchization / reverse hierarchy method> Although Lifting has been used above as an example of a method for layering and delayering attribute information, this technology can be applied to any technology for layering attribute information. In other words, the method for layering and delayering attribute information may be other than Lifting. Furthermore, the method for layering and delayering attribute information may be a scalable method such as that described in Non-Patent Document 3, or a non-scalable method.

[0226] <Control information> Control information related to the present technology described in each of the above embodiments may be transmitted from the encoding side to the decoding side. For example, control information (e.g., enabled_flag) that controls whether or not to permit (or prohibit) application of the above-described present technology may be transmitted. Also, for example, control information that specifies a range (e.g., an upper or lower limit of a block size, or both, a slice, a picture, a sequence, a component, a view, a layer, etc.) within which application of the above-described present technology is permitted (or prohibited) may be transmitted.

[0227] <Surroundings / neighborhood> In this specification, the positional relationship such as "nearby" or "surrounding" may include not only a spatial positional relationship but also a temporal positional relationship.

[0228] <Computer> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs constituting the software are installed on a computer. Here, the term "computer" includes computers built into dedicated hardware, and general-purpose personal computers, etc., that can execute various functions by installing various programs.

[0229] FIG. 41 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program.

[0230] In a computer 900 shown in FIG. 41, a CPU (Central Processing Unit) 901, a ROM (Read Only Memory) 902, and a RAM (Random Access Memory) 903 are interconnected via a bus 904.

[0231] An input / output interface 910 is also connected to the bus 904. To the input / output interface 910, an input unit 911, an output unit 912, a storage unit 913, a communication unit 914, and a drive 915 are connected.

[0232] The input unit 911 includes, for example, a keyboard, a mouse, a microphone, a touch panel, an input terminal, etc. The output unit 912 includes, for example, a display, a speaker, an output terminal, etc. The storage unit 913 includes, for example, a hard disk, a RAM disk, a non-volatile memory, etc. The communication unit 914 includes, for example, a network interface. The drive 915 drives removable media 921 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.

[0233] In a computer configured as above, the CPU 901 performs the above-described series of processes by, for example, loading a program stored in the storage unit 913 into the RAM 903 via the input / output interface 910 and the bus 904 and executing the program. The RAM 903 also stores data necessary for the CPU 901 to execute various processes as appropriate.

[0234] The program executed by the computer can be applied by recording it on removable media 921 such as package media, for example. In this case, the program can be installed in storage unit 913 via input / output interface 910 by inserting removable media 921 into drive 915.

[0235] This program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, digital satellite broadcasting, etc. In this case, the program can be received by the communication unit 914 and installed in the storage unit 913.

[0236] Alternatively, this program can be installed in advance in the ROM 902 or the storage unit 913 .

[0237] <Applicable targets of this technology> While the above describes the application of this technology to the encoding and decoding of point cloud data, this technology is not limited to these examples and can be applied to the encoding and decoding of 3D data of any standard. In other words, as long as it does not conflict with the above-described technology, various processes such as encoding and decoding methods and specifications of various data such as 3D data and metadata are arbitrary. Furthermore, as long as it does not conflict with the above-described technology, some of the above-described processes and specifications may be omitted.

[0238] Furthermore, although the encoding device 100 and the decoding device 200 have been described above as application examples of the present technology, the present technology can be applied to any configuration.

[0239] For example, this technology can be applied to various electronic devices, such as transmitters and receivers (e.g., television sets and mobile phones) used in satellite broadcasting, cable TV and other wired broadcasting, distribution over the Internet, and distribution to terminals via cellular communications, or devices (e.g., hard disk recorders and cameras) that record images on media such as optical disks, magnetic disks, and flash memories, or play images from these storage media.

[0240] Furthermore, for example, the present technology can also be implemented as a part of an apparatus, such as a processor (e.g., a video processor) as a system LSI (Large Scale Integration), a module (e.g., a video module) using multiple processors, a unit (e.g., a video unit) using multiple modules, or a set in which other functions are added to a unit (e.g., a video set).

[0241] Furthermore, for example, the present technology can also be applied to a network system configured with multiple devices. For example, the present technology may be implemented as cloud computing in which multiple devices share and collaborate on processing via a network. For example, the present technology may be implemented in a cloud service that provides image (video)-related services to any terminal, such as a computer, AV (Audio Visual) equipment, a portable information processing terminal, or an IoT (Internet of Things) device.

[0242] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all the components are contained in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.

[0243] <Fields and applications where this technology can be applied> Systems, devices, processing units, etc. to which the present technology is applied can be used in any field, such as transportation, medical care, crime prevention, agriculture, livestock farming, mining, beauty, factories, home appliances, weather, and nature monitoring. In addition, the applications thereof are also arbitrary.

[0244] <Other> In this specification, a "flag" refers to information for identifying multiple states, and includes not only information used to identify two states, true (1) or false (0), but also information capable of identifying three or more states. Therefore, the value that this "flag" can take may be, for example, two values, 1 / 0, or three or more values. In other words, the number of bits constituting this "flag" is arbitrary, and may be one bit or multiple bits. Furthermore, identification information (including flags) can be assumed not only to include the identification information in the bit stream, but also to include difference information of the identification information relative to certain reference information in the bit stream. Therefore, in this specification, "flag" and "identification information" include not only the information itself, but also difference information relative to the reference information.

[0245] Furthermore, various types of information (metadata, etc.) related to the coded data (bitstream) may be transmitted or recorded in any form as long as they are associated with the coded data. Here, the term "associate" means, for example, that one piece of data can be used (linked) when processing the other piece of data. In other words, data associated with each other may be combined into one piece of data or may be individual pieces of data. For example, information associated with coded data (image) may be transmitted over a transmission path separate from that of the coded data (image). Also, for example, information associated with coded data (image) may be recorded on a recording medium separate from that of the coded data (image) (or on a different recording area of ​​the same recording medium). Note that this "association" may refer to only a portion of the data, rather than the entire data. For example, an image and information corresponding to that image may be associated with each other in any unit, such as multiple frames, one frame, or a portion of a frame.

[0246] In this specification, terms such as "composite," "multiplex," "add," "integrate," "include," "store," "embed," "insert," and the like refer to combining multiple items into one, such as combining encoded data and metadata into one piece of data, and refer to one method of "associating" as described above.

[0247] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present technology.

[0248] For example, a configuration described as one device (or processing unit) may be divided and configured as multiple devices (or processing units). Conversely, configurations described above as multiple devices (or processing units) may be combined and configured as one device (or processing unit). Of course, configurations other than those described above may be added to the configuration of each device (or each processing unit). Furthermore, as long as the configuration and operation of the entire system are substantially the same, part of the configuration of one device (or processing unit) may be included in the configuration of another device (or other processing unit).

[0249] Furthermore, for example, the above-described program may be executed in any device, as long as the device has the necessary functions (functional blocks, etc.) and can obtain the necessary information.

[0250] Also, for example, each step of a single flowchart may be executed by one device, or may be shared and executed by multiple devices. Furthermore, when one step includes multiple processes, the multiple processes may be executed by one device, or may be shared and executed by multiple devices. In other words, multiple processes included in one step can be executed as multiple step processes. Conversely, processes described as multiple steps can be executed collectively as one step.

[0251] For example, the steps of a program executed by a computer may be executed in chronological order in the order described herein, or may be executed in parallel or individually at the required timing, such as when a call is made. In other words, as long as no contradiction occurs, the steps may be executed in an order different from the order described above. Furthermore, the steps of this program may be executed in parallel with the processing of another program, or may be executed in combination with the processing of another program.

[0252] Furthermore, for example, multiple technologies related to the present technology can be implemented independently and independently, as long as no contradiction occurs. Of course, any multiple technologies can also be implemented in combination. For example, part or all of the present technology described in any embodiment can be implemented in combination with part or all of the present technology described in another embodiment. Furthermore, part or all of any of the above-described present technologies can be implemented in combination with other technologies not described above.

[0253] The present technology can also be configured as follows. (1) A hierarchical structure generating unit generates a hierarchy of attribute information for each point of a point cloud representing a three-dimensional object as a set of points by recursively repeating classification of the reference points into prediction points for deriving a difference value between the attribute information and a prediction value of the attribute information, and reference points used to derive the prediction value, The layering unit sets the reference point based on the center of gravity of the points. Information processing device. (2) The hierarchical division unit sets a point that is closer to the center of gravity among the candidate points as the reference point. The information processing device described in (1). (3) The layering unit sets the reference point based on the center of gravity of points located within a predetermined range. An information processing device according to (1) or (2). (4) The layering unit sets the reference point based on a predetermined search order from among a plurality of candidates having substantially the same conditions with respect to the center of gravity. An information processing device according to any one of (1) to (3). (5) For attribute information for each point of a point cloud that represents a three-dimensional object as a set of points, a classification is performed recursively for the reference points into prediction points for deriving a difference value between the attribute information and a prediction value of the attribute information, and reference points used to derive the prediction value, and the reference points are set based on the center of gravity of the points when hierarchizing the attribute information. Information processing methods.

[0254] (6) A hierarchical structure generating unit generates a hierarchy of attribute information for each point of a point cloud that represents a three-dimensional object as a set of points, by recursively repeating classification of the reference points into predicted points for deriving a difference value between the attribute information and a predicted value of the attribute information, and reference points used to derive the predicted value, The layering unit sets the reference points based on a distribution pattern of points. Information processing device. (7) The layering unit sets the reference points based on table information that specifies points close to a center of gravity of the points for each distribution pattern of the points. (6) An information processing device according to the present invention. (8) The layering unit sets the reference points based on table information that specifies predetermined points for each distribution pattern of the points. An information processing device according to (6) or (7). (9) An encoding unit that encodes information about the table information is further provided. (8) An information processing device according to (8). (10) For attribute information for each point of a point cloud that represents a three-dimensional object as a set of points, a classification is recursively repeated for the reference points into prediction points that derive a difference value between the attribute information and a prediction value of the attribute information, and reference points that are used to derive the prediction value, and when hierarchizing the attribute information, the reference points are set based on the distribution pattern of the points. Information processing methods.

[0255] (11) A hierarchical layering unit that classifies attribute information for each point of a point cloud that represents a three-dimensional object as a set of points into a predicted point that derives a difference value between the attribute information and a predicted value of the attribute information and a reference point that is used to derive the predicted value, by recursively repeating the classification of the reference point. an encoding unit that encodes information regarding the setting of the reference point by the layering unit; An information processing device comprising: (12) The encoding unit encodes information regarding the setting of the reference point for all points. (11) An information processing device according to (11). (13) The encoding unit encodes information regarding the setting of the reference point for a point in a part of layers. (11) An information processing device according to (11). (14) The encoding unit further encodes information regarding the setting of the reference point for a point that satisfies a predetermined condition. (13) An information processing device according to (13). (15) For attribute information for each point of a point cloud that represents a three-dimensional object as a set of points, a classification is performed into a prediction point for deriving a difference value between the attribute information and a prediction value of the attribute information, and a reference point used to derive the prediction value, by recursively repeating the classification for the reference point, thereby hierarchizing the attribute information; Encoding information regarding the setting of said reference point Information processing methods.

[0256] (16) A hierarchical structure generating unit generates a hierarchy of attribute information for each point of a point cloud that represents a three-dimensional object as a set of points, by recursively repeating classification of the reference points into prediction points that derive a difference value between the attribute information and a prediction value of the attribute information, and reference points that are used to derive the prediction value, The layering unit alternately selects, for each layer, a point closer to the center of the bounding box and a point farther from the center of the bounding box from among the reference point candidates as the reference point. Information processing device. (17) The layering unit selects, from the candidates, a point that is closer to the center of the bounding box and a point that is farther from the center of the bounding box based on a search order according to a position within the bounding box. (16) An information processing device according to (16). (18) The search order is the order of distance from the center of the bounding box. (17) An information processing device according to (17). (19) The search order is set for each of the eight regions obtained by dividing the bounding box. (17) An information processing device according to (17). (20) When classifying attribute information for each point of a point cloud that represents a three-dimensional object as a set of points, into predicted points for deriving a difference value between the attribute information and a predicted value of the attribute information and reference points used to derive the predicted value, the attribute information is hierarchically classified by recursively repeating the classification for the reference points. In this way, points closer to the center of a bounding box and points farther from the center of the bounding box are alternately selected for each hierarchy from among the reference point candidates. Information processing methods. [Explanation of symbols]

[0257] 100 encoding device, 101 position information encoding unit, 102 position information decoding unit, 103 point cloud generation unit, 104 attribute information encoding unit, 105 bitstream generation unit, 111 hierarchical processing unit, 112 quantization unit, 113 encoding unit, 121 reference point setting unit, 122 reference relationship setting unit, 123 inversion unit, 124 weight value derivation unit, 200 decoding device, 201 encoded data extraction unit, 202 position information decoding unit, 203 attribute information decoding unit, 204 point cloud generation unit, 211 decoding unit, 212 inverse quantization unit, 213 inverse hierarchical processing unit< / lifting> < / octree>

Claims

1. Obtaining geometry data of a point cloud that represents a three-dimensional object as a set of points; setting a hierarchical structure based on the geometry data; classifying the points at each level of the hierarchical structure by a recursive process into either a predicted point or a reference point that is closer to the centroid of a limited number of points including the predicted point; A predicted value for the reference point is derived, and a difference value between the attribute data corresponding to the geometry data and the predicted value is derived based on the predicted point, thereby deriving the attribute data. Reverse layering section A decoding device comprising:

2. In the recursive processing, the inverse layering unit searches for the prediction point and the reference point from candidate points within an NxNxN (N is an integer equal to or greater than 2) voxel region of each layer. The decoding device according to claim 1 .

3. The inverse layering unit searches the prediction points and the reference points in Morton order in the recursive processing. The decoding device according to claim 2 .

4. In the recursive process, the inverse layer generation unit leaves the reference points in the current layer in a layer of lower resolution. The decoding device according to claim 1 .

5. A decoding device comprising: Obtaining geometry data of a point cloud that represents a three-dimensional object as a set of points; setting a hierarchical structure based on the geometry data; classifying the points at each level of the hierarchical structure by a recursive process into either a predicted point or a reference point that is closer to the centroid of a limited number of points including the predicted point; A predicted value for the reference point is derived, and a difference value between the attribute data corresponding to the geometry data and the predicted value is derived based on the predicted point, thereby deriving the attribute data. Decryption method.

6. In the recursive process, the predicted point and the reference point are searched for from candidate points within an NxNxN (N is an integer equal to or greater than 2) voxel region of each layer. The decoding method according to claim 5.

7. In the recursive process, the prediction points and the reference points are searched in Morton order. The decoding method according to claim 6.

8. In the recursive process, the reference points in the current layer are left in the lower resolution layer. The decoding method according to claim 5.

Citation Information

Patent Citations

  • Point cloud compression

    WO2019055772A1

  • Information processing device and method

    WO2019065298A1

  • Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device

    WO2019240286A1

  • Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device

    WO2019244931A1