Decoding device and decoding method
By using multiple hierarchical methods in the information processing device to layer and inversely layer the attribute information of point cloud data, the problem of decreasing encoding efficiency in the prior art is solved, and efficient resolution scalable decoding is achieved.
Patent Information
- Application Number
- CN202510460506.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-07-02
- Filing Date
- 2020-06-18
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art may lead to a reduced encoding efficiency when encoding point cloud data in three-dimensional shapes, especially when point cloud data is sparse.
By using multiple hierarchical methods in the information processing device to perform hierarchical and inverse hierarchical processing on the attribute information of the point cloud data, specifically, classifying the points into prediction points or reference points, using the attribute information of the reference point to derive the predicted value of the attribute information of the predicted point, and derive the difference between the attribute information of the predicted point and the predicted value. The device can use different hierarchical methods to process at different levels and reference points at the same level during the inverse hierarchical process to generate attribute information of the predicted points.
Through the combined use of multiple hierarchical methods, resolution scalable decoding can be supported while suppressing the reduction in encoding efficiency, improving the encoding efficiency and flexibility of point cloud data.
Smart Images

Figure CN120186352A_ABST
Abstract
Description
[0001] This application is a divisional application of a patent application with international application number PCT / JP2020 / 023989 filed on June 18, 2020, and national stage application number 202080038761.X in China, titled "Information Processing Apparatus and Method". Technical Field
[0002] The present disclosure relates to an information processing apparatus and method, and more particularly to an information processing apparatus and method capable of suppressing a decrease in coding efficiency. Background Art
[0003] For example, a conventional method for encoding 3D data representing a three-dimensional structure, such as point cloud, has been designed (for example, see Non-Patent Document 1). Point cloud data is composed of geometric data (also referred to as position information) and attribute data (also referred to as attribute information) for each point. Therefore, encoding of the point cloud is performed separately for the geometric data and the attribute data. As methods for encoding the attribute data, various methods have been proposed. For example, a method using a technique such as Lifting has been proposed (for example, see Non-Patent Document 2). In addition, a method capable of performing scalable decoding of the attribute data has also been proposed (for example, see Non-Patent Document 3).
[0004] [Citation List]
[0005] [Non-Patent Document]
[0006] [Non-Patent Document 1]
[0007] R. Mekuria, Student Member IEEE, K. Blom, P. Cesar., Member, IEEE, “Design, Implementation and Evaluation of a Point Cloud Codec for Tele-Immersive Video”, tcsvt_paper_submitted_february.pdf
[0008] [Non-Patent Document 2]
[0009] Khaled Mammou, Alexis Tourapis, Jungsun Kim, Fabrice Robinet, Valery Valentin, Yeping Su, “Lifting Scheme for Lossy Attribute Encoding in TMC1”, ISO / IEC JTC1 / SC29 / WG11 MPEG2018 / m42640, April 2018, San Diego, US [Non-Patent Document 3]
[0010] Ohji Nakagami, Satoru Kuma, “[G-PCC] Spatial scalability support for G-PCC”, ISO / IEC JTC1 / SC29 / WG11 MPEG2019 / m47352, March 2019, Geneva, CH Summary of the Invention
[0011] [Technical Problem]
[0012] However, each of these methods has its own characteristics, and all methods may not always be the best in all cases. For example, a method adopting the technology described in Non-Patent Document 3 can perform scalable decoding of attribute information that is difficult to perform in a method adopting Lifting described in Non-Patent Document 2. However, in the method adopting the technology described in Non-Patent Document 3, the overall encoding efficiency may be reduced compared to the method adopting Lifting described in Non-Patent Document 2.
[0013] In view of this situation, the present disclosure aims to suppress a reduction in encoding efficiency.
[0014] [Solution to the Problem]
[0015] An information processing apparatus according to an aspect of the present technology is an information processing apparatus that processes attribute information of each point in a point cloud representing an object of a three-dimensional shape as a set of points. The information processing apparatus includes: a layering unit configured to recursively repeat the following processing for reference points among points classified as prediction points or reference points to perform layering of the attribute information: deriving a difference between a predicted value of the attribute information of the prediction point derived using the attribute information of the reference point and the attribute information of the prediction point. The layering unit uses a first layering method for a first layer and uses a second layering method different from the first layering method for a second layer different from the first layer.
[0016] An information processing method according to one aspect of the present technology is an information processing method as follows. This information processing method is used for: for the attribute information of each point in a point cloud representing an object of a three-dimensional shape as a set of points, when performing hierarchical processing of the attribute information by recursively repeating the following processing for a reference point, multiple hierarchical methods are used to generate different levels according to the multiple hierarchical methods: classifying points into prediction points or reference points, deriving a predicted value of the attribute information of the prediction points using the attribute information of the reference points, and deriving the difference between the attribute information of the prediction points and the predicted value.
[0017] An information processing apparatus according to another aspect of the present technology is an information processing apparatus as follows. This information processing apparatus is used for processing the attribute information of each point in a point cloud representing an object of a three-dimensional shape as a set of points. The information processing apparatus includes: an inverse hierarchical unit configured to perform inverse hierarchical processing on the hierarchical attribute information obtained by recursively repeating the following processing for the reference points among the points classified as prediction points or reference points: deriving the difference between the predicted value of the attribute information of the prediction points derived using the attribute information of the reference points and the attribute information of the prediction points. Among them, the inverse hierarchical unit performs inverse hierarchical processing on the first level using the first hierarchical method, and performs inverse hierarchical processing on the second level different from the first level using the second hierarchical method different from the first hierarchical method.
[0018] An information processing method according to another aspect of the present technology is an information processing method as follows. This information processing method is used for: for the attribute information of each point in a point cloud representing an object of a three-dimensional shape as a set of points, when performing inverse hierarchical processing on the hierarchical attribute information obtained by recursively repeating the following processing for a reference point, multiple hierarchical methods are used to perform inverse hierarchical processing on different levels according to the multiple hierarchical methods: classifying points into prediction points or reference points, deriving a predicted value of the attribute information of the prediction points using the attribute information of the reference points, and deriving the difference between the attribute information of the prediction points and the predicted value.
[0019] An information processing apparatus according to still another aspect of the present technology is an information processing apparatus as follows. This information processing apparatus is used for processing the attribute information of each point in a point cloud representing an object of a three-dimensional shape as a set of points. The information processing apparatus includes: a hierarchical unit configured to perform hierarchical processing of the attribute information by recursively repeating the following processing for the reference points among the points classified as prediction points or reference points: deriving the difference between the predicted value of the attribute information of the prediction points derived using the attribute information of the reference points and the attribute information of the prediction points. Among them, the hierarchical unit derives the predicted values of the prediction points of multiple levels with reference to the reference points of the same level.
[0020] An information processing method according to another aspect of the present technology is an information processing method as follows. This information processing method is used for: for the attribute information of each point in a point cloud representing an object with a three-dimensional shape as a set of points, when performing hierarchical processing of the attribute information by recursively repeating the following processing for a reference point, the predicted values of the predicted points at multiple levels are derived by referring to the reference points at the same level: classifying the points into predicted points or reference points, deriving the predicted value of the attribute information of the predicted points using the attribute information of the reference points, and deriving the difference between the attribute information of the predicted points and the predicted value.
[0021] An information processing apparatus according to another aspect of the present technology is an information processing apparatus as follows. This information processing apparatus is used for processing the attribute information of each point in a point cloud representing an object with a three-dimensional shape as a set of points. The information processing apparatus includes: an inverse hierarchical unit configured to: perform inverse hierarchical processing on the hierarchical attribute information obtained by recursively repeating the following processing for the reference points among the points classified as predicted points or reference points: deriving the difference between the predicted value of the attribute information of the predicted points derived using the attribute information of the reference points and the attribute information of the predicted points. Wherein, when performing the inverse hierarchical processing, the inverse hierarchical unit derives the predicted values of the attribute information of the predicted points at multiple levels by referring to the attribute information of the reference points at the same level, and generates the attribute information of the predicted points by adding the derived predicted values and the difference.
[0022] An information processing method according to another aspect of the present technology is an information processing method as follows. This information processing method is used for: for the attribute information of each point in a point cloud representing an object with a three-dimensional shape as a set of points, when performing inverse hierarchical processing on the hierarchical attribute information obtained by recursively repeating the processing of classifying the points into predicted points or reference points, deriving the predicted value of the attribute information of the predicted points using the attribute information of the reference points, and deriving the difference between the attribute information of the predicted points and the predicted value for the reference points, the predicted values of the attribute information of the predicted points at multiple levels are derived by referring to the attribute information of the reference points at the same level, and the attribute information of the predicted points is generated by adding the derived predicted values and the difference.
[0023] In the information processing apparatus and method according to one aspect of the present technology, for the attribute information of each point in a point cloud representing an object with a three-dimensional shape as a set of points, when performing hierarchical processing of the attribute information by recursively repeating the following processing for a reference point, multiple hierarchical methods are used to generate different levels according to the multiple hierarchical methods: classifying the points into predicted points or reference points, deriving the predicted value of the attribute information of the predicted points using the attribute information of the reference points, and deriving the difference between the attribute information of the predicted points and the predicted value.
[0024] In an information processing apparatus and method according to another aspect of the present technology, for the attribute information of each point in a point cloud representing an object of a three-dimensional shape as a set of points, when performing inverse layering on the attribute information that is hierarchically layered by recursively repeating the following processing for a reference point, inverse layering is performed on different levels according to the plurality of layering methods using the plurality of layering methods: classifying a point as a prediction point or a reference point, deriving a predicted value of the attribute information of the prediction point using the attribute information of the reference point, and deriving a difference between the attribute information of the prediction point and the predicted value.
[0025] In an information processing apparatus and method according to still another aspect of the present technology, for the attribute information of each point in a point cloud representing an object of a three-dimensional shape as a set of points, when performing the layering of the attribute information by recursively repeating the following processing for a reference point, predicted values of prediction points at a plurality of levels are derived with reference to reference points at the same level: classifying a point as a prediction point or a reference point, deriving a predicted value of the attribute information of the prediction point using the attribute information of the reference point, and deriving a difference between the attribute information of the prediction point and the predicted value.
[0026] In an information processing apparatus and method according to still another aspect of the present technology, for the attribute information of each point in a point cloud representing an object of a three-dimensional shape as a set of points, when performing inverse layering on the attribute information that is hierarchically layered by recursively repeating the processing of classifying a point as a prediction point or a reference point, deriving a predicted value of the attribute information of the prediction point using the attribute information of the reference point, and deriving a difference between the attribute information of the prediction point and the predicted value for a reference point, predicted values of the attribute information of prediction points at a plurality of levels are derived with reference to the attribute information of reference points at the same level, and the attribute information of the prediction point is generated by adding the derived predicted value and the difference. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 Figure 1 is a diagram for describing an example of a state of Lifting that does not support scalable decoding.
[0028] Figure 2 Figure 2 is a diagram for describing an example of a state of Lifting that does not support scalable decoding.
[0029] Figure 3 Figure 3 is a diagram for describing an example of a state of Lifting that supports scalable decoding.
[0030] Figure 4 Figure 4 is a diagram for describing an example of a state of Lifting that supports scalable decoding.
[0031] Figure 5 Figure 5 is a diagram for describing the hierarchical method of attribute data.
[0032] Figure 6 Figure 6 is a diagram for describing an example of the state when both the Lifting method that supports scalable decoding and the Lifting method that does not support scalable decoding are applied.
[0033] Figure 7 Figure 7 is a diagram for describing the Lifting that supports scalable decoding.
[0034] Figure 8 Figure 8 is a diagram showing a reference relationship.
[0035] Figure 9 Figure 9 is a diagram for describing the hierarchical method of attribute data.
[0036] Figure 10 Figure 10 is a diagram showing a reference relationship.
[0037] Figure 11 Figure 11 is a block diagram showing an example of the main components of an encoding device.
[0038] Figure 12 Figure 12 is a block diagram showing an example of the main components of an attribute information encoding unit.
[0039] Figure 13 Figure 13 is a block diagram showing an example of the main components of a hierarchical processing unit.
[0040] Figure 14 Figure 14 is a flowchart for describing an example of an encoding processing flow.
[0041] Figure 15 Figure 15 is a flowchart for describing an example of an attribute information encoding processing flow.
[0042] Figure 16 Figure 16 is a flowchart for describing an example of a hierarchical processing flow.
[0043] Figure 17 Figure 17 is a block diagram showing an example of the main components of a decoding device.
[0044] Figure 18 Figure 18 It is a block diagram showing an example of the main components of an attribute information decoding unit.
[0045] Figure 19 Figure 19 It is a block diagram showing an example of the main components of an inverse layer processing unit.
[0046] Figure 20 Figure 20 It is a flowchart showing an example of a decoding process flow.
[0047] Figure 21 Figure 21 It is a flowchart showing an example of an attribute information decoding process flow.
[0048] Figure 22 Figure 22 It is a flowchart showing an example of an inverse layer processing process flow.
[0049] Figure 23 Figure 23 It is a flowchart showing an example of a scalable inverse layer processing process flow.
[0050] Figure 24 Figure 24 It is a flowchart showing an example of a layer processing process flow.
[0051] Figure 25 Figure 25 It is a flowchart showing an example of an inverse layer processing process flow.
[0052] Figure 26 Figure 26 It is a block diagram showing an example of the main components of a computer. Detailed Implementation Manner
[0053] Hereinafter, the manners for implementing the present disclosure (hereinafter referred to as implementation manners) will be described. The description will be made in the following order.
[0054] 1. Switching of Layer / Inverse Layer Methods
[0055] 2. Control of Reference Relationships
[0056] 3. First Embodiment (Encoding Device)
[0057] 4. Second Embodiment (Decoding Device)
[0058] 5. Third Embodiment (Encoding Device)
[0059] 6. Fourth Embodiment (Decoding Device)
[0060] 7. Supplementary
[0061] <Switching between layering / inverse layering methods>
[0062] <Literature such as those supporting technical details and technical terms>
[0063] The scope of the present technology disclosure includes not only the details described in the embodiments but also the details described in the following non-patent literatures known at the time of filing the application.
[0064] Non-patent literature 1: (as described above)
[0065] Non-patent literature 2: (as described above)
[0066] Non-patent literature 3: (as described above)
[0067] Non-patent literature 4: TELECOMMUNICATION STANDARDIZATION SECTOR OF ITU(International Telecommunication Union), “Advanced video coding for genericaudiovisualservices”, H.264, 04 / 2017
[0068] Non-patent literature 5: TELECOMMUNICATION STANDARDIZATION SECTOR OFITU(InternationalTelecommunication Union), “High efficiency video coding”, H.265, 12 / 2016
[0069] Non-patent literature 6: Jianle Chen,Elena Alshina,Gary J.Sullivan,Jens-Rainer,Jill Boyce, “Algorithm Description of Joint Exploration Test Model 4”, JVET-G1001_v1, Joint Video Exploration Team(JVET)of ITU-T SG 16WP 3andISO / IEC JTC1 / SC 29 / WG 11 7th Meeting:Torino,IT,13-21July 2017
[0070] That is, the details described in the above non-patent literature are also the basis for determining the support conditions. For example, although the quadtree block structure described in non-patent literature 5 and the QTBT (quadtree plus binary tree) block structure described in non-patent literature 6 are not explicitly described in the embodiments, they are considered to be included in the disclosure scope of the present technology and meet the support requirements of the claims. In addition, even if technical terms such as parsing, syntax, and semantics are not explicitly described in the embodiments, they are considered to be included in the disclosure scope of the present technology and meet the support requirements of the claims.
[0071] <Point cloud>
[0072] Conventionally, there is 3D data such as point clouds and meshes. A point cloud represents a three-dimensional structure based on the position information, attribute information, etc. of a point group. A mesh is composed of vertices, edges, and faces and uses polygons to represent and define a three-dimensional shape.
[0073] For example, in the case of a point cloud, a three-dimensional structure (an object with a three-dimensional shape) can be represented as a set of multiple points (a point group). That is, the data of the point cloud (also referred to as point cloud data) is composed of the geometric data (also referred to as position information) and attribute data (also referred to as attribute information) of each point in the point group. The attribute data can include any information. For example, color information, reflection information, normal information, etc. can be included in the attribute data. Therefore, the data structure is relatively simple, and an arbitrarily complex spatial structure can be represented with sufficient accuracy using a sufficient number of points.
[0074] <Quantifying position information using voxels>
[0075] Since the amount of such point cloud data is relatively large, in order to compress the data volume according to encoding, etc., an encoding method using voxels has been conceived. A voxel is a three-dimensional region used to quantify geometric data (position information).
[0076] That is, the three-dimensional region including the point cloud is divided into small three-dimensional regions called voxels, and it is indicated whether each voxel includes a point. By doing so, the position of each point is quantified in units of voxels. Therefore, the increase in the amount of information (usually a reduction in the amount of information) can be suppressed by converting the point cloud data into voxel data (also referred to as voxel data).
[0077] <Octree>
[0078] In addition, it has been conceived to use such voxel data for geometric data to construct an octree. An octree is a tree structure in which voxel data is structured. The value of each bit of the lowest node of the octree represents the presence or absence of a point in each voxel. For example, the value "1" represents a voxel including a point, and the value "0" represents a voxel not including a point. In the octree, 1 node corresponds to 8 voxels. That is, each node of the octree is composed of 8-bit data, and the 8 bits represent the presence or absence of a point in 8 voxels.
[0079] In addition, the upper-level nodes of the octree represent the presence or absence of a point in a region where the 8 voxels corresponding to the lower-level nodes belonging to the upper-level node are combined into one voxel. That is, the upper-level node is generated by combining the information of the voxels of the lower-level nodes. At the same time, in the case of a node with a value of "0", that is, when all 8 voxels corresponding to it do not include a point, the node is deleted.
[0080] By doing so, an octree composed of nodes with values not equal to "0" is constructed. That is, the octree can represent the presence or absence of a point in voxels at each resolution. According to the structuring into an octree and encoding, the position information can be decoded from the lowest resolution (highest level) to the required level (resolution), so that the point cloud data with the required resolution can be reconstructed. That is, the information can be easily decoded at any resolution without decoding the information at unnecessary levels (resolutions). In other words, scalability of voxels (resolutions) can be achieved.
[0081] In addition, as described above, since the resolution of voxels in a region without points can be reduced by omitting nodes with a value of "0", an increase in the amount of information can be further suppressed (usually the amount of information can be reduced).
[0082] <lifting>
[0083] On the other hand, when encoding attribute data (attribute information), assuming that geometric data (position information) including degradation due to encoding is known, the position relationship between points is used for encoding. As a method for encoding such attribute data, region adaptive hierarchical transform (RAHT) and a method using a transform called Lifting as described in Non-Patent Document 2 are conceived. By applying these techniques, the attribute data can be hierarchically structured like the octree of geometric data.
[0084] For example, in the case of Lifting described in Non-Patent Document 2, the attribute data of each point is encoded as a difference value relative to a predicted value derived using the attribute data of another point. In this case, the points are hierarchically structured, and the difference values are derived according to this hierarchical structure.
[0085] That is, for the attribute data of each point, the points are classified into predicted points or reference points, the predicted value of the attribute data of the predicted points is derived using the attribute data of the reference points, and the difference value between the attribute data of the predicted points and the predicted value is derived. Such processing is recursively repeated for the reference points to hierarchically structure the attribute data of each point.
[0086] In the case of Lifting described in Non-Patent Document 2, this hierarchical structuring is performed using geometric data and based on the distance between points. That is, points located within a predetermined distance range from a reference point are set as predicted points, and the predicted value of the predicted points is derived with reference to the reference point. For example, in Figure 1 , it can be assumed that point P5 is selected as the reference point. In this case, predicted points are searched for within a circular region centered on point P5 with a radius of R. In this case, since point P9 is located within this region, point P9 is set as a predicted point for deriving its predicted value with reference to reference point P5 (a predicted point with reference point P5).
[0087] According to such processing, for example, the difference values of points P7 to P9 represented by white circles, the difference values of points P1, P3, and P6 represented by diagonal patterns, and the difference values of points P0, P2, P4, and P5 represented by gray circles are derived as difference values at different levels.
[0088] However, this hierarchical structure is generated independently of the hierarchical structure of geometric data (e.g., octree) and is basically not associated with the hierarchical structure of geometric data. In order to reconstruct point cloud data, it is necessary to associate geometric data with attribute data, so it is necessary to decode geometric data and attribute data to the highest resolution (i.e., the lowest level).
[0089] For example, when decoding geometric data 1 at a specific resolution, as Figure 2 As shown, only part of the geometric data 1A can be decoded. However, in order to obtain attribute data with the same resolution, both the geometric data 1 and the attribute data 2 need to be decoded to the highest resolution (i.e., the lowest level). That is, the method using Lifting described in Non-Patent Document 2 does not support resolution scalable decoding.
[0090] <Hierarchical structure supporting scalable decoding>
[0091] On the other hand, the hierarchical structure described in Non-Patent Document 3 supports resolution scalable decoding. In the case of the method described in Non-Patent Document 3, hierarchical processing of the attribute data is performed such that the hierarchical structure of the attribute data is consistent with the hierarchical structure of the octree of the geometric data. That is, when there is a point in the region corresponding to the voxel of the geometric data (when there is attribute data corresponding to the point), reference points and prediction points are selected such that a point also exists in the voxel one level higher than the voxel (there is attribute data corresponding to the point). That is, the attribute information is hierarchically processed according to the hierarchical structure of the octree of the geometric data.
[0092] For example, in Figure 3 when a two-dimensional hierarchical structure is described for simplicity of description, voxels 10-1 to 10-4 are formed to be one level lower than a specific voxel 10, and voxels 10-4-1 to 10-4-4 are formed to be one level lower than voxel 10-4. A point 11-1 of the attribute data (indicating the attribute data corresponding to point 11-1, the same below) exists in voxel 10-1. A point 11-2 of the attribute data exists in voxel data 10-2, and a point 11-3 of the attribute data exists in voxel data 10-3. Points 11-4-1 to 11-4-4 of the attribute data exist in voxels 10-4-1 to 10-4-4.
[0093] In this case, for voxels 10-4-1 to 10-4-4, for example, point 11-4-1 is set as the reference point, and points 11-4-2 to 11-4-4 are set as the prediction points such that one point (attribute data) is retained in voxel 10-4 which is one level higher than voxels 10-4-1 to 10-4-4.
[0094] Similarly, for voxels 10-1 to 10-4, for example, point 11-1 is set as the reference point, and points 11-2, 11-3, and 11-4-1 are set as the prediction points such that one point is retained in voxel 10 which is one level higher than voxels 10-1 to 10-4.
[0095] In this way, the same hierarchical structure as that of the geometric data can also be achieved for the attribute data. In addition, the hierarchical structure of the attribute data can be associated with the hierarchical structure of the geometric data to support scalable decoding.
[0096] For example, when decoding the geometric data 21 at a specific resolution, as Figure 4 shown, only a part of the geometric data 21A can be decoded. In this case, since the hierarchical structures in the geometric data 21 and the attribute data 22 are associated with each other, only a part of the attribute data 22A corresponding to the geometric data 21A can also be decoded in the attribute data 22.
[0097] That is, the point cloud data can be easily reconstructed at the required resolution without performing decoding to the lowest level. In this way, the method using the technique described in Non-Patent Document 3 supports resolution scalable decoding.
[0098] However, this method may reduce the coding efficiency to be lower than that of the method described in Non-Patent Document 2. Especially when the points are in a sparse state, the reduction in coding efficiency may be greater.
[0099] <Switching of Hierarchical Methods in Hierarchy>
[0100] Accordingly, in the hierarchy of the attribute data, the hierarchical method of the attribute data can be switched at an intermediate level, as Figure 5 described in "Method 1" in the first row of the table shown.
[0101] For example, for the attribute information of each point in a point cloud representing an object of a three-dimensional shape as a set of points, when hierarchically processing the attribute information by recursively repeating the following processing for a reference point, different levels are generated using multiple hierarchical methods: classifying the points into predicted points or reference points, deriving a predicted value of the attribute information of the predicted points using the attribute information of the reference points, and deriving the difference between the attribute information of the predicted points and the predicted value.
[0102] For example, an information processing apparatus may include a layering unit that recursively repeats the following processing for a reference point to perform layering of attribute information for each point's attribute information in a point cloud representing an object with a three-dimensional shape as a set of points: classifying a point as a prediction point or a reference point, deriving a predicted value of the attribute information of the prediction point using the attribute information of the reference point, and deriving a difference between the attribute information of the prediction point and the predicted value. Here, the layering unit may perform layering using multiple layering methods and generate levels that differ according to the layering methods. In other words, an information processing apparatus that processes the attribute information of each point in a point cloud representing an object with a three-dimensional shape as a set of points may include a layering unit that recursively repeats the following processing for a reference point to perform layering of attribute information: deriving a difference between the predicted value of the attribute information of a prediction point, which is derived using the attribute information of a reference point among the points classified as prediction points or reference points, and the attribute information of the prediction point. Here, the layering unit may layer a first level using a first layering method and layer a second level different from the first level using a second layering method different from the first layering method.
[0103] By doing so, a hierarchical structure can be formed using multiple layering methods, and thus a method suitable for the characteristics of each level can be selected. Therefore, layering can be performed on more different types of attribute data by a more appropriate method to suppress a decrease in coding efficiency.
[0104] Meanwhile, the number of times of switching the layering method is arbitrary and can be once or more. In addition, the number of applied layering methods can be any complex number, such as three or more. Furthermore, the same layering method can be applied multiple times. For example, the layering methods can be switched among in the order of layering method A, layering method B, and layering method A at an intermediate level.
[0105] In addition, when switching between layering methods in this way, for example, it is possible to switch between a method that supports scalable decoding and a method that does not support scalable decoding, as Figure 5 described in "Method 1-1" in the second row of the table shown. In other words, certain levels of the hierarchical structure of the attribute data can be layered according to a method that supports scalable decoding.
[0106] For example, it is possible to switch the layering method (a method that supports scalable decoding) described in Non-Patent Document 3 and Lifting (a method that does not support scalable decoding) described in Non-Patent Document 2 at an intermediate level of the hierarchical structure. That is, the layering method described in Non-Patent Document 3 can be applied only to certain levels of the hierarchical structure of the attribute data.
[0107] By doing so, it is possible to generate only the layers that are likely to be scalable decoded using the hierarchical method described in Non-Patent Document 3, and the other layers can be generated using the Lifting described in Non-Patent Document 2. Therefore, compared with the case of generating all layers using the hierarchical method described in Non-Patent Document 3, a decrease in the encoding efficiency can be suppressed. In other words, it is possible to support scalable decoding while suppressing a decrease in the encoding efficiency.
[0108] In this case, the layers higher than a predetermined layer can be hierarchically divided according to a method that does not support scalable decoding, and the predetermined layer and the lower layers can be hierarchically divided according to a method that supports scalable decoding. For example, as Figure 5 described in "Method 1-1-1" in the third row of the table shown. For example, the layers higher than the predetermined layer including the first layer can be generated according to the first hierarchical method that does not support scalable decoding, and the layers lower than the predetermined layer including the second layer can be generated according to the second hierarchical method that supports scalable decoding.
[0109] In this case, it is possible to switch between the Lifting that does not support scalable decoding (the method described in Non-Patent Document 2) and the Lifting that supports scalable decoding (the method described in Non-Patent Document 3). For example, as Figure 5 described in "Method 1-1-1-1" in the fourth row of the table shown.
[0110] For example, the layers higher than the predetermined layer can be hierarchically divided according to the Lifting described in Non-Patent Document 2, and the predetermined layer and the lower layers can be hierarchically divided according to the hierarchical method (Lifting that supports scalable decoding) described in Non-Patent Document 3.
[0111] When applying a method that supports scalable decoding to higher layers where the points may be in a sparse state, the encoding efficiency may decrease. In addition, the resolution significantly decreases at the layers close to the highest layer, so it is less likely to perform decoding at that layer. That is, it is less likely to perform scalable decoding at higher layers.
[0112] Therefore, as described above, by applying a method that does not support scalable decoding at the layers higher than the predetermined layer in the hierarchical structure of the attribute data, it is possible to suppress a decrease in the encoding efficiency while suppressing a substantial limitation on scalable decoding. In other words, it is possible to support scalable decoding while suppressing a decrease in the encoding efficiency.
[0113] In this case, when decoding the geometric data 31 at a specific resolution, as Figure 6 As shown, only part of the geometric data 31A can be decoded. In this case, since the hierarchical structures in the geometric data 31 and the attribute data 32 correspond to each other, only part of the attribute data 32A and part of the attribute data 32B corresponding to the geometric data 31A in the attribute data 32 can also be decoded.
[0114] That is, the point cloud data can be easily reconstructed at the required resolution without performing decoding to the lowest level. However, in this case, the attribute data 32A is data hierarchically divided according to Lifting that does not support scalable decoding. That is, in this case, the higher level (stage L1) of the attribute data 32 is hierarchically divided according to Lifting that does not support scalable decoding, and the levels lower than the higher level (stage L2) are hierarchically divided according to Lifting that supports scalable decoding. Therefore, scalable decoding cannot be performed from stage 0 to stage (L1 - 1) starting from the top (both the geometric data 31 and the attribute data 32 need to be decoded to stage L1). However, as described above, since such a higher level has a low resolution, it is less likely to perform scalable decoding on it, and thus a reduction in coding efficiency can be suppressed while suppressing substantial limitations on scalable decoding.
[0115] <Other examples of switching methods>
[0116] At the same time, the level at which the hierarchical method is switched is arbitrary. In addition, the switched hierarchical method can be any method and is not limited to the above example (switching between a method that supports scalable decoding and a method that does not support scalable decoding). For example, it is possible to switch between using a hierarchical method based on distance sampling (such as Lifting described in Non-Patent Document 2) and a method of hierarchically dividing points by arranging each point (the attribute data corresponding to each point) in Morton order and sampling the points at equal intervals, as Figure 5 described in "Method 1-2" in the fifth row of the table shown.
[0117] For example, a hierarchical method based on distance sampling can be applied to levels higher than a predetermined level, and a hierarchical method of equal-interval sampling in Morton order can be applied to the predetermined level and lower levels. By doing so, as in the above example, a reduction in coding efficiency can be suppressed.
[0118] <Processing order>
[0119] In hierarchical division, the order of level processing (generation order) is arbitrary. For example, hierarchical division can be sequentially performed starting from the lowest level, as Figure 5 as described in "Method 1-3" in the sixth row of the table shown. For example, levels are generated one by one, and after all levels are generated, the levels are reversed. For example, level numbers can be assigned in the reverse order of the generation order. By doing so, levels can be generated from the lowest level to the highest level.
[0120] According to the layering, each point is classified as a reference point or a prediction point, and the reference points are processed again as points of a higher level. That is, the reference relationship between levels is constructed according to the layering. As described above, the above recursive processing can be performed by generating levels sequentially starting from the lowest level, so the reference relationship between levels can be constructed more easily.
[0121] <Control Information>
[0122] For example, control information regarding the layering of attribute information can be signaled (sent from the encoding side to the decoding side), as Figure 5 described in "Method 1-4" in the seventh row of the table shown.
[0123] The method of sending the control information is arbitrary. For example, the control information can be defined by syntax etc., included in the encoded data (bitstream) of the point cloud data (e.g., written in the header etc.) and sent. Additionally, different from the encoded data of the point cloud data, the control information can be sent as data associated with the encoded data of the point cloud data.
[0124] <scalable_enable_flag>
[0125] The content of the control information is arbitrary. For example, when it is possible to switch between a method that supports scalable decoding and a method that does not support scalable decoding at an intermediate level, control information indicating whether the layering method that supports scalable decoding can be applied can be sent to the decoder side. In this case, the control information can indicate whether the layering method that supports scalable decoding can be applied in any way.
[0126] For example, scalable_enable_flag can be sent as flag information. scale_enable_flag is flag information indicating whether the layering method that supports scalable decoding can be applied. This flag information indicates that the layering method that supports scalable decoding can be applied when its value is true (e.g., "1"), and indicates that the layering method that supports scalable decoding cannot be applied when its value is false (e.g., "0"). For example, it can be assumed that the case where scalable_enable_flag is omitted is equivalent to the case where the value of this flag information is false (e.g., "0").
[0127] In addition, the data unit for transmitting the control information (scalable_enable_flag) is arbitrary. For example, scalable_enable_flag can be transmitted for each sequence of point cloud data. In addition, for example, scalable_enable_flag can be transmitted for each piece of attribute data. For example, when there are multiple pieces of attribute data for geometric data, it can be indicated whether a hierarchical method supporting scalable decoding can be applied to each piece of attribute data by transmitting scalable_enable_flag for each piece of attribute data. That is, for example, only some attribute data can be made applicable to scalable decoding. Needless to say, scalable_enable_flag can be transmitted for each data unit other than these examples.
[0128] By transmitting the control information in this way, the decoding side can identify whether to perform inverse hierarchical processing corresponding to the hierarchical method applicable to scalable decoding based on this control information. Therefore, identification can be performed more easily. Thereby, an increase in the decoding processing load can be suppressed.
[0129] For example, the hierarchical method applicable to scalable decoding is not applied to the attribute data for which scalable_enable_flag is false, so no processing for scalable decoding is required. Therefore, based on the fact that scalable_enable_flag is false, the decoding device can omit all processing for scalable decoding. For example, it can also omit referring to the control information that represents the range of levels to which the hierarchical method applicable to scalable decoding is applied, which will be described later. In addition, it can also omit the determination of whether to use the hierarchical method applicable to scalable decoding during inverse hierarchical processing. In this way, an increase in the decoding processing load can be suppressed.
[0130] <scalable_enable_num_of_lod>
[0131] In addition, for example, control information representing the range of levels to which the hierarchical method applicable to scalable decoding can be applied can be transmitted to the decoding side. In this case, this control information can represent the range of levels to which the hierarchical method applicable to scalable decoding can be applied in any way. For example, the control information can represent a range, represent the level at which the hierarchical method is switched (the level serving as the boundary of the range), or represent the range of levels to which the hierarchical method applicable to scalable decoding is not applied.
[0132] For example, when the control information indicates the layer range to which a hierarchical method applicable to scalable decoding can be applied, the control information can indicate the upper and lower limits of the range, or can indicate the range by the size of the range (number of layers) and the layer number (identification number, starting from the highest layer 0 and increasing by 1 for each layer down) of the reference position of the range (e.g., start position, middle position, end position, etc.). Additionally, when the lower limit is equal to the lowest layer or the upper limit is equal to the highest layer, the range can be represented by the number of layers starting from the lower limit (lowest layer) or the upper limit (highest layer).
[0133] For example, the syntax element scalable_enable_num_of_lod can be sent. scalable_enable_num_of_lod is control information indicating the layer range to which a hierarchical method applicable to scalable decoding can be applied. Its value represents the layer range to which a hierarchical method applicable to scalable decoding can be applied by the number of layers starting from the lowest layer. That is, when scalable_enable_num_of_lod = N, the encoded data of the attribute data of the layers from the lowest layer to the Nth stage is applicable to scalable decoding, and the layers at the (N + 1)th stage and higher are not applicable to scalable decoding.
[0134] By sending such control information, the decoding side can easily and accurately determine the layers to which a hierarchical method applicable to scalable decoding can be applied based on the control information. Therefore, the encoded data to which multiple hierarchical methods have been applied can be correctly decoded. Therefore, a decrease in coding efficiency can be suppressed.
[0135] <Others>
[0136] Meanwhile, the control information is not limited to the above examples, and any information can be sent. Additionally, the number of control information sent is arbitrary, and multiple control information can be sent. For example, the above scalable_enable_flag and scalable_enable_num_of_lod can be sent. scalable_enable_num_of_lod can be sent only when scalable_enable_flag is true.
[0137] In any case, scalable_enable_num_of_lod is only referenced when scalable_enable_flag is true, and the layers to which a hierarchical method applicable to scalable decoding is applied are recognized on the decoding side.
[0138] <Switching of Hierarchical Method in Inverse Hierarchy>
[0139] Although the processing in encoding has been described above, the present technology can be applied to the decoding side. For example, in inverse layering, the attribute information layering method can be switched at an intermediate level (i.e., switched between inverse layering methods), as Figure 5 described in "Method 2" in the eighth row of the table shown.
[0140] For example, for the attribute information of each point in a point cloud that represents an object of a three-dimensional shape as a set of points, when performing inverse layering on the attribute information that has been layered by recursively repeating the following processing for reference points, multiple layering methods are used to perform inverse layering on different levels: classifying points as prediction points or reference points, deriving a predicted value of the attribute information of a prediction point using the attribute information of a reference point, and deriving the difference between the attribute information of the prediction point and the predicted value.
[0141] For example, an information processing apparatus may include an inverse layering unit that, for the attribute information of each point in a point cloud that represents an object of a three-dimensional shape as a set of points, performs inverse layering on the attribute information that has been layered by recursively repeating the following processing for reference points: classifying points as prediction points or reference points, deriving a predicted value of the attribute information of a prediction point using the attribute information of a reference point, and deriving the difference between the attribute information of the prediction point and the predicted value, where the inverse layering unit can use multiple layering methods to perform inverse layering on different levels. In other words, an information processing apparatus that processes the attribute information of each point in a point cloud that represents an object of a three-dimensional shape as a set of points may include an inverse layering unit that performs inverse layering on the attribute information that has been layered by recursively repeating the following processing for reference points among the points classified as prediction points or reference points: deriving the difference between the predicted value of the attribute information of a prediction point derived using the attribute information of a reference point and the attribute information of the prediction point, where the inverse layering unit can use a first layering method to perform inverse layering on a first level and can use a second layering method different from the first layering method to perform inverse layering on a second level different from the first level.
[0142] By doing so, inverse layering can be performed using multiple layering methods, and thus attribute data layered using multiple layering methods can be appropriately inverse-layered. That is, encoded data that has been encoded using multiple layering methods can be correctly decoded. Therefore, a decrease in encoding efficiency can be suppressed.
[0143] At the same time, the number of times of switching the layering method is arbitrary and can be one or more times. In addition, the number of applied layering methods can be any plural number, such as three or more. Furthermore, the same layering method can be applied multiple times. For example, at an intermediate level, the layering methods can be switched in the order of layering method A, layering method B, and layering method A.
[0144] In addition, as in the case of layering, it is possible to switch between a method applicable to scalable decoding and a method not applicable to scalable decoding. For example, as described in "Method 2-1" in the ninth row of the table shown in Figure 5 . In other words, some levels of the hierarchical structure of the attribute data can be inverse-layered according to the method applicable to scalable decoding.
[0145] For example, at the intermediate level of the hierarchical structure, it is possible to switch between the layering method described in Non-Patent Document 3 (a method supporting scalable decoding) and Lifting described in Non-Patent Document 2 (a method not supporting scalable decoding). That is, the layering method described in Non-Patent Document 3 can be applied only to the inverse layering of some levels of the hierarchical structure of the attribute data.
[0146] By doing so, it is possible to support scalable decoding while suppressing a decrease in coding efficiency.
[0147] In this case, it is possible to inverse-layer the levels higher than a predetermined level according to the method not applicable to scalable decoding, and inverse-layer the predetermined level and lower levels according to the method applicable to scalable decoding. For example, as described in "Method 2-1-1" in the tenth row of the table shown in Figure 5 . For example, it is possible to perform inverse layering on the levels higher than the predetermined level including the first level according to the first layering method not applicable to scalable decoding, and perform inverse layering on the levels lower than the predetermined level including the second level according to the second layering method applicable to scalable decoding.
[0148] In this case, for example, as described in "Method 2-1-1-1" in the eleventh row of the table shown in Figure 5 , it is possible to switch between Lifting not applicable to scalable decoding (the method described in Non-Patent Document 2) and Lifting applicable to scalable decoding (the method described in Non-Patent Document 3).
[0149] For example, it is possible to perform inverse layering on the levels higher than the predetermined level according to Lifting described in Non-Patent Document 2, and perform inverse layering on the predetermined level and lower levels according to the layering method (Lifting applicable to scalable decoding) described in Non-Patent Document 3.
[0150] By doing so, it is possible to suppress a decrease in coding efficiency while suppressing a substantial limitation on scalable decoding. In other words, it is possible to apply to scalable decoding while suppressing a decrease in coding efficiency.
[0151] <Other examples of switching methods>
[0152] Meanwhile, the levels of the switching hierarchical method are arbitrary. Additionally, the switching hierarchical method can be any method and is not limited to the above examples (switching between a method that supports scalable decoding and a method that does not support scalable decoding). For example, it is possible to switch between using a hierarchical method that samples according to distance (such as Lifting described in Non-Patent Document 2) and a method of hierarchically dividing points by arranging each point (the attribute data corresponding to each point) in Morton order and sampling the points at equal intervals, as Figure 5 described in "Method 2-2" in the twelfth row of the table shown.
[0153] For example, the hierarchical method that samples according to distance can be applied to the inverse hierarchical division of levels higher than a predetermined level, and the hierarchical method that samples at equal intervals in Morton order can be applied to the inverse hierarchical division of the predetermined level and lower levels. By doing so, as in the above example, a decrease in coding efficiency can be suppressed.
[0154] <Control Information>
[0155] For example, as Figure 5 described in "Method 2-3" in the thirteenth row of the table shown, inverse hierarchical division (switching between inverse hierarchical methods) can be performed based on the attribute data signaled (sent from the encoding side to the decoding side).
[0156] The content of this control information is arbitrary. For example, the control information can be the aforementioned scalable_enable_flag. Additionally, the control information can be, for example, the aforementioned scalable_enable_num_of_lod. Moreover, both of them are acceptable. For example, when scalable_enable_flag indicates that a hierarchical method applicable to scalable decoding can be applied, the level for switching the inverse hierarchical method can be identified for this attribute data based on scalable_enable_num_of_lod.
[0157] It goes without saying that examples other than these are also possible. By using such control information, it is possible to switch between hierarchical methods as in the case of encoding. Therefore, encoded data to which multiple hierarchical methods have been applied can be correctly decoded. Therefore, a decrease in coding efficiency can be suppressed.
[0158] <2. Control of Reference Relationships>
[0159] <Hierarchical Structure of Attribute Data>
[0160] However, since the attribute data is hierarchically divided by classifying its points into reference points or prediction points, nodes may not be formed at the lowest level when the points are sparse.
[0161] For example, in the region of the voxel 50 shown in A of Figure 7 , voxels 50-1 to 50-4 are formed at a level one level lower than the voxel 50, as shown in B of Figure 7 .
[0162] In the case of the layering of geometric data, as shown in A of Figure 7 , even at levels lower than the voxel 50, nodes corresponding to the unique point 51 existing in the voxel 50 are generated. For example, in Figure 7 B, a node of the point 51 is formed in the voxel 50-4. However, in the case of attribute data, when the unique point 51 is assigned to the level of the voxel 50, nodes cannot be assigned to voxels lower than the voxel 50.
[0163] Therefore, in the case of attribute data, the number of points at a specific level may be less than the number of points at levels higher than the specific level. For example, Figure 8 A shows an example of the hierarchical structure of attribute data. Figure 8 A shows examples of four levels from level of detail (LoD)=0 to LoD=3. The circles at each level represent points. In the case of this example, the number of points at LoD=2 is 6, while the number of points at LoD=3 lower than LoD=2 is 2. In this case, the reference relationship is formed according to this hierarchical structure and is thus configured as shown by the arrows in B of Figure 8 .
[0164] The number of levels and the level width (the number of points per level) have optimal values depending on the data. However, generally speaking, a configuration in which the number of points monotonically increases from a higher level to a lower level has the highest coding efficiency. As in the example of Figure 8 , in a configuration where the number of points at a lower level is less than the number of points at a higher level, the coding efficiency may decrease.
[0165] <Merging of Nodes at Multiple Levels>
[0166] Therefore, the reference relationship can be constructed by merging points (nodes) at multiple levels.
[0167] For example, for the attribute information of each point in a point cloud representing an object of a three-dimensional shape as a set of points, when the attribute information is hierarchically processed by recursively repeating the following process for a reference point, the predicted value of the predicted points at multiple levels is derived by referring to the reference points at the same level: classifying the points into predicted points or reference points, deriving the predicted value of the attribute information of the predicted points using the attribute information of the reference points, and deriving the difference between the attribute information of the predicted points and the predicted value.
[0168] For example, the information processing apparatus may include a layering unit that recursively repeats the following process for each piece of attribute information of each point in a point cloud representing an object of a three-dimensional shape as a set of points to perform layering of the attribute information: classifying a point as a prediction point or a reference point, deriving a predicted value of the attribute information of the prediction point using the attribute information of the reference point, and deriving a difference between the attribute information of the prediction point and the predicted value, where the layering unit may derive the predicted values of the prediction points at multiple levels with reference to the reference points at the same level. In other words, the information processing apparatus that processes the attribute information of each point in a point cloud representing an object of a three-dimensional shape as a set of points may include a layering unit that recursively repeats the following process for the reference points among the points classified as prediction points or reference points to perform layering of the attribute information: deriving a difference between the predicted value of the attribute information of the prediction point derived using the attribute information of the reference point and the attribute information of the prediction point, where the layering unit may perform derivation of the predicted values of the prediction points at multiple levels with reference to the reference points at the same level.
[0169] By doing so, a reference relationship with a configuration in which the number of points monotonically increases from a higher level to a lower level can be constructed, and thus a decrease in coding efficiency can be suppressed.
[0170] For example, the node (point) at the processing target level is merged into the node (point) at the level one higher to construct a reference relationship in the hierarchical structure of the attribute information, as Figure 9 described in "Method 3" in the first row of the table shown. For example, in the attribute data having a hierarchical structure as shown in A of Figure 8 , the points at the LoD = 3 level are merged with the points at the LoD = 2 level, which is the level immediately above, to construct a reference relationship. By doing so, a reference relationship with a configuration in which the number of points monotonically increases from a higher level to a lower level can be constructed, as Figure 10 shown.
[0171] In addition, when the number of nodes at the processing target level is greater than the number of nodes at the level one higher, the above process may be performed. For example, as Figure 9 described in "Method 3-1" in the second row of the table shown. That is, when the number of nodes at the processing target level is greater than the number of nodes at the level one higher, a reference relationship is constructed such that the nodes at the processing target level refer to the nodes at the level one higher. By doing so, a reference relationship with a configuration in which the number of points monotonically increases from a higher level to a lower level can be constructed more easily.
[0172] <Control Information>
[0173] For example, control information regarding the construction of the reference relationship (sent from the encoding side to the decoding side) may be signaled, as Figure 9 as described in "Method 3-2" in the third row of the table shown.
[0174] The method of sending the control information is arbitrary. For example, the control information can be defined by a syntax or the like, included in the encoded data (bitstream) of the point cloud data (for example, described in a header or the like) and sent. Additionally, different from the encoded data of the point cloud data, the control information can be sent as data associated with the encoded data of the point cloud data.
[0175] <merge_lower_lod_flag>
[0176] The content of the control information is arbitrary. For example, control information indicating whether to construct a reference relationship according to merging with other levels can be sent to the decoding side. In this case, the control information can indicate whether to construct a reference relationship according to merging with other levels in any way.
[0177] For example, merge_lower_lod_flag can be sent as flag information. merge_lower_lod_flag is flag information indicating whether to construct a reference relationship according to merging with the level one higher. For a point (node) of a level, when the value of merge_lower_lod_flag is true (for example, "1"), it indicates constructing a reference relationship according to merging with the level one higher, and when its value is false (for example, "0"), it indicates constructing a reference relationship without merging with the level one higher (usually only referring to the level one higher). For example, it can be assumed that the case of omitting merge_lower_lod_flag is equivalent to the case where the value of this flag information is false (for example, "0").
[0178] Meanwhile, this control information (merge_lower_lod_flag) is set for each level (level of detail (LoD)).
[0179] <Inverse hierarchical division>
[0180] Although the processing in encoding has been described above, this technology can be applied to the decoding side. For example, in inverse hierarchical division, based on constructing a reference relationship according to the merging of points (nodes) of multiple levels, reference points of the same level can be referred to reconstruct predicted points of multiple levels.
[0181] For example, for the attribute information of each point in a point cloud that represents an object with a three-dimensional shape as a set of points, when performing inverse stratification on the stratified attribute information obtained by recursively repeating the process of classifying points into predicted points or reference points, deriving a predicted value of the attribute information of a predicted point using the attribute information of a reference point, and deriving the difference between the attribute information of the predicted point and the predicted value for a reference point, the predicted value of the attribute information of predicted points at multiple levels is derived by referring to the attribute information of reference points at the same level, and the derived predicted value is added to the difference to generate the attribute information of the predicted point.
[0182] For example, an information processing apparatus may include an inverse stratification unit that performs inverse stratification on the stratified attribute information obtained by recursively repeating the process of classifying points into predicted points or reference points, deriving a predicted value of the attribute information of a predicted point using the attribute information of a reference point, and deriving the difference between the attribute information of the predicted point and the predicted value for each point in a point cloud that represents an object with a three-dimensional shape as a set of points. When performing inverse stratification, the inverse stratification unit may derive the predicted value of the attribute information of predicted points at multiple levels by referring to the attribute information of reference points at the same level, and add the derived predicted value to the difference to generate the attribute information of the predicted point. In other words, the information processing apparatus, which processes the attribute information of each point in a point cloud that represents an object with a three-dimensional shape as a set of points, may include an inverse stratification unit that performs inverse stratification on the stratified attribute information obtained by recursively repeating the process of deriving the difference between the predicted value of the attribute information of a predicted point derived using the attribute information of a reference point and the attribute information of the predicted point for reference points among the points classified as predicted points or reference points. When performing inverse stratification, the inverse stratification unit may derive the predicted value of the attribute information of predicted points at multiple levels by referring to the attribute information of reference points at the same level, and add the derived predicted value to the difference to generate the attribute information of the predicted point.
[0183] By doing so, it is possible to correctly inverse-stratify the attribute data in which a reference relationship with a configuration where the number of points monotonically increases from a higher level to a lower level is constructed, thereby suppressing a decrease in coding efficiency.
[0184] For example, it is possible to search for the reference destination of a node (point) at the processing target level at a level two levels higher, as Figure 9 described in "Method 4" in the fourth row of the table shown. By doing so, it is possible to correctly inverse-stratify the attribute data in which a reference relationship with a configuration where the number of points monotonically increases from a higher level to a lower level is constructed as Figure 10 shown. Therefore, a decrease in coding efficiency can be suppressed.
[0185] <Control Information>
[0186] For example, inverse layer division can be performed based on the signal-sent attribute data (sent from the encoding side to the decoding side), and (switching between inverse layer division methods is possible), as Figure 9 described in "Method 4-1" in the fifth row of the table shown.
[0187] The content of this control information is arbitrary. For example, the aforementioned merge_lower_lod_flag can be used. The merge_lower_lod_flag can also be control information indicating whether to refer to the same layer as other layers during inverse layer division. That is, during inverse layer division, the layer to be used as the reference destination can be identified based on the merge_lower_lod_flag. For example, the layer to be used as the reference destination for the processing target layer can be identified based on the merge_lower_lod_flag, and the predicted value of the attribute data of the predicted point of the processing target layer can be derived by referring to the attribute data of the reference point of the identified layer.
[0188] By using such control information, inverse layer division can be performed by the same method as that used during encoding. Therefore, a decrease in encoding efficiency can be suppressed.
[0189] <Case of not using control information>
[0190] Meanwhile, the inverse layer division method can be set without using the aforementioned control information. For example, on the decoding side, the number of points of the processing target layer can be compared with the number of points of the layer one level higher, and for a layer whose number of points is not sufficiently greater than that of the layer one level higher, the predicted value of the attribute information of the predicted point can be derived by referring to the attribute information of the reference point of the layer two levels higher. That is, the same determination process as that performed on the encoding side can be performed on the decoding side, and the reference relationship can be constructed based on the determination result, and the predicted value can be derived.
[0191] By doing so, inverse layer division can be performed by the same method as that used during encoding. Therefore, a decrease in encoding efficiency can be suppressed. In addition, the transmission of control information can be omitted, so a further decrease in encoding efficiency can be suppressed.
[0192] <3. First Embodiment>
[0193] <Encoding Device>
[0194] Next, the device applying the present technology described above in <1. Switching of Layer Division / Inverse Layer Division Method> will be described. Figure 11 is a block diagram showing an example of the configuration of an encoding device as an aspect of an information processing device applying the present technology. Figure 11 The encoding device 100 shown encodes point cloud (3D data). The encoding device 100 encodes the point cloud using the present technology described above in <1. Switching between hierarchical / inverse hierarchical methods>.
[0195] Meanwhile, Figure 11 main parts such as processing units and data flows are shown, but the processing units and data flows are not limited to Figure 11 the processing units and data flows shown in Figure 11 That is, Figure 11 processing units not shown as blocks in
[0196] and Figure 11 processing and data flows not shown as arrows, etc. in
[0197]
[0198] The position information encoding unit 101 encodes the geometric data (position information) of the point cloud (3D data) input to the encoding device 100. The encoding method can be any method corresponding to scalable decoding. For example, the position information encoding unit 101 hierarchically divides the geometric data to generate an octree and encodes the octree. In addition, for example, processes such as filtering and quantization for suppressing noise (denoising) can be performed. The position information encoding unit 101 provides the encoded data of the generated geometric data to the position information decoding unit 102 and the bitstream generation unit 105.
[0199] The position information decoding unit 102 acquires the encoded data of the geometric data provided by the position information encoding unit 101 and decodes the encoded data. The decoding method can be any method corresponding to the encoding performed by the position information encoding unit 101. For example, processes such as filtering and quantization for denoising can be performed. The position information decoding unit 102 provides the generated geometric data (decoding result) to the point cloud generation unit 103.
[0200] The point cloud generation unit 103 acquires the attribute data (attribute information) of the point cloud input to the encoding device 100 and the geometric data (decoding result) provided by the position information decoding unit 102. The point cloud generation unit 103 performs a process (recoloring process) of associating the attribute data with the geometric data (decoding result). The point cloud generation unit 103 provides the attribute data associated with the geometric data (decoding result) to the attribute information encoding unit 104.The attribute information encoding unit 104 acquires the geometric data (decoding result) and the attribute data provided by the point cloud generation unit 103. The attribute information encoding unit 104 encodes the attribute data using the geometric data (decoding result), and generates encoded data of the attribute data.
[0201] In this case, the attribute information encoding unit 104 encodes the attribute data using the present technique described above in <1. Switching of Hierarchical / Inverse Hierarchical Methods>. For example, the attribute information encoding unit 104 switches the hierarchical method of the attribute data at an intermediate level in the hierarchy of the attribute data. The attribute information encoding unit 104 provides the generated encoded data of the attribute data to the bitstream generation unit 105.
[0202] The bitstream generation unit 105 acquires the encoded data of the position information provided by the position information encoding unit 101. In addition, the bitstream generation unit 105 acquires the encoded data of the attribute information provided by the attribute information encoding unit 104. The bitstream generation unit 105 generates a bitstream including the encoded data. The bitstream generation unit 105 outputs the generated bitstream to the outside of the encoding device 100.
[0203] By adopting such a configuration, the encoding device 100 can form a hierarchical structure of the attribute data using multiple hierarchical methods. Therefore, hierarchical encoding can be performed on more different types of attribute data by a more appropriate method to suppress a decrease in encoding efficiency.
[0204] Meanwhile, these processing units (position information encoding unit 101 to bitstream generation unit 105) of the encoding device 100 have arbitrary configurations. For example, each processing unit can be configured as a logic circuit for implementing the above processing. In addition, each processing unit can include, for example, a central processing unit (CPU), a read-only memory (ROM), a random access memory (RAM), etc., and execute a program using these components to implement the above processing. Needless to say, each processing unit can have the above two configurations, implementing a part of the above processing according to a logic circuit, and implementing the other part of the processing by executing a program. The processing units can have independent configurations. For example, some processing units can implement the above partial processing according to a logic circuit, some other processing units can implement the above processing by executing a program, and still some other processing units can implement the above processing in both ways of a logic circuit and executing a program.
[0205] <Attribute Information Encoding Unit>
[0206] Figure 12 is a block diagram showing an example of the main components of the attribute information encoding unit 104 ( Figure 11 ). Meanwhile, Figure 12 shows main parts such as a processing unit and a data stream, but the processing unit and the data stream are not limited to Figure 12 the processing unit and the data stream shown in Figure 12 That is, the processing unit not shown as a block in Figure 12 and the processing and data stream not shown as an arrow or the like in
[0207] such as Figure 12 shown, the attribute information encoding unit 104 includes a hierarchical processing unit 111, a quantization unit 112, and an encoding unit 113.
[0208] The hierarchical processing unit 111 performs hierarchical processing on the attribute data. For example, the hierarchical processing unit 111 acquires the attribute data and the geometric data (decoding result) provided from the point cloud generation unit 103. The hierarchical processing unit 111 hierarchically processes the attribute data using the geometric data. In this case, the hierarchical processing unit 111 performs hierarchical processing using the present technology described above in <1. Switching of Hierarchical / Inverse Hierarchical Methods>. For example, the hierarchical processing unit 111 switches the hierarchical method of the attribute data at an intermediate level in the hierarchical processing of the attribute data. In other words, the hierarchical processing unit 111 hierarchically processes the attribute data using multiple hierarchical methods and generates different levels according to the hierarchical methods. The hierarchical processing unit 111 provides the hierarchically processed attribute data (difference) to the quantization unit 112.
[0209] In this case, the hierarchical processing unit 111 also generates control information regarding the hierarchical processing. The hierarchical processing unit 111 also provides the generated control information together with the attribute data (difference) to the quantization unit 112.
[0210] The quantization unit 112 acquires the attribute data (difference) and the control information provided from the hierarchical processing unit 111. The quantization unit 112 quantizes the attribute data (difference). The method for quantization is arbitrary. The quantization unit 112 provides the quantized attribute data (difference) and the control information to the encoding unit 113.
[0211] The encoding unit 113 acquires the quantized attribute data (difference) and the control information provided from the quantization unit 112. The encoding unit 113 encodes the quantized attribute data (difference) to generate encoded data of the attribute data. The method for encoding is arbitrary. In addition, the encoding unit 113 includes the control information in the generated encoded data. In other words, the encoding unit 113 generates encoded data of the attribute data including the control information. The encoding unit 113 provides the generated encoded data to the bitstream generation unit 105.
[0212] By performing layering as described above, the attribute information encoding unit 104 can form a hierarchical structure of attribute data using multiple layering methods. Therefore, more different types of attribute data can be layered by a more appropriate method to suppress a decrease in encoding efficiency.
[0213] Meanwhile, the processing units (the layer processing unit 111 to the encoding unit 113) have an arbitrary configuration. For example, each processing unit can be configured as a logic circuit for implementing the above processing. In addition, each processing unit can include, for example, a CPU, a ROM, a RAM, etc. and execute a program using these components to implement the above processing. Needless to say, each processing unit can have the above two configurations, implementing a part of the above processing according to the logic circuit and implementing another part of the processing by executing the program. The processing units can have an independent configuration. For example, some processing units can implement the above partial processing according to the logic circuit, some other processing units can implement the above processing by executing the program, and still some other processing units can implement the above processing in both ways of the logic circuit and executing the program.
[0214] <Layer processing unit>
[0215] Figure 13 is a block diagram showing an example of the main components of the layer processing unit 111 ( Figure 12 ). Meanwhile, Figure 13 shows the main parts such as the processing units and the data flow, but the processing units and the data flow are not limited to Figure 13 the processing units and the data flow shown in Figure 13 . That is, Figure 13 the processing units not shown as blocks in
[0216] Here, a description will be given assuming that the layer processing unit 111 switches between a layering method applicable to scalable decoding and a layering method not applicable to scalable decoding at an intermediate level. As Figure 13 shown, the layer processing unit 111 includes a control unit 121, a scalable layer processing unit 122, a non-scalable layer processing unit 123, an inversion unit 124, and a weighting unit 125.
[0217] The control unit 121 performs processing for controlling layering. For example, the control unit 121 acquires the attribute data and the geometric data (decoding result) provided by the point cloud generation unit 103. The control unit 121 provides the acquired attribute data and geometric data (decoding result) to the scalable layer processing unit 122.
[0218] In addition, the control unit 121 controls the scalable hierarchical processing unit 122 and the non-scalable hierarchical processing unit 123 to perform hierarchical processing. For example, the control unit 121 causes the present technology described above in "<1. Switching of Hierarchical / Inverse Hierarchical Methods>" to be used for performing hierarchical processing. That is, for example, the control unit 121 drives the scalable hierarchical processing unit 122, the non-scalable hierarchical processing unit 123, or both of them, and causes them to perform hierarchical processing according to their hierarchical methods.
[0219] In addition, the control unit 121 generates control information regarding the hierarchical processing of the attribute data and provides the control information to the quantization unit 112.
[0220] The scalable hierarchical processing unit 122 performs processing for hierarchical processing of the attribute data according to a method applicable to scalable decoding (for example, the hierarchical method described in Non-Patent Document 3). For example, the scalable hierarchical processing unit 122 acquires the attribute data and the geometric data (decoding result) provided from the control unit 121.
[0221] The scalable hierarchical processing unit 122 hierarchically processes the acquired attribute data using the acquired geometric data according to the control of the control unit 121 by a method applicable to scalable decoding. For example, the scalable hierarchical processing unit 122 generates a layer specified by the control unit 121 (for example, a layer that allows processing to be performed, and a layer that does not prohibit processing is also acceptable) using a method applicable to scalable decoding.
[0222] The scalable hierarchical processing unit 122 provides the attribute data of the generated layer, the attribute data that has not been hierarchically processed, the geometric data, etc. to the non-scalable hierarchical processing unit 123.
[0223] Meanwhile, when the control unit 121 does not permit hierarchical processing, the scalable hierarchical processing unit 122 may omit hierarchical processing. In this case, the scalable hierarchical processing unit 122 provides all the acquired attribute data and geometric data to the non-scalable hierarchical processing unit 123.
[0224] The non-scalable hierarchical processing unit 123 performs processing for hierarchical processing of the attribute data according to a method not applicable to scalable decoding (for example, Lifting described in Non-Patent Document 2). For example, the non-scalable hierarchical processing unit 123 acquires the attribute data and the geometric data (decoding result) provided from the scalable hierarchical processing unit 122.
[0225] The non-scalable layer processing unit 123 performs layer processing on the non-layered attribute data according to the control of the control unit 121, using the acquired geometric data, by a method not applicable to scalable decoding. For example, the non-scalable layer processing unit 123 uses a method not applicable to scalable decoding to generate the levels specified by the control unit 121 (for example, the levels that allow processing to be performed, and it is also possible to have levels that do not prohibit processing).
[0226] All levels of the attribute data are generated according to the layer processing performed by the scalable layer processing unit 122 and the non-scalable layer processing unit 123 (all attribute data is layered).
[0227] The non-scalable layer processing unit 123 provides the layered attribute data to the inversion unit 124.
[0228] At the same time, when the control unit 121 does not allow layer processing, the non-scalable layer processing unit 123 can omit layer processing. In this case, all the acquired attribute data is layered, and the scalable layer processing unit 122 provides this attribute data to the inversion unit 124.
[0229] The inversion unit 124 performs processing for inverting the levels. For example, the inversion unit 124 acquires the layered attribute data provided from the non-scalable layer processing unit 123. In this attribute data, the information of each level has been layered in the generation order.
[0230] The inversion unit 124 inverts the levels of the attribute data. For example, the inversion unit 124 assigns level numbers (numbers used to identify levels, which increase by 1 each time one level is decreased starting from the highest level 0 to reach the maximum value corresponding to the lowest level) to each level of the attribute data in the order opposite to the generation order, so that the generation order becomes the order from the lowest level to the highest level.
[0231] The inversion unit 124 provides the attribute data with inverted levels to the weighting unit 125.
[0232] The weighting unit 125 performs processing for weighting. For example, the weighting unit 125 acquires the attribute data provided from the inversion unit 124. The weighting unit 125 derives the weight value of the acquired attribute data. The method for deriving the weight value is arbitrary. The method for deriving the weight value can be changed at the levels generated by the scalable layer processing unit 122 and the levels generated by the non-scalable layer processing unit 123.
[0233] The weighting unit 125 provides the attribute data (difference) and the derived weight value to the quantization unit 112( Figure 12 )。In addition, the weighting unit 125 can provide the derived weight value as control information to the quantization unit 112 and send the control information to the decoding side.
[0234] The control unit 121 controls the scalable hierarchical processing unit 122 and the non-scalable hierarchical processing unit 123, and switches the attribute data hierarchical method at the intermediate level as described above in <1. Switching of Hierarchical / Inverse Hierarchical Methods>. By doing so, the hierarchical processing unit 111 can form a hierarchical structure of attribute data using multiple hierarchical methods. Therefore, hierarchical processing can be performed on more different types of attribute data by a more appropriate method to suppress a decrease in coding efficiency.
[0235] Meanwhile, the processing units (control unit 121 to weighting unit 125) have arbitrary configurations. For example, each processing unit can be configured as a logic circuit for implementing the above processing. In addition, each processing unit can include, for example, a CPU, ROM, RAM, etc. and execute a program using these components to implement the above processing. Needless to say, each processing unit can have the above two configurations, implement a part of the above processing according to the logic circuit, and implement the other part of the processing by executing a program. The processing units can have independent configurations. For example, some processing units can implement the above partial processing according to the logic circuit, some other processing units can implement the above processing by executing a program, and still some other processing units can implement the above processing in both ways of the logic circuit and executing a program.
[0236] <Flow of Encoding Processing>
[0237] Next, the processing performed by the encoding device 100 will be described. The encoding device 100 encodes the data of the point cloud by executing encoding processing. An example of the flow of this encoding processing will be described with reference to Figure 14 the flowchart of.
[0238] When the encoding processing starts, in step S101, the position information encoding unit 101 of the encoding device 100 encodes the geometric data (position information) of the input point cloud to generate encoded data of the geometric data.
[0239] In step S102, the position information decoding unit 102 decodes the encoded data of the geometric data generated in step S101 to generate position information.
[0240] In step S103, the point cloud generation unit 103 performs a recoloring process using the attribute data (attribute information) of the input point cloud and the geometric data (decoding result) generated in step S102 to associate the attribute data with the geometric data.
[0241] In step S104, the attribute information encoding unit 104 encodes the attribute data that has undergone the recoloring process in step S103 by performing an attribute information encoding process to generate encoded data of the attribute data. In this case, the attribute information encoding unit 104 encodes the attribute data using the present technology described above in <1. Switching between Hierarchical / Inverse Hierarchical Methods>. For example, the attribute information encoding unit 104 switches the hierarchical method of the attribute data at an intermediate level in the hierarchy of the attribute data. Details of the attribute information encoding process will be described later.
[0242] In step S105, the bitstream generation unit 105 generates and outputs a bitstream including the encoded data of the geometric data generated in step S101 and the encoded data of the attribute data generated in step S104.
[0243] When the process of step S105 ends, the encoding process ends.
[0244] In this way, the encoding device 100 can form a hierarchical structure of the attribute data using multiple hierarchical methods by performing the processing of each step. Therefore, hierarchical processing can be performed on more different types of attribute information by a more appropriate method to suppress a decrease in encoding efficiency.
[0245] <Flow of Attribute Information Encoding Process>
[0246] Next, with reference to Figure 15 the flowchart of Figure 14 an example of the flow of the attribute information encoding process performed in step S104 of
[0247] When the attribute information encoding process starts, in step S111, the hierarchical processing unit 111 of the attribute information encoding unit 104 hierarchically processes the attribute data by performing hierarchical processing and derives the difference of the attribute data for each point. In this case, the hierarchical processing unit 111 performs hierarchical processing using the present technology described above in <1. Switching between Hierarchical / Inverse Hierarchical Methods>. For example, the hierarchical processing unit 111 switches the hierarchical method of the attribute data at an intermediate level. Details of the hierarchical processing will be described later.
[0248] In step S112, the quantization unit 112 quantizes each difference derived in step S111.
[0249] In step S113, the encoding unit 113 encodes the differences quantized in step S112 to generate encoded data of the attribute data.
[0250] When the process of step S113 ends, the attribute information encoding process ends and the process returns to Figure 14 .
[0251] In this way, the attribute information encoding unit 104 can form a hierarchical structure of attribute data by performing the processing of each step using multiple hierarchical methods. Therefore, more different types of attribute information can be hierarchically processed by a more appropriate method to suppress the reduction of encoding efficiency.
[0252] <Flow of hierarchical processing>
[0253] Next, an example of the flow of the hierarchical processing performed in step S111 of Figure 16 will be described with reference to the flowchart of Figure 15 an example of the flow of the hierarchical processing performed in step S111 of
[0254] When the hierarchical processing starts, in step S121, the control unit 121 of the hierarchical processing unit 111 sets the application range of the method applicable to scalable decoding. The control unit 121 controls the hierarchy based on this setting.
[0255] In step S122, the scalable hierarchical processing unit 122 generates hierarchies within the application range of the method applicable to scalable decoding set in step S121 according to the method applicable to scalable decoding. That is, the scalable hierarchical processing unit 122 sets reference points and prediction points for each hierarchy encoded in a way that scalable decoding can be performed using the method applicable to scalable decoding, derives the predicted value of the attribute data of the prediction point using the attribute data of the reference point, and derives the difference between the attribute data of the prediction point and the predicted value.
[0256] In step S123, the non - scalable hierarchical processing unit 123 generates hierarchies that were not generated in step S122, that is, hierarchies outside the application range of the method applicable to scalable decoding set in step S121, according to the method not applicable to scalable decoding. That is, the non - scalable hierarchical processing unit 123 sets reference points and prediction points for each hierarchy encoded in a way that scalable decoding cannot be performed using the method not applicable to scalable decoding, derives the predicted value of the attribute data of the prediction point using the attribute data of the reference point, and derives the difference between the attribute data of the prediction point and the predicted value.
[0257] When all hierarchies of the attribute data are generated according to the processing of step S122 and step S123, in step S124, the inversion unit 124 inverts the hierarchies of the generated attribute data and assigns hierarchy numbers to each hierarchy in the reverse order of the generation order.
[0258] In step S125, the weighting unit 125 derives the weight value for the attribute data of each hierarchy.
[0259] In step S126, the control unit 121 generates hierarchical control information regarding the attribute data, provides the control information to the quantization unit 112, and causes the control information to be sent to the decoding side.
[0260] When the processing of step S126 ends, the processing returns to Figure 15 .
[0261] In this way, the hierarchical processing unit 111 can form a hierarchical structure of the attribute data using multiple hierarchical methods by performing the processing of each step. Therefore, hierarchical processing can be performed on more different types of attribute data by a more appropriate method to suppress a decrease in coding efficiency.
[0262] <4. Second Embodiment>
[0263] <Decoding Device>
[0264] Next, a device applying the present technology described above in <1. Switching of Hierarchical / Inverse Hierarchical Methods> will be described. Figure 17 is a block diagram showing an example of the configuration of a decoding device as an aspect to which the present technology is applied. Figure 17 The decoding device 200 shown decodes encoded data of point clouds (3D data). The decoding device 200 decodes the encoded data of point clouds using the present technology described above in <1. Switching of Hierarchical / Inverse Hierarchical Methods>.
[0265] Meanwhile, Figure 17 shows main parts such as processing units and data flows, but the processing units and data flows are not limited to Figure 17 the processing units and data flows shown therein. That is, processing units not shown as blocks in Figure 17 and processing and data flows not shown as arrows, etc. in Figure 17 may exist in the decoding device 200.
[0266] As Figure 17 shown, the decoding device 200 includes a decoding target LoD depth setting unit 201, an encoded data extraction unit 202, a position information decoding unit 203, an attribute information decoding unit 204, and a point cloud generation unit 205.
[0267] The decoding target LoD depth setting unit 201 performs processing for setting the depth of the level of detail (LoD) as the decoding target. For example, the decoding target LoD depth setting unit 201 sets the level to which the encoded data of the point cloud held in the encoded data extraction unit 202 is decoded. The method of setting the level depth as the decoding target is arbitrary.
[0268] For example, the decoding target LoD depth setting unit 201 can set the hierarchical depth based on instructions regarding the hierarchical depth from an external source such as a user or an application. Additionally, the decoding target LoD depth setting unit 201 can obtain the hierarchical depth to be decoded based on arbitrary information such as the output image and set the hierarchical depth.
[0269] For example, the decoding target LoD depth setting unit 201 can set the hierarchical depth to be decoded based on the viewpoint position, direction, viewing angle, movement of the viewpoint (movement, translation, tilt, and zoom), etc. of the two-dimensional image generated from the point cloud.
[0270] Meanwhile, the data unit for setting the hierarchical depth to be decoded is arbitrary. For example, the decoding target LoD depth setting unit 201 can set the hierarchical depth for all point clouds, set the hierarchical depth for each object, or set the hierarchical depth for each local region within an object. Needless to say, the hierarchical depth can also be set with data units other than these examples.
[0271] The encoded data extraction unit 202 acquires and holds the bitstream input to the decoding device 200. The encoded data extraction unit 202 extracts the encoded data of the geometric data (position information) and the attribute data (attribute information) from the highest level to the level specified by the decoding target LoD depth setting unit 201 from the held bitstream. The encoded data extraction unit 202 provides the extracted encoded data of the geometric data to the position information decoding unit 203. The encoded data extraction unit 202 provides the extracted encoded data of the attribute data to the attribute information decoding unit 204.
[0272] The position information decoding unit 203 acquires the encoded data of the geometric data provided by the encoded data extraction unit 202. The position information decoding unit 203 decodes the encoded data of the geometric data to generate geometric data (decoding result). This decoding method can be the same as that of the position information decoding unit 102 of the encoding device 100. The position information decoding unit 203 provides the generated geometric data (decoding result) to the attribute information decoding unit 204 and the point cloud generation unit 205.
[0273] The attribute information decoding unit 204 obtains the encoded data of the attribute data provided by the encoded data extraction unit 202. The attribute information decoding unit 204 obtains the geometric data (decoding result) provided by the position information decoding unit 203. The attribute information decoding unit 204 uses the position information (decoding result) to decode the encoded data of the attribute data according to the method of the present technology described above in <1. Switching between hierarchical / inverse hierarchical methods> to generate the attribute data (decoding result). For example, the attribute information decoding unit 204 switches the inverse hierarchical method of the attribute data at the intermediate level. The attribute information decoding unit 204 provides the generated attribute data (decoding result) to the point cloud generation unit 205.
[0274] The point cloud generation unit 205 obtains the geometric data (decoding result) provided by the position information decoding unit 203. The point cloud generation unit 205 obtains the attribute data (decoding result) provided by the attribute information decoding unit 204. The point cloud generation unit 205 generates a point cloud (decoding result) using the geometric data (decoding result) and the attribute data (decoding result). The point cloud generation unit 205 outputs the data of the generated point cloud (decoding result) to the outside of the decoding device 200.
[0275] By adopting such a configuration, the decoding device 200 can perform inverse hierarchical processing using multiple hierarchical methods, so that the attribute data hierarchically divided using multiple hierarchical methods can be correctly inverse hierarchically divided. That is, the encoded data that has been encoded using multiple hierarchical methods can be correctly decoded. Therefore, a decrease in encoding efficiency can be suppressed.
[0276] At the same time, the processing units (the decoding target LoD depth setting unit 201 to the point cloud generation unit 205) have arbitrary configurations. For example, each processing unit can be configured as a logic circuit for implementing the above processing. In addition, each processing unit can include, for example, a CPU, a ROM, a RAM, etc. and execute a program using these components to implement the above processing. It goes without saying that each processing unit can have the above two configurations, implement a part of the above processing according to the logic circuit, and implement the other part of the processing by executing a program. The processing units can have independent configurations. For example, some processing units can implement the above partial processing according to the logic circuit, some other processing units can implement the above processing by executing a program, and still some other processing units can implement the above processing in both ways of the logic circuit and executing a program.
[0277] <Attribute information decoding unit>
[0278] Figure 18 is a block diagram showing an example of the main components of the attribute information decoding unit 204 ( Figure 17 ). At the same time, Figure 18 shows main parts such as a processing unit and a data stream, but the processing unit and the data stream are not limited to Figure 18 the processing unit and the data stream shown in Figure 18 That is, a processing unit not shown as a block in Figure 18 and a process and a data stream not shown as an arrow or the like in
[0279] may exist in the attribute information decoding unit 204 as Figure 18 shown, the attribute information decoding unit 204 includes a decoding unit 211, an inverse quantization unit 212, and an inverse layer processing unit 213.
[0280] The decoding unit 211 performs a process of decoding the encoded data of the attribute data. For example, the decoding unit 211 acquires the encoded data of the attribute data provided from the attribute information decoding unit 204.
[0281] The decoding unit 211 decodes the encoded data of the attribute data to generate attribute data (decoding result). The decoding method may be a method corresponding to the encoding method performed by the encoding unit 113 of the encoding device 100 ( Figure 12 ). In addition, the generated attribute data (decoding result) corresponds to the attribute data before being encoded, and is the difference between the attribute data and its predicted value, and has been quantized, as described in the first embodiment. The decoding unit 211 provides the generated attribute data (decoding result) to the inverse quantization unit 212.
[0282] Meanwhile, when the encoded data of the attribute data includes control information about a weight value and control information about the layer of the attribute data, the decoding unit 211 also provides the control information to the inverse quantization unit 212.
[0283] The inverse quantization unit 212 performs a process of inverse quantization of the attribute data. The inverse quantization unit 212 acquires the attribute data (decoding result) and the control information provided from the decoding unit 211.
[0284] The inverse quantization unit 212 performs inverse quantization on the attribute data (decoding result). In this case, when control information about a weight value is provided from the decoding unit 211, the inverse quantization unit 212 also acquires the control information and performs inverse quantization on the attribute data (decoding result) based on the control information (using the weight value obtained based on the control information).
[0285] In addition, when control information about the layer of the attribute data is provided from the decoding unit 211, the inverse quantization unit 212 also acquires the control information.
[0286] The inverse quantization unit 212 supplies the inverse quantized attribute data (decoding result) to the inverse layer processing unit 213. Further, when the control information regarding the layer of the attribute data is acquired from the decoding unit 211, the inverse quantization unit 212 also supplies the control information to the inverse layer processing unit 213.
[0287] The inverse layer processing unit 213 acquires the inverse quantized attribute data (decoding result) supplied from the inverse quantization unit 212. As described above, the attribute data is a difference value. Further, the inverse layer processing unit 213 acquires the geometric data (decoding result) supplied from the position information decoding unit 203. The inverse layer processing unit 213 performs inverse layer processing on the acquired attribute data (difference value) using the geometric data, and the inverse layer processing is the inverse process of the layer processing performed by the layer processing unit 111 of the encoding device 100 ( Figure 12 ).
[0288] In this case, the inverse layer processing unit 213 performs inverse layer processing using the present technique described above in <1. Switching of layer / inverse layer methods>. That is, the inverse layer processing unit 213 switches the attribute information layer method at the intermediate level in the inverse layer (that is, switches between the inverse layer methods). The inverse layer processing unit 213 supplies the inverse layer processed attribute data as a decoding result to the point cloud generation unit 205 ( Figure 17 ).
[0289] By performing the inverse layer processing as described above, the attribute information decoding unit 204 can correctly perform inverse layer processing on the attribute data layered using a plurality of layer methods. That is, the encoded data that has been encoded using a plurality of layer methods can be correctly decoded. Therefore, a decrease in encoding efficiency can be suppressed.
[0290] Meanwhile, the processing units (the decoding unit 211 to the inverse layer processing unit 213) have arbitrary configurations. For example, each processing unit may be configured as a logic circuit for implementing the above processing. Further, each processing unit may include, for example, a CPU, a ROM, a RAM, etc. and execute a program using these components to implement the above processing. Needless to say, each processing unit may have the above two configurations, implement a part of the above processing according to the logic circuit, and implement the other part of the processing by executing the program. The processing units may have an independent configuration. For example, some processing units may implement the above partial processing according to the logic circuit, some other processing units may implement the above processing by executing the program, and still some other processing units may implement the above processing in both ways of the logic circuit and executing the program.
[0291] <Inverse layer processing unit>
[0292] Figure 19 shows the inverse layer processing unit 213 ( Figure 18 Block diagram of an example of the main components of (). At the same time, Figure 19 The main parts such as the processing unit and the data stream are shown, but the processing unit and the data stream are not limited to Figure 19 the processing unit and the data stream shown in Figure 19 That is, a processing unit not shown as a block in Figure 19 and a process and data stream not shown as an arrow, etc. in
[0293] Here, a description will be given on the assumption that the inverse layer processing unit 213 switches between an inverse layer method applicable to scalable decoding and an inverse layer method not applicable to scalable decoding at an intermediate level. As Figure 19 shown, the inverse layer processing unit 213 includes a control unit 221, a non-scalable inverse layer processing unit 222, and a scalable inverse layer processing unit 223.
[0294] The control unit 221 performs processing for controlling inverse layer. For example, the control unit 221 acquires the inverse quantized attribute data provided from the inverse quantization unit 212 ( Figure 18 ) and the control information on the inverse layer of the attribute data. In addition, the control unit 221 also acquires the geometric data (decoding result) provided from the position information decoding unit 203.
[0295] The control unit 221 uses the present technology described above in <1. Switching of layer / inverse layer method> to control the non-scalable inverse layer processing unit 222 and the scalable inverse layer processing unit 223, and causes them to perform layering.
[0296] For example, the control unit 221 determines whether to apply a method applicable to scalable decoding to the acquired attribute data based on the control information (e.g., scalable_enable_flag) indicating whether a layer method applicable to scalable decoding can be applied.
[0297] For example, if the control information indicates that a layer method applicable to scalable decoding cannot be applied, the control unit 221 provides the attribute data and the geometric data to the non-scalable inverse layer processing unit 222, and performs inverse layer for all levels of the attribute data using a method not applicable to scalable decoding.
[0298] In addition, if the control information indicates that a layer method applicable to scalable decoding can be applied, the control unit 221 refers to the control information (e.g., scalable_enable_num_of_lod) indicating the level range of applying the layer method applicable to scalable decoding.
[0299] In addition, the control unit 221 selects whether to use a method not applicable to scalable decoding or both a method not applicable to scalable decoding and a method applicable to scalable decoding to perform inverse layer splitting on the attribute data to be processed, based on the control information.
[0300] For example, if the layer of the attribute data to be processed does not include at least a part of the layer range for applying the layer splitting method applicable to scalable decoding indicated by the control information, the control unit 221 provides all the attribute data, geometric data, etc. to be processed to the non-scalable inverse layer splitting processing unit 222, and performs inverse layer splitting on all the layers of the attribute data according to the method not applicable to scalable decoding.
[0301] In addition, for example, if the layer of the attribute data to be processed includes at least a part of the layer range for applying the layer splitting method applicable to scalable decoding indicated by the control information, the control unit 221 provides the attribute data and geometric data of the layers included in this range to the scalable inverse layer splitting processing unit 223. In addition, the control unit 221 provides the attribute data and geometric data of the layers not included in this range to the non-scalable inverse layer splitting processing unit 222. That is, each inverse layer splitting processing unit performs inverse layer splitting by each method.
[0302] The non-scalable inverse layer splitting processing unit 222 performs inverse layer splitting on the attribute data provided by the control unit 221 using the geometric data provided by the control unit 221 by a method not applicable to scalable decoding. The non-scalable inverse layer splitting processing unit 222 provides the inversely layer-split attribute data (non-scalable layer) to the point cloud generation unit 205( Figure 17 )
[0303] The scalable inverse layer splitting processing unit 223 performs inverse layer splitting on the attribute data provided by the control unit 221 using the geometric data provided by the control unit 221 by a method applicable to scalable decoding. The scalable inverse layer splitting processing unit 223 provides the inversely layer-split attribute data (scalable layer) to the point cloud generation unit 205( Figure 17 )
[0304] In this way, the control unit 221 controls the non-scalable inverse layer splitting processing unit 222 and the scalable inverse layer splitting processing unit 223, and switches between the inverse layer splitting methods of the attribute data at the intermediate layer as described above in <1. Switching of Layer Splitting / Inverse Layer Splitting Method>. Therefore, the inverse layer splitting processing unit 213 can correctly perform inverse layer splitting on the attribute data split by multiple layer splitting methods. That is, the encoded data that has been encoded using multiple layer splitting methods can be correctly decoded. Therefore, a reduction in encoding efficiency can be suppressed.
[0305] Meanwhile, the processing units (control unit 221 to scalable inverse hierarchical processing unit 223) have an arbitrary configuration. For example, each processing unit may be configured as a logic circuit for implementing the above processing. In addition, each processing unit may include, for example, a CPU, a ROM, a RAM, etc. and execute a program using these components to implement the above processing. Needless to say, each processing unit may have the above two configurations, implementing a part of the above processing according to the logic circuit and implementing another part of the processing by executing a program. The processing units may have an independent configuration. For example, some processing units may implement the above partial processing according to the logic circuit, some other processing units may implement the above processing by executing a program, and still some other processing units may implement the above processing in both ways of the logic circuit and executing a program.
[0306] <Flow of decoding process>
[0307] Next, the processing performed by the decoding device 200 will be described. The decoding device 200 decodes the encoded data of the point cloud by performing a decoding process. The flow of this decoding process will be described with reference to Figure 20 the flowchart.
[0308] When the decoding process starts, in step S201, the LoD depth setting unit 201 of the decoding device 200 sets the LoD depth to be decoded (i.e., the hierarchical range as the decoding target).
[0309] In step S202, the encoded data extraction unit 202 acquires and holds the bitstream, and determines whether to decode the levels up to the level that can be scalably decoded based on the decoding target hierarchical range set in step S201 and the control information (e.g., scalable_enable_num_of_lod) representing the hierarchical range of the hierarchical method applicable to scalable decoding.
[0310] If the decoding target hierarchical range overlaps with the hierarchical range of the hierarchical method applicable to scalable decoding, and it is determined to decode the levels up to the level that can be scalably decoded (also referred to as the scalable level), the process proceeds to step S203.
[0311] In step S203, the encoded data extraction unit 202 extracts the encoded data of all levels within the decoding target hierarchical range. When the processing of S203 ends, the process proceeds to step S205.
[0312] In addition, in step S202, if the decoding target hierarchical range does not overlap with the hierarchical range of the hierarchical method applicable to scalable decoding, and it is determined not to decode the levels up to the scalable level, the process proceeds to step S204.
[0313] In step S204, the encoded data extraction unit 202 extracts the encoded data of all levels (also referred to as non-scalable levels) that cannot be decoded in a scalable manner. When the processing of S204 ends, the processing proceeds to step S205.
[0314] In step S205, the position information decoding unit 203 decodes the encoded data of the geometric data extracted in step S203 or step S204 to generate position information (decoding result).
[0315] In step S206, the attribute information decoding unit 204 decodes the encoded data of the attribute data extracted in step S203 or step S204 to generate attribute information (decoding result). In this case, the attribute information decoding unit 204 performs processing using the present technique described above in <1. Switching of Hierarchical / Inverse Hierarchical Methods>. For example, the attribute information decoding unit 204 switches between attribute information hierarchical methods at an intermediate level in the inverse hierarchy (i.e., switches between inverse hierarchical methods). Details of the attribute information decoding process will be described later.
[0316] In step S207, the point cloud generation unit 205 uses the geometric data (decoding result) generated in step S205 and the attribute data (decoding result) generated in step S206 to generate a point cloud (decoding result) and outputs the point cloud.
[0317] When the processing of step S207 ends, the decoding process ends.
[0318] By performing the processing of each step in this way, the decoding device 200 can perform inverse hierarchy using multiple hierarchical methods, and thus can correctly perform inverse hierarchy on the attribute data hierarchically divided using multiple hierarchical methods. That is, the encoded data that has been encoded using multiple hierarchical methods can be correctly decoded. Therefore, a decrease in encoding efficiency can be suppressed.
[0319] <Flow of Attribute Information Decoding Process>
[0320] Next, with reference to Figure 21 the flowchart of Figure 20 an example of the flow of the attribute information decoding process performed in step S206 of
[0321] When the attribute information decoding process starts, in step S211, the decoding unit 211 of the attribute information decoding unit 204 decodes the encoded data of the attribute data (attribute information) to generate attribute data (decoding result). This attribute data (decoding result) is quantized as described above.
[0322] In step S212, the inverse quantization unit 212 inverse quantizes the attribute data (decoding result) generated in step S211 by performing inverse quantization processing.
[0323] In step S213, the inverse hierarchical processing unit 213 inverse hierarchizes the attribute data (difference value) inverse quantized in step S212 by performing inverse hierarchical processing to derive the attribute data for each point. In this case, the inverse hierarchical processing unit 213 performs inverse hierarchical processing using the present technology described above in <1. Switching of Hierarchical / Inverse Hierarchical Methods>. For example, the inverse hierarchical processing unit 213 switches between the attribute information hierarchical methods at an intermediate level in the inverse hierarchy (i.e., switches between the inverse hierarchical methods). Details of the inverse hierarchical processing will be described later.
[0324] When the processing of step S213 ends, the attribute information decoding processing ends and the processing returns to Figure 20 .
[0325] By performing the processing of each step in this way, the attribute information decoding unit 204 can perform inverse hierarchical processing using multiple hierarchical methods, and thus can correctly inverse hierarchize the attribute data hierarchized using multiple hierarchical methods. That is, the encoded data that has been encoded using multiple hierarchical methods can be correctly decoded. Therefore, a decrease in encoding efficiency can be suppressed.
[0326] <Flow of Inverse Hierarchical Processing>
[0327] Next, an example of the flow of the inverse hierarchical processing performed in step S213 of Figure 22 will be described with reference to the flowchart of Figure 21 .
[0328] When the inverse hierarchical processing starts, in step S221, the control unit 221 of the inverse hierarchical processing unit 213 hierarchizes the attribute data (decoding result) using the geometric data (decoding result). That is, the control unit 221 performs the same processing as the hierarchical processing described in the first embodiment for the level to be decoded, so that the hierarchical structure of the attribute data is associated with the hierarchical structure of the geometric data. Therefore, the attribute data (decoding result) for each point is associated with the geometric data (decoding result). The following processing is performed based on this association relationship.
[0329] In step S222, the non-scalable inverse hierarchical processing unit 222 inverse hierarchizes the attribute data of the non-scalable level.
[0330] In step S223, the control unit 221 determines whether to decode the levels up to the scalable level. When it is determined to decode the levels up to the scalable level, the processing proceeds to step S224.
[0331] In step S224, the scalable inverse layer processing unit 223 performs inverse layer processing on the attribute data of the scalable layer through scalable inverse layer processing. Details of this scalable inverse layer processing will be described later.
[0332] When the processing of step S224 ends, the inverse layer processing ends, and the processing returns to Figure 21 .
[0333] In addition, when it is determined in step S223 that the layers up to the scalable layer are not to be decoded, the processing of step S224 is omitted, the inverse layer processing ends, and the processing returns to Figure 21 .
[0334] By performing the processing of each step in this way, the inverse layer processing unit 213 can perform inverse layer processing using multiple layer methods, so that the attribute data layered using multiple layer methods can be correctly inverse-layered. That is, the encoded data that has been encoded using multiple layer methods can be correctly decoded. Therefore, a reduction in encoding efficiency can be suppressed.
[0335] <Flow of Scalable Inverse Layer Processing>
[0336] Next, an example of the flow of the scalable inverse layer processing performed in step S224 of Figure 23 will be described with reference to the flowchart of Figure 22 .
[0337] When starting the scalable inverse layer processing, in step S231, the scalable inverse layer processing unit 223 sets the processing target LoD to the highest LoD of the scalable layer.
[0338] In step S232, the scalable inverse layer processing unit 223 derives the predicted value of the predicted point from the reference point based on the geometric data of the resolution of the processing target LoD.
[0339] In step S233, the scalable inverse layer processing unit 223 adds the predicted value derived in step S232 to the difference value to reconstruct the attribute information of the predicted point.
[0340] In step S234, the scalable inverse layer processing unit 223 merges the attribute data of the predicted point and the reference point.
[0341] In step S235, the scalable inverse layer processing unit 223 determines whether the processing target LoD is the lowest LoD of the decoding target layer range. When it is determined that the processing target LoD is not the lowest LoD of the decoding target layer range, the processing proceeds to step S236.
[0342] In step S236, the scalable inverse layer processing unit 223 moves (updates) the processing target LoD to the next level. When the processing of step S236 ends, the processing returns to S232 and the subsequent processing is repeated.
[0343] That is, the processing of steps S232 to S236 is performed for all levels that are the processing targets.
[0344] Then, when it is determined in step S235 that the processing target LoD has reached the lowest LoD within the decoding target level range, that is, when it is determined that all levels that are the decoding targets have been processed, the scalable inverse layer processing ends, and the processing returns to Figure 22 .
[0345] By performing the processing of each step in this way, the scalable inverse layer processing unit 223 can correctly perform inverse layer processing on the scalable levels that are the decoding targets. That is, the encoded data that has been encoded using multiple layer methods can be correctly decoded. Therefore, a reduction in encoding efficiency can be suppressed.
[0346] <5. Third Embodiment>
[0347] <Encoding Device>
[0348] Next, a device applying the present technology described above in <2. Control of Reference Relationships> will be described. In this case, the encoding device 100 has the same configuration as that in the first embodiment described with reference to Figure 11 and performs substantially the same processing. However, the attribute information encoding unit 104 encodes the attribute data using the present technology described above in <2. Control of Reference Relationships>. For example, the attribute information encoding unit 104 constructs a reference relationship by merging multiple levels (nodes).
[0349] <Attribute Information Encoding Unit>
[0350] In addition, even in this case, the attribute information encoding unit 104 has the same configuration as that in the first embodiment described with reference to Figure 12 and performs substantially the same processing. However, the layer processing unit 111 performs layer processing using the present technology described above in <2. Control of Reference Relationships>. For example, the layer processing unit 111 merges points (nodes) of multiple levels to construct a reference relationship.
[0351] <Layer Processing Unit>
[0352] Furthermore, in this case, the layer processing unit 111 also has the same configuration as that in the first embodiment described with reference to Figure 13 The configurations are the same as those in the first embodiment described and perform substantially the same processing. However, in this case, the hierarchical processing unit 111 may include at least one of a scalable hierarchical processing unit 122 and a non-scalable hierarchical processing unit 123. That is, the layering can be performed according to a method applicable to scalable decoding or a method not applicable to scalable decoding.
[0353] In addition, the control unit 121 controls the layering by using the method of the present technology described above in <2. Control of reference relationships>. For example, the control unit 121 merges points (nodes) of multiple levels to construct a reference relationship. For example, the control unit 121 derives the predicted values of the predicted points of multiple levels by referring to the reference points of the same level. For example, the control unit 121 compares the number of points (number of nodes) of the level that is the processing target as attribute data with the number of points (number of nodes) of the level one level higher. Then, if the number of points of the level that is the processing target is not sufficiently greater than the number of points of the level one level higher (if there is no increase in a predetermined number of points (for example, the number of points decreases)), the control unit 121 merges the points of the level that is the processing target into the level one level higher.
[0354] Furthermore, the control unit 121 generates control information (for example, merge_lower_lod_flag) indicating whether a reference relationship has been constructed for each level by merging with other levels, and provides the control information to the quantization unit 112 so that the control information is sent to the decoding side. When the encoding unit 113 obtains the control information through the quantization unit 112, the encoding unit 113 includes the control information in the encoded data of the attribute data.
[0355] By performing the layering as described above, the hierarchical processing unit 111 can construct a reference relationship with a configuration in which the number of points monotonically increases from a higher level to a lower level. Therefore, the encoding device 100 can suppress a decrease in encoding efficiency.
[0356] <Flow of encoding processing>
[0357] Next, the processing performed by the encoding device 100 in this case will be described. In this case, the encoding device 100 also performs the encoding processing through a process substantially the same as the process in the first embodiment described with reference to Figure 14 However, in step S104, the attribute information encoding unit 104 performs the processing by using the present technology described above in <2. Control of reference relationships>. For example, the attribute information encoding unit 104 merges points (nodes) of multiple levels to construct a reference relationship.
[0358] <Flow of attribute information encoding processing>
[0359] In addition, in this case, the attribute information encoding unit 104 performs the attribute information encoding process executed in step S104 in Figure 15 by a process substantially the same as the process in the first embodiment described with reference to Figure 14 . However, in step S111, the hierarchical processing unit 111 performs processing using the present technology described above in <2. Control of reference relationships>.
[0360] <Process flow of hierarchical processing>
[0361] A flowchart with reference to Figure 24 will describe an example of the process flow of hierarchical processing in this case.
[0362] When the hierarchical processing starts, in step S301, the control unit 121 compares the number of nodes in each level with the number of nodes in the level one level higher.
[0363] In step S302, the control unit 121 includes (merges) the nodes in the level where the number of nodes is not sufficiently greater than the number of nodes in the level one level higher (for example, the number of nodes decreases) into the nodes in the level one level higher.
[0364] In step S303, the control unit 121 controls the scalable hierarchical processing unit 122 or the non-scalable hierarchical processing unit 123 to set reference points for each level from the level one level higher. That is, a reference relationship between levels is constructed.
[0365] The processing in steps S304 to S306 is performed in the same manner as the processing in steps S124 to S126 ( Figure 16 ).
[0366] When the processing in step S306 ends, the hierarchical processing ends and the processing returns to Figure 15 .
[0367] By performing the processing as described above, the hierarchical processing unit 111 can construct a reference relationship with a configuration in which the number of points monotonically increases from a higher level to a lower level. Therefore, the encoding device 100 can suppress a decrease in encoding efficiency.
[0368] <6. Fourth embodiment>
[0369] <Decoding device>
[0370] Next, another example of a device applying the present technology described above in <2. Control of reference relationships> will be described. In this case, the decoding device 200 has the same as that with reference to Figure 17 The configurations in the second embodiment described are the same and perform substantially the same processing. However, the attribute information decoding unit 204 decodes the attribute data using the present technology described above in <2. Control of reference relationships>. For example, when performing inverse layering, the attribute information decoding unit 204 reconstructs prediction points at multiple levels by referring to reference points at the same level based on the construction of a reference relationship according to the merging of points (nodes) at multiple levels.
[0371] <Attribute information decoding unit>
[0372] Furthermore, in this case, the attribute information decoding unit 204 also has the same configuration as that in the second embodiment described with reference to Figure 18 and performs substantially the same processing. However, the inverse layering processing unit 213 performs layering using the present technology described above in <2. Control of reference relationships>. For example, when performing inverse layering, the inverse layering processing unit 213 reconstructs prediction points at multiple levels by referring to reference points at the same level based on the construction of a reference relationship according to the merging of points (nodes) at multiple levels.
[0373] <Inverse layering processing unit>
[0374] Furthermore, in this case, the inverse layering processing unit 213 also has the same configuration as that in the second embodiment described with reference to Figure 19 and performs substantially the same processing. However, in this case, the inverse layering processing unit 213 may include at least one of a non-scalable inverse layering processing unit 222 and a scalable inverse layering processing unit 223. That is, inverse layering can be performed according to a method applicable to scalable decoding or a method not applicable to scalable decoding.
[0375] Furthermore, the control unit 221 controls the inverse layering by using the method of the present technology described above in <2. Control of reference relationships>. For example, the control unit 221 reconstructs prediction points at multiple levels by referring to reference points at the same level based on the construction of a reference relationship according to the merging of points (nodes) at multiple levels. For example, the control unit 221 appropriately merges points at the level to be processed into the level one higher based on control information (e.g., merge_lower_lod_flag) indicating whether a reference relationship has been constructed according to the merging with other levels to construct a reference relationship.
[0376] By performing inverse layering as described above, the inverse layering processing unit 213 can correctly perform inverse layering on attribute data in which a reference relationship with a configuration where the number of points monotonically increases from a higher level to a lower level is constructed. Therefore, the decoding device 200 can suppress a decrease in coding efficiency.
[0377] <Flow of decoding processing>
[0378] Next, the processing performed by the decoding device 200 in this case will be described. In this case, the decoding device 200 also performs decoding processing through a process substantially the same as the process in the second embodiment described with reference to Figure 20 However, the attribute information decoding unit 204 performs processing using the present technology described above in <2. Control of reference relationships> in step S206. For example, the attribute information decoding unit 204 reconstructs predicted points at multiple levels by referring to reference points at the same level based on the construction of a reference relationship according to the merging of points (nodes) at multiple levels.
[0379] <Flow of attribute information decoding processing>
[0380] In addition, in this case, the attribute information decoding unit 204 performs the attribute information decoding processing executed in step S206 through a process substantially the same as the process in the second embodiment described with reference to Figure 21 However, the inverse hierarchical processing unit 213 performs processing using the present technology described above in <2. Control of reference relationships> in step S213. Figure 20
[0381] <Flow of inverse hierarchical processing>
[0382] An example of the flow of the inverse hierarchical processing in this case will be described with reference to the flowchart of Figure 25
[0383] When the inverse hierarchical processing starts, in step S401, the control unit 221 hierarchically divides the attribute data using the geometric data.
[0384] In step S402, the control unit 221 includes nodes at a level where the number of nodes is not sufficiently larger than the number of nodes at the level one level higher in the nodes at the level one level higher based on control information (e.g., merge_lower_lod_flag) indicating whether a reference relationship has been constructed according to the merging with other levels.
[0385] In step S403, the non-scalable inverse hierarchical processing unit 222 (or the scalable inverse hierarchical processing unit 223) derives the predicted value of the predicted point for each level using the reference point at the level one level higher.
[0386] In step S404, the non-scalable inverse hierarchical processing unit 222 (or the scalable inverse hierarchical processing unit 223) adds the predicted value to the difference for each level to reconstruct the attribute information of the predicted point.
[0387] In step S405, the non-scalable inverse layer processing unit 222 (or the scalable inverse layer processing unit 223) merges the attribute data of the prediction points and the reference points for each layer.
[0388] When the processing of step S405 ends, the inverse layer processing ends and the processing returns to Figure 21 .
[0389] By performing the processing described above, the inverse layer processing unit 213 can correctly inverse-layer the attribute data in which the reference relationship with a configuration where the number of points monotonically increases from a higher layer to a lower layer is constructed. Therefore, the decoding device 200 can suppress a decrease in coding efficiency.
[0390] <7. Supplementary>
[0391] <Layer / Inverse Layer Method>
[0392] Although Lifting has been illustrated above as the attribute information layer / inverse layer method, the present technology can be applied to any technology for layer-ing attribute information. That is, the attribute information layer / inverse layer method can be a method other than Lifting.
[0393] <Control Information>
[0394] Regarding the control information of the present technology described in each embodiment, it can be sent from the encoding side to the decoding side. For example, control information (e.g., enabled_flag) for controlling whether to allow (or prohibit) the application of the present technology described above can be sent. In addition, for example, control information for specifying the range in which the application of the present technology described above is allowed (or prohibited) (e.g., the upper limit or lower limit of the block size or both, slice, picture, sequence, component, view, layer, etc.) can be sent.
[0395] <Surrounding / Nearby>
[0396] At the same time, positional relationships such as "surrounding" and "nearby" in this specification include not only spatial positional relationships but also temporal positional relationships.
[0397] <Computer>
[0398] The above series of processes can be executed by hardware or software. When the series of processes are executed by software, the program configuring the software is installed on a computer. Here, the computer includes, for example, a computer built in dedicated hardware and a general-purpose personal computer installed with various programs to be able to execute various functions.
[0399] Figure 26 is a block diagram showing an example of the hardware configuration of a computer that executes the above series of processes according to a program.
[0400] In Figure 26 the computer 900 shown in Figure 26 , a central processing unit (CPU) 901, a read-only memory (ROM) 902, and a random access memory (RAM) 903 are connected to each other via a bus 904.
[0401] An input / output interface 910 is also connected to the bus 904. An input unit 911, an output unit 912, a storage unit 913, a communication unit 914, and a drive 915 are connected to the input / output interface 910.
[0402] The input unit 911 includes, for example, a keyboard, a mouse, a microphone, a touchpad, an input terminal, etc. The output unit 912 includes, for example, a display, a speaker, an output terminal, etc. The storage unit 913 includes, for example, a hard disk, a RAM disk, a non-volatile memory, etc. The communication unit 914 includes, for example, a network interface. The drive 915 drives a removable medium 921, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0403] In the computer configured as described above, for example, the CPU 901 executes the above-described series of processes by loading a program stored in the storage unit 913 into the RAM 903 via the input / output interface 910 and the bus 904 and executing the program. Data and the like required for the CPU 901 to execute various types of processes are also appropriately stored in the RAM 903.
[0404] A program executed by the computer can be recorded on, for example, a removable medium 921 such as a packaged medium for application. In this case, the program can be installed in the storage unit 913 via the input / output interface 910 by inserting the removable medium 921 into the drive 915.
[0405] In addition, the program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting. In such a case, the program can be received by the communication unit 914 and can be installed in the storage unit 913
[0406] In addition, the program can be pre-installed in the ROM 902 or the storage unit 913.
[0407] <Application target of the prior art>
[0408] Although the application of the present technology to the encoding / decoding of point cloud data has been described above, the present technology is not limited to such examples and can be applied to the encoding / decoding of any standard 3D data. That is, various types of processing such as encoding / decoding methods and the specifications of various types of data such as 3D data and metadata can be arbitrary as long as they do not conflict with the present technology described above. In addition, some of the above-mentioned processing and specifications can be omitted as long as they do not conflict with the present technology.
[0409] In addition, although the encoding device 100 and the decoding device 200 have been described above as examples of applying the present technology, the present technology can be applied to any configuration.
[0410] For example, the present technology can be applied to various electronic devices such as transmitters and receivers (e.g., TV receivers or cellular phones) in satellite broadcasting, cable broadcasting such as cable TV, distribution on the Internet, distribution to terminals according to cellular communication, etc.; or devices for recording images on media such as optical discs, magnetic disks, and flash memories or reproducing images from these storage media (e.g., hard disk recorders or cameras).
[0411] In addition, for example, the present technology can also be implemented as a part of components of a device, such as a processor (e.g., a video processor) used as a system large-scale integration (LSI) circuit, etc., a module (e.g., a video module) using multiple processors, etc., a unit (e.g., a video unit) using multiple modules, etc., or a collection of units having other additional functions (e.g., a video collection).
[0412] In addition, the present technology can also be applied to, for example, a network system formed by multiple devices. For example, the present technology can be implemented as cloud computing in which functions are shared and jointly processed by multiple devices via a network. For example, the present technology can also be implemented in a cloud service that provides services related to images (moving images) to any terminal such as a computer, an audio-visual (AV) device, a portable information processing terminal, an Internet of Things (IoT) device, etc.
[0413] Meanwhile, in this specification, a system means a collection of multiple components (devices, modules (parts), etc.), and it does not matter whether all components are arranged in a single housing. Therefore, multiple devices accommodated in different housings and connected via a network and a single device in which multiple modules are accommodated in one housing are both systems.
[0414] <Fields / Applications to which the Present Technology can be Applied>
[0415] Systems, devices, processing units, etc. to which the present technology is applied can be used in, for example, any fields such as transportation, medical care, crime prevention, agriculture, animal husbandry, mining, beauty, factories, household appliances, weather and natural monitoring. In addition, its applications are arbitrary.
[0416] <Others>
[0417] Meanwhile, in this specification, a "flag" is information for identifying multiple states, and includes not only information for identifying two states of true (1) and false (0), but also information capable of identifying three or more states. Therefore, a "flag" can have not only two values of 1 / 0, but also three or more values. That is, the number of bits constituting the "flag" is arbitrary and can be 1 bit or multiple bits. Additionally, it is assumed that identification information (including flags) not only has a form in which the identification information is included in the bit stream, but also has a form in which the difference information of the identification information with respect to a specific reference information is included in the bit stream. Therefore, in this specification, "flags" and "identification information" include not only the information itself, but also the difference information with respect to the reference information.
[0418] In addition, various types of information (such as metadata) regarding the encoded data (bit stream) can be sent or recorded in any form as long as they are associated with the encoded data. Here, the term "associated" means, for example, making other information available (linkable) when processing one piece of information. That is, the associated information can be collected as one piece of data or can be individual information. For example, information associated with the encoded data (image) can be transmitted on a transmission path different from the transmission path for the encoded data (image). Additionally, for example, information associated with the encoded data (image) can be recorded on a recording medium different from the recording medium for the encoded data (image) (or another recording area of the same recording medium). Meanwhile, such "association" can be for partial data rather than all data. For example, an image and the information corresponding to the image can be associated with each other in any unit such as multiple frames, one frame, or a part within a frame.
[0419] Meanwhile, in this specification, terms such as "constitute", "multiplex", "add", "integrate", "include", "store", "put", "introduce", "insert", etc. mean combining multiple things into one, for example, combining encoded data and metadata into one piece of data, and refer to the above-mentioned method of "association".
[0420] Furthermore, the embodiments of the present technology are not limited to the above embodiments, and various modifications can be made without departing from the gist of the present technology.
[0421] For example, a configuration described as a single device (or a processing unit) can be divided into a configuration of multiple devices (or processing units). Conversely, a configuration described as multiple devices (or processing units) in the above description can be configured together as a single device (or a processing unit). In addition, it goes without saying that components other than the above components can be added to the configuration of each device (or each processing unit). Moreover, as long as the configuration or operation of the entire system is substantially the same, a part of the configuration of one device (or one processing unit) can be included in the configuration of another device (or another processing unit).
[0422] In addition, for example, the above program can be executed in any device. In this case, the device can have the necessary functions (function blocks, etc.) capable of obtaining the necessary information.
[0423] In addition, for example, each step of the flowchart can be executed by one device or executed in a shared manner by multiple devices. In addition, in the case where a single step includes multiple processes, the multiple processes can be executed by one device or can be executed in a shared manner by multiple devices. In other words, the multiple processes included in a single step can be executed as processes of multiple steps. Conversely, the processes described as multiple steps can be executed together as a single step.
[0424] In addition, for example, a program executed by a computer can be executed such that the processing steps described in the program are executed chronologically in the order described in this specification, or executed in parallel or executed separately at a necessary timing, for example, in response to a call. That is, as long as there is no contradiction, each processing step can be executed in an order different from the above order. Moreover, the various processes of the steps described in the program can be executed in parallel with the processes of another program, or can be executed in combination with the processes of another program.
[0425] In addition, for example, as long as there is no contradiction, various technologies regarding the present technology can be implemented independently and separately. It goes without saying that any number of the present technologies can also be implemented in parallel. For example, part or all of the present technology described in any one embodiment can be implemented in combination with part or all of the present technology described in other embodiments. In addition, part or all of the above-described present technology can be implemented in combination with other technologies not described above.
[0426] Meanwhile, the present technology can also adopt the following configuration.
[0427] (1) An information processing device, comprising:
[0428] A hierarchical unit configured to perform hierarchical processing of the attribute information for each point in a point cloud representing an object of a three-dimensional shape as a set of points by recursively repeating the following processing for a reference point: classifying the points into prediction points or reference points, deriving a predicted value of the attribute information of the prediction points using the attribute information of the reference points, and deriving a difference between the attribute information of the prediction points and the predicted value, where
[0429] The hierarchical unit performs hierarchical processing using a plurality of hierarchical methods to generate levels different according to the hierarchical methods.
[0430] (2) The information processing apparatus according to (1), wherein the hierarchical unit generates levels higher than a predetermined level according to a hierarchical method not applicable to scalable decoding, and generates levels lower than the predetermined level according to a hierarchical method applicable to scalable decoding.
[0431] (3) The information processing apparatus according to (2), wherein the hierarchical unit generates each level in ascending order from the lowest level to the highest level.
[0432] (4) The information processing apparatus according to any one of (1) to (3), further comprising:
[0433] A generation unit configured to generate control information regarding the hierarchical processing of the attribute information performed by the hierarchical unit; and
[0434] An encoding unit configured to encode the attribute information hierarchically processed by the hierarchical unit and generate encoded data of the attribute information including the control information generated by the generation unit.
[0435] (5) An information processing method for: using a plurality of hierarchical methods to generate levels different according to the plurality of hierarchical methods when performing hierarchical processing of the attribute information for each point in a point cloud representing an object of a three-dimensional shape as a set of points by recursively repeating the following processing for a reference point: classifying the points into prediction points or reference points, deriving a predicted value of the attribute information of the prediction points using the attribute information of the reference points, and deriving a difference between the attribute information of the prediction points and the predicted value.
[0436] (6) An information processing apparatus including:
[0437] An inverse stratification unit configured to perform inverse stratification on attribute information that has been stratified by recursively repeating the following process for each point in a point cloud representing an object of a three-dimensional shape as a set of points: classifying the points as prediction points or reference points, deriving a predicted value of the attribute information of the prediction points using the attribute information of the reference points, and deriving the difference between the attribute information of the prediction points and the predicted value, where
[0438] The inverse stratification unit performs inverse stratification on different levels according to the stratification methods using a plurality of stratification methods.
[0439] (7) The information processing apparatus according to (6), wherein the inverse stratification unit performs inverse stratification on levels higher than a predetermined level according to a stratification method not applicable to scalable decoding, and performs inverse stratification on levels lower than the predetermined level according to a stratification method applicable to scalable decoding.
[0440] (8) The information processing apparatus according to (7), wherein the inverse stratification unit identifies the predetermined level based on first control information indicating a level range to which a stratification method applicable to scalable decoding is applied.
[0441] (9) The information processing apparatus according to (8), wherein the inverse stratification unit identifies the predetermined level based on the first control information for attribute information for which second control information indicating whether a stratification method applicable to scalable decoding can be applied indicates that the stratification method applicable to scalable decoding can be applied.
[0442] (10) The information processing apparatus according to any one of (6) to (9) further includes:
[0443] A decoding unit configured to decode encoded data of the attribute information to generate the attribute information, where
[0444] The inverse stratification unit performs inverse stratification on the attribute information obtained by decoding the encoded data by the decoding unit.
[0445] (11) An information processing method for, for attribute information of each point in a point cloud representing an object of a three-dimensional shape as a set of points, when performing inverse stratification on attribute information that has been stratified by recursively repeating the following process for a reference point, performing inverse stratification on different levels according to a plurality of stratification methods: classifying the points as prediction points or reference points, deriving a predicted value of the attribute information of the prediction points using the attribute information of the reference points, and deriving the difference between the attribute information of the prediction points and the predicted value.
[0446] (12) An information processing apparatus including:
[0447] A hierarchical unit configured to perform hierarchical processing of attribute information for each point in a point cloud representing a three-dimensional shape object as a set of points by recursively repeating the following processing for the reference points: classifying a point as a predicted point or a reference point, deriving a predicted value of the attribute information of the predicted point using the attribute information of the reference point, and deriving a difference between the attribute information of the predicted point and the predicted value, where,
[0448] the hierarchical unit refers to reference points at the same level and derives predicted values of predicted points at multiple levels.
[0449] (13) The information processing apparatus according to (12), wherein the hierarchical unit refers to reference points at a level two levels higher to calculate predicted values of predicted points at a level where the number of points is not sufficiently larger than the number of points at a level one level higher.
[0450] (14) The information processing apparatus according to (12) or (13), further comprising:
[0451] a generation unit configured to generate control information indicating whether a reference relationship has been constructed by merging with other levels; and
[0452] an encoding unit configured to encode the attribute information stratified by the hierarchical unit and generate encoded data of the attribute information including the control information generated by the generation unit.
[0453] (15) An information processing method for:
[0454] For the attribute information of each point in a point cloud representing a three-dimensional shape object as a set of points, when performing hierarchical processing of the attribute information by recursively repeating the following processing for the reference points, referring to reference points at the same level to derive predicted values of predicted points at multiple levels: classifying a point as a predicted point or a reference point, deriving a predicted value of the attribute information of the predicted point using the attribute information of the reference point, and deriving a difference between the attribute information of the predicted point and the predicted value.
[0455] (16) An information processing apparatus comprising:
[0456] An inverse hierarchical unit configured to perform inverse hierarchical processing of the attribute information stratified by recursively repeating the following processing for each point in a point cloud representing a three-dimensional shape object as a set of points: classifying a point as a predicted point or a reference point, deriving a predicted value of the attribute information of the predicted point using the attribute information of the reference point, and deriving a difference between the attribute information of the predicted point and the predicted value, where,
[0457] When performing inverse stratification, the inverse stratification unit derives predicted values of the attribute information of the predicted points of multiple levels with reference to the attribute information of the reference points of the same level, and generates the attribute information of the predicted points by adding the derived predicted values to the difference value.
[0458] (17) The information processing apparatus according to (16), wherein, when performing inverse stratification, the inverse stratification unit identifies the level that is the reference destination of the processing target level based on control information indicating whether to refer to the same level as other levels, and derives the predicted value of the attribute information of the predicted points of the processing target level with reference to the attribute information of the reference points of the identified level.
[0459] (18) The information processing apparatus according to (16) or (17), wherein the inverse stratification unit refers to the attribute information of the reference points of the level two levels higher to derive the predicted value of the attribute information of the predicted points for a level whose number of points is not sufficiently larger than the number of points of the level one level higher.
[0460] (19) The information processing apparatus according to any one of (16) to (18) further includes:
[0461] A decoding unit configured to decode the encoded data of the attribute information to generate the attribute information, wherein
[0462] The inverse stratification unit performs inverse stratification on the attribute information reconstructed by decoding the encoded data by the decoding unit.
[0463] (20) An information processing method for:
[0464] For the attribute information of each point in a point cloud representing an object of a three-dimensional shape as a set of points, when performing inverse stratification on the attribute information stratified by recursively repeating the process of classifying points into predicted points or reference points, deriving the predicted value of the attribute information of the predicted points using the attribute information of the reference points, and deriving the difference between the attribute information of the predicted points and the predicted value,
[0465] Derive the predicted values of the attribute information of the predicted points of multiple levels with reference to the attribute information of the reference points of the same level, and
[0466] Generate the attribute information of the predicted points by adding the derived predicted values to the difference value.
[0467] [List of Reference Signs]
[0468] 100 Encoding device
[0469] 101 Location Information Encoding Unit
[0470] 102 Location Information Decoding Unit
[0471] 103 Point Cloud Generation Unit
[0472] 104 Attribute Information Encoding Unit
[0473] 105 Bitstream Generation Unit
[0474] 111 Hierarchical Processing Unit
[0475] 112 Quantization Unit
[0476] 113 Encoding Unit
[0477] 121 Control Unit
[0478] 122 Scalable Hierarchical Processing Unit
[0479] 123 Non-Scalable Hierarchical Processing Unit
[0480] 124 Inversion Unit
[0481] 125 Weighting Unit
[0482] 200 Decoding Device
[0483] 201 Decoding Target LoD Depth Setting Unit
[0484] 202 Encoded Data Extraction Unit
[0485] 203 Location Information Decoding Unit
[0486] 204 Attribute Information Decoding Unit
[0487] 205 Point Cloud Generation Unit
[0488] 211 Decoding Unit
[0489] 212 Inverse Quantization Unit
[0490] 213 Inverse Hierarchical Processing Unit
[0491] 221 Control Unit
[0492] 222 Non-Scalable Inverse Hierarchical Processing Unit
[0493] 223 Scalable Inverse Hierarchical Processing Unit< / lifting>
Claims
1. A decoding device for representing a point cloud of a three-dimensional object, the decoding device comprising: A circuit configured to: Obtain a scalability enable flag for decoding an intermediate level of a hierarchical structure of geometric data; Based on the scalability enable flag, obtain hierarchical points of geometric data up to the intermediate level without decoding the highest resolution level of the hierarchical structure, wherein the number of nodes at the highest resolution level is less than the number of nodes at the intermediate level; Classify each of the hierarchical points into a prediction point or a reference point; and Based on the hierarchical points, perform inverse hierarchical processing on attribute data using recursive processing to make it correspond to the geometric data up to the intermediate level, the recursive processing including: Based on the hierarchical points and the reference points, derive a predicted value of the attribute data of the prediction point; For the reference points, derive a difference between the predicted value and the attribute data of the prediction point; and Add the predicted value and the difference.
2. The decoding device according to claim 1, wherein, The intermediate level includes a first intermediate level and a second intermediate level, The hierarchical structure sequentially includes the highest resolution level, the first intermediate level, and the second intermediate level, and The circuit is further configured to, for decoding the attribute data, make the highest resolution level reference the second intermediate level so that the number of nodes in the hierarchical structure increases monotonically from the high level to the low level.
3. A decoding method for representing a point cloud of a three-dimensional object, the decoding method comprising: Obtain a scalability enable flag for decoding an intermediate level of a hierarchical structure of geometric data; Based on the scalability enable flag, obtain hierarchical points of geometric data up to the intermediate level without decoding the highest resolution level of the hierarchical structure, wherein the number of nodes at the highest resolution level is less than the number of nodes at the intermediate level; Classify each of the hierarchical points into a prediction point or a reference point; and Based on the hierarchical points, perform inverse hierarchical processing on attribute data using recursive processing to make it correspond to the geometric data up to the intermediate level, the recursive processing including: Based on the hierarchical points and the reference points, derive a predicted value of the attribute data of the prediction point; For the reference points, derive a difference between the predicted value and the attribute data of the prediction point; and Add the predicted value and the difference.
4. The decoding method according to claim 3, wherein, The intermediate level includes a first intermediate level and a second intermediate level, The hierarchical structure sequentially includes the highest resolution level, the first intermediate level, and the second intermediate level, and The decoding method further includes: for decoding the attribute data, make the highest resolution level directly reference the second intermediate level so that the number of nodes in the hierarchical structure increases monotonically from the high level to the low level.