Information processing device and method
By classifying the points in the point cloud into prediction points and reference points and performing layering on the attribute information, the load problem caused by high-resolution decoding of position information in the existing technology is solved, and scalable and correct decoding of attribute information is achieved.
Patent Information
- Application Number
- CN202080020600.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-01
- Filing Date
- 2020-03-05
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2040-03-05
AI Technical Summary
When encoding attribute information of three-dimensional point cloud data, the existing technology needs to decode the position information at an unnecessary high resolution, resulting in an increased decoding processing load.
By classifying the points of the point cloud into prediction points and reference points, the attribute information is layered, the prediction points are selected so that they exist in the voxels of the layer one level higher than the voxels of the current layer, and the resolution of the attribute information and position information of the reference points is used to derive the prediction value during the layering.
Scalable decoding of attribute information is achieved, which reduces the decoding processing load and ensures that attribute information can be correctly decoded at intermediate resolutions.
Smart Images

Figure CN113557552B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an information processing apparatus and method, and more particularly to an information processing apparatus and method that enable easier scalable decoding of attribute information. Background Art
[0002] For example, methods for encoding 3D data representing three-dimensional structures, such as point clouds, have been developed (see, for example, Non-Patent Document 1). Point cloud data includes both position information and attribute information for each point. Therefore, point cloud encoding is performed on both position information and attribute information. As a method for encoding attribute information, for example, a technique called "lifting" has been proposed (see, for example, Non-Patent Document 2).
[0003] At the same time, scalable decoding of this type of encoded data for point clouds has been considered. For example, when representing a wide three-dimensional space using a point cloud, the amount of encoded data is very large. Decoding all of this data would unnecessarily increase the load. Therefore, scalable decoding is required to obtain the necessary data for the object at the required resolution.
[0004] Reference List
[0005] Non-patent literature
[0006] Non-patent document 1: R.Mekuria, Student Member IEEE, and K.Blom and P.Cesar, Members IEEE, "Design, Implementation and Evaluation of a Point Cloud Codec for Tele-Immersive Video",tcsvt_paper_submitted_february.pdf
[0007] Non-patent document 2: Khaled Mammou, Alexis Tourapis, Jungsun Kim, FabriceRobinet, Valery Valentin, and Yeping Su, "Lifting Scheme for Lossy AttributeEncoding in TMC1", ISO / IEC JTC1 / SC29 / WG11 MPEG2018 / m42640, April 2018, San Diego, US Summary of the Invention
[0008] Problems to be solved by the present invention
[0009] However, in the case of conventional encoding methods, the highest-resolution position information is required when decoding the encoded data of the attribute information. Therefore, the position information needs to be decoded at an unnecessary high resolution, which may increase the load of the decoding process.
[0010] The present disclosure has been made in view of such circumstances, and will enable easier scalable decoding of attribute information.
[0011] Solution to the problem
[0012] An information processing device according to one aspect of the present technology is an information processing device including a hierarchical unit that hierarchizes attribute information of a point cloud by recursively repeating a process of classifying each point of a point cloud representing a three-dimensional object into a predicted point and a reference point, the predicted point being a point where a difference between the attribute information and the predicted value is left, and the reference point being a point whose attribute information was referenced during derivation of the predicted value. During the hierarchical process, the hierarchical unit selects the predicted point so that the point also exists in a voxel of a layer one level higher than the layer to which the voxel containing the point of the current layer belongs.
[0013] An information processing method according to one aspect of the present technology is an information processing method including: performing layering on attribute information of a point cloud representing a three-dimensional object by recursively repeating a process of classifying points of a point cloud representing a three-dimensional object into predicted points and reference points for reference points, wherein the predicted points are points that leave a difference between the attribute information and the predicted value, and the reference points are points to which the attribute information is referenced during the derivation of the predicted value; and during the layering, selecting the predicted points in such a manner that the points also exist in voxels of a layer one level higher than the layer to which the voxels containing the points of the current layer belong.
[0014] An information processing device according to another aspect of the present technology is an information processing device including an inverse hierarchical unit that performs inverse hierarchical processing on attribute information of a point cloud representing a three-dimensional object, the attribute information having been hierarchically layered by recursively repeating a process of classifying points of the point cloud into predicted points and reference points, the predicted points being points at which a difference between the attribute information and a predicted value is left, the reference points being points to which the attribute information was referenced during derivation of the predicted value. During the inverse hierarchical processing, the inverse hierarchical unit derives, in each level, a predicted value of the attribute information of the predicted point using the attribute information of the reference point and position information of a resolution of a level that is not the lowest level of the point cloud, and derives the attribute information of the predicted point using the predicted value and the difference.
[0015] An information processing method according to another aspect of the present technology is an information processing method including: performing inverse layering on attribute information of a point cloud representing a three-dimensional object, the attribute information having been layered by recursively repeating a process of classifying points of the point cloud into predicted points and reference points on reference points, wherein the predicted points are points that leave a difference between the attribute information and the predicted value, and the reference points are points to which the attribute information is referenced during the derivation of the predicted value; and during inverse layering, in each level, the predicted value of the attribute information of the predicted point is derived using the attribute information of the reference point and position information of the resolution of a level that is not the lowest level of the point cloud, and the attribute information of the predicted point is derived using the predicted value and the difference.
[0016] According to another aspect of the present technology, an information processing device is an information processing device including: a generation unit that generates information indicating a hierarchical structure of position information of a point cloud representing a three-dimensional object; and a hierarchical unit that hierarchizes attribute information of the point cloud based on the information generated by the generation unit so as to associate the attribute information with the hierarchical structure of the position information.
[0017] According to another aspect of the present technology, an information processing device is an information processing device including: a generation unit that generates information indicating a hierarchical structure of position information of a point cloud representing a three-dimensional object; and a de-hierarchical unit that de-hierarchically performs attribute information of the point cloud based on the information generated by the generation unit so as to associate the hierarchical structure of the attribute information with the hierarchical structure of the position information.
[0018] In an information processing device and method according to one aspect of the present technology, attribute information of a point cloud is hierarchized by recursively repeating a process of classifying each point of a point cloud representing a three-dimensional object into predicted points (points where a difference between attribute information and predicted values is left), and reference points (points whose attribute information was referenced during the derivation of predicted values). During the hierarchization, predicted points are selected so that the points also exist in voxels of a layer one level higher than the layer to which the voxel containing the point of the current layer belongs.
[0019] In an information processing device and method according to another aspect of the present technology, with respect to attribute information of a point cloud representing a three-dimensional object, a process of classifying each point of the point cloud into predicted points and reference points is recursively repeated for reference points, so that inverse stratification is performed on attribute information that has been stratified, the predicted points being points where a difference between attribute information and predicted values is left, and the reference points being points where attribute information was referenced during derivation of the predicted values. During the inverse stratification, a predicted value of the attribute information of the predicted point is derived at each stratum using the attribute information of the reference point and position information of a resolution of a stratum that is not the lowest stratum of the point cloud, and the attribute information of the predicted point is derived using the predicted value and the difference.
[0020] In an information processing device according to another aspect of the present technology, information indicating a hierarchical structure of position information of a point cloud representing a three-dimensional object is generated, and based on the generated information, attribute information of the point cloud is hierarchized to be associated with the hierarchical structure of the position information.
[0021] In an information processing device according to another aspect of the present technology, information indicating a hierarchical structure of position information of a point cloud representing a three-dimensional object is generated, and based on the generated information, attribute information of the point cloud is inversely hierarchicalized, and the hierarchical structure of the attribute information is considered to be associated with the hierarchical structure of the position information. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 This is a diagram for explaining an example of the hierarchical structure of position information.
[0023] Figure 2 It is a diagram for explaining an example of "lifting".
[0024] Figure 3 It is a diagram for explaining an example of "lifting".
[0025] Figure 4 It is a diagram for explaining an example of "lifting".
[0026] Figure 5 is a diagram for explaining an example of quantization.
[0027] Figure 6 is a diagram for explaining an example method for hierarchizing attribute information.
[0028] Figure 7 This is a diagram for explaining an example of the hierarchy of attribute information.
[0029] Figure 8 This is a diagram for explaining an example of inverse hierarchical structure of attribute information.
[0030] Figure 9 This is a diagram for explaining an example of the hierarchy of attribute information.
[0031] Figure 10 This is a diagram for explaining an example of the hierarchy of attribute information.
[0032] Figure 11 This is a diagram for explaining an example of inverse hierarchical structure of attribute information.
[0033] Figure 12 is a diagram for explaining an example of a quantization method.
[0034] Figure 13 is a diagram for explaining an example of quantization.
[0035] Figure 14 is a block diagram showing a general example configuration of an encoding device.
[0036] Figure 15 is a block diagram showing a typical example configuration of an attribute information encoding unit.
[0037] Figure 16 is a block diagram showing a general example configuration of a hierarchical processing unit.
[0038] Figure 17 This is a flowchart for explaining an example flow in encoding processing.
[0039] Figure 18 This is a flowchart for explaining an example flow in attribute information encoding processing.
[0040] Figure 19 This is a flowchart for explaining an example flow in hierarchical processing.
[0041] Figure 20 is a flowchart for explaining an example flow in quantization processing.
[0042] Figure 21 This is a flowchart for explaining an example flow in hierarchical processing.
[0043] Figure 22 is a flowchart for explaining an example flow in quantization processing.
[0044] Figure 23 is a block diagram showing a typical example configuration of a decoding device.
[0045] Figure 24 is a block diagram showing a typical example configuration of an attribute information decoding unit.
[0046] Figure 25 is a block diagram showing a general example configuration of a reverse layering processing unit.
[0047] Figure 26 This is a flowchart for explaining an example flow in a decoding process.
[0048] Figure 27 This is a flowchart for explaining an example flow in attribute information decoding processing.
[0049] Figure 28 This is a flowchart for explaining an example flow in inverse quantization processing.
[0050] Figure 29 This is a flowchart for explaining an example flow in the de-hierarchical processing.
[0051] Figure 30 This is a flowchart for explaining an example flow in inverse quantization processing.
[0052] Figure 31 is a diagram showing an example of an octree to which position information of DCM is applied.
[0053] Figure 32 is a diagram showing an example of an octree to be referred to in the processing of attribute information.
[0054] Figure 33 is a diagram showing an example of an octree in encoding attribute information.
[0055] Figure 34 is a diagram showing an example of an octree in decoding attribute information.
[0056] Figure 35 is a diagram showing an example of an octree in encoding attribute information.
[0057] Figure 36 is a diagram showing an example of an octree in decoding attribute information.
[0058] Figure 37 is a diagram illustrating an example of derivation of quantization weights.
[0059] Figure 38 is a diagram showing an example of a hierarchical structure of position information.
[0060] Figure 39 is a diagram showing an example of labeling.
[0061] Figure 40 is a diagram showing an example of a hierarchical structure of position information.
[0062] Figure 41 is a diagram showing an example of labeling.
[0063] Figure 42 is a diagram showing an example of labeling.
[0064] Figure 43 is a block diagram showing a general example configuration of an encoding device.
[0065] Figure 44 This is a flowchart for explaining an example flow in encoding processing.
[0066] Figure 45 This is a flowchart for explaining an example flow in hierarchical processing.
[0067] Figure 46 This is a flowchart for explaining an example flow in a decoding process.
[0068] Figure 47 This is a flowchart for explaining an example flow in the de-hierarchical processing.
[0069] Figure 48 is a block diagram showing a typical example configuration of a computer. DETAILED DESCRIPTION
[0070] The following is a description of modes for implementing the present disclosure (these modes are hereinafter referred to as embodiments). Note that the description will be given in the following order.
[0071] 1. Scalable decoding
[0072] 2. First embodiment (encoding device)
[0073] 3. Second embodiment (decoding device)
[0074] 4.DCM
[0075] 5. Quantization weights
[0076] 6.LoD Generation
[0077] 7. Notes
[0078] <1. Scalable Decoding>
[0079] <Documents supporting technical content and terminology, etc.>
[0080] The scope disclosed in the present technology includes not only the contents disclosed in the embodiments but also the contents disclosed in the following non-patent documents known at the time of filing.
[0081] Non-Patent Document 1: (as mentioned above)
[0082] Non-Patent Document 2: (as mentioned above)
[0083] Non-patent document 3: TELECOMMUNICATION STANDARDIZATION SECTOR OF ITU (International Telecommunication Union), "Advanced video coding for generic audiovisual services", H.264,04 / 2017
[0084] Non-patent document 4: TELECOMMUNICATION STANDARDIZATION SECTOR OF ITU (International Telecommunication Union), "High efficiency video coding", H.265,12 / 2016
[0085] Non-Patent Document 5: Jianle Chen, Elena Alshina, Gary J. Sullivan, Jens-Rainer, and Jill Boyce, “Algorithm Description of Joint Exploration Test Model 4”, JVET-G1001_v1, Joint Video Exploration Team (JVET) of ITU-T SG 16WP 3 and ISO / IEC JTC1 / SC 29 / WG 11 7th Meeting: Torino, IT, July 13-21, 2017
[0086] Non-Patent Literature 6: Sebastien Lasserre and David Flynn, “[PCC] Inference of amode using point location direct coding in TMC3”, ISO / IEC JTC1 / SC29 / WG11MPEG2018 / m42239, January 2018, Gwangju, Korea
[0087] That is, the contents disclosed in the non-patent documents listed above are also the basis for determining the support requirements. For example, even when the quadtree block structure disclosed in non-patent document 4 and the quadtree plus binary tree (QTBT) block structure disclosed in non-patent document 4 are not directly disclosed in the embodiment, these structures are within the scope of the present technology and meet the support requirements of the claims. In addition, for example, technical terms such as parsing, grammar, and semantics are within the scope of the disclosure of the present technology and meet the support requirements of the claims even when these technical terms are not directly described.
[0088] <Point Cloud>
[0089] There are already 3D data such as point clouds and meshes. Point clouds represent three-dimensional structures and have point location information, attribute information, etc. Meshes are formed by vertices, edges, and planes and use polygonal representation to define three-dimensional shapes.
[0090] For example, in the case of a point cloud, a three-dimensional structure (three-dimensional object) is represented as a collection of a large number of points. Specifically, the point cloud data (also called point cloud data) consists of the positional information and attribute information of each point in the point cloud. For example, attribute information includes color, reflectivity, and general information. Therefore, the data structure is relatively simple, and any desired three-dimensional structure can be represented with sufficiently high accuracy using a sufficient number of points.
[0091] <Octtree>
[0092] For example, as disclosed in Non-Patent Document 1, since the amount of point cloud data is relatively large, compression through encoding is considered. For example, the position information of the point cloud is first encoded (decoded), and then the attribute information is encoded (decoded). During encoding, the position information is quantized using voxels, and then the position information is further hierarchically formed into an octree.
[0093] Voxels are small areas of a predetermined size that divide a three-dimensional space. Position information is corrected so that a point is provided for each voxel. Figure 1 As shown in A of FIG. 1 , for example, the region 10 as a two-dimensional region is divided by voxels 11-1 represented by squares, and the position information of each point 12-1 is corrected so that a point 12-1 represented by a black circle is provided for each voxel 11-1. That is, the position information is quantized according to the size of the voxel. Note that in Figure 1 In A, although only one voxel 11-1 is represented by a reference numeral, Figure 1 All squares in region 10 of A are voxels 11-1. Similarly, Figure 1 All black circles shown in A are points 12-1.
[0094] In an octree, a voxel is divided into two in each of the x, y, and z directions (meaning the voxel is divided into eight), forming voxels of the next lower level (also called LoD). In other words, two voxels aligned in each of the x, y, and z directions (eight voxels in total) are combined to form voxels of the next higher level (LoD). Position information is then quantized using voxels within each level.
[0095] For example, in Figure 1 In the case shown in A, the region is two-dimensional. Therefore, Figure 1 As shown in FIG. 1B , four voxels 11-1 arranged vertically and horizontally are integrated to form a higher-level voxel 11-2 indicated by a thick line. The position information is then quantized using the voxel 11-2. That is, when the point 12-1 ( Figure 1 When A) exists in voxel 11-2, its position information is corrected so that point 12-1 is converted into point 12-2 corresponding to voxel 11-2. Figure 1 In B, although only one voxel 11-1 is represented by a reference numeral, Figure 1 All squares indicated by dotted lines in region 10 in B are voxels 11-1. Figure 1 In B, although only one voxel 11-2 is represented by a reference numeral, Figure 1All squares indicated by thick lines in region 10 in B are voxels 11-2. Figure 1 All black circles shown in B are points 12-2.
[0096] Likewise, Figure 1 As shown in FIG. 1C , four voxels 11-2 arranged vertically and horizontally are integrated to form a higher-level voxel 11-3 indicated by a thick line. The position information is then quantized using the voxel 11-3. That is, when the point 12-2 ( Figure 1 When B) exists in voxel 11-3, its position information is corrected so that point 12-2 is converted into point 12-3 corresponding to voxel 11-3. Figure 1 In C, although only one voxel 11-2 is represented by a reference numeral, Figure 1 All squares indicated by dotted lines in region 10 in C are voxels 11-2. Figure 1 In C, although only one voxel 11-3 is represented by a reference numeral, Figure 1 All squares indicated by thick lines in region 10 in C are voxels 11-3. Figure 1 All black circles shown in C are points 12-3.
[0097] Likewise, Figure 1 As shown in FIG. 1D , the four voxels 11-3 arranged vertically and horizontally are integrated to form a higher-level voxel 11-4 indicated by a thick line. The position information is then quantized using the voxel 11-4. That is, when the point 12-3 ( Figure 1 When C) exists in voxel 11-4, its position information is corrected so that point 12-3 is converted into point 12-4 corresponding to voxel 11-4. Figure 1 In D, although only one voxel 11-3 is represented by a reference numeral, Figure 1 All squares indicated by dotted lines in region 10 in D are voxels 11 - 3 .
[0098] In this way, the position information is hierarchical according to the hierarchical structure of voxels.
[0099] <Ascension>
[0100] On the other hand, when attribute information is encoded, the positional relationship between points is used for encoding, where it is known that the positional information includes degradation caused by encoding. As a method for encoding such attribute information, a method using a region-adaptive hierarchical transform (RAHT) or a transform called "lifting" disclosed in Non-Patent Document 2 is considered. By using these techniques, attribute information can be hierarchized like an octree of position information.
[0101] For example, in the case of "lifting", the attribute information of each point is encoded as the difference from the predicted value derived by using the attribute information of other points. Then the points from which the difference (ie, predicted value) is derived are hierarchically selected.
[0102] For example, in the two-dimensional structure Figure 2 In the hierarchy shown in Figure A, points P7, P8, and P9, indicated by white circles, are selected as prediction points among the circled points (P0 to P9) and predicted values are derived at these points. The other points, P0 to P6, are selected as reference points, which are points whose attribute information is referenced during the derivation of the predicted values. The predicted value for prediction point P7 is then derived by referencing the attribute information of reference points P0 and P1. The predicted value for prediction point P8 is derived by referencing the attribute information of reference points P2 and P3. The predicted value for prediction point P9 is derived by referencing the attribute information of reference points P4 to P6. That is, in this hierarchy, the difference values for each of points P7 to P9 are obtained.
[0103] like Figure 2 As shown in B, in the next higher level, the pair is selected as Figure 2 The reference points (P0 to P6) in the hierarchy (lower level) shown in A are executed with Figure 2 The classification (ranking) between the prediction points and the reference points is the same as in the case of the hierarchy shown in A.
[0104] For example, in Figure 2 In Figure 1, points P1, P3, and P6, indicated by gray circles, are selected as prediction points, and points P0, P2, P4, and P5, indicated by black circles, are selected as reference points. The predicted value for prediction point P1 is then derived by referring to the attribute information of reference points P0 and P2. The predicted value for prediction point P3 is derived by referring to the attribute information of reference points P2 and P4. The predicted value for prediction point P6 is derived by referring to the attribute information of reference points P4 and P5. In other words, in this hierarchy, the difference between each of points P1, P3, and P6 is obtained.
[0105] like Figure 2 As shown in C, in the next higher level, the pair is selected as Figure 2 Classification (sorting) is performed on the points (P0, P2, P4, and P5) of the reference points in the hierarchy (lower one level) shown in B.
[0106] Since such classification is recursively repeated for reference points at the next lower level, the attribute information is hierarchical.
[0107] <Point Classification>
[0108] The process of classifying (sorting) points in this "boosting" process is now described in more detail. As described above, in "boosting," points are classified in order from low to high levels. In each level, the points are arranged in Morton order. Next, the point at the top of the sequence of points arranged in Morton order is selected as a reference point. Next, a search is performed for points located near the reference point (nearby points), and the detected points (nearby points) are set as predicted points (also called index points).
[0109] For example, Figure 3 As shown, a point is searched in a circle 22 with a radius R, with the processing target reference point 21 as the center. The radius R is set in advance for each level. Figure 3 In the case of the illustrated example, points 23 - 1 to 23 - 4 are detected, and the points 23 - 1 to 23 - 4 are set as predicted points.
[0110] Next, the remaining points are classified in a similar manner to the above. That is, among the points that have not been selected as reference points or prediction points at this stage, the point at the top of the Morton order is selected as a reference point, and the points near this reference point are detected and set as prediction points.
[0111] The above process is repeated until all points have been classified and the hierarchy is complete. The processing target then moves to the next higher hierarchy. The above process is then repeated for that hierarchy. That is, the points selected as reference points in the next lower hierarchy are arranged in Morton order and classified as reference points and prediction points as described above. As the above process is repeated, the attribute information is layered.
[0112] <Derivation of predicted values>
[0113] Furthermore, in the case of the above-mentioned “boosting”, the predicted value of the attribute information of the prediction point is derived by using the attribute information of the reference points around the prediction point. For example, Figure 4 As shown, the predicted value of the prediction point Q(i, j) is derived by referring to the attribute information of reference points P1 to P3. In this case, as shown in the following expression (1), the attribute information of each reference point is weighted using a quantization weight (α(P, Q(i, j))) that depends on the inverse of the distance between the prediction point and the reference point, and then integrated. Thus, the predicted value is derived. Here, A(P) represents the attribute information of point P.
[0114] [Mathematical formula 1]
[0115]
[0116] The highest resolution (which is the lowest level) position information is used when deriving the distance between the predicted point and its reference point.
[0117] <Quantification>
[0118] After being layered as described above, the attribute information is further quantized and encoded. Figure 5 As shown in the example shown, the attribute information (difference) of each point is weighted according to the hierarchical structure. Figure 5 As shown, for each point, the quantization weight W is derived by using the quantization value of the lower level. Note that the quantization weight can also be used for "boosting" (layering of attribute information) to improve compression efficiency.
[0119] <hierarchy mismatch>
[0120] As described above, the method for hierarchizing attribute information differs from the method used in the case of the octree of position information. Therefore, there is no guarantee that the hierarchical structure of attribute information matches the hierarchical structure of position information. Therefore, there is a possibility that at intermediate resolutions (levels higher than the lowest level), points in attribute information do not match points in position information, and even using position information at intermediate resolutions, attribute information cannot be correctly predicted (i.e., inverse hierarchization of attribute information is difficult).
[0121] In other words, in order to perform correct decoding (reverse layering) on attribute information, it is necessary to decode the position information down to the lowest layer regardless of the current layer. Therefore, the load of the decoding process may increase.
[0122] <Position Mismatch>
[0123] Furthermore, as described above, position information is quantized using voxels, so point position information may vary depending on the layer. Consequently, it may be impossible to accurately predict attribute information from intermediate-resolution position information (i.e., it may be difficult to de-layer attribute information).
[0124] In other words, in order to perform correct decoding (reverse layering) on attribute information, it is necessary to decode the position information down to the lowest layer regardless of the current layer. Therefore, the load of the decoding process may increase.
[0125] <Attribute Information Hierarchy: Method 1>
[0126] Considering the above situation, the prediction points are selected so that the attribute information is retained in the voxels of the level (LoD) one level higher than the level to which the voxel containing the point belongs (method 1), as Figure 6 The first row of the table is shown in .
[0127] Specifically, with respect to attribute information of a point cloud representing a three-dimensional object, the process of recursively classifying each point of the point cloud into predicted points and reference points is repeated for each reference point, so that the attribute information is layered. Predicted points are points where the difference between the attribute information and the predicted value is left, and reference points are points where the attribute information was referenced during the derivation of the predicted value. During the layering, predicted points are selected so that the point also exists in voxels of a layer one level higher than the layer to which the voxel containing the point in the current layer belongs.
[0128] For example, the information processing device includes a stratification unit that stratifies attribute information of a point cloud by recursively repeating a process of classifying each point of a point cloud representing a three-dimensional object into a predicted point and a reference point, wherein the predicted point is a point where a difference between the attribute information and the predicted value is left, and the reference point is a point whose attribute information was referenced during the derivation of the predicted value. During the stratification, the stratification unit selects the predicted point so that the point also exists in a voxel of a layer one level higher than the layer to which the voxel containing the point of the current layer belongs.
[0129] For example, Figure 7 The encoding of the two-dimensional interpretation is performed as shown. For example, Figure 7 As shown in FIG. 1A , position information is set so that points 102-1 to 102-9 are placed in voxel 101-1 of a predetermined layer in a predetermined spatial region 100. Note that when it is not necessary to distinguish voxels 101-1 to 101-3 from each other, voxels 101-1 to 101-3 are referred to as voxels 101. Furthermore, when it is not necessary to distinguish points 102-1 to 102-9 from each other, points 102-1 to 102-9 are referred to as points 102.
[0130] At this level, if Figure 7 As shown in B in FIG, points 102-1 to 102-9 are classified into predicted points and reference points so that points also exist in voxel 101-2 of a level higher than the level of voxel 101-1 in which points 102-1 to 102-9 exist. Figure 7 In the example shown in FIG. 8B , points 102 - 3 , 102 - 5 , and 102 - 8 indicated by white circles are set as prediction points, and the other points are set as reference points.
[0131] Likewise, in the next higher level, the point 102 is classified into a predicted point and a reference point so that a point ( Figure 7 C in Figure 7 In the example shown in FIG. 4C , points 102 - 1 , 102 - 4 , and 102 - 7 indicated by gray circles are set as prediction points, and the other points are set as reference points.
[0132] In this way, hierarchical layering is performed so that there is a point 102 in each voxel 101-3 where there is a point 102 in a lower level, as shown in FIG. Figure 7 As shown in D of FIG. Such processing is performed in each level. That is, a hierarchical structure similar to that of an octree can be achieved.
[0133] like Figure 8 As shown, for example, with Figure 7 Decoding is performed in the reverse order shown. Figure 8 As shown in FIG. 1A , the position information is set so that the points 102 - 2 , 102 - 6 and 102 - 9 are placed in the voxel 101 - 3 of the predetermined level in the predetermined spatial region 100 (with Figure 7 A state similar to the state shown in D).
[0134] At the next lower level, e.g. Figure 8 As shown in FIG. 1B , the predicted values of the points 102-1, 102-4, and 102-7 are derived using the attribute information of the points 102-2, 102-6, and 102-9 of each voxel 101-3, and the predicted values are added to the difference values, thereby restoring the attribute information of the point 102 of each voxel 101-2 (the difference between the predicted values and the ... Figure 7 A state similar to the state shown in C).
[0135] Furthermore, at the next lower level, e.g. Figure 8 As shown in FIG. 3 , the attribute information of the points 102-1, 102-2, 102-4, 102-6, 102-7, and 102-9 of each voxel 101-2 is also used to derive the predicted values of the points 102-3, 102-5, and 102-8, and the predicted values are added to the difference values, thereby restoring the attribute information (with Figure 7 A state similar to the state shown in B).
[0136] In this way, the attribute information of the point 102 of each voxel 101-1 is restored, such as Figure 8 As shown in D (with Figure 7 That is, as in the case of the octree, attribute information of each level can be restored by using attribute information in a higher level.
[0137] In this way, the hierarchical structure of attribute information can be linked to the hierarchical structure of position information. As a result, even at intermediate resolutions, the position information corresponding to each piece of attribute information can be obtained. Therefore, the attribute information can be correctly decoded by decoding the intermediate resolution position information and attribute information. Consequently, scalable decoding of attribute information can be performed more easily.
[0138] <Attribute Information Hierarchy: Method 1-1>
[0139] When deriving the prediction value using the above method 1, the prediction value can be derived using the location information of the resolution of the current level (LoD) during both encoding and decoding, such as Figure 6 The second row from the top of the table shown is mentioned (Method 1-1).
[0140] For example, with respect to attribute information of a point cloud representing a three-dimensional object, a process of classifying each point of the point cloud into predicted points and reference points can be recursively repeated for reference points, thereby performing inverse stratification of the stratified attribute information. The predicted points are points where a difference between the attribute information and the predicted value is left, and the reference points are points where the attribute information was referenced during the derivation of the predicted value. During the inverse stratification, at each level, a predicted value of the attribute information of the predicted point can be derived using the attribute information of the reference point and position information of a resolution of a level other than the lowest level of the point cloud, and the attribute information of the predicted point can be derived using the predicted value and the difference.
[0141] Furthermore, for example, the information processing device may include a de-stratification unit that performs de-stratification on attribute information of a point cloud representing a three-dimensional object, the attribute information having been stratified by recursively repeating a process of classifying each point of the point cloud into a predicted point and a reference point, the predicted point being a point where a difference between the attribute information and the predicted value is left, the reference point being a point to which the attribute information was referenced during derivation of the predicted value. In the de-stratification, the de-stratification unit may derive a predicted value of the attribute information of the predicted point in each hierarchy by using the attribute information of the reference point and position information of a resolution of a hierarchy that is not the lowest hierarchy of the point cloud, and derive the attribute information of the predicted point by using the predicted value and the difference.
[0142] Furthermore, for example, in reverse hierarchical layering, the predicted value may be derived using the attribute information of the reference point and the position information of the reference point and the prediction point at the resolution of the current layer.
[0143] As mentioned above Figure 4 As explained, the derivation of the predicted value of the predicted point is performed using the attribute information of the reference point weighted according to the distance between the predicted point and the reference point.
[0144] In either encoding or decoding, distance can be derived using the position information of the current resolution level being processed. That is, within a layer, the predicted value of a prediction point can be derived using attribute information of a reference point weighted by the distance based on the position information of the current resolution level in the point cloud.
[0145] In the case of encoding, Figure 7In B, the distance between the predicted point and the reference point is derived using the position information of the resolution of voxel 101-1, and Figure 7 In C, the distance between the prediction point and the reference point is derived using the position information of the resolution of voxel 101-2. Similarly, in the case of decoding, Figure 8 In B, the distance between the predicted point and the reference point is derived using the position information of the resolution of voxel 101-2, and Figure 8 In C, the distance between the predicted point and the reference point is derived using the position information of the resolution of the voxel 101 - 1 .
[0146] In this way, the same distance can be derived in both encoding and decoding. That is, the same prediction value can be derived in both encoding and decoding. Therefore, even when the attribute information is decoded at an intermediate resolution, the reduction in accuracy due to errors in distance calculation (the increase in error due to decoding) can be reduced. More specifically, when the attribute information is decoded at an intermediate resolution, decoding can be performed more accurately than in the case of method 1-1' described later. Compared with the case of method 1-2 described later, the increase in encoding / decoding load can also be reduced to a smaller amount.
[0147] Note that in decoding (inverse layering) ( Figure 8 ), the correspondence between the position information and the attribute information of each point at the intermediate resolution is unclear. Figure 8 In B, it is unclear which voxel 101-2 contains which point 102. Therefore, as in the case of layering, attribute information can be layered using position information (layering down to the intermediate resolution to be decoded), so that at each layer level, attribute information for each point is associated with position information. In this way, more accurate inverse layering can be performed.
[0148] <Attribute Information Hierarchy: Method 1-1'>
[0149] Note that when deriving the prediction value using the above method 1, the position information of the resolution of the current level (LoD) can be used to derive the prediction value during decoding, such as Figure 6 The third row from the top of the table shown in the figure (Method 1-1') is mentioned. That is, in encoding, the position information of the highest resolution (resolution of the lowest level) is known, and therefore, the prediction value can be derived using the position information of the highest resolution (more specifically, the distance between the reference point and the prediction point to be used when deriving the prediction value).
[0150] That is, the predicted value of the predicted point can be derived using the attribute information of the reference point weighted according to the distance based on the position information of the lowest level of resolution in the point cloud.
[0151] For example, if Figure 9 The hierarchy shown in A (with Figure 7 If the level is the same as the level shown in A of FIG) is the lowest level, the position information of the resolution of the voxel 101-1 is used to perform the derivation of the predicted values of the points 102-3, 102-5 and 102-8, as shown in FIG. Figure 9 This is similar to Figure 7 The situation shown in B.
[0152] However, in this case, as Figure 9 As shown in FIG. 3C , the position information of the resolution of the voxel 101 is also used to perform the derivation of the predicted values of the points 102 - 1 , 102 - 4 and 102 - 7 , as shown in FIG. Figure 9 As shown in B. That is, Figure 9 As shown in FIG. 4D , the position information of the points 102 - 2 , 102 - 6 , and 102 - 9 of each voxel 101 - 3 also has the resolution of the voxel 101 - 1 .
[0153] In this way, the prediction accuracy can be made higher than that in the case of method 1-1, and therefore, the encoding efficiency can be made higher than that in the case of method 1-1. However, at the time of decoding, as Figure 9 As shown in FIG. 1 , the position information of the points 102-2, 102-6, and 102-9 of each voxel 101-3 cannot have the resolution of the voxel 101-1. Figure 8 Decoding is performed in the same manner as in the case shown (method 1-1). That is, the distance between the reference point and the prediction point is derived by using position information having different resolutions for encoding and decoding.
[0154] <Attribute Information Hierarchy: Methods 1-2>
[0155] In addition, when deriving the predicted value using the above method 1, a virtual reference point can be used to derive the predicted value, such as Figure 6 The last row of the table shown in FIG1 is referred to as (method 1-2). That is, in the hierarchy, all points in the current level can be classified as prediction points, reference points can be set in the voxels in the next higher level, and the attribute information of the reference points can be used to derive the prediction values of the respective prediction points.
[0156] In addition, for example, in reverse layering, the predicted values of each prediction point can be derived using the attribute information of the reference point, the position information of the reference point at the resolution of a layer one layer higher than the current layer, and the position information of the prediction point at the resolution of the current layer.
[0157] For example, Figure 10 The encoding shown performs a two-dimensional interpretation. Figure 10 A shows the Figure 7 The hierarchical level shown is the same, and points 111-1 to 111-9 similar to points 102-1 to 102-9 exist in each voxel 101-1 in the spatial region 100. Note that when it is not necessary to distinguish the points 111-1 to 111-9 from each other, the points 111-1 to 111-9 are referred to as points 111.
[0158] In this case, all points 111-1 to 111-9 are set as prediction points, which is the same as Figure 7 The situation shown is different. Figure 10 As shown in B, in a hierarchy one level higher than voxel 101-1, virtual points 112-1 to 112-7 of voxel 101-2 are set as reference points in voxel 101-2 to which voxel 101-1 containing the predicted point (point 111) belongs. Note that when it is not necessary to distinguish points 112-1 to 112-7 from each other, points 112-1 to 112-7 are referred to as points 112.
[0159] The attribute information of the newly set virtual point 112 is derived by performing a recoloring process using the attribute information of the nearby point 111. Note that the position of the point 112 is any appropriate position in the voxel 101-2 and may be different from Figure 10 The center position is shown in B.
[0160] like Figure 10 As shown in FIG. 8B , the predicted values of the respective prediction points (points 111 ) are derived using these reference points (points 112 ).
[0161] Then, at the next higher level, Figure 10 The point 112 set in B is set as the prediction point, such as Figure 10 As shown in C. Figure 10 In a manner similar to that shown in the case of B, in a hierarchy one level higher than the voxel 101-2, the virtual points 113-1 to 113-3 of the voxel 101-3 are set as reference points in the voxel 101-3 to which the voxel 101-2 containing the predicted point (point 112) belongs. Note that when it is not necessary to distinguish the points 113-1 to 113-3 from each other, the points 113-1 to 113-3 are referred to as points 113.
[0162] In a manner similar to that in the case of the point 112, attribute information of the newly set virtual point 113 is derived through the recoloring process. Figure 10 As shown in FIG. 3C , the predicted values of the respective prediction points (points 112 ) are derived by using these reference points (points 113 ).
[0163] In this way, hierarchical layering is performed so that a point 113 exists in each voxel 101-3 where a point 111 exists in a lower level. Figure 10 As shown in D. Such processing is performed in each level. That is, a layering similar to that of an octree can be achieved.
[0164] like Figure 11 As shown, for example, with Figure 10 The decoding is performed in the reverse order of the order shown. Figure 11 As shown in FIG. 1A , the position information is set so that the points 113 - 1 to 113 - 3 are placed in the voxel 101 - 3 of the predetermined level in the predetermined spatial region 100 (similar to FIG. Figure 10 The state shown in D)
[0165] At the next lower level, e.g. Figure 11 As shown in FIG. 1B , the attribute information of the points 113-1 to 113-3 of each voxel 101-3 is used to derive the predicted values of the points 112-1 to 112-7, and the predicted values are added to the difference values so that the attribute information of the point 112 of each voxel 101-2 is restored (similar to FIG. Figure 10 The state shown in C).
[0166] Furthermore, at the next lower level, e.g. Figure 11 As shown in FIG. 1C, the attribute information of the points 112-1 to 112-7 of each voxel 101-2 is used to derive the predicted values of the points 111-1 to 111-9, and the predicted values are added to the difference values so that the attribute information is restored (similar to Figure 10 of the state shown in B).
[0167] In this way, the attribute information of the point 111 of each voxel 101-1 is restored, such as Figure 11 As shown in D (similar to Figure 10 That is, as in the case of the octree, attribute information of each level can be restored by using attribute information of a higher level.
[0168] In addition, in this case, prediction is performed using virtual points that have been recolored. Therefore, the prediction accuracy can be made higher than that in the case of Method 1-1 and Method 1-1'. That is, the encoding efficiency can be made higher than that in the case of Method 1-1 and Method 1-1'. In addition, in the case of this method, the distance between the reference point and the prediction point can be derived using position information of the same resolution in both encoding and decoding. Therefore, decoding can be performed more accurately than in the case of Method 1-1'. However, in this case, a virtual point is set and a recoloring process is performed. Therefore, the load increases accordingly.
[0169] <Hierarchy of Attribute Information: Combination>
[0170] Note that any of the above methods can be selected and adopted. In this case, the encoding side selects and adopts a method, and information indicating which method has been selected is sent from the encoding side to the decoding side. Based on this information, the decoding side is only required to adopt the same method as the encoding side.
[0171] <Refer to lower layers when deriving quantization weights>
[0172] At the same time, as described in "Quantization," the quantization weight W used for quantizing and layering the attribute information (difference) is derived for each point using the quantization weights of the lower layers. Therefore, during decoding, in order to inverse quantize the attribute information at the intermediate resolution, the quantization weights of the lower layers are required. This makes scalable decoding difficult. In other words, to decode the attribute information at the desired intermediate resolution, all layers must be decoded. This can increase the decoding load.
[0173] <Quantization weights: Method 1>
[0174] In view of the above situation, instead of using the quantization weight of each point mentioned above, the quantization weight of each level (LoD) is used to perform quantization and inverse quantization of attribute information (method 1), such as Figure 12 As described in the first row of the table shown, that is, the difference value of each point of each level generated as described above can be quantized and inversely quantized using the quantization weight of each level.
[0175] For example, Figure 13 As shown, the quantization weight of the level LoD (N) is set to C0, and the attribute information (differences) of all points of the level LoD (N) are quantized and inversely quantized using the quantization weight C0. In addition, the quantization weight of the next higher level LoD (N-1) is set to C1, and the attribute information (differences) of all points of the level LoD (N-1) are quantized and inversely quantized using the quantization weight C1. In addition, the quantization weight of the next higher level LoD (N-2) is set to C2, and the attribute information (differences) of all points of the level LoD (N-2) are quantized and inversely quantized using the quantization weight C2.
[0176] Note that these quantization weights C0, C1, and C2 are set independently of each other. When using quantization weights for individual points, it is necessary to consider how to weight points at the same level. This can not only complicate processing but also make it difficult to maintain the independence of the quantization weights for each level, for example because information from lower levels is required. On the other hand, it is easier to derive the quantization weights for each level independently of each other.
[0177] Since quantization weights are independent between layers, quantization and inverse quantization can be performed independently of any lower layer. That is, since there is no need to refer to any lower layer information as described above, scalable decoding can be achieved more easily.
[0178] Note that the quantization weights at higher levels may be larger (in Figure 13 In the example shown, C0≤C1≤C2). Generally, hierarchical attribute information is more important because higher-level attribute information affects the attribute information of a large number of points (or is directly or indirectly used to derive the predicted values of a large number of points). Therefore, by using larger quantization weights to quantize and inversely quantize the higher-level attribute information (difference), the reduction in encoding and decoding accuracy can be reduced.
[0179] <Quantization Weight: Method 1-1>
[0180] Moreover, when method 1 is adopted, the function for deriving the quantization weight (quantization weight derivation function) can be shared between encoding and decoding, such as Figure 12 As described in the second row from the top of the table (Method 1-1), the quantization weights for each layer can be derived using the same quantization weight derivation function in both encoding and decoding. This allows the same quantization weights to be derived in both encoding and decoding. Consequently, the loss in accuracy due to quantization and inverse quantization can be reduced.
[0181] Note that the quantization weight derivation function can be any type of function. That is, for each level, the quantization weight to be derived using the quantization weight derivation function is any appropriate quantization weight. For example, the quantization weight derivation function can be a function that obtains larger quantization weights for higher levels (a function that increases monotonically with the level). Moreover, the rate of increase of the quantization weights for each level can be constant or not. In addition, the range of the quantization weights to be derived using the quantization weight derivation function can also be any appropriate range.
[0182] <Quantization Weight: Method 1-1-1>
[0183] Again, in that case, predefined functions can be shared in advance, e.g. Figure 12 As described in the third row from the top of the table (Method 1-1-1), in other words, the same predetermined quantization weight derivation function can be used to derive the quantization weights for each layer in both encoding and decoding. This arrangement makes it easier to share the quantization weight derivation function between encoding and decoding.
[0184] <Quantization Weight: Method 1-1-2>
[0185] In addition, if Figure 12 As described in the fourth row from the top of the table shown, parameters for defining a quantization weight derivation function (e.g., information or coefficients specifying a function) can be shared (method 1-1-2). That is, the quantization weight derivation function can be defined at the time of encoding, and the parameters defining the quantization weight derivation function can be sent (included in the bitstream) to the decoding side. That is, the quantized difference and the parameters defining the function for deriving the quantization weights of each level can be encoded to generate encoded data. In addition, the encoded data can be decoded to obtain the quantized difference and the parameters defining the function for deriving the quantization weights of each level. With this arrangement, the quantization weight derivation function used at the time of encoding can be easily restored based on the parameters included in the bitstream. Therefore, the same quantization weight derivation function as in the case of encoding can be easily adopted at the time of decoding.
[0186] <Quantization Weights: Methods 1-2>
[0187] In addition, when using method 1, the quantization weights can be shared between encoding and decoding, such as Figure 12 As described in the fifth row from the top of the table shown (Method 1-2). That is, in encoding and decoding, quantization and inverse quantization can be performed using the same quantization weight. By doing so, the accuracy reduction caused by quantization and inverse quantization can be reduced. Note that the quantization weights of each layer can have any appropriate value.
[0188] <Quantization Weight: Method 1-2-1>
[0189] Also, in that case, predetermined quantization weights can be shared in advance, e.g. Figure 12 As described in the sixth row from the top of the table (method 1-2-1). That is, quantization and inverse quantization can be performed using the same predetermined quantization weights in encoding and decoding. With this arrangement, quantization weights can be more easily shared between encoding and decoding.
[0190] <Quantization Weight: Method 1-2-2>
[0191] Moreover, if Figure 12As mentioned in the seventh row from the top of the table shown, the quantization weight can be sent from the encoding side to the decoding side (method 1-2-2). That is, the quantization weight can be set at the time of encoding, and the quantization weight can be sent (included in the bit stream) to the decoding side. That is, the quantized difference values and quantization weights of each level can be encoded to generate encoded data. The encoded data can then be decoded to obtain the quantized difference values and quantization weights of each level. With this arrangement, the quantization weight included in the bit stream can be used for inverse quantization at the time of decoding. Therefore, the quantization weight used in the quantization at the time of encoding can be more easily used in the inverse quantization at the time of decoding.
[0192] <Quantization weights: Method 2>
[0193] Note that the quantization weights of each point can be used in quantization and inverse quantization. Figure 12 As described in the eighth row from the top of the table, quantization weights for each point can be defined and used in encoding (quantization), and the quantization weights for each point (included in the bitstream) can be sent to the decoding side (method 2). With this arrangement, quantization and inverse quantization can be performed using quantization weights that are more suitable for the attribute information (difference value) of each point. Therefore, it is possible to reduce the reduction in encoding efficiency.
[0194] <Quantization Weight: Method 2-1>
[0195] When using method 2, the quantization weights of all defined points can be sent (included in the bitstream), such as Figure 12 As described in the ninth row from the top of the table shown (method 2-1). With this arrangement, the quantization weights included in the bitstream can be used without any changes during decoding (inverse quantization). Therefore, inverse quantization can be performed more easily.
[0196] <Quantization Weight: Method 2-2>
[0197] Similarly, when method 2 is adopted, the difference between the quantization weight of each defined point and the predetermined value can be sent (included in the bit stream), such as Figure 12 As described in the bottom row of the table shown (Method 2-2). This predetermined value may be any appropriate value. A value different from each quantization weight may be set as the predetermined value, or the difference between the quantization weights may be transmitted. With this arrangement, for example, compared to the case of Method 2-1, the increase in the amount of information related to the quantization weights can be reduced to a smaller value, and the reduction in encoding efficiency can be reduced to a smaller value.
[0198] <Quantization weight: combination>
[0199] Note that any of the above methods can be selected and adopted. In this case, the encoding side selects and adopts a method, and information indicating which method has been selected is sent from the encoding side to the decoding side. Based on this information, the decoding side is only required to adopt the same method as the encoding side.
[0200] <Applying quantized weights to layers>
[0201] Quantization weights can be used to layer attribute information. For example, in the case of "lifting," quantization weights are used to update the difference values derived at each layer, and the updated difference values are used to update the attribute information of the point set as the reference point. This reduces the reduction in compression efficiency. When performing such an update process in the layered attribute information, the quantization weights described above for each layer can be used.
[0202] With this arrangement, it is not necessary to refer to any lower level information, and therefore scalable decoding can be more easily achieved.
[0203] <2. First embodiment>
[0204] <Encoding device>
[0205] Next, the apparatus according to the present technology described in <1. Scalable decoding> above is described. Figure 14 : is a block diagram illustrating an example configuration of an encoding device as an embodiment of an information processing device to which the present technology is applied. Figure 14 The encoding device 200 shown is a device that encodes a point cloud (3D data). The encoding device 200 encodes the point cloud by using the present technology described in <1. Scalable Decoding> above.
[0206] Notice, Figure 14 The main components and aspects, such as processing units and data flows, are shown, but Figure 14 Not all components and aspects are necessarily shown. That is, in the encoding device 200, there may be Figure 14 The processing units are shown as blocks, or there may be processing units that are not shown in Figure 14 The process or data flow is shown as arrows or the like.
[0207] like Figure 14 As shown, the encoding device 200 includes a position information encoding unit 201, a position information decoding unit 202, a point cloud generation unit 203, an attribute information encoding unit 204 and a bit stream generation unit 205.
[0208] The position information encoding unit 201 encodes the position information of the point cloud (3D data) input to the encoding device 200. The encoding method used here can be any appropriate method. For example, processing such as filtering and quantization for reducing noise (denoising) can be performed. The position information encoding unit 201 provides the generated encoded data of the position information to the position information decoding unit 202 and the bitstream generation unit 205.
[0209] The position information decoding unit 202 obtains the encoded data of the position information provided by the position information encoding unit 201 and decodes the encoded data. The decoding method adopted here can be any appropriate method compatible with the encoding performed by the position information encoding unit 201. For example, processing such as filtering and inverse quantization for noise reduction can be performed. The position information decoding unit 202 provides the generated position information (decoding result) to the point cloud generation unit 203.
[0210] The point cloud generation unit 203 acquires the attribute information of the point cloud input to the encoding device 200 and the position information (decoded result) provided by the position information decoding unit 202. The point cloud generation unit 203 performs processing (recoloring processing) to associate the attribute information with the position information (decoded result). The point cloud generation unit 203 provides the attribute information associated with the position information (decoded result) to the attribute information encoding unit 204.
[0211] The attribute information encoding unit 204 receives the position information (decoded result) and attribute information provided by the point cloud generation unit 203. Using the position information (decoded result), the attribute information encoding unit 204 encodes the attribute information using the encoding method of the present technology described above in <1. Scalable Decoding>, and generates encoded data of the attribute information. The attribute information encoding unit 204 provides the generated encoded data of the attribute information to the bitstream generation unit 205.
[0212] The bitstream generation unit 205 obtains the encoded data of the position information provided by the position information encoding unit 201. The bitstream generation unit 205 also obtains the encoded data of the attribute information provided by the attribute information encoding unit 204. The bitstream generation unit 205 generates a bitstream including these encoded data. The bitstream generation unit 205 then outputs the generated bitstream to the outside of the encoding device 200.
[0213] With this configuration, the encoding device 200 can associate the hierarchical structure of attribute information with the hierarchical structure of position information. Similarly, with this configuration, quantization and inverse quantization can be performed independently of any lower layers. That is, the encoding device 200 can further facilitate scalable decoding of attribute information.
[0214] Note that these processing units (position information encoding unit 201 to bit stream generating unit 205) all have any appropriate configuration. For example, each processing unit can be formed with a logic circuit that implements the above-mentioned processing. Alternatively, each processing unit can, for example, include a central processing unit (CPU), a read-only memory (ROM), a random access memory (RAM), etc., and use these components to execute a program to perform the above-mentioned processing. Each processing unit can certainly have two configurations, and perform some of the above-mentioned processing by a logic circuit, and perform other processing by executing a program. The configurations of each processing unit can be independent of each other. For example, one processing unit can perform some of the above-mentioned processing using a logic circuit, while other processing units perform the above-mentioned processing by executing a program. In addition, some other processing units can perform the above-mentioned processing using a logic circuit and by executing a program.
[0215] <Attribute Information Encoding Unit>
[0216] Figure 15 The attribute information encoding unit 204 ( Figure 14 ) is a block diagram of a typical example configuration. Note that Figure 15 The main components and aspects are shown, such as processing units and data flows, but Figure 15 Not all components and aspects are necessarily shown. That is, in the attribute information encoding unit 204, there may be Figure 15 The processing units are not shown as blocks, or may exist in Figure 15 Not shown are processes or data flows as arrows or the like.
[0217] like Figure 15 As shown, the attribute information encoding unit 204 includes a hierarchical processing unit 211, a quantization unit 212, and an encoding unit 213.
[0218] The layering processing unit 211 performs processing related to the layering of attribute information. For example, the layering processing unit 211 obtains the attribute information and position information (decoding results) provided to the attribute information encoding unit 204. The layering processing unit 211 uses the position information to layer the attribute information. In doing so, the layering processing unit 211 performs layering using the present technique described in <1. Scalable Decoding> above. In other words, the layering processing unit 211 generates attribute information having a layered structure similar to that of the position information (correlating the layered structure of the attribute information with the layered structure of the position information). The layering processing unit 211 provides the layered attribute information (difference) to the quantization unit 212.
[0219] The quantization unit 212 receives the attribute information (difference value) supplied from the hierarchical processing unit 211. The quantization unit 212 quantizes the attribute information (difference value). In doing so, the quantization unit 212 performs quantization using the present technique described in <1. Scalable Decoding> above. That is, as described in <1. Scalable Decoding>, the quantization unit 212 quantizes the attribute information (difference value) using the quantization weights for each hierarchical level. The quantization unit 212 supplies the quantized attribute information (difference value) to the encoding unit 213. Note that the quantization unit 212 also supplies information regarding the quantization weights used in quantization to the encoding unit 213, as necessary.
[0220] The encoding unit 213 obtains the quantized attribute information (difference value) provided by the quantization unit 212. The encoding unit 213 encodes the quantized attribute information (difference value) and generates encoded data of the attribute information. The encoding method used here can be any appropriate method. Note that when information about quantization weights is provided from the quantization unit 212, the encoding unit 213 also incorporates information (e.g., such as quantization weight derivation function definition parameters and quantization weights) into the encoded data. The encoding unit 213 provides the generated encoded data to the bitstream generation unit 205.
[0221] By performing the hierarchical layering and quantization described above, the attribute information encoding unit 204 can associate the hierarchical structure of the attribute information with the hierarchical structure of the position information. Furthermore, quantization can be performed independently of any lower layers. Furthermore, inverse quantization can be performed independently of any lower layers. In other words, the attribute information encoding unit 204 can further facilitate scalable decoding of the attribute information.
[0222] Note that these processing units (hierarchical processing unit 211 to encoding unit 213) have any appropriate configuration. For example, each processing unit can be formed with a logic circuit that implements the above-mentioned processing. In addition, each processing unit can also include, for example, a CPU, ROM, RAM, etc., and use them to execute programs to perform the above-mentioned processing. Each processing unit can certainly have two configurations, and use logic circuits to perform some of the above-mentioned processing, and perform other processing by executing programs. The configurations of each processing unit can be independent of each other. For example, one processing unit can perform some of the above-mentioned processing using logic circuits, while other processing units perform the above-mentioned processing by executing programs. In addition, some other processing units can perform the above-mentioned processing using logic circuits and by executing programs.
[0223] <Hierarchical Processing Unit>
[0224] Figure 16 is a diagram showing the hierarchical processing unit 211 ( Figure 15 ) is a block diagram of a typical example configuration. Note that Figure 16The main components and aspects are shown, such as processing units and data flows, but Figure 16 Not all components and aspects are necessarily shown. That is, in the layered processing unit 211, there may be Figure 16 The processing units are not shown as blocks, or may exist in Figure 16 Not shown are processes or data flows as arrows or the like.
[0225] like Figure 16 As shown, the hierarchical processing unit 211 includes individual hierarchical processing units 221-1, individual hierarchical processing units 221-2, .... The individual hierarchical processing unit 221-1 processes the lowest hierarchical level (N) of the attribute information. The individual hierarchical processing unit 221-2 processes the hierarchical level (N-1) that is one level higher than the lowest hierarchical level (N) of the attribute information. When there is no need to distinguish between these individual hierarchical processing units 221-1, 221-2, ..., they are referred to as individual hierarchical processing units 221. That is, the hierarchical processing unit 211 includes the same number of individual hierarchical processing units 221 as the number of hierarchical levels of attribute information that can be processed.
[0226] The individual level processing unit 221 performs processing related to the stratification of attribute information of the current level (current LoD) as the processing target. For example, the individual level processing unit 221 classifies each point of the current level into a prediction point of the current level and a reference point (a prediction point of a higher level), and derives a prediction value and a difference value. In doing so, the individual level processing unit 221 performs processing using the present technology described in <1. Scalable Decoding> above.
[0227] The individual level processing unit 221 includes a point classification unit 231 , a prediction unit 232 , an arithmetic unit 233 , an updating unit 234 , and an arithmetic unit 235 .
[0228] The point classification unit 231 classifies each point of the current level into a prediction point and a reference point. In doing so, as described above in <1. Scalable Decoding>, the point classification unit 231 selects a prediction point (in other words, a reference point) so that the point also exists in a voxel of a level higher than the level to which the voxel containing the point of the current level belongs. That is, the point classification unit 231 selects a prediction point (in other words, a reference point) by referring to the above. Figure 6 The point classification unit 231 classifies the points using any of the methods described in the table shown in . The point classification unit 231 supplies the attribute information H(N) of the point selected as the prediction point to the arithmetic unit 233. The point classification unit 231 also supplies the attribute information L(N) of the point selected as the reference point to the prediction unit 232 and the arithmetic unit 235.
[0229] Note that the point classification unit 231 , the prediction unit 232 , and the updating unit 234 share the position information of each point.
[0230] The prediction unit 232 uses the attribute information L(N) and position information of the point selected as the reference point, which are provided from the point classification unit 231, to derive a predicted value P(N) of the attribute information of the predicted point. In doing so, the prediction unit 232 derives the predicted value P(N) using the method of the present technology described above in <1. Scalable Decoding>. The prediction unit 232 supplies the derived predicted value P(N) to the arithmetic unit 233.
[0231] For each prediction point, the arithmetic unit 233 subtracts the prediction value P(N) supplied from the prediction unit 232 from the attribute information H(N) supplied from the point classification unit 231, thereby deriving a difference value D(N). The arithmetic unit 233 supplies the derived difference value D(N) for each prediction point to the quantization unit 212 ( Figure 15 ). The arithmetic unit 233 also provides the difference value D(N) of each derived prediction point to the updating unit 234.
[0232] The updating unit 234 obtains the difference value D(N) of the predicted point supplied from the arithmetic unit 233. The updating unit 234 updates the difference value D(N) by performing a predetermined operation and derives the coefficient U(N). The arithmetic method employed herein may be any appropriate method, but for example, the arithmetic operation shown in the following expression (2) may be performed using the quantization weight W to be used for quantization.
[0233] [Mathematical formula 2]
[0234]
[0235] In doing so, the updating unit 234 independently derives the quantization weight w to be used in the arithmetic operation for each level, as described above <1. Scalable decoding>. That is, the point classification unit 231 uses the quantization weight w to be used in the arithmetic operation by adopting the above reference Figure 12 The arithmetic operation is performed using the quantization weight w derived by one of the methods described in the table shown in The updating unit 234 supplies the derived coefficient U(N) to the operation unit 235 .
[0236] The arithmetic unit 235 obtains the attribute information L(N) of the reference point provided from the point classification unit 231. The arithmetic unit 235 also obtains the coefficient U(N) provided from the updating unit 234. The arithmetic unit 235 updates the attribute information L(N) by adding the coefficient U(N) to the attribute information L(N) (derives the updated attribute information L′(N)). As a result, the reduction in compression efficiency can be reduced. The arithmetic unit 235 provides the updated attribute information L′(N) to the individual level processing unit 221 (the point classification unit 231) whose processing target is the next higher level (N-1). For example, in Figure 16In the case shown, the arithmetic unit 235 of the individual level processing unit 221-1 provides the updated attribute information L'(N) to the individual level processing unit 221-2. The individual level processing unit 221 provided with the updated attribute information L'(N) performs similar processing, where the processing target is a point (i.e., the next higher level is the current level).
[0237] By performing the layering as described above, the layering processing unit 211 can associate the hierarchical structure of the attribute information with the hierarchical structure of the position information. In addition, the layering can be performed using quantization weights derived independently of any lower layers. That is, the layering processing unit 211 can further facilitate scalable decoding of the attribute information.
[0238] Note that these processing units (point classification unit 231 to arithmetic unit 235) have any appropriate configuration. For example, each processing unit can be formed with a logic circuit that implements the above-mentioned processing. In addition, each processing unit can also include, for example, a CPU, ROM, RAM, etc., and use them to execute programs to perform the above-mentioned processing. Each processing unit can of course have two configurations, and perform some of the above-mentioned processing by logic circuits, and perform other processing by executing programs. The configurations of the various processing units can be independent of each other. For example, one processing unit can perform some of the above-mentioned processing using logic circuits, while other processing units perform the above-mentioned processing by executing programs. In addition, some other processing units can perform the above-mentioned processing using logic circuits and by executing programs.
[0239] <Encoding Process Flow>
[0240] Next, the processing to be performed by the encoding device 200 is described. The encoding device 200 encodes the data of the point cloud by performing the encoding processing. Figure 17 The flowchart shown describes an example flow in the encoding process.
[0241] When the encoding process starts, in step S201 , the position information encoding unit 201 of the encoding device 200 encodes the input position information of the point cloud and generates encoded data of the position information.
[0242] In step S202 , the position information decoding unit 202 decodes the encoded data of the position information generated in step S201 , and generates position information.
[0243] In step S203 , the point cloud generation unit 203 performs a recoloring process using the attribute information of the input point cloud and the position information (decoding result) generated in step S202 , and associates the attribute information with the position information.
[0244] In step S204, the attribute information encoding unit 204 performs attribute information encoding processing to encode the attribute information that has undergone the recoloring processing in step S203 and generates encoded attribute information data. In doing so, the attribute information encoding unit 204 utilizes the present technology described in <1. Scalable Decoding> above. The attribute information encoding processing will be described in detail later.
[0245] In step S205 , the bit stream generation unit 205 generates and outputs a bit stream containing the encoded data of the position information generated in step S201 and the encoded data of the attribute information generated in step S204 .
[0246] When the processing in step S205 is completed, the encoding process ends.
[0247] By performing the processing in each step in the manner described above, the encoding device 200 can associate the hierarchical structure of attribute information with the hierarchical structure of position information. Furthermore, with this configuration, quantization and inverse quantization can be performed independently of any lower hierarchical levels. In other words, the encoding device 200 can further facilitate scalable decoding of attribute information.
[0248] <Flow of Attribute Information Coding Process>
[0249] Next, refer to Figure 18 The flowchart shown in Figure 17 An example flow of the attribute information encoding process performed in step S204.
[0250] When the attribute information encoding process begins, in step S221, the layering processing unit 211 of the attribute information encoding unit 204 performs a layering process to layer the attribute information and derive the difference in attribute information at each point. In doing so, the layering processing unit 211 employs the present technique described in <1. Scalable Decoding> above to perform the layering. The layering process will be described in detail later.
[0251] In step S222, the quantization unit 212 quantizes each difference value derived in step S221 by performing a quantization process. In doing so, the quantization unit 212 performs quantization using the present technology described in <1. Scalable Decoding> above. The quantization process will be described in detail later.
[0252] In step S223 , the encoding unit 213 encodes the difference value quantized in step S222 , and generates encoded data of attribute information.
[0253] When the processing in step S223 is completed, the attribute information encoding processing ends, and the processing returns to Figure 17 .
[0254] By performing the processing in each step in the manner described above, the attribute information encoding unit 204 can associate the hierarchical structure of the attribute information with the hierarchical structure of the position information. Furthermore, quantization can be performed independently of any lower hierarchical levels. Furthermore, inverse quantization can be performed independently of any lower hierarchical levels. In other words, the attribute information encoding unit 204 can further facilitate scalable decoding of the attribute information.
[0255] <Flow of layered processing>
[0256] Next, refer to Figure 19 The flowchart shown in Figure 18 An example process of layered processing performed in step S221 in FIG. Figure 6 The case of method 1-1 is described in the second row from the top of the shown table.
[0257] When the hierarchical processing starts, in step S241 , the individual hierarchy processing unit 221 of the hierarchical processing unit 211 sets the lowest hierarchy (LoD) as the current hierarchy (current LoD) that is a processing target.
[0258] In step S242, the point classification unit 231 selects a prediction point so that a point also exists in a voxel of a layer one level higher than the layer to which the voxel containing the point of the current layer belongs. In other words, the point classification unit 231 selects a reference point through this process.
[0259] For example, the point classification unit 231 sets reference points one by one for voxels in a layer one level higher than the layer to which each voxel containing the point of the current layer belongs. Note that when there are multiple candidate points, the point classification unit 231 selects one of them. The selection method adopted herein may be any suitable method. For example, a point with the most points in its vicinity may be selected as a reference point. After setting the reference point, the point classification unit 231 next sets points other than the reference point as predicted points. In this way, predicted points are selected so that points also exist in voxels in a layer one level higher than the layer to which the voxel containing the point of the current layer belongs.
[0260] In step S243 , the point classification unit 231 sets reference points corresponding to the respective prediction points based on the position information of the resolution of the current hierarchy.
[0261] For example, the point classification unit 231 uses the position information at the current level of resolution to search for prediction points (neighboring prediction points) within a predetermined range from the reference point. The point classification unit 231 associates the prediction point with the reference point so that the detected neighboring prediction points will reference that reference point. Since the point classification unit 231 performs this process for each reference point, all prediction points are associated with the reference point. In other words, a reference point corresponding to each prediction point is set (to be referenced by each prediction point).
[0262] In step S244, the prediction unit 232 derives the predicted value of each prediction point. That is, based on the correspondence relationship set by the process in step S243, the prediction unit 232 uses the attribute information of the reference point corresponding to each prediction point to derive the predicted value of the attribute information of the prediction point.
[0263] For example, as mentioned above Figure 4 As described above, the prediction unit 232 weights the attribute information of each reference point corresponding to the prediction point using a quantization weight (α(P, Q(i, j))) corresponding to the inverse of the distance between the prediction point and the reference point, and integrates the attribute information by performing the arithmetic operation shown in Expression (1) shown above. In doing so, the prediction unit 232 uses the position information of the resolution of the current layer to derive the distance between the prediction point and the reference point and adopts this distance.
[0264] In step S245 , the arithmetic unit 233 derives difference values for the respective prediction points by subtracting the prediction value derived in step S244 from the attribute information of the prediction point selected in step S242 .
[0265] In step S246, the updating unit 234 derives the quantization weights to be used in the updating process. The updating unit 234 uses Figure 12 The quantization weights are derived using one of the methods described in the table shown in FIG.
[0266] In step S247, the updating unit 234 updates the difference value derived in step S245 using the quantization weight derived in step S246. For example, the updating unit 234 performs this updating by performing the arithmetic operation expressed by the above-mentioned expression (2).
[0267] In step S248 , the operation unit 235 updates the attribute information of the reference point set in step S243 using the difference value updated in step S247 .
[0268] In step S249, the individual level processing unit 221 determines whether the current level (current LoD) is the highest level. Note that the highest level in the hierarchical structure of attribute information is set independently of the level of position information and, therefore, may be different from the highest level in the hierarchical structure of position information. If it is determined that the current level is not the highest level, the process proceeds to step S250.
[0269] In step S250, the individual hierarchy processing unit 221 sets the next higher hierarchy (LoD) as the current hierarchy (current LoD) that is the processing target. When the processing in step S250 is completed, the processing returns to step S242. That is, the processing in each of steps S242 to S250 is performed for each hierarchy.
[0270] Furthermore, if it is determined in step S249 that the current level is the highest level, the hierarchical processing ends, and the processing returns to the Figure 18 .
[0271] By performing the processing in each step in the manner described above, the hierarchical processing unit 211 can associate the hierarchical structure of the attribute information with the hierarchical structure of the position information. Furthermore, the hierarchical structure can be performed using quantization weights derived independently of any lower hierarchical level. In other words, the hierarchical processing unit 211 can further facilitate scalable decoding of the attribute information.
[0272] <Quantization Process>
[0273] Next, refer to Figure 20 The flowchart shown describes Figure 18 In this paper, the example process of the quantization process to be performed in step S222 of Figure 12 The case of method 1-1-2 described in the fourth row from the top of the shown table will be described.
[0274] When the quantization process starts, in step S271 , the quantization unit 212 defines a quantization weight derivation function.
[0275] In step S272 , the quantization unit 212 derives the quantization weight of each level (LoD) using the quantization weight derivation function defined in step S271 .
[0276] In step S273 , the quantization unit 212 quantizes the difference value of the attribute information using the quantization weight of each LoD derived in step S272 .
[0277] In step S274, the quantization unit 212 provides the quantization weight derivation function definition parameters that define the quantization weight derivation function used in quantization weight derivation to the encoding unit 213, and causes the encoding unit 213 to perform encoding. That is, the quantization unit 212 sends the quantization weight derivation function definition parameters to the decoding side.
[0278] When the processing in step S274 is completed, the quantization processing ends, and the processing returns to Figure 18 .
[0279] By performing the quantization process described above, the quantization unit 212 can perform quantization that is not dependent on any lower layer. In addition, it can perform inverse quantization that is not dependent on any lower layer. That is, the quantization unit 212 can further facilitate scalable decoding of attribute information.
[0280] <Flow of layered processing>
[0281] Note that when using Figure 6 When the method 1-1' is described in the third row from the top of the table shown, basically follow Figure 19 Hierarchical processing is performed in a similar manner to the process in the flowchart shown in FIG. However, in the process in step S243, the point classification unit 231 sets the reference point corresponding to each prediction point based on the position information of the lowest level of resolution (i.e., the highest resolution). Furthermore, in the process in step S244, the prediction unit 232 uses the position information of the lowest level of resolution (i.e., the highest resolution) to derive the distance between the prediction point and the reference point, and derives the prediction value using a quantization weight that depends on the inverse of the distance.
[0282] By doing so, the hierarchical processing unit 211 can derive a predicted value of a predicted point using the attribute information of the reference point weighted according to the distance based on the position information of the lowest level of resolution in the point cloud.
[0283] Next, refer to Figure 21 The flowchart shown in Figure 6 An example flow of hierarchical processing in the case of method 1-2 described in the fourth row from the top of the table shown will be described.
[0284] In this case, the hierarchical processing unit 211 is Figure 19 The processing in step S301 and each of steps S303 to S310 is performed in a manner similar to the processing in step S241 and each of steps S243 to S250 in FIG.
[0285] Furthermore, in step S302, the point classification unit 231 sets all points in the current hierarchy as predicted points and also sets virtual reference points so that points also exist in voxels in a hierarchy one level higher than the hierarchy to which the voxel containing the point in the current hierarchy belongs. The point classification unit 231 performs recoloring processing using the attribute information of the existing points (predicted points) and generates attribute information for the set virtual reference points.
[0286] In step S303 , the point classification unit 231 sets reference points corresponding to the respective prediction points among the virtual reference points set in step S302 based on the position information of the resolution of the current hierarchy.
[0287] By performing the processing in each step in the manner described above, the hierarchical processing unit 211 can associate the hierarchical structure of the attribute information with the hierarchical structure of the position information. Furthermore, the hierarchical structure can be performed using quantization weights derived independently of any lower hierarchical level. In other words, the hierarchical processing unit 211 can further facilitate scalable decoding of the attribute information.
[0288] <Quantization Process>
[0289] Next, refer to Figure 22 The flowchart shown in Figure 12 An example flow of the quantization process in the case of method 1-2-2 is described in the seventh row from the top of the table shown.
[0290] When the quantization process starts, in step S321 , the quantization unit 212 derives a quantization weight for each level (LoD).
[0291] In step S322 , the quantization unit 212 quantizes the difference value of the attribute information using the quantization weight of each LoD derived in step S321 .
[0292] In step S323, the quantization unit 212 supplies the quantization weight used for quantization to the encoding unit 213, and causes the encoding unit 213 to perform encoding. That is, the quantization unit 212 causes the quantization weight to be transmitted to the decoding side.
[0293] When the processing in step S323 is completed, the quantization processing ends, and the processing returns to Figure 18 .
[0294] By performing the quantization process described above, the quantization unit 212 can perform quantization that is not dependent on any lower layer. In addition, it can perform inverse quantization that is not dependent on any lower layer. That is, the quantization unit 212 can further facilitate scalable decoding of attribute information.
[0295] <3. Third embodiment>
[0296] <Decoding Device>
[0297] Figure 23 : is a block diagram showing an example configuration of a decoding device as an embodiment of an information processing device to which the present technology is applied. Figure 23 The decoding device 400 shown is a device that decodes coded data of a point cloud (3D data). The decoding device 400 decodes the coded data of the point cloud by adopting the present technology described above in <1. Scalable Decoding>.
[0298] Notice, Figure 23 The main components and aspects such as processing units and data flow are shown, but Figure 23Not all components and aspects are necessarily shown. That is, in the decoding device 400, there may be Figure 23 The processing units shown as blocks in FIG, or there may be processing units not shown in FIG. Figure 23 The flow of processing or data is indicated by arrows etc.
[0299] like Figure 23 As shown, the decoding device 400 includes a decoding target LoD depth setting unit 401, a coded data extraction unit 402, a position information decoding unit 403, an attribute information decoding unit 404 and a point cloud generation unit 405.
[0300] The decoding target LoD depth setting unit 401 performs processing related to setting the depth of the layer (LoD) as the decoding target. For example, the decoding target LoD depth setting unit 401 determines the layer to which the encoded data of the point cloud stored in the encoded data extraction unit 402 is decoded. The method used herein to set the depth of the layer as the decoding target may be any appropriate method.
[0301] For example, the decoding target LoD depth setting unit 401 may also perform setting based on an instruction related to the layer depth, the instruction being sent from an external source such as a user or an application. Alternatively, the decoding target LoD depth setting unit 401 may also calculate and set the depth of the layer to be decoded based on desired information such as an output image.
[0302] For example, the decoding target LoD depth setting unit 401 may set the depth of the layer to be decoded based on the viewpoint position, direction, viewing angle, viewpoint movement (motion, translation, tilt, or zoom), etc. of the two-dimensional image to be generated from the point cloud.
[0303] Note that the data unit for setting the depth of the layer as the decoding target can be any appropriate data unit. For example, the decoding target LoD depth setting unit 401 can set the layer depth of the entire point cloud, can set the layer depth of each object, and can set the layer depth of each partial area within the object. Of course, in addition to these examples, the layer depth can also be set for each data unit.
[0304] The coded data extraction unit 402 acquires and stores the bitstream input to the decoding device 400. The coded data extraction unit 402 extracts the coded data of the position information and attribute information of the hierarchies from the highest hierarchy to the hierarchy specified by the decoding target LoD depth setting unit 401 from the stored bitstream. The coded data extraction unit 402 supplies the extracted coded data of the position information to the position information decoding unit 403. The coded data extraction unit 402 supplies the extracted coded data of the attribute information to the attribute information decoding unit 404.
[0305] The position information decoding unit 403 obtains the encoded data of the position information provided by the encoded data extraction unit 402. The position information decoding unit 403 decodes the encoded data of the position information and generates position information (decoding result). The decoding method adopted herein may be any appropriate method similar to the method used in the case of the position information decoding unit 202 of the encoding device 200. The position information decoding unit 403 provides the generated position information (decoding result) to the attribute information decoding unit 404 and the point cloud generation unit 405.
[0306] The attribute information decoding unit 404 obtains the encoded data of the attribute information provided by the encoded data extraction unit 402. The attribute information decoding unit 404 obtains the position information (decoding result) provided by the position information decoding unit 403. The attribute information decoding unit 404 uses this position information (decoding result) to decode the encoded data of the attribute information using the method described above in <1. Scalable Decoding> of the present technology, and generates attribute information (decoding result). The attribute information decoding unit 404 provides the generated attribute information (decoding result) to the point cloud generation unit 405.
[0307] The point cloud generation unit 405 receives the position information (decoded result) provided by the position information decoding unit 403. The point cloud generation unit 405 also receives the attribute information (decoded result) provided by the attribute information decoding unit 404. The point cloud generation unit 405 generates a point cloud (decoded result) using the position information (decoded result) and the attribute information (decoded result). The point cloud generation unit 405 outputs the generated point cloud (decoded result) data to the outside of the decoding device 400.
[0308] With this configuration, decoding device 400 can correctly derive a predicted value for attribute information and correctly perform inverse layering of the attribute information associated with the hierarchical structure of position information. This allows decoding device 400 to perform inverse quantization independently of any lower layers. In other words, decoding device 400 can more easily perform scalable decoding of attribute information.
[0309] Note that these processing units (decoding target LoD depth setting unit 401 to point cloud generation unit 405) have any appropriate configuration. For example, each processing unit can be formed with a logic circuit that implements the above-mentioned processing. In addition, each processing unit can also include, for example, a CPU, ROM, RAM, etc., and use them to execute programs to perform the above-mentioned processing. Each processing unit can of course have two configurations, and use logic circuits to perform some of the above-mentioned processing, and perform other processing by executing programs. The configurations of the various processing units can be independent of each other. For example, one processing unit can perform some of the above-mentioned processing using logic circuits, while other processing units perform the above-mentioned processing by executing programs. In addition, some other processing units can perform the above-mentioned processing using logic circuits and by executing programs.
[0310] <Attribute Information Decoding Unit>
[0311] Figure 24 The attribute information decoding unit 404 ( Figure 23 ) is a block diagram of a typical example configuration. Note that Figure 24 The main components and aspects such as processing units and data flow are shown, but Figure 24 Not all components and aspects are necessarily shown. That is, in the attribute information decoding unit 404, there may be Figure 24 The processing units shown as blocks in the Figure 24 Processing or data flow not shown as arrows etc.
[0312] like Figure 24 As shown, the attribute information decoding unit 404 includes a decoding unit 421, an inverse quantization unit 422, and an inverse layering processing unit 423.
[0313] The decoding unit 421 obtains the encoded data of the attribute information provided to the attribute information decoding unit 404. The decoding unit 421 decodes the encoded data of the attribute information and generates attribute information (decoding result). The decoding method used herein may be the same as that of the encoding unit 213 ( Figure 15 ) is compatible with the encoding method used. Furthermore, the generated attribute information (decoded result) corresponds to the attribute information before encoding, is the difference between the attribute information and the predicted value as described in the first embodiment, and has been quantized. The decoding unit 421 supplies the generated attribute information (decoded result) to the inverse quantization unit 422.
[0314] Note that when the encoded data of the attribute information includes information related to the quantization weight (e.g., quantization weight derivation function definition parameters and quantization weight), the decoding unit 421 also generates information related to the quantization weight (decoding result). The decoding unit 421 provides the generated information related to the quantization weight (decoding result) to the inverse quantization unit 422. The decoding unit 421 can also use the generated information related to the quantization weight (decoding result) to decode the encoded data of the attribute information.
[0315] The inverse quantization unit 422 acquires the attribute information (decoding result) supplied from the decoding unit 421. The inverse quantization unit 422 inversely quantizes the attribute information (decoding result). In doing so, the inverse quantization unit 422 performs inverse quantization using the present technique described above in <1. Scalable Decoding>, which is performed by the quantization unit 212 ( Figure 15 ) is an inverse process of quantization performed by . That is, the inverse quantization unit 422 inversely quantizes the attribute information (decoding result) using the quantization weights of each hierarchy as described above in <1. Scalable decoding>.
[0316] Note that when the information on the quantization weight (decoding result) is supplied from the decoding unit 421, the inverse quantization unit 422 also acquires the information on the quantization weight (decoding result). In this case, the inverse quantization unit 422 inversely quantizes the attribute information (decoding result) using the quantization weight specified by the information on the quantization weight (decoding result) or the quantization weight derived based on the information on the quantization weight (decoding result).
[0317] The inverse quantization unit 422 supplies the inversely quantized attribute information (decoding result) to the inverse layering processing unit 423 .
[0318] The inverse layering processing unit 423 acquires the inversely quantized attribute information (decoded result) supplied from the inverse quantization unit 422. As described above, this attribute information is a difference value. The inverse layering processing unit 423 also acquires the position information (decoded result) supplied from the position information decoding unit 403. The inverse layering processing unit 423 uses the position information (decoded result) to perform inverse layering of the acquired attribute information (difference value), which is performed by the layering processing unit 211 ( Figure 15) performs inverse processing of the layering performed by . In doing so, the inverse layering processing unit 423 adopts the present technology described above in <1. Scalable decoding> to perform inverse layering. That is, the inverse layering processing unit 423 uses the position information to derive a prediction value, and restores the attribute information by adding the prediction value to the difference value. In this way, the inverse layering processing unit 423 uses the position information to inversely layer the attribute information having a hierarchical structure similar to that of the position information. The inverse layering processing unit 423 provides the inversely layered attribute information as a decoding result to the point cloud generation unit 405 ( Figure 23 ).
[0319] By performing inverse quantization and inverse layering as described above, attribute information decoding unit 404 can correctly derive a predicted value for the attribute information and correctly perform inverse layering on the hierarchical structure of attribute information associated with the hierarchical structure of position information. Furthermore, inverse quantization can be performed independently of any lower hierarchical levels. In other words, decoding apparatus 400 can more easily perform scalable decoding of attribute information.
[0320] Note that these processing units (decoding unit 421 to inverse layered processing unit 423) have any appropriate configuration. For example, each processing unit can be formed with a logic circuit that implements the above-mentioned processing. In addition, each processing unit can also include, for example, a CPU, ROM, RAM, etc., and use them to execute a program to perform the above-mentioned processing. Each processing unit can certainly have two configurations, and utilize logic circuits to perform some of the above-mentioned processing, and perform other processing by executing a program. The configurations of each processing unit can be independent of each other. For example, one processing unit can utilize logic circuits to perform some of the above-mentioned processing, while other processing units perform the above-mentioned processing by executing a program. In addition, some other processing units can utilize logic circuits and perform the above-mentioned processing by executing a program.
[0321] <Reverse Layering Processing Unit>
[0322] Figure 25 is a diagram showing the reverse layering processing unit 423 ( Figure 24 ) is a block diagram of a typical example configuration. Note that Figure 25 The main components and aspects such as processing units and data flow are shown, but Figure 25 Not all components and aspects are necessarily shown. That is, in the inverse layering processing unit 423, there may be Figure 25 The processing units shown as blocks in the Figure 25 Processing or data flow not shown as arrows etc.
[0323] like Figure 25As shown, the inverse hierarchical processing unit 423 includes an individual hierarchical processing unit 441-1, an individual hierarchical processing unit 441-2, .... The individual hierarchical processing unit 441-1 performs processing on the highest hierarchical level (0) of the attribute information. The individual hierarchical processing unit 441-2 performs processing on the hierarchical level (1) one level lower than the highest hierarchical level of the attribute information. When there is no need to distinguish these individual hierarchical processing units 441-1, 441-2, ..., they are referred to as individual hierarchical processing units 441. That is, the inverse hierarchical processing unit 423 includes the same number of individual hierarchical processing units 441 as the number of hierarchical levels of the attribute information that can be processed.
[0324] The individual level processing unit 441 performs processing related to inverse layering of the attribute information (difference value) of the current level (current LoD) as the processing target. For example, the individual level processing unit 441 derives a predicted value for the attribute information of each point in the current level, and adds the predicted value to the difference value to derive the attribute information of each point in the current level. In doing so, the individual level processing unit 441 performs processing using the present technology described above in <1. Scalable Decoding>.
[0325] The individual hierarchy processing unit 441 includes a prediction unit 451 , an arithmetic unit 452 , and a merge processing unit 453 .
[0326] The prediction unit 451 obtains the attribute information L(N) of the reference point of the current layer, provided from the inverse quantization unit 422 or the individual layer processing unit 441 of the previous layer. The prediction unit 451 also obtains the position information (decoding result) provided from the position information decoding unit 403. The prediction unit 451 uses the position information and the attribute information to derive the predicted value P(N) of the attribute information of the prediction point. In doing so, the prediction unit 451 derives the predicted value P(N) using the method of the present technology described above in <1. Scalable Decoding>. The prediction unit 451 provides the derived predicted value P(N) to the arithmetic unit 452.
[0327] The arithmetic unit 452 obtains the difference value D(N) of the attribute information of the prediction point of the current hierarchy supplied from the inverse quantization unit 422. The arithmetic unit 452 also obtains the predicted value P(N) of the attribute information of the prediction point of the current hierarchy supplied from the prediction unit 451. The arithmetic unit 452 adds the difference value D(N) to the predicted value P(N) to generate (restore) the attribute information H(N) of the prediction point. The arithmetic unit 452 supplies the generated attribute information H(N) of the prediction point to the merge processing unit 453.
[0328] The merge processing unit 453 obtains the attribute information L(N) of the reference point of the current hierarchy, supplied from the inverse quantization unit 422 or the individual hierarchy processing unit 441 of the previous level. The merge processing unit 453 also obtains the attribute information H(N) of the prediction point of the current hierarchy, supplied from the arithmetic unit 452. The merge processing unit 453 merges the obtained attribute information L(N) of the reference point and the attribute information H(N) of the prediction point. Thus, attribute information of all points in the hierarchy from the highest level to the current level is generated. The merge processing unit 453 supplies the generated attribute information of all points as the attribute information L(N+1) of the reference point of the hierarchy one level lower than the current level to the individual hierarchy processing unit 441 of the next level (the prediction unit 451 and the merge processing unit 453).
[0329] Note that, in the case of the individual hierarchy processing unit 441 which is the lowest hierarchy of the current hierarchy to be decoded, the merging processing unit 453 supplies the attribute information of all points generated as the decoding result to the point cloud generation unit 405 ( Figure 23 ).
[0330] By performing inverse layering as described above, the inverse layering processing unit 423 can correctly derive the predicted value of the attribute information and perform correct inverse layering on the hierarchical structure of the attribute information associated with the hierarchical structure of the position information. That is, the decoding device 400 can more easily perform scalable decoding of the attribute information.
[0331] Note that these processing units (prediction unit 451 to merging processing unit 453) have any appropriate configuration. For example, each processing unit can be formed with a logic circuit that implements the above-mentioned processing. In addition, each processing unit can also include, for example, a CPU, ROM, RAM, etc., and use them to execute programs to perform the above-mentioned processing. Each processing unit can of course have two configurations, and use logic circuits to perform some of the above-mentioned processing, and perform other processing by executing programs. The configurations of the various processing units can be independent of each other. For example, one processing unit can perform some of the above-mentioned processing using logic circuits, while other processing units perform the above-mentioned processing by executing programs. In addition, some other processing units can perform the above-mentioned processing using logic circuits and by executing programs.
[0332] <Decoding Process Flow>
[0333] Next, the processing performed by the decoding device 400 is described. The decoding device 400 decodes the encoded data of the point cloud by performing the decoding processing. Figure 26 The flowchart shown describes an example flow of the decoding process.
[0334] When the decoding process starts, in step S401 , the decoding target LoD depth setting unit 401 of the decoding device 400 sets the depth of the LoD to be decoded.
[0335] In step S402 , the encoded data extraction unit 402 acquires and saves the bit stream, and extracts the encoded data of the position information and the attribute information down to the LoD depth set in step S401 .
[0336] In step S403 , the position information decoding unit 403 decodes the encoded data of the position information extracted in step S402 , and generates position information (decoding result).
[0337] In step S404, the attribute information decoding unit 404 decodes the encoded data of the attribute information extracted in step S402 and generates attribute information (decoding result). In doing so, the attribute information decoding unit 404 uses the present technology described above in <1. Scalable Decoding> to perform processing. The attribute information decoding process will be described in detail later.
[0338] In step S405 , the point cloud generation unit 405 generates and outputs a point cloud (decoded result) using the position information (decoded result) generated in step S403 and the attribute information (decoded result) generated in step S404 .
[0339] When the processing in step S405 is completed, the processing result is decoded.
[0340] By performing the processing of each of the above steps, decoding device 400 can correctly derive the predicted value of the attribute information and correctly perform inverse layering of the hierarchical structure of the attribute information associated with the hierarchical structure of the position information. In doing so, decoding device 400 can also perform inverse quantization that is independent of any lower layers. In other words, decoding device 400 can more easily perform scalable decoding of the attribute information.
[0341] <Flow of attribute information decoding process>
[0342] Next, refer to Figure 27 The flowchart shown describes Figure 26 An example flow of the attribute information decoding process performed in step S404.
[0343] When the attribute information decoding process starts, in step S421, the decoding unit 421 of the attribute information decoding unit 404 decodes the encoded data of the attribute information and generates attribute information (decoding result). This attribute information (decoding result) has been quantized as described above.
[0344] In step S422, the inverse quantization unit 422 performs inverse quantization processing to inversely quantize the attribute information (decoding result) generated in step S421. In doing so, the inverse quantization unit 422 employs the present technique described above in <1. Scalable Decoding> to perform inverse quantization. The inverse quantization processing will be described in detail later. Note that, as described above, the attribute information to be inversely quantized is the difference between the attribute information and the predicted value.
[0345] In step S423, the inverse layering processing unit 423 performs inverse layering to inversely layer the attribute information (difference value) inversely quantized in step S422, and derives attribute information for each point. In doing so, the inverse layering processing unit 423 employs the present technique described above in <1. Scalable Decoding> to perform inverse layering. The inverse layering processing will be described in detail later.
[0346] When the processing in step S423 is completed, the attribute information decoding processing ends, and the processing returns to Figure 26 .
[0347] By performing the processing of each step described above, the attribute information decoding unit 404 can correctly derive the predicted value of the attribute information and correctly perform inverse quantization on the hierarchical structure of the attribute information associated with the hierarchical structure of the position information. In this way, the decoding device 400 can also perform inverse quantization that is independent of any lower hierarchical structure. In other words, the decoding device 400 can more easily perform scalable decoding of the attribute information.
[0348] <Flow of Inverse Quantization Processing>
[0349] Next, refer to Figure 28 The flowchart shown describes Figure 27 In this paper, the inverse quantization process is performed in step S422. Figure 12 The case of method 1-1-2 described in the fourth row from the top of the shown table will be described.
[0350] When the inverse quantization process starts, in step S441 , the inverse quantization unit 422 acquires quantization weight derivation function definition parameters generated by decoding the encoded data.
[0351] In step S442 , the inverse quantization unit 422 defines a weight derivation function based on the function definition parameters acquired in step S441 .
[0352] In step S443 , the inverse quantization unit 422 derives the quantization weight for each level (LoD) using the weight derivation function defined in step S442 .
[0353] In step S444 , the inverse quantization unit 422 inversely quantizes the attribute information (decoding result) using the quantization weight derived in step S443 .
[0354] When the processing in step S444 is completed, the inverse quantization processing ends, and the processing returns to Figure 27 .
[0355] By performing the inverse quantization process as described above, the inverse quantization unit 422 can perform inverse quantization that does not depend on any lower layer. That is, the decoding device 400 can more easily perform scalable decoding of attribute information.
[0356] <Flow of De-stratification Processing>
[0357] Next, refer to Figure 29 The flowchart shown in Figure 27 In this paper, the reverse layering process is performed in step S423. Figure 6 The case of method 1-1 described in the second row from the top of the shown table will be described.
[0358] When the inverse layering process begins, in step S461, the inverse layering unit 423 uses the position information (decoding result) to layer the attribute information (decoding result). Specifically, the inverse layering unit 423 performs a process similar to the layering process described in the first embodiment, associating the hierarchical structure of the attribute information at the layer to be decoded with the hierarchical structure of the position information. Consequently, the attribute information (decoding result) at each point is associated with the position information (decoding result). Based on this correspondence, the following process is executed.
[0359] In step S462 , the individual hierarchy processing unit 441 of the inverse hierarchical processing unit 423 determines the highest hierarchy (LoD) as the current hierarchy (current LoD) that is a processing target.
[0360] In step S463, the prediction unit 451 derives a predicted value for each prediction point. That is, the prediction unit 451 derives a predicted value for the attribute information of the prediction point using the attribute information of the reference point corresponding to each prediction point.
[0361] For example, as mentioned above Figure 4As described, the prediction unit 451 weights the attribute information of each reference point corresponding to the prediction point using a quantization weight (α(P, Q(i, j))) corresponding to the inverse of the distance between the prediction point and the reference point, and integrates the attribute information by performing the arithmetic operation shown in the above expression (1). In this way, the prediction unit 451 derives a predicted value. In doing so, the prediction unit 451 uses the position information of the resolution of the current layer (current LoD) to derive and adopt the distance between the prediction point and the reference point.
[0362] In step S464 , the arithmetic unit 452 restores the attribute information of the predicted point by adding the predicted value derived in step S463 and the difference value of the predicted point.
[0363] In step S465 , the merging processing unit 453 merges the derived attribute information of the prediction point and the attribute information of the reference point to generate attribute information of all points in the current hierarchy.
[0364] In step S466, the individual hierarchy processing unit 441 determines whether the current hierarchy (current LoD) is the lowest hierarchy. If it is determined that the current hierarchy is not the lowest hierarchy, the process proceeds to step S467.
[0365] In step S467, the individual level processing unit 441 sets the next lower level (LoD) as the current level (current LoD) that is the processing target. When the processing in step S467 is completed, the processing returns to step S463. That is, the processing in each step from step S463 to step S467 is performed for each level.
[0366] Furthermore, if it is determined in step S466 that the current level is the lowest level, the reverse layering process ends and the process returns to step S466. Figure 27 .
[0367] By performing the processing of each of the above steps, the de-layering processing unit 423 can correctly derive the predicted value of the attribute information and correctly perform de-layering on the hierarchical structure of the attribute information associated with the hierarchical structure of the position information. In other words, the decoding device 400 can more easily perform scalable decoding of the attribute information.
[0368] Note that when using Figure 6 When the method 1-1' described in the third row from the top of the table shown in Figure 6 When the method 1-2 described in the fourth row from the top of the table shown is basically the same as Figure 29 The reverse hierarchical processing is performed by a process similar to the process in the flowchart shown. Therefore, this article does not describe these cases.
[0369] <Flow of Inverse Quantization Processing>
[0370] Next, refer to Figure 30 The flowchart shown in Figure 12 An example flow of the inverse quantization process of method 1-2-2 described in the seventh row from the top of the table shown will be described.
[0371] When the inverse quantization process starts, in step S481 , the inverse quantization unit 422 acquires the quantization weight of each level (LoD), which has been generated by decoding the encoded data.
[0372] In step S482 , the inverse quantization unit 422 inversely quantizes the attribute information (decoding result) using the quantization weight acquired in step S481 .
[0373] When the processing in step S482 is completed, the inverse quantization processing ends, and the processing returns to Figure 27 .
[0374] By performing the inverse quantization process as described above, the inverse quantization unit 422 can perform inverse quantization that does not depend on any lower layer. That is, the decoding device 400 can more easily perform scalable decoding of attribute information.
[0375] <4.DCM>
[0376] <DCM for location information>
[0377] As described above, the position information is converted into an octree. In recent years, as disclosed in Non-Patent Document 6, for example, a direct coding mode (DCM) has been proposed for encoding the relative distance from a node to each leaf in an octree having only a certain number of leaves or less (a node at the lowest level in the octree, or a point with the highest resolution).
[0378] For example, when converting voxel data into an octree, if a processing target node satisfies predetermined conditions and is determined to be sparse, the DCM is used to calculate and encode the relative distance from the processing target node to each leaf (in each of the x, y, and z directions) that directly or indirectly belongs to the processing target node.
[0379] Note that a "directly owned node" refers to a node that is suspended from another node in the tree structure. For example, a node that directly belongs to a processing target node is a node that belongs to the processing target node and is one level lower than the processing target node (a so-called child node). At the same time, an "indirectly owned node" refers to a node that is suspended from another node via another node in the tree structure. For example, a node that indirectly belongs to a processing target node is a node that belongs to the processing target node via another node and is at least two levels lower than the processing target node (a so-called grandchild node).
[0380] For example, Figure 31 As shown, when it is determined that node n0 (LoD=n) is sparse during conversion to an octree, leaf p0 belonging to (hanging on) node n0 is identified, and the difference (in position) between node n0 and leaf p0 is derived. That is, encoding of position information of nodes of intermediate resolution (nodes at levels (LoD) between LoD=n and LoD=1) is skipped.
[0381] Similarly, when it is determined that node n2 (LoD=m) is sparse, leaf p2 belonging to (hanging on) node n2 is identified, and the difference (in position) between node n2 and leaf p2 is derived. That is, encoding of position information of nodes of intermediate resolution (nodes at levels (LoD) between LoD=m and LoD=1) is skipped.
[0382] By adopting DCM in this way, the generation of sparse nodes can be skipped, thereby reducing the increase in processing load for conversion to an octree.
[0383] <Hierarchy of DCM attribute information>
[0384] The layering of attribute information can be done in Figure 32 The node formation of the intermediate resolution shown is performed under the assumption that the position information is in the octree, so that when this DCM is adopted, scalable decoding can also be performed on the attribute information. That is, assuming that the DCM is not adopted, encoding and decoding according to this technology can be performed on the attribute information.
[0385] exist Figure 32 In the example case shown, the layering of attribute information is performed, assuming that there are nodes of intermediate resolution (node n0', node n0", and node n0'" indicated by a white circle between node n0 and leaf p0), and a node of intermediate resolution (node n2') indicated by a white circle between node n2 and leaf p2. That is, attribute information corresponding to these nodes is generated.
[0386] In this way, attribute information of nodes whose position information is not encoded according to DCM is also encoded. Therefore, attribute information of any desired level (resolution) can be easily restored to the highest resolution without being decoded. In other words, even when using DCM, easier scalable decoding of attribute information can be achieved.
[0387] For example, in the case of the encoding device 200, the attribute information encoding unit 204 ( Figure 14 ) for those with Figure 31 The hierarchical structure of the location information similar to the example shown is hierarchical into the attribute information Figure 32The hierarchical structure is similar to the example shown, and then the attribute information is quantized and encoded. That is, the hierarchical processing unit 211 ( Figure 15 ) generates Figure 32 The example shown is similar to the hierarchical structure of attribute information.
[0388] In addition, in the case of the decoding device 400, for example, the attribute information decoding unit 404 ( Figure 23 ) Decode the encoded data of the attribute information of the desired resolution (to the desired level). As described above, in the encoded data, the attribute information has Figure 32 Therefore, when decoding to the highest resolution is not required, the attribute information decoding unit 404 (the inverse layer processing unit 423 ( Figure 24 )) can reverse the hierarchy and restore attribute information at any desired level.
[0389] <Node resolution setting 1>
[0390] Note that when position information is referenced during processing related to such attribute information, nodes of intermediate resolutions whose encoding of position information is skipped by DCM can be represented by, for example, the resolution of the hierarchy.
[0391] <In the case of encoding>
[0392] For example, a case where the level LoD=m is a processing target in encoding of point cloud data will now be described. The position information encoding unit 201 applies DCM to sparse nodes and generates a point cloud having the same Figure 31 The position information of the hierarchical structure similar to the example shown in is encoded. That is, the encoded data of the position information does not include the position information of the node n0 ″ having LoD=m associated with the node n0 .
[0393] On the other hand, Figure 33 In the example shown, the attribute information encoding unit 204 generates attribute information of the node n0" of LoD=m that is associated with the node n0 but skipped by the DCM. In doing so, the attribute information encoding unit 204 (hierarchical processing unit 211) uses the resolution of the hierarchy (LoD=m) to represent the node n0". That is, the attribute information encoding unit 204 (hierarchical processing unit 211) generates attribute information corresponding to the node n0" of the resolution of the hierarchy (LoD=m).
[0394] In this case, the method for generating attribute information corresponding to the node n0" at the resolution of the level (LoD=m) may be any appropriate method. For example, attribute information corresponding to the node n0 or the leaf p0 may be used to generate attribute information corresponding to the node n0". In this case, for example, the attribute information of the node n0 or the leaf p0 may be converted based on the resolution difference between the node n0 or the leaf p0 and the node n0" to generate the attribute information corresponding to the node n0". Alternatively, for example, the attribute information of a node located nearby may be used to generate the attribute information corresponding to the node n0".
[0395] In doing so, the encoding device 200 can encode the attribute information corresponding to the node n0 ″ of the resolution of the hierarchy (LoD=m).
[0396] <In case of decoding>
[0397] Next, a description is given of a case where the level LoD=m is a processing target in the decoding of the encoded data of the point cloud data. The encoded data of the attribute information encoded in the above manner includes the attribute information corresponding to the node n0" of the resolution of the level (LoD=m). Therefore, the attribute information decoding unit 404 can easily restore the attribute information corresponding to the node n0" of the resolution of the level (LoD=m) by decoding the encoded data of the attribute information to the level (LoD=m).
[0398] That is, even when DCM is adopted, the decoding apparatus 400 can more easily decode attribute information in a scalable manner.
[0399] Note that, regarding the position information, the position information corresponding to the node n0" of the level (LoD=m) is not encoded. Therefore, Figure 34 In the example shown, the position information decoding unit 403 can generate position information corresponding to the node n0" at the resolution of the level (LoD=m) by restoring the position information corresponding to the leaf p0 and converting the resolution from the final resolution (highest resolution) to the resolution of the level (LoD=m).
[0400] <Node resolution setting 2>
[0401] Furthermore, for example, a node of an intermediate resolution whose encoding of position information is skipped by DCM may be represented with the highest resolution, which is the resolution of a leaf (LoD=1).
[0402] <In the case of encoding>
[0403] For example, a case where level LoD=m is a processing target in encoding of point cloud data is now described. In this case, as Figure 35In the example shown, the attribute information encoding unit 204 generates attribute information of the node n0" of LoD=m that is associated with the node n0 but skipped by the DCM. In doing so, the attribute information encoding unit 204 (hierarchical processing unit 211) uses the resolution of the leaf (LoD=l) to represent the node n0". That is, the attribute information encoding unit 204 (hierarchical processing unit 211) generates attribute information corresponding to the node n0" of the resolution of the leaf (LoD=l) (that is, the same attribute information as the attribute information corresponding to the leaf p0).
[0404] In doing so, the encoding device 200 can encode the attribute information corresponding to the node n0 ″ of the resolution of the leaf (LoD=1).
[0405] <In case of decoding>
[0406] Next, a case is described where the level LoD=m is a processing target in the decoding of the encoded data of the point cloud data. The encoded data of the attribute information encoded in the above manner includes the attribute information corresponding to the node n0" at the resolution of the leaf (LoD=1). Therefore, the attribute information decoding unit 404 can easily restore the attribute information corresponding to the node n0" at the resolution of the leaf (LoD=1) by decoding the encoded data of the attribute information to the level (LoD=m).
[0407] That is, even when DCM is adopted, the decoding apparatus 400 can more easily decode attribute information in a scalable manner.
[0408] Note that as for the position information, the node n0" of the level (LoD=m) is not encoded. Therefore, Figure 36 In the illustrated example, the position information decoding unit 403 may restore the position information corresponding to the leaf p0 and may use the restored position information as the position information corresponding to the node n0 ″ of the resolution of the leaf (LoD=1).
[0409] <5. Quantitative Weights>
[0410] <Example 1 of Quantization Weight Derivation Function>
[0411] In <1. Scalable Decoding>, refer to Figure 12 and other figures describe examples of quantization weights. For example, in <Quantization Weight: Method 1-1-1>, an example of quantization weight derivation using a predetermined quantization weight derivation function is described. For example, the quantization weight derivation function used to derive the quantization weight may be any appropriate function, and may be, for example, Figure 37 The functions of Expression (11) and Expression (12) shown in A. Expression (11) and Expression (12) are also shown below.
[0412] [Mathematical formula 3]
[0413]
[0414] [Formula 4]
[0415] if(i == LoDCount)
[0416] QuantizationWeight[i]=1
[0417] …(12)
[0418] In expressions (11) and (12), LoDCount represents the lowest level (highest resolution level) of the hierarchical structure of attribute information. In addition, QuantizationWeight[i] represents the quantization weight (the quantization weight is a value set for each level) of level i (the (i-1)th level from the highest level (lowest resolution level)) in the hierarchical structure. In the expressions, pointCount represents the total number of nodes in the hierarchical structure, that is, the total number of points. In addition, predictorCount[i] represents the number of prediction points of level i, that is, the number of points from which prediction values are derived in level i. The integrated value of the number of prediction points of each level from level 0 (highest level) to level LoDCount (lowest level) (the rightmost numerator of expression (11)) indicates the total number of points (pointCount).
[0419] As shown in Expression (11), in this case, the ratio between the total number of points of level i and the number of prediction points is set in the quantization weight (QuantizationWeight[i]) of level i. However, as shown in Expression (12), in the case of the lowest level (i == LoDCount), "1" is set in the quantization weight (QuantizationWeight[i]).
[0420] By using such a quantization weight derivation function to derive the quantization weight (QuantizationWeight[i]), the encoding device 200 and the decoding device 400 can derive the quantization weight independently of any lower layer. Therefore, as described above in <Quantization Weight: Method 1>, quantization and inverse quantization can be performed independently of any lower layer. That is, since it is not necessary to refer to information related to any lower layer as described above, the encoding device 200 and the decoding device 400 can more easily achieve scalable decoding.
[0421] Furthermore, using the quantization weight derivation function (Expressions (11) and (12)), the encoding device 200 and the decoding device 400 can derive quantization weights such that higher layers have larger quantization weights. Therefore, it is possible to reduce the reduction in encoding and decoding accuracy.
[0422] Note that, as described above in <Quantization Weight: Method 1-1-1>, the encoding device 200 and the decoding device 400 can share the quantization weight derivation function (Expression (11) and Expression (12)). By doing so, the encoding device 200 and the decoding device 400 can derive the same quantization weights. Furthermore, since it is not necessary to transmit information for sharing the quantization weight derivation function, it is possible to reduce a decrease in encoding efficiency and an increase in the load of the encoding and decoding processes.
[0423] In addition, as described above in <Quantization Weight: Method 1-1-2>, the encoding device 200 can define the quantization weight derivation function (Expression (11) and Expression (12)) at the time of encoding. In this case, the encoding device 200 can transmit the defined quantization weight derivation function or the parameters (quantization weight derivation function definition parameters) from which the decoding device 400 can derive the quantization weight derivation function to the decoding device 400.
[0424] For example, the encoding device 200 may transmit such quantization weight derivation parameters or such quantization weight derivation function definition parameters included in a bitstream including the encoded data of the point cloud to the decoding device 400 .
[0425] For example, in the above reference Figure 20 In the quantization process described in the flowchart shown, the quantization unit 212 defines the above expression (11) or expression (12) as a quantization weight derivation function (step S271). Then, the quantization unit 212 uses the quantization weight derivation function to derive the quantization weight of each level (LoD) (step S272). Then, the quantization unit 212 quantizes the difference in attribute information using the derived quantization weight of each LoD (step S273). Then, the quantization unit 212 provides the quantization weight derivation function or the quantization weight derivation function definition parameter to the encoding unit 213, and causes the encoding unit 213 to perform encoding (step S274). That is, the quantization unit 212 sends the quantization weight derivation function or the quantization weight derivation function definition parameter to the decoding side.
[0426] Note that the encoding device 200 may transmit the quantization weight derivation function or the quantization weight derivation function definition parameters to the decoding device 400 as data or a file independent of the bit stream including the encoded data of the point cloud.
[0427] Since such a quantization weight derivation function or parameter is transmitted from the encoding device 200 to the decoding device 400 in the above manner, the encoding device 200 and the decoding device 400 can share the quantization weight derivation function (Expression (11) or Expression (12)). Therefore, the encoding device 200 and the decoding device 400 can derive the same quantization weight.
[0428] Furthermore, as described above in <Quantization Weight: Method 1-2>, the encoding device 200 and the decoding device 400 may share the quantization weight (QuantizationWeight[i]) derived using the quantization weight derivation function (Expression (11) and Expression (12)). For example, as described above in <Quantization Weight: Method 1-2-2>, the encoding device 200 may derive the quantization weight using the quantization weight derivation function during encoding and transmit the quantization weight to the decoding device 400.
[0429] For example, the decoding apparatus 200 may transmit such quantization weights included in a bitstream including encoded data of a point cloud to the decoding apparatus 400 .
[0430] For example, in the above reference Figure 22 In the quantization process described in the flowchart shown, the quantization unit 212 uses the above quantization weight derivation function (Expression (11) and Expression (12)) to derive the quantization weight for each level (LoD) (step S321). Then, the quantization unit 212 quantizes the difference value of the attribute information using the derived quantization weight for each LoD (step S322). Then, the quantization unit 212 provides the quantization weight used in quantization to the encoding unit 213, and causes the encoding unit 213 to perform encoding (step S323). That is, the quantization unit 212 transmits the quantization weight to the decoding side.
[0431] Note that the encoding device 200 may transmit the quantization weights to the decoding device 400 as data or a file independent of a bit stream including encoded data of a point cloud.
[0432] Since the quantization weights are transmitted from the encoding device 200 to the decoding device 400 in the above manner, the encoding device 200 and the decoding device 400 can share the quantization weights. Therefore, the encoding device 200 and the decoding device 400 can easily derive the quantization weights independently of any lower layers. Therefore, the encoding device 200 and the decoding device 400 can more easily implement scalable decoding. In addition, since the decoding device 400 does not need to use a quantization weight derivation function to derive the quantization weights, the increase in the decoding process load can be reduced.
[0433] <Example 2 of Quantization Weight Derivation Function>
[0434] For example, the quantization weight derivation function can be as follows Figure 37 The function of expression (13) shown in B. Expression (13) is also shown below.
[0435] [Formula 5]
[0436]
[0437] In expression (13), as in the case of expressions (11) and (12), LoDCount represents the lowest level of the hierarchical structure of attribute information. In addition, QuantizationWeight[i] represents the quantization weight of level i in the hierarchical structure. In addition, predictorCount[i] represents the number of prediction points of level i, that is, the number of points of predicted values derived in level i. The integrated value of the number of prediction points of each level from level i (processing target level) to level LoDCount (lowest level) (the rightmost numerator of expression (13)) indicates the number of points below level i.
[0438] As shown in Expression (13), in this case, the ratio between the number of points below level i and the number of prediction points of level i is set in the quantization weight of level i (QuantizationWeight[i]).
[0439] When the quantization weight derivation function of Expression (13) is adopted, the encoding device 200 and the decoding device 400 can derive the quantization weight independently of any lower layer, as in the case of the quantization weight derivation functions of Expressions (11) and (12). Therefore, as described above in <Quantization Weight: Method 1>, quantization and inverse quantization can be performed independently of any lower layer. That is, as described above, since there is no need to refer to information related to any lower layer, the encoding device 200 and the decoding device 400 can more easily achieve scalable decoding.
[0440] Furthermore, when the quantization weight derivation function of Expression (13) is employed, the encoding device 200 and the decoding device 400 can derive the quantization weight so that the higher layer has a larger quantization weight, as in the case of employing the quantization weight derivation function of Expression (11) and Expression (12). Therefore, it is possible to reduce the reduction in the accuracy of encoding and decoding.
[0441] The value of the quantization weight (QuantizationWeight[i]) is set based on the following concept: the more frequently a point is referenced by other points, the more important it is (the greater the impact on subjective image quality), such as Figure 5As shown. In addition, usually, the prediction points of higher levels are likely to be referenced more frequently by other points. Therefore, a larger quantization weight is set for the prediction points of higher levels. Figure 12 As described, when quantization weights are set for respective levels (LoDs), based on this concept, a larger quantization weight is set for a higher LoD.
[0442] In the case of a point cloud with sufficiently dense points, or when there are basically other points near each point, the number of points generally increases as the resolution increases. Generally, the degree of increase in the number of points is high enough, and the number of predicted points (predictorCount[i]) increases monotonically toward the lowest level. Therefore, whether the quantization weight derivation function of Expression (11) and Expression (12) or the quantization weight derivation function of Expression (13) is adopted, the higher the level, the larger the quantization weight (QuantizationWeight[i]) of level i tends to be.
[0443] On the other hand, in the case of a point cloud with sparse points, or when the probability of other points existing near each point is small, the number of points does not usually increase significantly with increasing resolution. Typically, in higher levels of the hierarchy, a large proportion of points are likely to be designated as predicted points, while in lower levels, the number of predicted points is smaller.
[0444] When the quantization weight derivation functions of Expressions (11) and (12) are used, the larger the quantization weight (QuantizationWeight[i]) of level i, the smaller the number of prediction points (predictorCount[i]) of the level i. Therefore, there is a possibility that a higher level, where a larger proportion of points are designated as prediction points, has a smaller quantization weight than a lower level. In other words, the correspondence between the quantization weight and the number of times the prediction point is referenced may be weakened, and the degradation of subjective image quality due to quantization may increase.
[0445] On the other hand, the number of points below level i increases monotonically toward the highest level. Therefore, when the quantization weight derivation function of Expression (13) is adopted, the quantization weight (QuantizationWeight[i]) of level i is likely to be larger as the level is higher. In other words, it is easy to maintain the tendency of setting larger quantization weights for prediction points (higher levels) that are frequently referenced by other points. Therefore, compared with the case of adopting the weight derivation functions of Expressions (11) and (12), the degradation of subjective image quality due to quantization can be made smaller.
[0446] Note that, as described above in <Quantization Weight: Method 1-1-1>, the encoding device 200 and the decoding device 400 can share the quantization weight derivation function (Expression (13)) in advance. By doing so, the encoding device 200 and the decoding device 400 can derive the same quantization weights. Furthermore, since there is no need to transmit information about the shared quantization weight derivation function, it is possible to reduce a decrease in encoding efficiency and an increase in the load of the encoding and decoding processes.
[0447] In addition, as described above in <Quantization Weight: Method 1-1-2>, the encoding device 200 can define the quantization weight derivation function (Expression (13)) at the time of encoding. In this case, the encoding device 200 can transmit the defined quantization weight derivation function or the parameters (quantization weight derivation function definition parameters) from which the decoding device 400 can derive the quantization weight derivation function to the decoding device 400.
[0448] For example, the encoding device 200 may transmit such a quantization weight derivation function or such a quantization weight derivation function definition parameter included in a bitstream including the encoded data of the point cloud to the decoding device 400 .
[0449] For example, in the above reference Figure 20 In the quantization process described in the flowchart shown, the quantization unit 212 defines the above expression (13) as a quantization weight derivation function (step S271). Then, the quantization unit 212 uses the quantization weight derivation function to derive the quantization weight of each level (LoD) (step S272). Then, the quantization unit 212 quantizes the difference in attribute information using the derived quantization weight of each LoD (step S273). Then, the quantization unit 212 provides the quantization weight derivation function or the quantization weight derivation function definition parameter to the encoding unit 213, and causes the encoding unit 213 to perform encoding (step S274). That is, the quantization unit 212 sends the quantization weight derivation function or the quantization weight derivation function definition parameter to the decoding side.
[0450] Note that the encoding device 200 may transmit the quantization weight derivation function or the quantization weight derivation function definition parameters to the decoding device 400 as data or a file independent of the bit stream including the encoded data of the point cloud.
[0451] Since such a quantization weight derivation function or parameter is transmitted from the encoding device 200 to the decoding device 400 in the above manner, the encoding device 200 and the decoding device 400 can share the quantization weight derivation function (Expression (13)). Therefore, the encoding device 200 and the decoding device 400 can derive the same quantization weight.
[0452] Furthermore, as described above in <Quantization Weight: Method 1-2>, the encoding device 200 and the decoding device 400 may share the quantization weight (QuantizationWeight[i]) derived using the quantization weight derivation function (Expression (13)). For example, as described above in <Quantization Weight: Method 1-2-2>, the encoding device 200 may derive the quantization weight using the quantization weight derivation function during encoding and transmit the quantization weight to the decoding device 400.
[0453] For example, the encoding device 200 may transmit such quantization weights included in a bitstream including the encoded data of the point cloud to the decoding device 400 .
[0454] For example, in the above reference Figure 22 In the quantization process described in the flowchart shown, the quantization unit 212 uses the above quantization weight derivation function (Expression (13)) to derive the quantization weight for each level (LoD) (step S321). The quantization unit 212 then quantizes the difference in attribute information using the derived quantization weight for each LoD (step S322). The quantization unit 212 then provides the encoding unit 213 with the quantization weight used in quantization and causes the encoding unit 213 to perform encoding (step S323). That is, the quantization unit 212 transmits the quantization weight to the decoding side.
[0455] Note that the encoding device 200 may transmit the quantization weights to the decoding device 400 as data or a file independent of a bit stream including encoded data of a point cloud.
[0456] Since such quantization weights are transmitted from the encoding device 200 to the decoding device 400 in the manner described above, the encoding device 200 and the decoding device 400 can share the quantization weights. Therefore, the encoding device 200 and the decoding device 400 can easily derive the quantization weights independently of any lower layers. Therefore, the encoding device 200 and the decoding device 400 can more easily implement scalable decoding. Furthermore, since the decoding device 400 does not need to use a quantization weight derivation function to derive the quantization weights, the increase in the decoding process load can be reduced.
[0457] <6.LoD Generation>
[0458] <LoD Generation During Decoding of Position Information>
[0459] As mentioned above Figure 6As described in <1. Scalable decoding> and other figures, the attribute information is hierarchically arranged to have a hierarchical structure similar to the hierarchical structure of the position information. Therefore, scalable decoding of the attribute information can be performed. Then, prediction points (in other words, reference points) for each level are selected so that the hierarchical structure of the attribute information is set. This process of setting the hierarchical structure is called LoD generation processing or subsampling processing.
[0460] As described above in <2. First embodiment> and <3. Second embodiment>, this LoD generation processing is performed during the encoding / decoding of the attribute information. For example, in the case of the encoding device 200, the point classification unit 231 of the attribute information encoding unit 204 performs this LoD generation processing (step S242). In addition, in the case of the decoding device 400, the inverse hierarchical processing unit 423 of the attribute information decoding unit 404 performs this LoD generation processing (step S461).
[0461] In examples other than this example, for example, the LoD generation processing can be performed when decoding the encoded data of the position information. As the encoded data of the position information is decoded, the hierarchical structure of the position information (which point belongs to which level) becomes clear. Therefore, in the LoD generation processing, information indicating this hierarchical structure of the position information is generated. Based on this information, the attribute information is then hierarchically arranged or inversely hierarchically arranged to be associated with the hierarchical structure of the position information. Since the above processing is performed, the hierarchical structure of the position information can be applied to the attribute information.
[0462] Then, such hierarchical arrangement and inverse hierarchical arrangement are applied to the encoding and decoding of the attribute information. That is, encoding and decoding are performed in which the hierarchical structure of the attribute information is associated with the hierarchical structure of the position information. Therefore, scalable decoding of the attribute information can be performed. In addition, redundant processing can be reduced, and an increase in the load of the attribute information encoding / decoding processing can be reduced.
[0463] Note that in the case of the encoding device 200, this LoD generation processing can be performed during the encoding of the position information.
[0464] <Example of LoD generation processing>
[0465] Now, such LoD generation processing will be described. For example, as the encoded data of the position information is decoded, a hierarchical structure such as an octree is restored. Figure 38 An example of the restored hierarchical structure of the position information is shown. The node 611 indicated by a circle in this hierarchical structure represents the point after hierarchical arrangement. Note that Figure 38 only a part of the hierarchical structure (some nodes 611) is shown. Although Figure 38Nodes 611 - 1 to 611 - 12 are shown in FIG, but when there is no need to distinguish the nodes from each other, they are referred to as node 611 .
[0466] In the direction from the highest level (LoDIdx=0) to the lowest level, position information is decoded for each level. Figure 38 In the example case shown, node 611-1 appears in the highest level (LoDIdx=0). In the next lower level (LoDIdx=1), as the resolution increases, the point corresponding to node 611-1 is split, and a new node 611-2 appears. In the next lower level (LoDIdx=2), the point corresponding to node 611-1 is split, so that new nodes 611-3 and 611-4 appear. The point corresponding to node 611-2 is split, so that new nodes 611-5 to 611-7 appear.
[0467] In the next lower level (LoDIdx=3), the point corresponding to node 611-1 is split, resulting in a new node 611-7. In the next lower level (LoDIdx=4), the point corresponding to node 611-1 is split, resulting in new nodes 611-8 to 611-10. The point corresponding to node 611-7 is split, resulting in new nodes 611-11 and 611-12.
[0468] Each node 611 appearing sequentially in this manner is labeled to indicate the level (LoDIdx) by, for example, using a list of data structures. Figure 39 Shows the Figure 38 The example shown is an example of labeling. This labeling is performed in each level in the order from the highest level to the lowest level in the hierarchical structure of the position information, and a label indicating the value of the level is attached to each newly appearing node.
[0469] For example, after the highest level (LoDIdx=0) is processed and the node 611-1 appears, a label "0" is appended to the node 611-1 in the list, as shown in FIG. Figure 39 As shown in the first row from the top (row "1") in FIG.
[0470] After the next lower level is processed (LoDIdx=1) and node 611-2 newly appears, a label "1" is appended to node 611-2 in the list, as shown in FIG. Figure 39 Since the node 611 - 1 is already labeled, no label is attached to the node 611 - 1 at this level.
[0471] Furthermore, after the next lower level is processed (LoDIdx=2) and nodes 611-3 to 611-6 newly appear, a label "2" is appended to each of nodes 611-3 to 611-6 in the list, as shown in FIG. Figure 39 As shown in the third row from the top (row "3") in . Since nodes 611-1 and 611-2 are already labeled, no label is attached to nodes 611-1 and 611-2 at this level. At the same time, in the hierarchical structure (octree), the nodes (points) are arranged in Morton order. Similarly, in the list, the labels of the nodes 611 are arranged in Morton order. Therefore, in Figure 39 In the example case shown, the label order is "0,2,2,1,2,2,...".
[0472] Furthermore, after the next lower level is processed (LoDIdx=3) and the node 611-7 newly appears, a label "3" is appended to the node 611-7 in the list, as shown in FIG. Figure 39 As shown in the fourth row from the top (row "4") in . Since node 611-1 has already been labeled, no label is attached to node 611-1 at this level. Since node 611-7 appears immediately after node 611-1 in Morton order, Figure 39 In the example case shown, the label order is "0,3,...".
[0473] Furthermore, after the next lower level is processed (LoDIdx=4) and nodes 611-8 to 611-12 newly appear, a label "4" is appended to nodes 611-8 to 611-12 in the list, as shown in FIG. Figure 39 As shown in the fifth row from the top (row "5") in . Since nodes 611-1 and 611-7 are already labeled, no labels are attached to nodes 611-1 and 611-7 at this level. In Morton order, nodes 611-8 to 611-10 appear immediately after node 611-1, and nodes 611-11 and 611-12 appear immediately after node 611-7. Therefore, in Figure 39 In the example case shown, the label sequence is "0,4,4,4,3,4,4,...".
[0474] Since the attribute information is hierarchically mapped and encoded / decoded with reference to the list generated as described above, the attribute information can be encoded / decoded using a hierarchical structure similar to that of the location information. Therefore, scalable decoding of the attribute information can be achieved. Furthermore, redundant processing can be reduced, and the increase in the load of the attribute information encoding / decoding process can be minimized.
[0475] <Case 1 including DCM>
[0476] A case where the DCM is included in the hierarchical structure of the location information is now described. Figure 40 An example of the hierarchical structure of the location information in this case is shown. Figure 40 In the hierarchical structure shown, the point of the node 611-21 that appears at the level LoDIdx = 1 is divided at the level LoDIdx = 3. That is, a new node 611-22 appears at the level LoDIdx = 3. Furthermore, DCM is applied to this node 611-21. DCM is also applied to the node 611-22.
[0477] Such nodes to which DCM is applied may be marked using a list different from the list for nodes to which DCM is not applied. Figure 41 The method shown in Figure 40 The example in the example performs the markup. Figure 41 As shown, the nodes without DCM are Figure 39 , are labeled using List A in a manner similar to the exemplary case shown in . On the other hand, nodes 611-21 and 611-22 to which DCM is applied are labeled using List B which is different from List A. Note that List A and List B are names given to the respective lists for convenience of explanation and may have any names in practice.
[0478] After processing the highest level (LoDIdx=0), the next lower level (LoDIdx=1) is processed, and the node 611-21 to which DCM is applied newly appears. Figure 41 In List B, shown to the right of the dotted line in [ 1 ], a label "1" is attached to node 611-21. DCM is also applied to node 611-22, and in LoDIdx = 1, the relative distance from node 611-21 is encoded. Therefore, a label is attached to node 611-22 at this level. However, since node 611-22 is separated from node 611-21 in LoDIdx = 3, a label "3" is attached to node 611-22.
[0479] Note that the label on node 611-21 and the label on node 611-22 are arranged in the lowest level (highest resolution) Morton order according to the positions of the two nodes. Figure 40 In the illustrated example, the labels are arranged in the order of "1, 3" because the node 611-21 and the node 611-22 are arranged in this order.
[0480] For other nodes 611 to which DCM is applied, labels are set in a similar manner in List B. The label strings in List A and List B are then used for encoding / decoding attribute information. At this stage, the label strings in List A and List B are merged and arranged in Morton order. That is, the labels are arranged to correspond to the nodes arranged in Morton order at the lowest level (highest resolution).
[0481] In addition, the labels in List B (labels on nodes to which DCM is applied) are processed to remove overlapping points. For node 611 to which DCM is applied, the node 611 to be separated is also marked when it first appears. For example, Figure 41 In the illustrated case, marking is performed on both the node 611 - 21 and the node 611 - 22 in the level LoDIdx=1.
[0482] However, depending on the decoding target level, there may be cases where nodes are not separated. For example, when Figure 40 In the example shown in FIG, when decoding is performed at level (resolution) LoDIdx = 2, nodes 611-21 and 611-22 are not separated. That is, at level LoDIdx = 2, the point corresponding to label "1" (node 611-21) and the point corresponding to label "3" (node 611-22) overlap. Therefore, no label is required for node 611-22. In view of this, the labels of the corresponding points that overlap are removed as described above.
[0483] That is, this process of removing overlapping points is performed only on the labels in List B (labels on nodes to which DCM is applied) (this does not apply to the labels in List A).
[0484] As described above, the label strings corresponding to the nodes to which DCM is not applied and the label strings corresponding to the nodes to which DCM is applied are merged, the individual labels are arranged in Morton order, and a label string from which overlapping points are appropriately removed is also used as the result of the LoD generation process to encode and decode the attribute information.
[0485] Since the attribute information is hierarchically mapped and encoded / decoded with reference to the list generated as described above, the attribute information can be encoded / decoded using a hierarchical structure similar to that of the location information. Therefore, when the DCM is included in the hierarchical structure of the location information, scalable decoding of the attribute information can also be achieved. Furthermore, redundant processing can be reduced, and the increase in the load of the attribute information encoding / decoding process can be reduced.
[0486] <Case 2 including DCM>
[0487] Alternatively, when DCM is included in the hierarchical structure of the location information, the nodes to which DCM is applied may not be separated from the nodes to which DCM is not applied. That is, marking may be performed using the same list. Figure 41 , Figure 42 Shows the Figure 40 An example of marking up the hierarchical structure shown in . Figure 42 In the example shown, Figure 39 In the example case shown in , tagging is performed using a single list.
[0488] First, after the highest level is processed (LoDIdx=0) and the node 611-1 appears, a label "0" is appended to the node 611-1 in the list ( Figure 42 The first row from the top (row "1")) of Figure 39 The example situation shown in .
[0489] After the next lower level is processed (LoDIdx=1) and the node 611-21 to which DCM is applied newly appears, a label "1" is attached to the node 611-21 in the list. DCM is also applied to the node 611-22, and the relative distance from the node 611-21 is encoded in LoDIdx=1. Therefore, a label is attached to the node 611-22 in this level. However, since the node 611-22 is separated from the node 611-21 in LoDIdx=3, a label "3" is attached to the node 611-22 ( Figure 42 The second row from the top (row "2") in the .
[0490] Note that the label on node 611-21 and the label on node 611-22 are arranged in the lowest level (highest resolution) Morton order according to the positions of the two nodes, as shown in Figure 41 The situation shown. Figure 40 In the illustrated example, the labels are arranged in the order of "1, 3" because the node 611-21 and the node 611-22 are arranged in this order.
[0491] However, in Figure 42 In the example case shown, each time a marking operation is performed on a node to which DCM is applied, overlapping points are removed. Figure 40 In the example in , when decoding is performed in the level (resolution) LoDIdx=2, the point corresponding to the label "1" (node 611-21) and the point corresponding to the label "3" (node 611-22) overlap, and therefore, the label "3" is removed. In other words, in Figure 42 In the second row from the top (row "2"), only the label "1" is attached.
[0492] On the other hand, when Figure 40 When decoding is performed at the level (resolution) LoDIdx=4 in the example of , the point corresponding to the label "1" (node 611-21) and the point corresponding to the label "3" (node 611-22) do not overlap, and therefore label removal is not performed. In other words, in Figure 42 In the second row from the top (row “2”), label “1” and label “3” are added.
[0493] Furthermore, after the next lower level is processed (LoDIdx=2) and the new node 611 appears, a label "2" is appended to these points in the list ( Figure 42 Likewise, after the next lower level is processed (LoDIdx=3) and a new node 611 appears, a label "3" is appended to these points in the list ( Figure 42 Likewise, after the next lower level is processed (LoDIdx=4) and new node 611 appears, the label "4" is appended to these points in the list ( Figure 42 Like node 611, these labels are arranged in Morton order.
[0494] A label string generated as described above and including labels corresponding to nodes to which DCM is not applied and labels corresponding to nodes to which DCM is applied is used as a result of the LoD generation process to encode and decode attribute information. In the case of this method, the individual labels are arranged in Morton order in the generated label string. Therefore, this label string can be used for encoding and decoding attribute information without the need for Figure 41 The example case shown arranges and merges the individual labels in Morton order.
[0495] Since the attribute information is hierarchically mapped and encoded / decoded with reference to the list generated as described above, the attribute information can be encoded / decoded using a hierarchical structure similar to that of the location information. Therefore, this method also enables scalable decoding of the attribute information. Furthermore, redundant processing can be reduced, and the increase in the load of the attribute information encoding / decoding process can be minimized.
[0496] <Encoding device>
[0497] Figure 43 2 is a block diagram showing an example configuration of the encoding device 200 in this case. Figure 43 The illustrated encoding device 200 encodes a point cloud by applying the present technology described above in <6. LoD Generation> to encoding.
[0498] Notice, Figure 43 The main components and aspects such as processing units and data flows are shown, but Figure 43 Not all components and aspects are necessarily shown. That is, in the encoding device 200, there may be Figure 43 A processing unit is not shown as a block, or may be present Figure 43 Not shown are processes or data flows as arrows or the like.
[0499] Figure 43 The encoding device 200 shown in FIG. 2 includes a position information encoding unit 201 to a bit stream generating unit 205. Figure 14 However, when decoding the encoded data of the position information, the position information decoding unit 202 uses the above reference Figures 38 to 42 The described list performs LoD generation processing (subsampling processing).
[0500] For example, when the hierarchical structure of the location information does not include any node to which DCM is applied, the location information decoding unit 202 decodes the location information by referring to the above. Figure 39 Furthermore, for example, when the hierarchical structure of the position information includes nodes to which DCM is applied, the position information decoding unit 202 performs LoD generation processing by referring to the above. Figure 41 or Figure 42 The described method performs a LoD generation process.
[0501] The position information decoding unit 202 then supplies the generated list (label string) to the attribute information encoding unit 204. The attribute information encoding unit 204 encodes the attribute information using the LoD generation processing result (label string). Specifically, the attribute information encoding unit 204 uses the label string (the hierarchical structure of the position information) as the hierarchical structure of the attribute information and performs encoding. By doing so, the attribute information encoding unit 204 can hierarchically separate each set of attribute information into a hierarchical structure similar to the hierarchical structure of the position information without performing LoD generation processing (subsampling processing).
[0502] Therefore, the encoding device 200 can achieve scalable decoding of attribute information. The encoding device 200 can also reduce redundant processing and an increase in the load of the attribute information encoding process.
[0503] <Encoding Process Flow>
[0504] Now refer to Figure 44 The flowchart shown in FIG6 illustrates an example flow of the encoding process in this case. When the encoding process starts, in step S601 , the position information encoding unit 201 of the encoding device 200 encodes the position information of the input point cloud and generates encoded data of the position information.
[0505] In step S602, the position information decoding unit 202 decodes the encoded data of the position information generated in step S601 and generates position information. In doing so, as described above with reference to Figures 38 to 42 As described, the position information decoding unit 202 performs LoD generation processing using the list, and performs labeling on each node.
[0506] In step S603 , the point cloud generation unit 203 performs recoloring processing using the attribute information of the input point cloud and the position information (decoding result) generated in step S602 , and associates the attribute information with the position information.
[0507] In step S604, the attribute information encoding unit 204 performs attribute information encoding processing to encode the attribute information that has been subjected to the recoloring processing in step S603, and generates encoded data of the attribute information. In doing so, the attribute information encoding unit 204 encodes the attribute information using the result (label character string) of the LoD generation processing performed in step S602. In this case, the attribute information is also encoded with the same Figure 18 The attribute information encoding process is performed in a process similar to the process shown in the flowchart in . However, the hierarchical process to be performed in step S221 will be described in detail later.
[0508] In step S605 , the bit stream generation unit 205 generates and outputs a bit stream including the encoded data of the position information generated in step S601 and the encoded data of the attribute information generated in step S604 .
[0509] When the processing in step S605 is completed, the encoding process ends.
[0510] By performing the processing in each step in the above manner, the encoding device 200 can associate the hierarchical structure of the attribute information with the hierarchical structure of the position information. That is, the encoding device 200 can also facilitate scalable decoding of the attribute information. The encoding device 200 can also reduce redundant processing and the increase in the encoding process load.
[0511] <Flow of layered processing>
[0512] In this case, Figure 18 The process of the flowchart in the Figure 44 However, in step S221 of the attribute information encoding process, the layered processing is performed in the following flow. Figure 45 The flowchart shown describes an example flow of the hierarchical processing in this case.
[0513] When the hierarchical processing starts, in step S621, the individual hierarchy processing unit 221 of the hierarchical processing unit 211 sets the lowest hierarchy (LoD) as the current hierarchy (current LoD) as the processing target. That is, the lowest hierarchy (LoD) is set as the processing target.
[0514] In step S622, the point classification unit 231 classifies the points based on the Figure 44 The predicted point is selected based on the list marked in the LoD generation process in step S602. That is, the point classification unit 231 constructs a hierarchical structure similar to the hierarchical structure of the position information.
[0515] With the steps S243 to S250 ( Figure 19 ) performs the processing in steps S623 to S630 in a manner similar to the processing in ).
[0516] By performing the processing in each step in the above manner, the hierarchical processing unit 211 can associate the hierarchical structure of the attribute information with the hierarchical structure of the position information. That is, the hierarchical processing unit 211 can also facilitate scalable decoding of the attribute information. The hierarchical processing unit 211 can also reduce redundant processing and the increase in the load of hierarchical processing.
[0517] <Decoding Device>
[0518] The configuration of the decoding device 400 in this case is also similar to Figure 23 The configuration of the example shown. That is, Figure 23 The decoding device 400 in the embodiment can decode the bit stream generated by encoding the point cloud by the present technology described above in <6. LoD generation>. However, in this case, when decoding the encoded data of the position information, the position information decoding unit 403 uses the above reference Figures 38 to 42 That is, the position information decoding unit 403 performs the LoD generation process (subsampling process) by a method similar to that employed by the position information decoding unit 202 .
[0519] The position information decoding unit 403 then supplies the generated list (label string) to the attribute information decoding unit 404. The attribute information decoding unit 404 uses the LoD generation processing result (label string) to decode the encoded attribute information. Specifically, the attribute information decoding unit 404 uses the label string (the hierarchical structure of the position information) as the hierarchical structure of the attribute information and performs decoding. This allows the attribute information decoding unit 404 to perform de-hierarchical processing on each set of attribute information as a hierarchical structure similar to the hierarchical structure of the position information, without having to perform LoD generation processing (subsampling processing).
[0520] Therefore, the decoding apparatus 400 can realize scalable decoding of attribute information. For the encoded data of attribute information, the decoding apparatus 400 can also reduce redundant processing and an increase in the load of the decoding process.
[0521] Note that in this case, the attribute information decoding unit 404 also has Figure 24 The configuration of the example case shown in is similar to that of the example case shown in . In addition, the inverse layering processing unit 423 has the same Figure 25 The configuration of the example case shown is similar to the configuration.
[0522] <Decoding Process Flow>
[0523] Now refer to Figure 46 The flowchart shown in FIG6 illustrates an example flow in the decoding process in this case. When the decoding process starts, in step S651 , the decoding target LoD depth setting unit 401 of the decoding apparatus 400 sets the depth of the LoD to be decoded.
[0524] In step S652 , the encoded data extracting unit 402 acquires and saves the bit stream, and extracts the encoded data of the position information and attribute information of the hierarchies ranging from the highest hierarchy LoD to the LoD depth set in step S651 .
[0525] In step S653, the position information decoding unit 403 decodes the encoded data of the position information extracted in step S652 and generates position information (decoding result). In doing so, the position information decoding unit 403 performs LoD generation processing using the list and performs labeling on each node, as described above with reference to Figures 38 to 42 described.
[0526] In step S654, the attribute information decoding unit 404 decodes the encoded data of the attribute information extracted in step S652 and generates attribute information (decoding result). In doing so, the attribute information decoding unit 404 decodes the encoded data of the attribute information using the result (label character string) of the LoD generation process performed in step S653. In this case, the attribute information is also decoded with the same Figure 27 The attribute information decoding process is performed in a similar manner to the process shown in the flowchart in . However, the inverse layering process to be performed in step S423 will be described in detail later.
[0527] In step S655 , the point cloud generation unit 405 generates and outputs a point cloud (decoded result) using the position information (decoded result) generated in step S653 and the attribute information (decoded result) generated in step S654 .
[0528] When the processing in step S655 is completed, the decoding process ends.
[0529] By performing the processing in each step in the above manner, decoding device 400 can associate the hierarchical structure of attribute information with the hierarchical structure of position information. That is, decoding device 400 can more easily perform scalable decoding of attribute information. Decoding device 400 can also reduce redundant processing and the increase in the decoding processing load.
[0530] <Flow of De-stratification Processing>
[0531] In this case, Figure 27 The process flow in the flowchart is similar to the process execution Figure 46 However, in step S423 of the attribute information decoding process, the inverse layering process is performed in the following flow. Figure 47 The flowchart shown describes an example flow of the inverse layering process in this case.
[0532] When the inverse hierarchical processing starts, in step S661, the individual hierarchy processing unit 441 of the inverse hierarchical processing unit 423 sets the highest hierarchy (LoD) as the current hierarchy (current LoD) as the processing target. That is, the highest hierarchy (LoD) is set as the processing target.
[0533] In step S662, the prediction unit 451 calculates the Figure 46 The list of markers in the LoD generation process of step S653 and the position information of the current LoD resolution are used to derive the predicted value of the predicted point from the reference point.
[0534] For example, the prediction unit 451 identifies the reference point of the current LoD based on the list or by considering a hierarchical structure of attribute information similar to the hierarchical structure of position information, and derives a predicted value of the predicted point of the current LoD based on the position information of the reference point and the resolution of the current LoD. The prediction value deriving method adopted herein may be any appropriate method. For example, the prediction unit 451 may derive a predicted value using a method similar to the method in the case of the processing in step S463.
[0535] With each step S464 to step S467 ( Figure 29 ) is performed in a manner similar to the processing in step S663 to step S666.
[0536] By performing the processing in each step in the above manner, the inverse layering processing unit 423 can associate the hierarchical structure of the attribute information with the hierarchical structure of the position information. That is, the decoding device 400 can more easily perform scalable decoding of the attribute information. The inverse layering processing unit 423 can also reduce redundant processing and the increase in the load of the inverse layering processing.
[0537] <7. Notes>
[0538] <Stratification and Destratification Methods>
[0539] In the above description, "boosting" has been explained as an example of a method for stratifying and destratifying attribute information. However, the present technology can be applied to any technology for stratifying attribute information. In other words, the method for stratifying and destratifying attribute information can be a method other than "boosting."
[0540] <Control Information>
[0541] Control information according to the present technology described in each of the above embodiments can be transmitted from the encoding side to the decoding side. For example, control information (e.g., enabled_flag) for controlling whether to enable (or disable) the application of the above-described present technology can be transmitted. In addition, for example, control information specifying the range within which the application of the above-described present technology is enabled (or disabled) (e.g., the upper and / or lower limits of block size, slice, picture, sequence, component, view, layer, etc.) can be transmitted.
[0542] <nearby / nearby>
[0543] Note that in this specification, positional relationships such as “near” and “adjacent” may include not only spatial positional relationships but also temporal positional relationships.
[0544] <Computer>
[0545] The above series of processes can be executed by hardware or by software. When the series of processes are to be executed by software, the program forming the software is installed in the computer. Here, the computer may be a computer incorporated into dedicated hardware, or may be, for example, a general-purpose personal computer that can execute various functions when various programs are installed therein.
[0546] Figure 48 : is a block diagram illustrating an example configuration of hardware of a computer that executes the above-described series of processes according to a program.
[0547] exist Figure 48 In the illustrated computer 900 , a central processing unit (CPU) 901 , a read-only memory (ROM) 902 , and a random access memory (RAM) 903 are interconnected via a bus 904 .
[0548] An input / output interface 910 is also connected to the bus 904. An input unit 911, an output unit 912, a storage unit 913, a communication unit 914, and a drive 915 are connected to the input / output interface 910.
[0549] The input unit 911 is formed of, for example, a keyboard, a mouse, a microphone, a touch panel, an input terminal, etc. The output unit 912 is formed of, for example, a display, a speaker, an output terminal, etc. The storage unit 913 is formed of, for example, a hard disk, a RAM disk, a nonvolatile memory, etc. The communication unit 914 is formed of, for example, a network interface. The drive 915 drives a removable medium 921 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0550] In the computer having the above configuration, the CPU 901 loads a program stored in the storage unit 913 into the RAM 903 via the input / output interface 910 and the bus 904, for example, and executes the program, so that the above-described series of processes are performed. The RAM 903 also stores data required for the CPU 901 to perform various processes, etc., as needed.
[0551] For example, the program to be executed by the computer may be recorded on the removable medium 921 as a package medium or the like to be used. In this case, when the removable medium 921 is mounted on the drive 915 , the program can be installed in the storage unit 913 via the input / output interface 910 .
[0552] Alternatively, the program may be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting. In this case, the program may be received by the communication unit 914 and installed in the storage unit 913.
[0553] Furthermore, the program may be installed in the ROM 902 or the storage unit 913 in advance.
[0554] <Objectives of applying this technology>
[0555] While the present technology has been described so far as being applied to the encoding and decoding of point cloud data, it is not limited to these examples and can be applied to the encoding and decoding of 3D data of any standard. Specifically, any processing such as encoding and decoding, and any specifications for various data types such as 3D data and metadata, can be employed as long as they do not conflict with the present technology described above. Furthermore, some of the aforementioned processing and specifications may be omitted as long as they do not conflict with the present technology.
[0556] Furthermore, in the above description, the encoding device 200 and the decoding device 400 have been described as example applications of the present technology, but the present technology can be applied to any desired configuration.
[0557] For example, the present technology can be applied to various electronic devices such as transmitters and receivers in satellite broadcasting (such as television receivers or portable telephone devices), wired broadcasting such as cable television, distribution via the Internet, distribution to terminals via cellular communications, etc., as well as devices that record images on media (such as optical disks, magnetic disks, and flash memories) and reproduce images from these storage media (such as hard disk recorders or cameras).
[0558] In addition, the present technology can also be implemented as a component of a device, such as a processor used as a system LSI (large-scale integration) etc. (e.g., a video processor), a module using multiple processors etc. (e.g., a video module), a unit using multiple modules etc. (e.g., a video unit), or a component having other functions added to the unit (e.g., a video component).
[0559] Furthermore, for example, the present technology can also be applied to a network system formed by multiple devices. For example, the present technology can be implemented as cloud computing in which multiple devices share and jointly process data via a network. For example, the present technology can be implemented in a cloud service that provides image (video) related services to any type of terminal, such as computers, audio-visual (AV) devices, portable information processing terminals, and IoT (Internet of Things) devices.
[0560] Note that in this specification, a system refers to an assembly of multiple components (devices, modules (parts), etc.), and not all of these components need to be housed in the same housing. Therefore, multiple devices housed in different housings and interconnected via a network form a system, and a single device comprising multiple modules housed in a single housing also constitutes a system.
[0561] <Fields and uses where this technology can be applied>
[0562] Systems, devices, processing units, etc. using this technology can be used in any appropriate field, such as transportation, medical care, crime prevention, agriculture, animal husbandry, mining, beauty, factories, household appliances, meteorology, or nature observation. In addition, this technology can also be used for any appropriate purpose.
[0563] <Other aspects>
[0564] Note that in this specification, a "flag" is information for identifying a plurality of states, and includes not only information for identifying two states of true (1) or false (0), but also information for identifying three or more states. Therefore, for example, the value that the "flag" can have can be two values of "1" and "0", or three or more values. That is, the "flag" can be formed by an arbitrary number of bits, and can be formed by one bit or a plurality of bits. In addition, for identification information (including a flag), not only identification information but also difference information of the identification information relative to the reference information can be included in the bit stream. Therefore, in this specification, "flag" and "identification information" include not only the information but also difference information relative to the reference information.
[0565] In addition, various information (such as metadata) about the encoded data (bit stream) can be transmitted or recorded in any mode associated with the encoded data. Here, the term "association" means that other data (or links to other data) can be used, for example, when processing the data. That is, mutually related data can be integrated into one piece of data or be regarded as separate data. For example, information associated with the encoded data (image) can be transmitted via a transmission path different from that of the encoded data (image). In addition, for example, information associated with the encoded data (image) can be recorded in a recording medium different from that of the encoded data (image) (or in a different recording area of the same recording medium). Note that this "association" can apply to certain data, not the entire data. For example, an image and information corresponding to the image can be associated with each other for any appropriate unit (for example, for multiple frames, each frame, or a portion of each frame).
[0566] Note that in this specification, the terms "combine", "multiplex", "add", "integrate", "include", "store", "contain", "merge", "insert", etc. mean combining multiple objects into one, such as combining encoded data and metadata into one piece of data, and mean the above-mentioned "association" method.
[0567] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications may be made to them without departing from the scope of the present technology.
[0568] For example, any configuration described above as one device (or one processing unit) may be divided into multiple devices (or processing units). Conversely, any configuration described above as multiple devices (or processing units) may be combined into one device (or one processing unit). Furthermore, of course, components other than those described above may be added to the configuration of each device (or each processing unit). Furthermore, certain components of a device (or processing unit) may be incorporated into the configuration of another device (or processing unit), as long as the configuration and functionality of the overall system remain substantially the same.
[0569] In addition, for example, the above program can be executed in any device. In this case, the device only needs to have necessary functions (functional blocks, etc.) so that necessary information can be obtained.
[0570] Furthermore, for example, a single device may execute each step in a flowchart, or multiple devices may execute each step. Furthermore, when a single step includes multiple processes, the multiple processes may be executed by a single device or by multiple devices. In other words, the multiple processes included in a single step may be executed as processes in multiple steps. Conversely, processes described as multiple steps may be executed collectively as a single step.
[0571] Furthermore, for example, the program executed by the computer may be a program for executing the processing in the steps in chronological order according to the sequence described in this specification, or may be a program for executing processing in parallel or executing processing when necessary (for example, when called). That is, as long as there is no contradiction, the processing in each step may be executed in an order different from the order described above. Furthermore, the processing in the steps according to this program may be executed in parallel with processing according to another program, or may be executed in conjunction with processing according to another program.
[0572] Furthermore, for example, each of the various techniques according to the present technology can be implemented independently, as long as there are no contradictions. Of course, combinations of some of the various techniques according to the present technology can also be implemented. For example, part or all of the present technology described in one embodiment can be implemented in combination with part or all of the present technology described in another embodiment. Furthermore, part or all of the above-described present technology can be implemented in combination with some other technology not described above.
[0573] Note that the present technology can also be implemented in the configuration described below.
[0574] (1) An information processing device comprising:
[0575] a hierarchical unit that performs hierarchical processing on attribute information of a point cloud representing a three-dimensional object by recursively repeating, for reference points, a process of classifying points of the point cloud into predicted points and reference points, the predicted points being points where a difference between the attribute information and the predicted value is left, and the reference points being points to which the attribute information is referenced during derivation of the predicted value,
[0576] During the layering, the layering unit selects the prediction point in such a manner that a point also exists in a voxel of a layer one level higher than the layer to which the voxel containing the point of the current layer belongs.
[0577] (2) The information processing device according to (1), wherein
[0578] The layered unit
[0579] deriving a predicted value of the prediction point using attribute information of the reference point weighted according to the distance between the prediction point and the reference point, and
[0580] The difference value is derived using the derived predicted value.
[0581] (3) The information processing device according to (2), wherein
[0582] The hierarchical unit derives a predicted value of the predicted point using attribute information of the reference point weighted according to a distance based on position information of a lowest-level resolution of the point cloud.
[0583] (4) The information processing device according to (2), wherein
[0584] The hierarchical unit derives a predicted value of the predicted point using attribute information of the reference point weighted according to a distance of position information based on a resolution of a current level of the point cloud.
[0585] (5) The information processing device according to (2), wherein
[0586] The hierarchical unit classifies all points in the current layer as the prediction points, sets reference points in voxels of a layer one level higher, and derives prediction values of the prediction points using attribute information of the reference points.
[0587] (6) The information processing device according to any one of (1) to (5), further comprising:
[0588] A quantization unit quantizes the difference value of each point of each level generated by the hierarchical unit using a quantization weight of each level.
[0589] (7) The information processing device according to (6), wherein
[0590] The quantization unit quantizes the difference value using a quantization weight of each level, the quantization weight having a larger value at a higher level.
[0591] (8) The information processing device according to (6), further comprising:
[0592] An encoding unit that encodes the difference value that has been quantized by the quantization unit and a parameter defining a function for deriving a quantization weight for each hierarchy, and generates encoded data.
[0593] (9) The information processing device according to (6), further comprising:
[0594] An encoding unit that encodes the difference value that has been quantized by the quantization unit and the quantization weight of each hierarchy, and generates encoded data.
[0595] (10) An information processing method comprising:
[0596] performing stratification on attribute information of a point cloud representing a three-dimensional object by recursively repeating a process of classifying points of the point cloud into predicted points and reference points, the predicted points being points where a difference between the attribute information and a predicted value is left, and the reference points being points to which the attribute information was referenced during derivation of the predicted value; and
[0597] During the layering, the prediction point is selected in such a way that a point also exists in a voxel of a level one level higher than the level to which the voxel containing the point of the current level belongs.
[0598] (11) An information processing device comprising:
[0599] an inverse stratification unit that performs inverse stratification on attribute information of a point cloud representing a three-dimensional object, the attribute information having been stratified by recursively repeating a process of classifying points of the point cloud into predicted points and reference points, the predicted points being points that leave a difference between the attribute information and a predicted value, the reference points being points that referenced the attribute information during derivation of the predicted value,
[0600] During the inverse layering, in each layer, the inverse layering unit uses the attribute information of the reference point and the position information of the resolution of the layer that is not the lowest layer of the point cloud to derive a predicted value of the attribute information of the predicted point, and uses the predicted value and the difference to derive the attribute information of the predicted point.
[0601] (12) The information processing device according to (11), wherein
[0602] The inverse hierarchical unit derives the prediction value using attribute information of the reference point and position information of a resolution of a current hierarchy of the reference point and the prediction point.
[0603] (13) The information processing device according to (11), wherein
[0604] The inverse hierarchical unit derives the prediction value using attribute information of the reference point, position information of the reference point at a resolution of a hierarchy one level higher than the current hierarchy, and position information of the prediction point at the resolution of the current hierarchy.
[0605] (14) The information processing device according to any one of (11) to (13), wherein
[0606] The inverse hierarchical unit derives a predicted value of the prediction point using attribute information of the reference point weighted according to a distance between the prediction point and the reference point.
[0607] (15) The information processing device according to any one of (11) to (14), wherein
[0608] The inverse hierarchical unit hierarchizes the attribute information using the position information, and associates the attribute information of each point with the position information in each hierarchical level.
[0609] (16) The information processing device according to any one of (11) to (15), further comprising:
[0610] An inverse quantization unit inversely quantizes the difference value of each point of each level that has been quantized using the quantization weight of each level.
[0611] (17) The information processing device according to (16), wherein
[0612] The inverse quantization unit inversely quantizes the quantized difference value using a quantization weight of each level, the quantization weight having a larger value at a higher level.
[0613] (18) The information processing device according to (16), further comprising:
[0614] A decoding unit that decodes the encoded data and obtains the quantized difference value and parameters defining a function for deriving a quantization weight for each layer.
[0615] (19) The information processing device according to (16), further comprising:
[0616] A decoding unit decodes the encoded data and obtains the quantized difference value and the quantization weight of each layer.
[0617] (20) An information processing method comprising:
[0618] performing inverse stratification on attribute information of a point cloud representing a three-dimensional object, the attribute information having been stratified by recursively repeating, for reference points, a process of classifying points of the point cloud into predicted points and reference points, the predicted points being points that leave a difference between the attribute information and a predicted value, the reference points being points that referenced the attribute information during derivation of the predicted value; and
[0619] During the inverse layering, in each layer, the attribute information of the reference point and the position information of the resolution of the layer that is not the lowest layer of the point cloud are used to derive a predicted value of the attribute information of the predicted point, and the attribute information of the predicted point is derived using the predicted value and the difference.
[0620] (21) An information processing device comprising:
[0621] a generating unit that generates information indicating a hierarchical structure of position information of a point cloud representing a three-dimensional object; and
[0622] A hierarchical unit hierarchizes the attribute information of the point cloud based on the information generated by the generating unit so as to associate the attribute information with the hierarchical structure of the position information.
[0623] (22) The information processing device according to (21), wherein
[0624] The generation unit generates information of the position information obtained by decoding encoded data of the position information.
[0625] (23) The information processing device according to (21) or (22), wherein:
[0626] The generation unit generates the information by marking each node of the hierarchical structure of the position information using a list.
[0627] (24) The information processing device according to (23), wherein
[0628] The generation unit marks each newly appearing node with a value indicating the level in order from the highest level to the lowest level in the hierarchical structure of the position information.
[0629] (25) The information processing device according to (24), wherein
[0630] The generating unit marks each newly appearing node in Morton order in each level in the hierarchical structure of the position information.
[0631] (26) The information processing device according to (24) or (25), wherein
[0632] When a node to which a direct coding mode (DCM) is applied newly appears, the generation unit performs marking on the node and performs marking on another node in a lower layer that is divided from the node and to which the DCM is applied.
[0633] (27) The information processing device according to (26), wherein
[0634] The generation unit performs marking on nodes to which the DCM is applied using a list different from a list used to mark nodes to which the DCM is not applied.
[0635] (28) The information processing device according to (26) or (27), wherein
[0636] When points corresponding to a plurality of labels overlap, the generation unit retains one of the plurality of labels and deletes the other labels of the plurality of labels.
[0637] (29) The information processing device according to any one of (21) to (28), wherein
[0638] The hierarchical unit selects prediction points of each hierarchy of the attribute information based on the information so as to associate the attribute information with the hierarchical structure of the position information.
[0639] (30) An information processing method comprising:
[0640] generating information indicating a hierarchical structure of position information of a point cloud representing a three-dimensional object; and
[0641] Based on the generated information, the attribute information is hierarchized so as to be associated with the hierarchical structure of the position information.
[0642] (31) An information processing device comprising:
[0643] a generating unit that generates information indicating a hierarchical structure of position information of a point cloud representing a three-dimensional object; and
[0644] A de-hierarchical unit de-hierarchically performs de-hierarchical processing on the attribute information of the point cloud based on the information generated by the generating unit so as to associate the hierarchical structure of the attribute information with the hierarchical structure of the position information.
[0645] (32) The information processing device according to (31), wherein
[0646] The generation unit generates information of the position information obtained by decoding encoded data of the position information.
[0647] (33) The information processing device according to (31) or (32), wherein:
[0648] The generation unit generates the information by marking each node of the hierarchical structure of the position information using a list.
[0649] (34) The information processing device according to (33), wherein
[0650] The generation unit marks each newly appearing node with a value indicating the level in order from the highest level to the lowest level in the hierarchical structure of the position information.
[0651] (35) The information processing device according to (34), wherein
[0652] The generating unit marks each newly appearing node in Morton order in each level in the hierarchical structure of the position information.
[0653] (36) The information processing device according to (34) or (35), wherein
[0654] When a node to which a direct coding mode (DCM) is applied newly appears, the generation unit performs marking on the node and performs marking on another node in a lower layer that is divided from the node and to which the DCM is applied.
[0655] (37) The information processing device according to (36), wherein
[0656] The generation unit performs marking on nodes to which the DCM is applied using a list different from a list used to mark nodes to which the DCM is not applied.
[0657] (38) The information processing device according to (36) or (37), wherein
[0658] When points corresponding to a plurality of labels overlap, the generation unit retains one of the plurality of labels and deletes the other labels of the plurality of labels.
[0659] (39) The information processing device according to any one of (31) to (38), wherein
[0660] The inverse hierarchical unit selects prediction points of each hierarchy of the attribute information based on the information so as to associate the attribute information with the hierarchical structure of the position information.
[0661] (40) An information processing method comprising:
[0662] generating information indicating a hierarchical structure of position information of a point cloud representing a three-dimensional object; and
[0663] Based on the generated information, attribute information of the point cloud is hierarchized, wherein the attribute information is associated with a hierarchical structure of the position information.
[0664] (41) An information processing device comprising:
[0665] a hierarchical unit that hierarchizes attribute information of a point cloud representing a three-dimensional object so as to associate the attribute information with a hierarchical structure of position information of the point cloud; and
[0666] A quantization unit quantizes the attribute information hierarchized by the hierarchical unit using a quantization weight of each level.
[0667] (42) The information processing device according to (41), wherein
[0668] The quantization unit quantizes the attribute information using a quantization weight formed by a ratio between the number of points in the point cloud and the number of points in a processing target layer.
[0669] (43) The information processing device according to (42), wherein
[0670] The quantization unit quantizes the attribute information of the lowest level using a quantization weight of a value "1".
[0671] (44) The information processing device according to (41), wherein
[0672] The quantization unit quantizes the attribute information using a quantization weight formed by a ratio between the number of points in a hierarchy below a processing target hierarchy and the number of points in the processing target hierarchy.
[0673] (45) An information processing method comprising:
[0674] hierarchical attribute information of a point cloud representing a three-dimensional object so as to associate the attribute information with a hierarchical structure of position information of the point cloud; and
[0675] The quantization weight of each level is used to quantize the layered attribute information.
[0676] (46) An information processing device comprising:
[0677] an inverse quantization unit that inversely quantizes attribute information of a point cloud using a quantization weight of each level, the attribute information having been hierarchically associated with a hierarchical structure of position information of the point cloud representing a three-dimensional object and having been quantized using the quantization weight; and
[0678] The inverse layering unit inversely layers the attribute information inversely quantized by the inverse quantization unit.
[0679] (47) The information processing device according to (46), wherein
[0680] The inverse quantization unit inversely quantizes the attribute information using a quantization weight formed by a ratio between the number of points in the point cloud and the number of points in a processing target layer.
[0681] (48) The information processing device according to (47), wherein
[0682] The inverse quantization unit inversely quantizes the attribute information of the lowest level using a quantization weight of a value "1".
[0683] (49) The information processing device according to (46), wherein
[0684] The inverse quantization unit inversely quantizes the attribute information using a quantization weight formed by a ratio between the number of points in a hierarchy below a processing target hierarchy and the number of points in the processing target hierarchy.
[0685] (50) The information processing device according to any one of (46) to (49), wherein
[0686] The inverse quantization unit acquires a quantization weight used to quantize the attribute information, and inversely quantizes the attribute information using the acquired quantization weight.
[0687] (51) An information processing method comprising:
[0688] hierarchical attribute information of a point cloud representing a three-dimensional object so as to associate the attribute information with a hierarchical structure of position information of the point cloud; and
[0689] The quantization weight of each level is used to quantize the layered attribute information.
[0690] Reference Signs List
[0691] 100 Space Area
[0692] 101 voxels
[0693] 102 points
[0694] 111 points
[0695] 112 points
[0696] 113 points
[0697] 200 Encoding Device
[0698] 201 Position information encoding unit
[0699] 202 Position information decoding unit
[0700] 203 Point Cloud Generation Unit
[0701] 204 Attribute Information Coding Unit
[0702] 205 Bitstream Generation Unit
[0703] 211 layered processing units
[0704] 212 Quantitative Units
[0705] 213 coding units
[0706] 221 Individual Level Processing Unit
[0707] 231 taxa
[0708] 232 prediction units
[0709] 233 Arithmetic Unit
[0710] 234 Update Unit
[0711] 235 Arithmetic Unit
[0712] 400 Decoding Device
[0713] 401 Decoding target LoD depth setting unit
[0714] 402 Encoded Data Extraction Unit
[0715] 403 Position Information Decoding Unit
[0716] 404 Attribute Information Decoding Unit
[0717] 405 Point Cloud Generation Unit
[0718] 421 Decoding Unit
[0719] 422 Inverse Quantization Unit
[0720] 423 Inverse Layer Processing Unit
[0721] 441 Individual Level Processing Unit
[0722] 451 prediction unit
[0723] 452 Arithmetic Unit
[0724] 453 Merge Processing Unit
Claims
1. An information processing device, comprising: a hierarchical unit that performs hierarchical processing on attribute information of a point cloud representing a three-dimensional object by recursively repeating, for reference points, a process of classifying points of the point cloud into predicted points, which are points where a difference between the attribute information and the predicted value is left, and reference points, which are points to which the attribute information is referenced during derivation of the predicted value; as well as a quantization unit that quantizes the difference value of each point of each level generated by the hierarchical unit using a quantization weight of each level, During the layering, the layering unit selects the prediction point in such a manner that a point also exists in a voxel of a layer one level higher than the layer to which the voxel containing the point of the current layer belongs.
2. The information processing device according to claim 1, wherein The layered unit deriving a predicted value of the prediction point using attribute information of the reference point weighted according to the distance between the prediction point and the reference point, and The difference value is derived using the derived predicted value.
3. The information processing device according to claim 2, wherein: The hierarchical unit derives a predicted value of the predicted point using attribute information of the reference point weighted according to a distance based on position information of a lowest-level resolution of the point cloud.
4. The information processing device according to claim 2, wherein: The hierarchical unit derives a predicted value of the predicted point using attribute information of the reference point weighted according to a distance of position information based on a resolution of a current level of the point cloud. The information processing apparatus according to claim 1 , wherein: The quantization unit quantizes the difference value using a quantization weight of each level, the quantization weight having a larger value at a higher level.
6. The information processing apparatus according to claim 1, further comprising: An encoding unit that encodes the difference value that has been quantized by the quantization unit and a parameter defining a function for deriving a quantization weight for each hierarchy, and generates encoded data.
7. The information processing apparatus according to claim 1, further comprising: An encoding unit that encodes the difference value that has been quantized by the quantization unit and the quantization weight of each hierarchy, and generates encoded data.
8. An information processing method, comprising: performing stratification on attribute information of a point cloud representing a three-dimensional object by recursively repeating a process of classifying points of the point cloud into predicted points and reference points, the predicted points being points where a difference between the attribute information and a predicted value is left, and the reference points being points to which the attribute information was referenced during derivation of the predicted value; and During the layering, the prediction point is selected in such a way that a point also exists in a voxel of a level one level higher than the level to which the voxel containing the point of the current level belongs; as well as The difference value of each point at each level is quantized using the quantization weight of each level.
9. An information processing device comprising: an inverse hierarchical unit that performs inverse hierarchical processing on attribute information of a point cloud representing a three-dimensional object, the attribute information having been hierarchical by recursively repeating, for reference points, a process of classifying points of the point cloud into predicted points and reference points, the predicted points being points that leave a difference between the attribute information and a predicted value, the reference points being points to which the attribute information was referenced during derivation of the predicted value; as well as an inverse quantization unit that inversely quantizes the difference value of each point of each level that has been quantized using the quantization weight of each level, During the inverse layering, in each layer, the inverse layering unit uses the attribute information of the reference point and the position information of the resolution of the layer that is not the lowest layer of the point cloud to derive a predicted value of the attribute information of the predicted point, and uses the predicted value and the difference to derive the attribute information of the predicted point.
10. The information processing apparatus according to claim 9, wherein: The inverse hierarchical unit derives the prediction value using attribute information of the reference point and position information of a resolution of a current hierarchy of the reference point and the prediction point.
11. The information processing apparatus according to claim 9, wherein: The inverse hierarchical unit derives a predicted value of the prediction point using attribute information of the reference point weighted according to a distance between the prediction point and the reference point.
12. The information processing apparatus according to claim 9, wherein: The inverse hierarchical unit hierarchizes the attribute information using the position information, and associates the attribute information of each point with the position information in each hierarchical level.
13. The information processing apparatus according to claim 9, wherein: The inverse quantization unit inversely quantizes the quantized difference value using a quantization weight of each level, the quantization weight having a larger value at a higher level.
14. The information processing apparatus according to claim 9, further comprising: A decoding unit that decodes the encoded data and obtains the quantized difference value and parameters defining a function for deriving a quantization weight for each layer.
15. The information processing apparatus according to claim 9, further comprising: A decoding unit decodes the encoded data and obtains the quantized difference value and the quantization weight of each layer.
16. An information processing method, comprising: performing inverse stratification on attribute information of a point cloud representing a three-dimensional object, the attribute information having been stratified by recursively repeating, for reference points, a process of classifying points of the point cloud into predicted points and reference points, the predicted points being points where a difference between the attribute information and a predicted value is left, the reference points being points to which the attribute information was referenced during derivation of the predicted value; as well as During the inverse layering, in each layer, using the attribute information of the reference point and position information of a resolution of a layer that is not a lowest layer of the point cloud to derive a predicted value of the attribute information of the predicted point, and using the predicted value and the difference value to derive the attribute information of the predicted point; The difference value of each point of each level that has been quantized is inversely quantized using the quantization weight of each level.
17. An information processing device comprising: a generating unit that generates information indicating a hierarchical structure of position information of a point cloud representing a three-dimensional object; a hierarchical unit that hierarchizes the attribute information of the point cloud based on the information generated by the generating unit so as to associate the attribute information with the hierarchical structure of the position information; as well as A quantization unit quantizes the attribute information layered by the layering unit using a quantization weight for each layer.
18. An information processing device comprising: a generating unit that generates information indicating a hierarchical structure of position information of a point cloud representing a three-dimensional object; as well as a de-hierarchical unit that de-hierarchically performs de-hierarchical processing on the attribute information of the point cloud based on the information generated by the generating unit so as to associate a hierarchical structure of the attribute information with a hierarchical structure of the position information; as well as An inverse quantization unit inversely quantizes attribute information of a point cloud using a quantization weight of each level, the attribute information having been hierarchically associated with a hierarchical structure of position information of the point cloud representing a three-dimensional object and having been quantized using the quantization weight.
Citation Information
Patent Citations
Hierarchical point cloud compression
US20190081638A1