Point cloud decoding device, point cloud decoding method, and program

The point cloud decoding device enhances coding efficiency by excluding upper layer intra prediction scaling and decoding the 0th layer scaling factor, addressing inefficiencies in conventional RAHT-based methods for intra prediction determination.

WO2025150284A1PCT designated stage expired Publication Date: 2025-07-17KDDI CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/041960
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-12
Filing Date
2024-11-27
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Conventional point cloud decoding methods using RAHT for determining intra prediction of AC coefficients are inefficient due to reliance on adjacent nodes, leading to inaccurate determination of effective intra prediction.

Method used

A point cloud decoding device and method that decodes a value indicating up to which upper layer intra prediction scaling is excluded, and when the value is '0', decodes the scaling factor of the 0th layer, incorporating an attribute information decoding unit to enhance coding efficiency.

Benefits of technology

Improves the coding efficiency of attribute information by accurately determining the effectiveness of intra prediction, thereby optimizing the decoding process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024041960_17072025_PF_FP_ABST
    Figure JP2024041960_17072025_PF_FP_ABST
Patent Text Reader

Abstract

A point cloud decoding device 200 according to the present invention is provided with an attribute information decoding unit 2060 that decodes a value indicating up to which upper layer inter-prediction scaling is not to be applied and decodes a scaling factor for the 0th layer when the value is "0".
Need to check novelty before this filing date? Find Prior Art

Description

Point group decoding device, point group decoding method and program

[0001] The present invention relates to a point group decoding device, a point group decoding method, and a program.

[0002] As a conventional technique, in decoding attribute information using RAHT, a method is known in which, in order to determine whether intra prediction of AC coefficients can be performed, the condition is that the number of adjacent nodes for a node to be processed is equal to or greater than a threshold.

[0003] G-PCC codec description, ISO / IEC JTC1 / SC29 / WG7 N00271G-PCC 2nd Edition codec description, ISO / IEC JTC1 / SC29 / WG7 N00506

[0004] However, the conventional technology has a problem in that it is only based on the condition of adjacent nodes and is not possible to precisely determine whether intra prediction is likely to be effective.

[0005] Therefore, the present invention has been made in consideration of the above-mentioned problems, and aims to provide a point cloud decoding device, a point cloud decoding method, and a program that can improve the coding efficiency of attribute information coding.

[0006] A first feature of the present invention is that it is a point cloud decoding device that includes an attribute information decoding unit that decodes a value indicating up to which upper layer inter-prediction scaling should not be applied, and decodes the scaling factor for the 0th layer if the value is "0."

[0007] A second feature of the present invention is that it is a point cloud decoding method, which includes a step of decoding a value indicating up to which upper layer inter-prediction scaling is not applied, and a step of decoding the scaling factor of the 0th layer if the value is "0".

[0008] A third feature of the present invention is a program that causes a computer to function as a point cloud decoding device, wherein the point cloud decoding device is provided with an attribute information decoding unit that decodes a value indicating up to which upper layer inter-prediction scaling should not be applied, and decodes the scaling factor of the 0th layer if the value is "0".

[0009] According to the present invention, it is possible to provide a point cloud decoding device, a point cloud decoding method, and a program that can improve the coding efficiency of attribute information coding.

[0010] FIG. 1 is a diagram showing an example of the configuration of a point cloud processing system 10 according to an embodiment. FIG. 2 is a diagram showing an example of functional blocks of a point cloud decoding device 200 according to an embodiment. FIG. 3 is a diagram showing an example of the configuration of encoded data (bit stream) received by a geometric information decoding unit 2010 of the point cloud decoding device 200 according to an embodiment. FIG. 4 is a diagram showing an example of the syntax configuration of GPS2011. FIG. 5 is an example of the configuration of encoded data (bit stream) received by an attribute information decoding unit 2060 of the point cloud decoding device 200 according to an embodiment. FIG. 6 is an example of the syntax configuration of APS2611 shown in FIG. 5. FIG. 7 is an example of the syntax configuration of ASH2612 shown in FIG. 5. FIG. 8 is a flowchart showing an example of the processing of the RAHT unit 2080. FIG. 9 is a flowchart showing an example of the processing of step S704. FIG. 10 is a flowchart showing an example of the intra prediction processing of step S28112. FIG. 11 is a diagram showing the relationship between a node to be decoded and adjacent nodes in a higher layer. FIG. 12 is a diagram showing the relationship between a node to be decoded and adjacent nodes in the subnode hierarchy. FIG. 13 is a diagram showing an example of the inter prediction process of step S28111. FIG. 14 is an example of a syntax configuration when raht_filter_taps is decoded based on raht_inter_skip_layers. FIG. 15 is a flowchart showing an example of the slice data decoding process of step S505. FIG. 16 is a flowchart showing an example of the intra prediction of step S703 described above. FIG. 17 is a flowchart showing an example of the process of the tree merging unit 2020 of the point group decoding device 200 according to an embodiment. FIG. 18 is a flowchart showing an example of the process of predicting position information in step S604. FIG. 19 is a diagram for explaining an example of the process of the tree merging unit 2020 of the point group decoding device 200 according to an embodiment. FIG. 20 is a diagram showing an example of functional blocks of the point group encoding device 100 according to an embodiment.

[0011] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that the components in the following embodiments can be appropriately replaced with existing components, etc., and various variations, including combinations with other existing components, are possible. Therefore, the description of the following embodiments does not limit the content of the invention described in the claims.

[0012] First Embodiment Hereinafter, a point cloud processing system 10 according to a first embodiment of the present invention will be described with reference to Figures 1 to 20. Figure 1 is a diagram showing a point cloud processing system 10 according to this embodiment.

[0013] As shown in FIG. 1 , the point cloud processing system 10 includes a point cloud encoding device 100 and a point cloud decoding device 200 .

[0014] The point cloud encoding device 100 is configured to generate encoded data (bitstream) by encoding an input point cloud signal, and the point cloud decoding device 200 is configured to generate an output point cloud signal by decoding the bitstream.

[0015] The input point cloud signal and the output point cloud signal are composed of position information and attribute information of each point in the point cloud, such as color information and reflectance of each point.

[0016] Here, the bit stream may be transmitted from the point group encoding device 100 to the point group decoding device 200 via a transmission path. Alternatively, the bit stream may be stored in a storage medium and then provided from the point group encoding device 100 to the point group decoding device 200.

[0017] (Point Group Decoding Device 200) The point group decoding device 200 according to this embodiment will be described below with reference to Fig. 2. Fig. 2 is a diagram showing an example of functional blocks of the point group decoding device 200 according to this embodiment.

[0018] As shown in Figure 2, the point cloud decoding device 200 has a geometric information decoding unit 2010, a tree synthesis unit 2020, an approximate surface synthesis unit 2030, a geometric information reconstruction unit 2040, an inverse coordinate transformation unit 2050, an attribute information decoding unit 2060, an inverse quantization unit 2070, an RAHT unit 2080, an LoD calculation unit 2090, an inverse lifting unit 2100, an inverse color transformation unit 2110, and a frame buffer 2120.

[0019] The geometric information decoding unit 2010 is configured to receive as input a bit stream relating to geometric information (geometric information bit stream) from among the bit streams output from the point group encoding device 100, and to decode the syntax.

[0020] The decoding process is, for example, a context-adaptive binary arithmetic decoding process. Here, for example, the syntax includes control data (flags and parameters) for controlling the decoding process of the position information.

[0021] The tree synthesis unit 2020 is configured to receive as input the control data decoded by the geometric information decoding unit 2010 and an occupancy code indicating at which node in the tree (described later) the point group exists, and to generate tree information indicating in which area in the decoding target space the point exists.

[0022] The decoding process of the occupancy code may be performed within the tree synthesis unit 2020.

[0023] This process divides the decoding target space into rectangular parallelepipeds, determines whether a point exists in each rectangular parallelepiped by referring to the occupancy code, divides the rectangular parallelepiped in which the point exists into multiple rectangular parallelepipeds, and refers to the occupancy code. By recursively repeating this process, tree information can be generated.

[0024] Here, when decoding such an occupancy code, inter prediction, which will be described later, may be used.

[0025] In this embodiment, a method called "Octree" can be used, which recursively performs octree division on the above-mentioned rectangular parallelepiped as a cube, and a method called "QtBt" can be used, which performs quadtree division and binary tree division in addition to octree division. Whether or not to use "QtBt" is transmitted from the point cloud encoding device 100 as control data.

[0026] Alternatively, when the use of predictive geometry coding is specified by the control data, the tree synthesis unit 2020 is configured to decode the coordinates of each point based on an arbitrary tree configuration determined by the point cloud encoding device 100.

[0027] The approximate surface synthesis unit 2030 is configured to generate approximate surface information using the tree information generated by the tree synthesis unit 2020, and to decode the point cloud based on the approximate surface information.

[0028] For example, when decoding three-dimensional point cloud data of an object, if the point cloud is densely distributed on the surface of the object, approximate surface information is used to represent the area where the point cloud exists by approximating it with a small plane, rather than decoding each individual point cloud.

[0029] Specifically, the approximate surface synthesis unit 2030 can generate approximate surface information and decode the point cloud using a technique called "Trisoup," for example. A specific example of the "Trisoup" process will be described later. Furthermore, when decoding a sparse point cloud acquired by Lidar or the like, this process can be omitted.

[0030] The geometric information reconstruction unit 2040 is configured to reconstruct the geometric information (position information in the coordinate system assumed by the decoding process) of each point of the point cloud data to be decoded based on the tree information generated by the tree synthesis unit 2020 and the approximate surface information generated by the approximate surface synthesis unit 2030.

[0031] The inverse coordinate transformation unit 2050 is configured to receive the geometric information reconstructed by the geometric information reconstruction unit 2040 as input, transform it from the coordinate system assumed by the decoding process to the coordinate system of the output point cloud signal, and output position information.

[0032] The frame buffer 2120 is configured to store, as a reference frame, the geometric information reconstructed by the geometric information reconstruction unit 2040 as an input. The stored reference frame is read from the frame buffer 2130 and used as the reference frame when the tree synthesis unit 2020 performs inter-prediction of a temporally different frame.

[0033] Here, which reference frame at which time point is to be used for each frame may be determined based on control data transmitted as a bit stream from the point cloud encoding device 100, for example.

[0034] The attribute information decoding unit 2060 is configured to receive as input a bit stream relating to attribute information (attribute information bit stream) from among the bit streams output from the point group encoding device 100, and to decode the syntax.

[0035] The decoding process is, for example, a context-adaptive binary arithmetic decoding process, where the syntax includes, for example, control data (flags and parameters) for controlling the decoding process of the attribute information.

[0036] Moreover, the attribute information decoding unit 2060 is configured to decode the quantized residual information from the decoded syntax.

[0037] The inverse quantization unit 2070 is configured to perform inverse quantization processing based on the quantized residual information decoded by the attribute information decoding unit 2060 and the quantization parameter, which is one of the control data decoded by the attribute information decoding unit 2060, to generate inverse quantized residual information.

[0038] The dequantized residual information is output to either the RAHT unit 2080 or the LoD calculation unit 2090, depending on the characteristics of the point group to be decoded. Which unit the information is output to is specified by control data decoded by the attribute information decoding unit 2060.

[0039] The RAHT unit 2080 is configured to receive as input the inverse-quantized residual information generated by the inverse quantization unit 2070 and the geometric information generated by the geometric information reconstruction unit 2040, and to decode the attribute information of each point using a type of Haar transform (inverse Haar transform in the decoding process) called RAHT (Region Adaptive Hierarchical Transform). The decoded information is the direct current component (DC coefficient) and alternating current component (AC coefficient) of the attribute information generated by using the RAHT in the encoding process, and is converted into attribute information by using the inverse transform of the RAHT in the decoding process. As a specific process of the RAHT, for example, the method described in Non-Patent Document 1 can be used.

[0040] The LoD calculation unit 2090 is configured to receive the geometric information generated by the geometric information reconstruction unit 2040 as an input and generate an LoD (Level of Detail).

[0041] LoD is information for defining a reference relationship (a point to be referenced and a point to be referenced) to realize predictive coding, such as predicting attribute information of another point from attribute information of another point and encoding or decoding the prediction residual.

[0042] In other words, LoD is information that defines a hierarchical structure in which each point contained in geometric information is classified into multiple levels, and the attributes of points belonging to lower levels are encoded or decoded using the attribute information of points belonging to higher levels.

[0043] As a specific method for determining the LoD, for example, the method described in Non-Patent Document 1 mentioned above may be used.

[0044] The inverse lifting unit 2100 is configured to decode attribute information of each point based on the hierarchical structure defined by the LoD, using the LoD generated by the LoD calculation unit 2090 and the inverse-quantized residual information generated by the inverse quantization unit 2070. As a specific process of inverse lifting, for example, the method described in the above-mentioned Non-Patent Document 1 can be used.

[0045] The inverse color conversion unit 2110 is configured to perform inverse color conversion processing on the attribute information output from the RAHT unit 2080 or the inverse lifting unit 2100 when the attribute information to be decoded is color information and color conversion has been performed on the point cloud encoding device 100 side. Whether or not to perform such inverse color conversion processing is determined by the control data decoded by the attribute information decoding unit 2060.

[0046] The point cloud decoding device 200 is configured to decode and output attribute information of each point in the point cloud through the above processing. (Geometric Information Decoding Unit 2010) The control data decoded by the geometric information decoding unit 2010 will be described below with reference to Figures 3 and 4.

[0047] FIG. 3 shows an example of the structure of coded data (bit stream) received by the geometric information decoding unit 2010.

[0048] First, the bitstream may include a GPS 2011. The GPS 2011 is also called a geometry parameter set and is a set of control data related to decoding of geometric information. A specific example will be described later. Each GPS 2011 includes at least GPS ID information for identifying each GPS 2011 when multiple GPS 2011 exist.

[0049] Second, the bitstream may include GSH2012A / 2012B. GSH2012A / 2012B is also called a geometry slice header or geometry data unit header, and is a collection of control data corresponding to a slice, which will be described later. Hereinafter, the term "slice" will be used, but "slice" can also be interpreted as "data unit." Specific examples will be described later. GSH2012A / 2012B includes at least GPS ID information for specifying the GPS2011 corresponding to each GSH2012A / 2012B.

[0050] Third, the bitstream may include slice data 2013A / 2013B following the GSH 2012A / 2012B. The slice data 2013A / 2013B includes data that encodes geometric information.

[0051] As described above, the bit stream is configured such that one GSH 2012A / 2012B and one GPS 2011 correspond to each slice data 2013A / 2013B.

[0052] As described above, the GPS id information is used to specify which GPS 2011 to refer to in the GSH 2012A / 2012B, so a common GPS 2011 can be used for multiple slice data 2013A / 2013B.

[0053] In other words, it is not necessary to transmit the GPS 2011 for each slice. For example, as shown in Fig. 3, the bit stream may be configured such that the GPS 2011 is not coded immediately before the GSH 2012B and slice data 2013B.

[0054] 3 is merely an example. As long as the GSH 2012A / 2012B and the GPS 2011 correspond to each slice data 2013A / 2013B, elements other than those described above may be added as components of the bit stream.

[0055] For example, as shown in Fig. 3, the bitstream may include a sequence parameter set (SPS) 2001. Similarly, when transmitted, the bitstream may be shaped into a configuration different from that shown in Fig. 3. Furthermore, the bitstream may be combined with a bitstream decoded by an attribute information decoding unit 2060 (described later) and transmitted as a single bitstream.

[0056] FIG. 4 shows an example of the syntax configuration of GPS2011.

[0057] Note that the syntax names explained below are merely examples. If the syntax functions explained below are similar, the syntax names may be different.

[0058] The GPS 2011 may include GPS id information (gps_geom_parameter_set_id) for identifying each GPS 2011.

[0059] 4 indicates how each syntax element is coded. ue(v) indicates an unsigned zeroth-order exponential-Golomb code, and u(1) indicates a 1-bit flag.

[0060] The GPS 2011 may include a flag (geom_tree_type) for controlling the tree type in the tree synthesis unit 2020 .

[0061] For example, when the value of geom_tree_type is "1", it may be defined that predictive geometry coding is used, and when the value of geom_tree_type is "0", it may be defined that Octree is used.

[0062] The GPS 2011 may include a flag (geom_angular_enabled) for controlling whether or not the tree synthesis unit 2020 performs processing in angular mode.

[0063] For example, when the value of geom_angular_enabled is "1", it may be defined that predictive geometry coding processing is performed as angular mode, and when the value of geom_angular_enabled is "0", it may be defined that predictive geometry coding processing is not performed as angular mode.

[0064] The GPS 2011 may include a flag (ptree_ang_azimuth_scaling_enabled) for controlling whether or not the adaptive azimuth angle quantization mode is in the angular mode in the tree synthesis unit 2020. The adaptive azimuth angle quantization mode is a mode in which adaptive azimuth angle quantization is performed according to the radius.

[0065] For example, when the value of ptree_ang_azimuth_scaling_enabled is "1", it may be defined that adaptive quantization of the azimuth angle according to the radius is performed, and when the value of ptree_ang_azimuth_scaling_enabled is "0", it may be defined that adaptive quantization of the azimuth angle according to the radius is not performed.

[0066] It may also be used as a flag to control whether or not to use the predictor list in predictor calculation (selection) in angular mode.

[0067] For example, if the value of ptree_azimuth_scaling_enabled is "1", it may be defined that a predictor list is used in the calculation of such a predictor, and if the value of ptree_ang_azimuth_scaling_enabled is "0", it may be defined that a predictor list is not used in the calculation of such a predictor.

[0068] The GPS 2011 may include a value (ptree_ang_azimuth_step_minus1) related to the laser rotation speed for use in calculating the predicted value of the azimuth angle in the tree synthesis unit 2020 in angular mode.

[0069] (Tree Merging Unit 2020) An example of the operation of the tree merging unit 2020 will be described below with reference to FIGS.

[0070] 17 is a flowchart showing an example of processing in the tree merging unit 2020. Note that the following describes an example in which trees are merged using "Predictive coding."

[0071] Predictive coding is also called predictive geometry coding, predictive geometry, or predictive tree.

[0072] Predictive geometry coding is a method for decoding the position information of the point cloud data by decoding the residual of the position information predicted based on an arbitrary tree structure determined by the point cloud encoding device 100 and the position information of the point cloud data, and adding the two together.

[0073] As shown in FIG. 17, in step S501, the tree synthesis unit 2020 determines whether to use inter prediction based on the value of interprediction_enabled_flag.

[0074] If inter prediction is to be used, the tree synthesis unit 2020 proceeds to step S502; if inter prediction is not to be used, the tree synthesis unit 2020 proceeds to step S505.

[0075] In step S502 , the tree synthesis unit 2020 obtains a reference frame from the frame buffer 2120 .

[0076] The frame buffer 2120 may store one previously decoded frame, and a decoded frame may be added to the frame buffer 2120 each time the decoding of one or a predetermined number of frames is completed. After obtaining the reference frame, the tree synthesis unit 2020 proceeds to step S503.

[0077] In step S503, the tree synthesis unit 2020 determines whether or not to perform global motion compensation based on the global_motion_enabled_flag.

[0078] If global motion compensation is to be performed, the tree compositing unit 2020 proceeds to step S504; if global motion compensation is not to be performed, the tree compositing unit 2020 proceeds to step S505.

[0079] In step S504, the tree synthesis unit 2020 performs global motion compensation on the reference frame obtained in step S502.

[0080] Global motion compensation is a process for correcting a global positional shift for each frame, and applies rotation and translation to all points in the reference frame or to a group of points within a specified range based on the global motion vector decoded by the geometric information decoding unit 2010. After global motion compensation, the tree synthesis unit 2020 proceeds to step S505.

[0081] In step S505, the tree merging unit 2020 decodes the slice data. Specific processing in step S505 will be described later. After decoding the slice data, the tree merging unit 2020 proceeds to step S506.

[0082] In step S506, the tree merging unit 2020 ends the process.

[0083] The processes of steps S503 and S504, that is, the determination and execution of global motion compensation, may be performed during the slice data decoding process of step S505.

[0084] FIG. 15 is a flowchart showing an example of the slice data decoding process in step S505 described above.

[0085] As shown in FIG. 15, in step S1601, the tree synthesis unit 2020 determines whether decoding of the position information of all point cloud data included in the slice has been completed.

[0086] This process can be performed, for example, by transmitting information indicating the number of point cloud data contained in the slice to the GSH, and comparing this number of point cloud data with the number of data already processed to determine whether processing of all points has been completed.

[0087] If the decoding of the position information of all point cloud data has been completed, the operation proceeds to step S1613, where the processing ends. If the decoding of the position information of all point cloud data has not been completed, the operation proceeds to step S1602.

[0088] In step S1602, the tree merging unit 2020 sets the parent node of the node to be decoded (node ​​to be processed) of the point cloud data.

[0089] For example, the tree synthesis unit 2020 decodes the number of child nodes of each node to be decoded, and stores the indexes of the nodes to be decoded for the number of child nodes.

[0090] When the tree synthesis unit 2020 processes a node to be decoded after a certain node, it may refer to the array of indexes of the node, obtain one index stored at the end of the array, and set the node of the obtained index as the parent node of the node to be decoded.

[0091] After the setting of the parent node is completed, the operation proceeds to step S1603.

[0092] In step S1603, the tree merging unit 2020 determines whether to perform processing in angular mode.

[0093] For example, the tree synthesis unit 2020 can refer to the value of the above-mentioned geom_angular_enabled to determine whether to perform processing in angular mode.

[0094] If processing is to be performed in angular mode, the operation proceeds to step S1604, and if processing is not to be performed in angular mode, the operation proceeds to step S1610.

[0095] In step S1604, the tree synthesis unit 2020 decodes the predictor information and spherical coordinate residuals to be used in step S1605. Here, the spherical coordinate residuals indicate the residuals of the radius, azimuth angle, and laser ID. Once this decoding is complete, the operation proceeds to step S1605.

[0096] In step S1605, the tree synthesis unit 2020 predicts the position information based on the predictor information decoded in step S1604. Here, the predictor information is a predictor index or a prediction mode. A specific method for predicting the position information will be described later.

[0097] After the prediction of the position information is completed, the operation proceeds to step S1606.

[0098] In step S1606, the tree synthesis unit 2020 reconstructs the spherical coordinates by adding the decoded spherical coordinate residual and the predictor.

[0099] After the reconfiguration is completed, the operation proceeds to step S1607.

[0100] In step S1607, the tree synthesis unit 2020 reconstructs orthogonal integer coordinates. In this process, the tree synthesis unit 2020 can convert spherical coordinates into orthogonal integer coordinates based on the reconstructed spherical coordinates. A specific method for this can be realized, for example, by the technique described in Non-Patent Document 1.

[0101] After the reconstruction of the orthogonal integer coordinates is completed, the operation proceeds to step S1608.

[0102] In step S1608, the tree synthesis unit 2020 decodes the orthogonal integer coordinate residual.

[0103] After the decoding of the orthogonal integer coordinate residual is completed, the operation proceeds to step S1609.

[0104] In step S1609, the tree synthesis unit 2020 reconstructs the original coordinates by adding the decoded orthogonal integer coordinate residual and the reconstructed orthogonal integer coordinates.

[0105] After the reconstruction of the original coordinates is completed, the operation returns to step S1601.

[0106] In step S1610, the tree synthesis unit 2020 predicts the position information. Specifically, the tree synthesis unit 2020 selects a predictor and sets the selected predictor as the predicted value of the position information.

[0107] For example, the tree synthesis unit 2020 may select a predictor based on the decoded predictor mode from among a plurality of predictors calculated based on a tree structure.

[0108] After the prediction of the position information is completed, the operation proceeds to step S1611.

[0109] In step S1611, the tree synthesis unit 2020 decodes the orthogonal integer coordinate residual.

[0110] After the decoding of the orthogonal integer coordinate residual is completed, the operation proceeds to step S1612.

[0111] In step S1612, the tree synthesis unit 2020 reconstructs the original coordinates by adding the residual of the orthogonal integer coordinates decoded in step S1611 to the position information predicted in step S1610.

[0112] After the reconstruction of the original coordinates is completed, the operation returns to step S1601.

[0113] FIG. 18 is a flowchart showing an example of the process of predicting location information in step S1605 described above.

[0114] As shown in FIG. 18, in step S701, the tree synthesis unit 2020 decodes the predictor flag.

[0115] Here, the slice data may include a flag indicating a predictor to be used for each node. For example, the slice data may include flags similar to those described in Non-Patent Documents 1 and 2, such as a flag indicating whether the predictor is an inter predictor or an intra predictor, or an index of the inter predictor. Alternatively, the slice data may include other flags, which will be described later.

[0116] After decoding the predictor flag, the tree synthesis unit 2020 proceeds to step S702.

[0117] In step S702, the tree synthesis unit 2020 determines whether to use an inter predictor based on the flag decoded in step S701.

[0118] If an inter-predictor is to be used, the tree synthesis unit 2020 proceeds to step S704; if an inter-predictor is not to be used, the tree synthesis unit 2020 proceeds to step S703.

[0119] In step S703, the tree synthesis unit 2020 performs intra prediction on the coordinates of the node to be processed.

[0120] Here, in such intra prediction, the tree synthesis unit 2020 configures a predictor based on the coordinates of the parent or ancestor node of the node to be processed (for example, the parent node of the parent node) and predicts the coordinates of the node to be processed.

[0121] In the process of step S703, first, the tree synthesis unit 2020 determines the type of predictor to be used for prediction.

[0122] For example, the tree synthesis unit 2020 may determine whether or not the adaptive azimuth angle quantization mode is selected based on the value of ptree_ang_azimuth_scaling_enabled, and determine the type of predictor to be used.

[0123] The tree synthesis unit 2020 may select a predictor to use based on the decoded prediction mode from among multiple predictors calculated using a tree structure, for example, in the case of an adaptive azimuth angle quantization mode.

[0124] Alternatively, in the adaptive azimuth angle quantization mode, the tree synthesis unit 2020 may store the position information of the decoded nodes as predictors in a list, and refer to the list for the predictor that corresponds to the index of the decoded predictor to select the predictor to use.

[0125] Once the tree synthesis unit 2020 has determined the type of predictor to be used, it uses that predictor as the predicted value of the position information.

[0126] After the intra prediction is completed, the tree synthesis unit 2020 proceeds to step S705.

[0127] In step S704, the tree synthesis unit 2020 performs inter prediction on the coordinates of the node to be processed.

[0128] In such inter prediction, the tree synthesis unit 2020 selects a node corresponding to the node to be processed from the reference frame as a predictor, and sets the coordinates of the selected predictor as the predicted values ​​of the coordinates of the node to be processed. A method for selecting a predictor from the reference frame will be described later.

[0129] After the inter prediction is completed, the tree synthesis unit 2020 proceeds to step S705.

[0130] In step S705, the tree merging unit 2020 ends the process of step S1605.

[0131] 19 is a diagram showing an example of the process of selecting a predictor from a reference frame in step S704. It should be noted that the example in FIG. 19 assumes that the angular mode is used. In the angular mode, the point of the parent node of the node to be processed can be considered to have been decoded immediately before or before that.

[0132] In FIG. 19, a node having the same laser ID and a large azimuth angle as the parent node of the node to be processed is searched for in the reference frame, and the two with the smallest azimuth angles are set as predictor 1 and predictor 2, respectively.

[0133] For example, the tree synthesis unit 2020 may perform bidirectional prediction. An example of the operation of the tree synthesis unit 2020 when performing bidirectional prediction will be described below.

[0134] First, the tree synthesis unit 2020 may group frames to be processed into a fixed number, and change the processing order within each group.

[0135] For example, the tree synthesis unit 2020 regards eight frames as one group and processes frames with intra-group frame indices of 0 to 7 in the order 0, 7, 1, 2, 3, 4, 5, 6.

[0136] Here, the intra-group frame index is a number assigned to each frame to be processed in the group.

[0137] Furthermore, there may be two reference frames for inter prediction for each frame to be processed, and the reference frames may be frames that are in the future in time series.

[0138] The intra-group frame index order pattern and the frame to which each intra-group frame index refers may be decoded as a flag included in APS 2611 or ASH 2612.

[0139] Here, the intra-group frame index order pattern refers to the pattern of the order of the frame indexes within the group.

[0140] When performing bidirectional prediction, the tree synthesis unit 2020 may, for example, search two reference frames for nodes that have the same laser ID and a large azimuth angle as the parent node of the node being processed, and may use the two with the smallest azimuth angles from each reference frame as predictors, generating a total of four predictors.The tree synthesis unit 2020 may then use one of these predictors as a predictor based on the decoded predictor index.

[0141] Furthermore, the tree synthesis unit 2020 may prepare a list of reference frames for the frames referenced by each intra-group frame index, and select a frame to reference from this list based on the value of the decoded intra-list index.

[0142] In addition, the tree synthesis unit 2020 may prepare two lists of reference frames, one for past frames and one for future frames in chronological order from the processing frame, and may update these lists at the timing of processing each frame.

[0143] Furthermore, the tree synthesis unit 2020 may fix and hard-code the frames that each intra-group frame index refers to for each intra-group frame index order pattern.

[0144] For example, when performing bi-prediction, the tree synthesis unit 2020 may create one predictor from the two selected frames based on the predictor index of each decoded reference frame.

[0145] Specifically, the tree synthesis unit 2020 may use the linear prediction values ​​of two frames as a predictor. That is, the tree synthesis unit 2020 may use the average value of the predictors of the two reference frames as a predictor. Here, the tree synthesis unit 2020 may predict the azimuth angle and radius. The tree synthesis unit 2020 may use a value quantized in units of rotation speed for the azimuth angle. For example, the tree synthesis unit 2020 may assign a weight according to the distance between the reference frame and the frame to be processed.

[0146] In the example of Figure 19, the tree synthesis unit 2020 searches the reference frame for nodes that have the same laser ID and a large azimuth angle as the parent node of the node to be processed, and selects the two with the smallest azimuth angles as predictor 1 and predictor 2, respectively.

[0147] FIG. 16 is a flowchart showing an example of the intra prediction in step S703 described above.

[0148] As shown in FIG. 16, in step S1701, the tree synthesis unit 2020 determines whether or not the adaptive azimuth angle quantization mode is selected based on the value of ptree_ang_azimuth_scaling_enabled.

[0149] If the adaptive azimuth angle quantization mode is selected, the operation proceeds to step S1702. On the other hand, if the adaptive azimuth angle quantization mode is not selected, the operation proceeds to step S1703.

[0150] In step S1702, the tree synthesis unit 2020 decodes the predictor index. After the decoding of the predictor index is completed, the operation proceeds to step S1704.

[0151] In step S1703, the tree synthesis unit 2020 decodes the prediction mode. After the prediction mode has been decoded, the operation proceeds to step S1704.

[0152] In step S1704, the tree synthesis unit 2020 decodes the number of azimuth angle steps. After the decoding of the number of azimuth angle steps is completed, the operation proceeds to step S1705.

[0153] In step S1705, the tree synthesis unit 2020 decodes the spherical coordinate residual. The tree synthesis unit 2020 may perform this decoding using the method described in Non-Patent Document 2. After the decoding is complete, the operation proceeds to step S1706, where the process ends.

[0154] (Attribute Information Decoding Unit 2060) Hereinafter, the control data decoded by the attribute information decoding unit 2060 will be described with reference to FIGS.

[0155] FIG. 5 shows an example of the structure of coded data (bit stream) received by the attribute information decoding unit 2060. FIGS. 6 and 7 show examples of the syntax structure of APS 2611 and ASH 2612.

[0156] Note that the syntax names explained below are merely examples. If the syntax functions explained below are similar, the syntax names may be different.

[0157] The APS 2611 may include APS id information (aps_geom_parameter_set_id) for identifying each APS 2611.

[0158] 4 indicates how each syntax element is coded. ue(v) indicates an unsigned zeroth-order exponential-Golomb code, and u(1) indicates a 1-bit flag.

[0159] The APS 2611 may include a flag (attr_coding_type) for controlling whether the inverse quantization unit 2070 outputs the inverse quantized residual information to the RAHT unit 2080 or the LoD calculation unit 2090 .

[0160] For example, when the value of attr_coding_type is “1”, it may be defined to be output to the LoD calculation unit 2090 , and when the value of attr_coding_type is “0”, it may be defined to be output to the RAHT unit 2080 .

[0161] The APS 2611 may include a flag (raht_prediction_enabled) for controlling whether or not the RAHT unit 2080 predicts attribute information.

[0162] For example, when the value of raht_prediction_enabled is "1", it may be defined that attribute information is predicted, and when the value of raht_prediction_enabled is "0", it may be defined that attribute information is not predicted.

[0163] The APS 2611 may include a value (raht_prediction_threshold0) indicating a threshold value for the number of adjacent nodes of a grandparent node, which is used to determine whether or not to perform intra prediction of attribute information in the RAHT unit 2080. Here, the grandparent node refers to the parent node of the parent node of the node to be processed.

[0164] The APS 2611 may include a value (raht_prediction_threshold1) indicating a threshold value for the number of adjacent nodes of a parent node, which is used to determine whether or not to perform intra prediction of attribute information in the RAHT unit 2080.

[0165] APS2611 may include values ​​(raht_prediction_intra_elegibility_threshold0) and (raht_prediction_intra_elegibility_threshold1) indicating the threshold of the value obtained by dividing or subtracting the predicted value of the DC coefficient of the node to be processed, which is used to determine whether or not to perform intra-prediction of attribute information in the RAHT unit 2080, from the DC coefficient obtained by RAHT conversion.

[0166] The APS 2611 may include a flag (raht_subnode_prediction_enable_flag) that controls whether or not the RAHT unit 2080 uses subnodes to predict attribute information.

[0167] For example, if the value of raht_subnode_prediction_enable_flag is "1", it may be defined that subnodes are used to predict attribute information, and if the value of raht_subnode_prediction_enable_flag is "0", it may be defined that subnodes are not used to predict attribute information.

[0168] The APS 2611 may include weight parameters (raht_prediction_weights) used when the RAHT unit 2080 performs intra prediction of attribute information.

[0169] For example, the value of intra_prediction_weights may be defined according to the manner in which the node to be decoded is adjacent to the adjacent node used for intra prediction.

[0170] The APS 2611 may include a flag (raht_inter_prediction_enabled) for controlling whether or not the RAHT unit 2080 performs inter prediction of attribute information.

[0171] For example, when the value of raht_inter_prediction_enabled is "1", it may be defined that prediction of attribute information is performed, and when the value of raht_inter_prediction_enabled is "0", it may be defined that prediction of attribute information is not performed.

[0172] The APS 2611 may include a value (raht_inter_prediction_depth_minus1) indicating a layer for which inter prediction of attribute information is enabled in the RAHT unit 2080.

[0173] For example, if raht_inter_prediction_depth_minus1 is "N-1", inter prediction may be enabled in up to the top N layers of the Octree structure.

[0174] The APS 2611 may include a value (raht_send_inter_filters) indicating whether to transmit a scaling factor in inter prediction of the attribute information.

[0175] For example, when raht_send_inter_filters is "1", it may be defined that a scaling factor in inter prediction of attribute information is transmitted, and when raht_send_inter_filters is "0", it may be defined that a scaling factor in inter prediction of attribute information is not transmitted.

[0176] The APS 2611 may include a value (rhat_inter_skip_layers) indicating how many upper layers from the root node of the Octree are to be excluded from the application of scaling in inter prediction for the attribute information. Here, the root node refers to a node in the slice that has never been subjected to Octree division.

[0177] For example, when rf_inter_skip_layers is "3", it may be defined that inter prediction is not applied to the first to third layers.

[0178] The APS 2611 may include a value (raht_enable_code_layer) indicating whether to transmit an applicability mode of inter prediction for each layer. Alternatively, the APS 2611 may include raht_enable_code_layer when raht_prediction_enabled is “1” and raht_inter_prediction_enabled is “1”.

[0179] For example, if raht_enable_code_layer is '1', it may be defined that the inter prediction applicability mode for each layer is transmitted, and if raht_enable_code_layer is '0', it may be defined that the inter prediction applicability mode for each layer is not transmitted.

[0180] The APS 2611 may include a flag (biPredictionPrediod) indicating the method of predicting attribute information in the RAHT section 2080 .

[0181] For example, when biPredictionPrediod is "0", the prediction method of the attribute information may be defined as "no prediction" or "intra prediction", when biPredictionPrediod is "1", the prediction method of the attribute information may be defined as "no prediction", "intra prediction", or "inter prediction", and when biPredictionPrediod is "2", the prediction method of the attribute information may be defined as "no prediction", "intra prediction", "inter prediction", or "bidirectional prediction".

[0182] Note that the prediction method of the attribute information being "no prediction" means that the RAHT unit 2080 does not predict AC coefficients, and decoded AC coefficients are used as they are for inverse RAHT.

[0183] The APS 2611 may include a value (raht_send_inter_filters_intra) indicating whether to transmit a scaling factor in intra prediction of the attribute information.

[0184] For example, when raht_send_inter_filters_intra is "1", it may be defined that a scaling factor in intra prediction of attribute information is transmitted, and when raht_send_inter_filters_intra is "0", it may be defined that a scaling factor in intra prediction of attribute information is not transmitted.

[0185] When raht_inter_prediction_enabled is "1" and raht_enable_code_layer is "1", ASH2612 may include a value (layer_code_depth) indicating the number of modes (raht_attr_layer_code_mode) for determining whether inter prediction is applicable for each layer, which will be described later.

[0186] Alternatively, ASH 2612 may include layer_code_depth if either raht_enable_code_layer or raht_send_inter_filters is "1".

[0187] Alternatively, for example, ASH2612 may include layer_code_depth when only raht_send_inter_filters is "1".

[0188] Alternatively, layer_code_depth may be defined as a value obtained by subtracting 1 from the number of layers of the frame, or may be used by adding 1 after decoding.

[0189] Alternatively, if layer_code_depth is "0", layer_code_depth may be used as "0", and if it is other than "0", 1 may be subtracted from it after decoding.

[0190] Alternatively, layer_code_depth may be set to be equal to the smaller of the number of layers of the slice minus 1 and raht_inter_prediction_depth_minus1. When raht_enable_code_layer is “1”, ASH2612 may include the number of layers of layer_code_depth, and may include an inter prediction applicability mode (raht_attr_layer_code_mode) for each layer.

[0191] For example, in each layer, if inter prediction is applied, it may be defined as "1", and if inter prediction is not applied, it may be defined as "0".

[0192] Alternatively, raht_attr_layer_code_mode may be configured as a 3-bit field, and the flags indicated by the respective bits may be defined as follows:

[0193] The first bit may be defined as a value indicating "no prediction" or "prediction", with "no prediction" defined when the first bit is "0" and "prediction" defined when the first bit is "1".

[0194] The second bit may be defined as a value indicating the "prediction method," with "intra prediction" defined when the second bit is "0" and "inter prediction" defined when the second bit is "1."

[0195] The third bit may be defined as a value indicating the "method of inter prediction," with "inter prediction" defined when the third bit is "0," and "bidirectional prediction" defined when the third bit is "1."

[0196] Furthermore, the number of bits of raht_attr_layer_code_mode to be decoded may be determined according to the value of biPredictionPrediod.

[0197] For example, when biPredictionPrediod is "0", only the first bit of raht_attr_layer_code_mode may be decoded.

[0198] Note that, for example, when biPredictionPrediod is "0" and raht_attr_layer_code_mode is "1", the prediction method may be defined as intra prediction.

[0199] When biPredictionPrediod is "1", only the first and second bits of raht_attr_layer_code_mode may be decoded.

[0200] When biPredictionPrediod is "2", the first, second, and third bits of raht_attr_layer_code_mode may be decoded.

[0201] When raht_send_inter_filters is "1", the ASH 2612 may include residuals (raht_filter_taps) from the scaling factors equal to the number of scaling factors in inter prediction (num_filter_taps).

[0202] As shown in FIG. 7, raht_filter_taps may be decoded when raht_attr_layer_code_mode[i+raht_inter_skip_layers-1] is "1" in decoding raht_filter_taps[i].

[0203] The initial value of raht_filter_taps may be defined as "0".

[0204] Furthermore, when raht_attr_layer_code_mode[i+raht_inter_skip_layers-1] is "0", raht_filter_taps[i] may be set to the initial value "0".

[0205] Furthermore, in decoding raht_filter_taps[i], if raht_inter_skip_layers is "0", the initial value "0" may be set when i is "0".

[0206] FIG. 14 shows an example of a syntax configuration when raht_filter_taps is decoded based on raht_inter_skip_layers.

[0207] The following describes only the differences from the syntax configuration described in Fig. 7. When decoding raht_filter_taps[i], if raht_inter_skip_layers is "0", raht_filter_taps may be decoded even when i is "0".

[0208] num_filter_taps may be derived based on syntax that specifies the decoded layer to which inter prediction is applied.

[0209] An example of a method for deriving num_filter_taps will be described below.

[0210] num_filter_taps may be included in ASH2612 when raht_enable_code_layer is "0", or may be derived in the following manner when raht_enable_code_layer is "1".

[0211] For example, num_filter_taps may be derived based on a value (raht_inter_skip_layers) indicating how many top layers are exempt from inter-prediction scaling, a value (raht_inter_prediction_depth_minus1) indicating the number of effective layers for inter-prediction, and a value (layer_code_depth) indicating the number of raht_attr_layer_code_modes.

[0212] Here, the number of effective layers for inter prediction is a numerical value indicating a threshold value of layers to which inter prediction is applied. For example, the number of effective layers for inter prediction may be a value obtained by adding 1 to raht_inter_prediction_depth_minus1, and when raht_inter_prediction_depth_minus1 is "N-1", the number of effective layers for inter prediction may be defined as "N".

[0213] Specifically, for example, if the number of layers of the frame is greater than the number of effective layers for inter prediction, num_filter_taps may be obtained by subtracting a value indicating how many top layers are not to be subject to inter prediction scaling from the number of effective layers for inter prediction; if the value indicating the number of raht_attr_layer_code_mode is smaller than the number of effective layers for inter prediction, num_filter_taps may be obtained by subtracting a value indicating how many top layers are not to be subject to inter prediction scaling from the value indicating the number of raht_attr_layer_code_mode.

[0214] That is, num_filter_taps may be derived as follows:

[0215]

[0216] Alternatively, num_filter_taps may be derived as follows, regardless of the values ​​of raht_inter_skip_layers and raht_inter_prediction_depth_minus1.

[0217] ​num_filter_taps = layer_code_depth - raht_inter_skip_layers - 1 Alternatively, num_filter_taps may be derived as num_filter_taps = layer_code_depth - raht_inter_skip_layers, and in this case, when decoding raht_filter_taps, if raht_attr_layer_code_mode[i + raht_inter_skip_layers] is "1", raht_filter_taps[i] may be decoded.

[0218] Furthermore, the attribute information decoding unit 2060 may derive the number of scaling factors using the applicability mode of inter prediction for each layer.

[0219] Specifically, the attribute information decoding unit 2060 may count the layers to which inter prediction is applied, based on, for example, the inter prediction applicability mode for each layer.

[0220] However, the attribute information decoding unit 2060 may exclude from the count layers to which inter prediction scaling is not applicable, based on a value indicating up to which uppermost layers inter prediction scaling is not applicable.

[0221] Alternatively, the attribute information decoding unit 2060 may decode the scaling factor only if the layer is a layer to which inter prediction is applied, based on the inter prediction applicability mode for each layer.

[0222] However, the attribute information decoding unit 2060 may not decode layers to which inter prediction scaling is not applicable, based on a value indicating up to which uppermost layers inter prediction scaling is not applicable. When raht_send_inter_filters_intra is "1", the ASH 2612 may include the number of scaling factors in intra prediction (num_filter_taps_intra).

[0223] The ASH 2612 may include the residuals of the scaling factors (rhat_filter_taps_intra) in the number specified by num_filter_taps_intra.

[0224] num_filter_taps_intra may be derived based on syntax that specifies the layer to which decoded intra prediction is applied.

[0225] Although the above description has been given of an example in which the above information is decoded by the APS 2611, this information may be included in the ASH 2612 or the SPS 2601. In other words, this information may be included in any of the headers.

[0226] (RAHT Unit 2080) An example of the processing of the RAHT unit 2080 will be described with reference to FIGS.

[0227] FIG. 8 is a flowchart showing an example of processing by the RAHT unit 2080.

[0228] 8, in step S28001, the RAHT unit 2080 recursively divides the nodes into octrees until they reach a predetermined size, using a technique called Octree. After the division is complete, the operation proceeds to step S28002.

[0229] In step S28002, the RAHT unit 2080 counts the total number of points belonging to the layer below the node for each node divided by the Octree.

[0230] Specifically, the RAHT unit 2080 sequentially scans the nodes in a certain layer and records the number of points belonging to each node. Next, the RAHT unit 2080 adds up the numbers of points recorded in the child nodes of each node in the node one layer above to calculate the number of points belonging to each node.

[0231] The RAHT unit 2080 repeats the above scanning from the bottom layer to the top layer. The total number of acquired points is used as a weight for the inverse transformation of the RAHT in step S28005, which will be described later. After this calculation is completed, the operation proceeds to step S28003.

[0232] In step S28003, the RAHT unit 2080 decodes the DC coefficients of the nodes belonging to the highest layer of the Octree. Alternatively, the RAHT unit 2080 may calculate the DC coefficients by predicting the DC coefficients using intra prediction and decoding and adding up the prediction residuals of the DC coefficients.

[0233] After completing the decoding of the DC coefficient, the RAHT unit 2080 calculates the attribute value Aroot of the root node using the total number of points belonging to the root node acquired in step S28002, wroot, and the decoded DC coefficient DCroot, using the following formula.

[0234] After this calculation is completed, the operation proceeds to step S28004.

[0235] In step S28004, the RAHT unit 2080 determines whether or not the decoding of the attribute information of all nodes included in the layer has been completed.

[0236] If not completed, the operation proceeds to step S28005; if completed, the operation proceeds to step S28007.

[0237] In step S28005, the RAHT unit 2080 decodes the AC coefficients. Details will be described later. After the decoding is completed, the operation proceeds to step S28006.

[0238] In step S28006, the RAHT unit 2080 calculates attribute values ​​using the inverse transform of the RAHT based on the total number of points belonging to the lower layer of each aggregated node, the decoded AC coefficients, and the DC coefficients calculated from the nodes in the upper layer using the method described below.

[0239] Here, the inverse transformation of the RAHT is performed in units of 8 nodes (2×2×2) divided by the Octree.

[0240] Specifically, attribute value A 1 , A 2 , ...A k is the DC coefficient DC of a node that holds k subnodes, and the AC coefficient AC 1 , A.C. 2 , ...ACk-1 and the total number of points belonging to the lower layer of each subnode, w = w 1 , w 2 ,...w k is used to obtain the following equation (1).

[0241] Here, T(w) -1 is a matrix used for the inverse transformation of the RAHT, and can be generated by the method described in Non-Patent Document 1, for example.

[0242] It is assumed that this conversion process is performed repeatedly in the order from the higher-level nodes to the lower-level nodes,

[0243] is used as a DC coefficient in the inverse transform of the RAHT of each sub-node. After this transform process is completed, the operation proceeds to step S28004.

[0244] In step S28007, the RAHT unit 2080 determines whether the decoding of nodes in all layers has been completed.

[0245] If not, the operation moves to the next lower hierarchical level and proceeds to step S28004. If completed, the operation proceeds to step S28008 and ends the process.

[0246] FIG. 9 is a flowchart showing an example of the process in step S28004.

[0247] 9, in step S28101, the RAHT unit 2080 determines whether to predict AC coefficients. When making this determination, the RAHT unit 2080 may refer to the value of raht_prediction_enabled and use this value.

[0248] The RAHT unit 2080 may decode a flag indicating whether or not AC coefficient prediction is to be performed in the current node to be processed, and use the value of the flag.

[0249] The flag may be decoded for each node or for each layer. The flag may be decoded only if the value of raht_prediction_enabled is "1", which indicates that prediction is enabled. The flag may be included in slice data.

[0250] If the result of the determination is that AC coefficients are not to be predicted, the operation proceeds to step S28102, and if AC coefficients are to be predicted, the operation proceeds to steps S28103 and S28104.

[0251] In step S28102, the RAHT unit 2080 decodes the AC coefficients. After the decoding is completed, the operation proceeds to step S28113, where the process ends.

[0252] In step S28107, the RAHT unit 2080 determines whether inter prediction is enabled.

[0253] The RAHT unit 2080 may refer to the value of raht_inter_prediction_enabled and use this value for such determination.

[0254] If the result of the determination is that inter prediction is enabled, the operation proceeds to step S28109, and if the result is that inter prediction is disabled, the operation proceeds to step S28112.

[0255] In step S28109, the RAHT unit 2080 determines whether the depth of the layer including the node to be processed is equal to or less than a threshold. The RAHT unit 2080 may refer to the value of raht_inter_prediction_depth_minus1 and use this value as the threshold.

[0256] If the result of the determination is that the depth is equal to or less than the threshold, the operation proceeds to step S28110, and if the depth is greater than the threshold, the operation proceeds to step S28104.

[0257] In step S28110, the RAHT unit 2080 determines whether or not to perform inter prediction on the AC coefficients of the node to be processed.

[0258] The RAHT unit 2080 may make this determination by checking whether inter prediction is possible, and if so, not performing inter prediction. This will be described in detail later.

[0259] The RAHT unit 2080 may decode a flag indicating whether or not to perform inter prediction on the AC coefficients of the node to be processed, and use the value of the flag to make the determination. The flag may be decoded for each node or for each layer. The flag may be decoded and the determination may be made only when it is determined that inter prediction is feasible. The flag may be included in the slice data.

[0260] This flag may refer to layer_attr_layer_code_mode and use its value. This value may be referenced when the depth of the layer containing the target node is smaller than layer_code_depth and when the depth of the layer containing the target node is greater than the layer of the root node.

[0261] That is, this value may be referenced when depth-1<layer_code_depth and depth-1≧0.

[0262] Here, depth is a value that is defined as "0" at the root node level and is counted up as the level becomes deeper.

[0263] If not, it may be determined that inter prediction is not feasible.

[0264] If it is determined that inter prediction is feasible, the operation proceeds to step S28111, and if it is determined that inter prediction is not feasible, the operation proceeds to step S28104.

[0265] In step S28111, the RAHT unit 2080 performs inter prediction of the AC coefficients of the node to be processed. Specific details will be described later.

[0266] In step S28104, the RAHT unit 2080 determines whether or not to perform intra prediction of the AC coefficients of the node to be processed.

[0267] For example, the RAHT unit 2080 may determine whether the number of adjacent nodes of the parent node and grandparent node of the node to be processed is greater than or equal to a threshold, and may determine to perform intra prediction if the number is greater than or equal to the threshold, and may determine not to perform intra prediction if the number is less than the threshold (i.e., if it is determined that the accuracy of intra prediction of AC coefficients is not high).

[0268] The RAHT unit 2080 may refer to the value of raht_prediction_threshold0 described above and use this value as the threshold for the adjacent node of the grandparent node, or may refer to the value of raht_prediction_threshold1 described above and use this value as the threshold for the adjacent node of the parent node.

[0269] Alternatively, the RAHT unit 2080 may perform an additional determination for a processing target node for which intra prediction is determined to be performed in the determination using raht_prediction_threshold0 and raht_prediction_threshold1 described above.

[0270] That is, in step S28104, the RAHT unit 2080 determines the effect of intra prediction of AC coefficients of attribute values ​​using RAHT. In other words, in step S28104, the RAHT unit 2080 determines whether the accuracy of intra prediction of AC coefficients of attribute values ​​using RAHT is high.

[0271] For example, the RAHT unit 2080 may use the DC coefficient to determine whether or not to perform intra prediction (that is, whether or not the accuracy of intra prediction of AC coefficients is high).

[0272] Specifically, the RAHT unit 2080 may determine to perform intra prediction if the value obtained by dividing the DC coefficient obtained in step S28006 by the predicted value of the DC coefficient of the node to be processed is within a threshold range (i.e., if it is determined that the accuracy of the intra prediction of the AC coefficient is high), and may determine not to perform intra prediction if the value is outside the threshold range (i.e., if it is determined that the accuracy of the intra prediction of the AC coefficient is not high).

[0273] Alternatively, the RAHT unit 2080 may determine to perform intra prediction if the value obtained by subtracting the DC coefficient obtained in step S28006 from the predicted value of the DC coefficient of the node to be processed is within a threshold range (i.e., if it is determined that the accuracy of the intra prediction of the AC coefficient is high), and may determine not to perform intra prediction if the value is outside the threshold range (i.e., if it is determined that the accuracy of the intra prediction of the AC coefficient is not high).

[0274] The RAHT unit 2080 may refer to the values ​​of raht_prediction_intra_eligibility_threshold0 and raht_prediction_intra_eligibility_threshold1 described above and use these values ​​as thresholds.

[0275] Specifically, the RAHT unit 2080 may determine to perform intra prediction when the value obtained by dividing the DC coefficient obtained in step S28006 by the predicted value of the DC coefficient of the node to be processed, or the value obtained by subtracting the DC coefficient obtained in step S28006 from the predicted value of the DC coefficient of the node to be processed, is greater than or equal to the value of raht_prediction_intra_eligibility_threshold0 and less than or equal to the value of raht_prediction_intra_eligibility_threshold1 (i.e., when it is determined that the accuracy of intra prediction of AC coefficients is high).

[0276] Here, the predicted value of the DC coefficient is a value obtained simultaneously when the predicted value of the attribute value is converted into an AC coefficient in step S28207 described later, and can be obtained by performing the same processing in step S28104.

[0277] If it is determined that intra prediction is not to be performed, the operation proceeds to step S28102, and if it is determined that intra prediction is to be performed, the operation proceeds to step S28112.

[0278] In step S28112, the RAHT unit 2080 performs intra prediction of the AC coefficients of the node to be processed. Specific details will be described later.

[0279] In step S28103, the RAHT 2080 decodes the AC coefficient residuals. After the decoding is completed, the operation proceeds to step S28105.

[0280] In step S28105, the RAHT unit 2080 adds the residuals of the decoded AC coefficients to the predicted AC coefficients to reconstruct the AC coefficients. After the reconstruction is completed, the operation proceeds to step S28106, where the process ends.

[0281] The conditional branch in step S28109 may be omitted.

[0282] In the inter prediction process of step S28111, a process equivalent to the intra prediction process of step S28112 may also be performed, and prediction may be performed by combining the results of the inter prediction and the intra prediction. Specific details will be described later.

[0283] FIG. 10 is a flowchart showing an example of the intra prediction process in step S28112.

[0284] As shown in FIG. 10, in step S28201, the RAHT unit 2080 determines whether or not to perform intra prediction using adjacent nodes in the subnode hierarchy.

[0285] The RAHT unit 2080 may refer to the value of raht_subnode_prediction_enable_flag and use this value for the determination.

[0286] If the RAHT unit 2080 does not use adjacent nodes in the subnode hierarchy, it performs intra prediction using only adjacent nodes in the upper hierarchy.

[0287] Here, the adjacent nodes in the higher hierarchy are, of the 19 nodes in total, the six nodes adjacent to the parent node of the node to be decoded on the face, the 12 nodes adjacent to the edge, and the parent node itself, the three nodes adjacent to the face of the node to be decoded, the three nodes adjacent to the edge, and the seven nodes of the parent node itself.

[0288] FIG. 11 is a diagram showing the relationship between a decoding target node and adjacent nodes in a higher layer.

[0289] When using adjacent nodes in the subnode hierarchy, the RAHT unit 2080 performs intra prediction using adjacent nodes in the upper hierarchy and adjacent nodes in the subnode hierarchy.

[0290] Here, an adjacent node in the subnode hierarchy is a subnode of an adjacent node in a higher hierarchy, which has a face or an edge adjacent to the decoding target node and has already been decoded.

[0291] FIG. 12 is a diagram showing the relationship between a node to be decoded and adjacent nodes in the subnode hierarchy.

[0292] If the determination result is that intra prediction is to be performed without using adjacent nodes in the subnode hierarchy, this operation proceeds to step S28202, and if intra prediction is to be performed using adjacent nodes in the subnode hierarchy, this operation proceeds to step S28204.

[0293] In step S28202, the RAHT unit 2080 acquires the attribute value of the adjacent node in the upper layer. After acquiring the attribute value of the adjacent node in the upper layer, the operation proceeds to step S28203.

[0294] In step S28203, the RAHT unit 2080 predicts the attribute value of the node to be decoded.

[0295] The RAHT unit 2080 calculates the attribute values ​​attr of the k adjacent nodes in the higher hierarchy. i and the weight w according to the type of adjacent node i i The attribute value attr may be predicted using the following formula:

[0296] Here, the RAHT unit 2080 uses the weight w i Depending on whether the adjacent node i is a face adjacent node in a higher layer, an edge adjacent node in a higher layer, or a parent node, a hard-coded value may be used, or the weight w may be calculated from the value of raht_prediction_weights. i may be calculated.

[0297] After the prediction of the attribute value is completed, the operation proceeds to step S28207.

[0298] In step S28204, the RAHT unit 2080 acquires the attribute value of the adjacent node in the upper layer.

[0299] Here, the targets for acquiring attribute values ​​are adjacent nodes in a higher hierarchy whose subnodes have not yet been decoded, or adjacent nodes in a higher hierarchy whose subnodes have been decoded but whose faces or edges are not adjacent to the node to be decoded.

[0300] After the attribute value has been acquired, the operation proceeds to step S28205.

[0301] In step S28205, the RAHT unit 2080 acquires the attribute value of the adjacent node in the subnode hierarchy. After acquiring the attribute value of the adjacent node in the subnode hierarchy, the operation proceeds to step S28206.

[0302] In step S28206, the RAHT unit 2080 predicts the attribute value of the node to be decoded.

[0303] The RAHT unit 2080 calculates the attribute values ​​attr of the acquired k adjacent nodes in the upper layer and the adjacent nodes in the subnode layer. i and the weight w according to the type i of the adjacent node i The attribute value attr may be predicted using the following formula:

[0304] Here, the RAHT unit 2080 uses the weight w i Depending on whether the adjacent node i is a face adjacent node in a higher layer, an edge adjacent node in a higher layer, a parent node, a face adjacent node in a subnode layer, or an edge adjacent node in a subnode layer, a hard-coded value may be used, or the weight w may be calculated by referring to the value of raht_prediction_weights. i may be calculated.

[0305] After the attribute value prediction is completed, the operation proceeds to step S28207.

[0306] In step S28207, the RAHT unit 2080 converts the predicted attribute values ​​into AC coefficients. The AC coefficients are generated by performing RAHT on the predicted attribute values. For example, the RAHT unit 2080 may use the method described in Non-Patent Document 1 as the conversion method.

[0307] The RAHT unit 2080 calculates predicted AC coefficients of the transformed AC coefficients. intra For the scaling factor α intra By α intra You can double it.

[0308] AC pred = α intra ×AC intra Here, the coefficient α intra can be any real number. intra may be decoded for each node or for each layer. intra The coefficient α may be decoded as a syntax included in the APS 2611 or the ASH 2612, or may be included in the slice data. intra may be hard-coded.

[0309] For example, the coefficient α intra is defined as follows using the depth of the hierarchy, and the coefficient α intra Instead of α intra ' may be decrypted.

[0310] α intra = 1 + α intra ' x 2 -depth For example, an integer β may be defined to range from integer a to integer b, and the RAHT unit 2080 may decode the integer β.

[0311] The RAHT unit 2080 uses a coefficient α intra Regarding the integer β, it may be calculated as a value obtained by adding the integer c to the decoded integer β and then dividing the result by the integer c, as follows:

[0312] α intra =(β+c) / c Here, the RAHT unit 2080 may decode the integer β using exponential-Golomb coding.

[0313] Alternatively, for example, if the value of the decoded raht_filter_taps_intra is “X”, the RAHT unit 2080 subtracts X from 128, shifts the result of the subtraction to the right by 7 bits, and sets the result as the scaling factor α intra may be used for inter prediction as

[0314] For example, when the value of rf_filter_taps_intra is “0”, the scaling factor α in the inter prediction of the attribute information is intra The value of may be defined as "1" obtained by subtracting 0 from 128 and shifting the result of the subtraction 7 bits to the right.

[0315] The RAHT unit 2080 may, for example, refer to raht_attr_layer_code_mode, and if it is determined that intra prediction is applied to the node to be processed, scale the intra prediction value using the decoded value of raht_filter_taps_intra.

[0316] After the conversion of the AC coefficients is completed, the operation proceeds to step S28208, where the process ends.

[0317] FIG. 13 is a diagram showing an example of the inter prediction process in step S28111.

[0318] The RAHT unit 2080 predicts the AC coefficients of the target node using information about a reference node, which is the corresponding node in a reference frame. Here, the information about the reference node may be its attribute value or AC coefficient. The reference frame may also refer to another decoded frame, and the information about the reference frame may be included in the previous frame buffer 2120.

[0319] The RAHT unit 2080 may apply the same Octree structure as the frame to be processed to the reference frame. In such a case, a node may be set at a position where there is no point. Such a node is called an empty node. If the reference node is an empty node, the RAHT unit 2080 may disable inter prediction in step S28110.

[0320] The RAHT unit 2080 may apply an Octree to the reference frame independently of the current frame and set an Octree structure different from that of the current frame. In such a case, a node may not necessarily exist at the same position as in the current frame. If a reference node is not found at a position corresponding to the current node, the RAHT unit 2080 may disable inter prediction in step S28143.

[0321] If the reference node is an empty node, or if the reference node cannot be found, the RAHT unit 2080 may estimate and interpolate the information of the reference node using information of nodes in nearby positions within the reference frame.

[0322] For example, the RAHT unit 2080 may estimate and interpolate the average value of the attribute values ​​or AC coefficients of the adjacent nodes, the nearest nodes, or the k nearest nodes relative to the reference node position as the attribute value or AC coefficient of the reference node, respectively.

[0323] The RAHT unit 2080 may apply the above-mentioned interpolation only to a specific layer or layers.

[0324] If it is determined that not decoding the AC coefficients of the attribute values ​​results in better coding efficiency, the RAHT unit 2080 may skip decoding the AC coefficients of the attribute values ​​of the hierarchical nodes below the node to be processed.

[0325] Specifically, if the number of nodes to be decoded within a parent node including the node to be processed is two or less, or if the value of the decoded AC coefficient is equal to or less than a threshold, or if the number of nodes to be decoded is two or less and the value of the decoded AC coefficient is equal to or less than a threshold, the RAHT unit 2080 may determine that it is more efficient to not decode the AC coefficient of the attribute value, and may skip decoding the AC coefficients of the hierarchical nodes subordinate to the node to be processed.

[0326] Here, the threshold value may be a hard-coded value, or may be a value that is used by referring to the value of raht_prediction_skip_threshold.

[0327] Furthermore, the skipping of decoding of AC coefficients in layers below the node to be processed may be applied only to a specific layer and thereafter.

[0328] The RAHT unit 2080 may predict the AC coefficients of the node to be processed from, for example, the attribute values ​​of the reference node.

[0329] Specifically, the RAHT unit 2080 calculates the value Attr of the decoded attribute value of the reference node. inter The predicted value Attr of the attribute value of the node to be processed is calculated using pred and the predicted value Attr pred By applying RAHT to the AC coefficients of the node to be processed, the predicted value AC pred It may also be possible to ask for

[0330] Attr pred = Attr inter AC pred =RAHT(Attr pred The RAHT unit 2080 may predict the AC coefficients of the node to be processed directly from the AC coefficients of the reference node, for example.

[0331] Specifically, the RAHT unit 2080 calculates the AC coefficient values ​​AC of the reference node using the RAHT in the reference frame. inter is calculated, and the calculated value is used as the predicted value AC of the AC coefficient of the node to be processed. pred It may also be possible to use the following.

[0332] AC pred =AC inter The RAHT unit 2080 may obtain the AC coefficients of the reference node by recording the AC coefficients of each node of the reference frame in the frame buffer 2120 and referring to the values ​​in the frame buffer 2120. In this case, if there are no AC coefficients of the reference node in the frame buffer 2120, the RAHT unit 2080 may determine in step S28110 that inter prediction is not executable.

[0333] The RAHT unit 2080 inter and A.C. inter may be multiplied by a scaling factor α.

[0334] Attr pred = αAttr inter Or AC pred = αAC inter The coefficient α may be any real number. The coefficient α may be decoded for each node or for each layer. The coefficient α may be decoded as syntax included in the APS 2611 or the ASH 2612, or may be included in the slice data.

[0335] For example, the coefficient α may be defined using the depth of the hierarchy as follows, and α′ may be decoded instead of the coefficient α.

[0336] α = 1 + α' 2 -depth For example, the integer β may be defined as an integer ranging from integer a to integer b, and β may be decoded. The coefficient α may be calculated by adding an integer c to the decoded β and then dividing the result by the integer c, as follows:

[0337] α=(β+c) / c The integer β may be decoded using an exponential-Golomb code.

[0338] Alternatively, for example, if the value of the decoded raht_filter_taps is "X", the RAHT unit 2080 may subtract X from 128, shift the result of the subtraction to the right by 7 bits, and use the result as the scaling factor α for inter prediction.

[0339] For example, when the value of raht_filter_taps is "0", the value of the scaling factor α for inter prediction in the attribute information may be defined as "1" obtained by subtracting 0 from 128 and shifting the result of the subtraction 7 bits to the right.

[0340] The RAHT unit 2080 determines whether to apply inter-prediction, for example, based on syntax specifying the layer to which inter-prediction is applied, and if it determines that inter-prediction is to be applied to such layer, it may scale the inter-prediction value using the decoded value of raht_filter_taps.

[0341] On the other hand, if it is determined that inter prediction scaling should not be applied to the layer, the RAHT unit 2080 does not need to scale the inter prediction.

[0342] Specifically, the RAHT unit 2080 may determine to scale the inter-predicted value if the depth of the layer containing the node to be processed is equal to or less than the number of effective layers for inter-prediction, and if the depth of the layer containing the node to be processed is equal to or greater than a value indicating how many upper layers are exempt from inter-prediction scaling.

[0343] Here, the RAHT unit 2080 may refer to the value of raht_inter_prediction_depth_minus1 and use this value as the number of effective layers for inter prediction.

[0344] Furthermore, the RAHT unit 2080 may refer to the value of raht_inter_skip_layers and use that value as the value indicating up to which upper layer the scaling of inter prediction is not applied.

[0345] Alternatively, the RAHT unit 2080 may determine to scale the inter-predicted value, for example, when the depth of the layer containing the node to be processed is equal to or less than the number of effective layers for inter-prediction, and the depth of the layer containing the node to be processed is equal to or greater than a value indicating how many upper layers are exempt from inter-prediction scaling, and it is determined that inter-prediction is applied to the layer containing the node to be processed.

[0346] Here, the RAHT unit 2080 may refer to the value of the above-mentioned raht_attr_layer_code_mode in the layer that includes the node to be processed, and determine whether or not inter prediction is to be applied based on this value.

[0347] Alternatively, the RAHT unit 2080 may determine to scale the inter-predicted value if the depth of the layer containing the node to be processed is equal to or greater than a value indicating how many upper layers are exempt from inter-prediction scaling.

[0348] Here, the RAHT unit 2080 may refer to the value of raht_inter_skip_layers and use the value as the value indicating up to which uppermost layer inter prediction scaling is not applied.

[0349] Although the above description has been given of the case where there is one scaling factor for each layer, for example, even when a scaling factor is transmitted for each frequency index idx of the AC coefficients, the number of scaling factors to be decoded can be derived by multiplying the number of scaling factors calculated above by the number of scaling factors for each layer. The number of scaling factors may be, for example, 7.

[0350] For example, the AC coefficients AC parent and the inter predicted value AC when the parent node is decoded. parent_inter may be calculated as follows using

[0351] α = AC parent / AC parent_inter For example, the RAHT unit 2080 calculates the AC coefficients AC of N neighboring nodes of the node to be decoded. neighbor1 , A.C. neighbor2 , ..., AC neighborN and the inter-predicted value AC when each adjacent node is decoded. neighbor_inter1 , A.C. neighbor_inter2 , ..., AC neighbor_interN and may be used to calculate α so as to minimize the cost.

[0352] The cost may be, for example, the sum of the AC coefficients of each adjacent node and the squared error of the predictor of the AC coefficients. The adjacent nodes may be, for example, only nodes adjacent to a face, or nodes adjacent to a face and nodes adjacent to an edge.

[0353] The RAHT unit 2080 may perform a similar operation in the inter prediction of DC coefficients in step S28003.

[0354] DC pred = αDC inter Here, the DC coefficient of the reference node is DC inter and the predicted value of the DC coefficient of the root node ispred Let's say.

[0355] Furthermore, the RAHT unit 2080 may calculate predicted values ​​of attribute values ​​or AC coefficients by combining inter prediction and intra prediction.

[0356] For example, an example in which the RAHT unit 2080 requests a prediction of an attribute value will be shown below.

[0357] Attr pred =W inter ・Attr inter +W intra ・Attr intra Here, Attr inter and Attr intra are the inter prediction and intra prediction of the attribute value, respectively. inter and W intra are the weights of inter prediction and intra prediction, respectively.

[0358] W inter and W intra may be determined depending on the depth of the layer to be processed so that the deeper the layer, the more importance is placed on intra prediction. For example, W inter =1-depth / N W intra =depth / N, where N is the maximum value of the depth of a layer for which inter prediction is enabled. The combination of inter prediction and intra prediction may be enabled only in a specific layer. For example, the combination of inter prediction and intra prediction may be enabled only when M<depth<N. M may be any real number less than N, and may be decoded as header information such as APS.

[0359] For example, the RAHT unit 2080 may perform bidirectional prediction. An example of the operation of the RAHT unit 2080 when performing bidirectional prediction will be described below.

[0360] First, the RAHT unit 2080 groups frames to be processed into a fixed number of groups, and processes the frames in a different order within each group.

[0361] For example, the RAHT unit 2080 may regard eight frames as one group and process frames with intra-group frame indexes from 0 to 7 in the order 0, 7, 1, 2, 3, 4, 5, 6.

[0362] Here, the intra-group frame index is a number assigned to each frame to be processed in the group.

[0363] Furthermore, there may be two reference frames for inter prediction for each frame to be processed, and the reference frames may be frames that are in the future in time series.

[0364] The intra-group frame index order pattern and the frames to which the intra-group frame indexes refer may be decoded as flags included in APS 2611 or ASH 2612.

[0365] Furthermore, the RAHT unit 2080 may refer to the value of the above-mentioned biPredictionPrediod and use this value to decode the intra-group frame index order pattern and the frame referenced by the intra-group frame index.

[0366] Here, the intra-group frame index order pattern refers to the pattern of the order of the frame indexes within the group.

[0367] Alternatively, raht_attr_layer_code_mode may include a flag indicating whether intra prediction, inter prediction, no prediction, or bidirectional prediction is to be performed for each layer.

[0368] Furthermore, when there are multiple intra-group frame index order patterns in bidirectional prediction, the above-mentioned raht_attr_layer_code_mode may include as many patterns as there are variations.

[0369] Furthermore, the RAHT unit 2080 may prepare a list of reference frames and select a frame to be referenced by each intra-group frame index from the list based on the value of the decoded index in the list.

[0370] In addition, the RAHT unit 2080 may prepare two lists of such reference frames, one for past frames and one for future frames in chronological order from the frame to be processed, and may update these lists at the timing of processing each frame.

[0371] Furthermore, the RAHT unit 2080 may fix and hard-code the frames that each intra-group frame index refers to for each intra-group frame index order pattern.

[0372] Furthermore, the RAHT unit 2080 may decode the above-mentioned raht_attr_layer_code_mode for each slice or for each layer.

[0373] (Point Cloud Encoding Device 100) The point cloud encoding device 100 according to this embodiment will be described below with reference to Fig. 20. Fig. 20 is a diagram showing an example of functional blocks of the point cloud encoding device 100 according to this embodiment.

[0374] As shown in FIG. 20 , the point cloud encoding device 100 includes a coordinate transformation unit 1010, a geometric information quantization unit 1020, a tree analysis unit 1030, an approximate surface analysis unit 1040, a geometric information encoding unit 1050, a geometric information reconstruction unit 1060, a color conversion unit 1070, an attribute transfer unit 1080, an RAHT unit 1090, an LoD calculation unit 1100, a lifting unit 1110, an attribute information quantization unit 1120, an attribute information encoding unit 1130, and a frame buffer 1140.

[0375] The coordinate transformation unit 1010 is configured to perform transformation processing from the three-dimensional coordinate system of the input point cloud to any different coordinate system. For example, the coordinate transformation may involve rotating the input point cloud to transform the x, y, and z coordinates of the input point cloud into any s, t, and u coordinates. Alternatively, as a variation of the transformation, the coordinate system of the input point cloud may be used as is.

[0376] The geometric information quantization unit 1020 is configured to quantize the position information of the input point group after coordinate transformation and remove points with overlapping coordinates. Note that when the quantization step size is 1, the position information of the input point group and the position information after quantization match. In other words, when the quantization step size is 1, it is equivalent to not performing quantization.

[0377] The tree analysis unit 1030 is configured to receive position information of the quantized point group as input, and to generate an occupancy code indicating at which node in the encoding target space a point exists, based on a tree structure described below.

[0378] In this process, the tree analysis unit 1030 is configured to recursively divide the encoding target space into rectangular parallelepipeds to generate a tree structure.

[0379] If a point exists within a rectangular parallelepiped, a tree structure can be generated by recursively dividing the rectangular parallelepiped into multiple rectangular parallelepipeds until the rectangular parallelepiped reaches a predetermined size. Each such rectangular parallelepiped is called a node. Each rectangular parallelepiped generated by dividing a node is called a child node, and the occupancy code is a value of 0 or 1 indicating whether or not the point is contained within the child node.

[0380] As described above, the tree analysis unit 1030 is configured to generate an occupancy code while recursively dividing a node until it reaches a predetermined size.

[0381] In this embodiment, a method called "Octree" can be used, which recursively performs octree division on the above-mentioned rectangular parallelepiped, always treating it as a cube, and a method called "QtBt" can be used, which performs quadtree division and binary tree division in addition to octree division.

[0382] Here, whether or not to use “QtBt” is transmitted to the point cloud decoding device 200 as control data.

[0383] Alternatively, predictive geometry coding using an arbitrary tree structure may be specified. In this case, the tree analysis unit 1030 determines the tree structure, and the determined tree structure is transmitted to the point cloud decoding device 200 as control data.

[0384] For example, the tree-structured control data may be configured so that it can be decoded according to the procedures described with reference to FIGS.

[0385] The approximate surface analysis unit 1040 is configured to generate approximate surface information using the tree information generated by the tree analysis unit 1030 .

[0386] For example, when decoding three-dimensional point cloud data of an object, if the point cloud is densely distributed on the surface of the object, approximate surface information is used to represent the area where the point cloud exists by approximating it with a small plane, rather than decoding each individual point cloud.

[0387] Specifically, the approximate surface analysis unit 1040 may be configured to generate approximate surface information using, for example, a method called "Trisoup." Furthermore, when decoding a sparse point cloud acquired by Lidar or the like, this process can be omitted.

[0388] The geometric information encoding unit 1050 is configured to generate a bitstream (geometric information bitstream) by encoding syntax such as the occupancy code generated by the tree analysis unit 1030 and the approximate surface information generated by the approximate surface analysis unit 1040. Here, the bitstream may include, for example, the syntax described in FIG. 4 .

[0389] The encoding process is, for example, a context-adaptive binary arithmetic coding process. Here, for example, the syntax includes control data (flags and parameters) for controlling the decoding process of the position information.

[0390] The geometric information reconstruction unit 1060 is configured to reconstruct the geometric information of each point of the point cloud data to be encoded (the coordinate system assumed by the encoding process, i.e., the position information after coordinate transformation in the coordinate transformation unit 1010) based on the tree information generated by the tree analysis unit 1030 and the approximate surface information generated by the approximate surface analysis unit 1040.

[0391] The frame buffer 1140 is configured to receive the geometric information reconstructed by the geometric information reconstruction unit 1060 as an input and store it as a reference frame.

[0392] The stored reference frame is read from the frame buffer 1140 and used as a reference frame when the tree analysis unit 1030 performs inter-prediction of a temporally different frame.

[0393] Here, which reference frame to use for each frame may be determined based on, for example, the value of a cost function representing encoding efficiency, and information on the reference frame to be used may be transmitted to the point cloud decoding device 200 as control data.

[0394] The color conversion unit 1070 is configured to perform color conversion when the input attribute information is color information. The color conversion does not necessarily have to be performed, and whether or not the color conversion process is to be performed is coded as part of the control data and transmitted to the point cloud decoding device 200.

[0395] The attribute transfer unit 1080 is configured to correct the attribute values ​​so as to minimize distortion of the attribute information, based on the position information of the input point cloud, the position information of the point cloud after reconstruction by the geometric information reconstruction unit 1060, and the attribute information after color change by the color conversion unit 1070. As a specific correction method, for example, the method described in Non-Patent Document 1 can be applied.

[0396] The RAHT unit 1090 is configured to receive as input the attribute information transferred by the attribute transfer unit 1080 and the geometric information generated by the geometric information reconstruction unit 1060, and to generate residual information for each point using a type of Haar transform called RAHT (Region Adaptive Hierarchical Transform).

[0397] The information to be decoded is the direct current component (DC coefficient) and alternating current component (AC coefficient) of the attribute information generated by using RAHT in the encoding process, and in the decoding process, it is converted into attribute information by using the inverse transform of RAHT.

[0398] As a specific example of the RAHT process, the method described in Non-Patent Document 1 can be used.

[0399] The LoD calculation unit 1100 is configured to receive the geometric information generated by the geometric information reconstruction unit 1060 as an input and generate an LoD (Level of Detail).

[0400] LoD is information for defining a reference relationship (a point to be referenced and a point to be referenced) to realize predictive coding, such as predicting attribute information of another point from attribute information of another point and encoding or decoding the prediction residual.

[0401] In other words, LoD is information that defines a hierarchical structure in which each point contained in geometric information is classified into multiple levels, and the attributes of points belonging to lower levels are encoded or decoded using the attribute information of points belonging to higher levels.

[0402] As a specific method for determining the LoD, for example, the method described in Non-Patent Document 1 mentioned above may be used.

[0403] The lifting unit 1110 is configured to generate residual information by a lifting process using the LoD generated by the LoD calculation unit 1100 and the attribute information after attribute transfer by the attribute transfer unit 1080 .

[0404] As a specific lifting process, for example, the method described in Non-Patent Document 1 above may be used.

[0405] The attribute information quantization unit 1120 is configured to quantize the residual information output from the RAHT unit 1090 or the lifting unit 1110. Here, a quantization step size of 1 is equivalent to no quantization being performed.

[0406] The attribute information encoding unit 1130 is configured to perform encoding processing using the quantized residual information, etc. output from the attribute information quantization unit 1120 as syntax, and to generate a bit stream related to the attribute information (attribute information bit stream).

[0407] The encoding process is, for example, a context-adaptive binary arithmetic coding process, where the syntax includes, for example, control data (flags and parameters) for controlling the decoding process of the attribute information.

[0408] Through the above processing, the point cloud encoding device 100 is configured to perform encoding processing using the position information and attribute information of each point in a point cloud as input, and to output a geometry information bit stream and an attribute information bit stream.

[0409] According to this embodiment, a DC coefficient is used to determine whether or not to apply intra-prediction to AC coefficients, and if it is determined in advance that the accuracy of intra-prediction is high, intra-prediction is performed, whereas if it is determined that the accuracy of intra-prediction is not high, no prediction is performed, thereby reducing the amount of code for the AC coefficients to be decoded and improving coding efficiency.

[0410] Furthermore, according to this embodiment, by scaling the intra-predicted attribute values ​​or the AC coefficients obtained by RAHTing those attribute values, prediction accuracy is improved, the residual to be decoded is reduced, and coding efficiency is improved.

[0411] Furthermore, the above-described point group encoding device 100 and point group decoding device 200 may be realized as a program that causes a computer to execute each function (each process).

[0412] In each of the above embodiments, the present invention has been described using the example of applying it to the point cloud encoding device 100 and the point cloud decoding device 200, but the present invention is not limited to such an example and can be similarly applied to a point cloud encoding / decoding system having the functions of the point cloud encoding device 100 and the point cloud decoding device 200.

[0413] According to this embodiment, for example, it is possible to improve the overall service quality in video communication, which makes it possible to contribute to Goal 9 of the Sustainable Development Goals (SDGs) led by the United Nations, which is to "Develop resilient infrastructure, promote sustainable industrialization and foster innovation."

[0414] 10... Point cloud processing system 100... Point cloud encoding device 1010... Coordinate transformation unit 1020... Geometric information quantization unit 1030... Tree analysis unit 1040... Approximate surface analysis unit 1050... Geometric information encoding unit 1060... Geometric information reconstruction unit 1070... Color conversion unit 1080... Attribute transfer unit 1090... RAHT unit 1100... LoD calculation unit 1110... Lifting unit 1120... Attribute information quantization unit 1130... Attribute information encoding unit 200... Point cloud decoding device 2010... Geometric information decoding unit 2020... Tree synthesis unit 2030... Approximate surface synthesis unit 2040... Geometric information reconstruction unit 2050... Inverse coordinate transformation unit 2060... Attribute information decoding unit 2070... Inverse quantization unit 2080... RAHT unit 2090...LoD calculation unit 2100...inverse lifting unit 2110...inverse color conversion unit

Claims

1. A point cloud decoding device, comprising: an attribute information decoding unit that decodes a value indicating up to which upper layer the scaling of inter prediction is not applied, and decodes a scaling factor of the 0th layer when the value is "0".

2. A point cloud decoding method, comprising: a step of decoding a value indicating up to which upper layer the scaling of inter prediction is not applied; and a step of decoding a scaling factor of the 0th layer when the value is "0".

3. A program for causing a computer to function as a point cloud decoding device, wherein the point cloud decoding device decodes a value indicating up to which upper layer the scaling of inter prediction is not applied, and comprises an attribute information decoding unit that decodes a scaling factor of the 0th layer when the value is "0".