Point cloud decoding device, point cloud decoding method, and program
Patent Information
- Application Number
- JP2023112555
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-07-07
- Publication Date
- 2025-06-13
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a point group decoding device, a point group decoding method, and a program. [Background technology]
[0002] Conventionally, a method is known in which the AC coefficient of an intra-predicted attribute value is added to a residual of the decoded AC coefficient to reconstruct the AC coefficient, and the attribute value is decoded by an inverse RAHT.
[0003] Also, a technique is known in which smoothing is performed on AC coefficients of intra-predicted attribute values based on predicted values of adjacent nodes. [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] G-PCC codec description, ISO / IEC JTC1 / SC29 / WG7 N00271 [Non-Patent Document 2] G-PCC 2nd Edition codec description, ISO / IEC JTC1 / SC29 / WG7 N00506 Summary of the Invention [Problem to be solved by the invention]
[0005] However, in the conventional technology, when outliers are included in the smoothing process, there is a problem that the smoothing process is significantly affected by the outliers.
[0006] Therefore, the present invention has been made in consideration of the above-mentioned problems, and aims to provide a point cloud decoding device, a point cloud decoding method, and a program that can improve the encoding efficiency of attribute information encoding. [Means for solving the problem]
[0007] A first feature of the present invention is a point cloud decoding device comprising a RAHT unit that performs smoothing using clipping, using attribute values intra-predicted at each subnode within the same parent node as the node to be decoded.
[0008] A second feature of the present invention is a point cloud decoding method comprising the steps of performing smoothing using clipping, using attribute values intra-predicted at each subnode within the same parent node as the node to be decoded.
[0009] A third feature of the present invention is a program for causing a computer to function as a point cloud decoding device, the point cloud decoding device comprising a RAHT unit that performs smoothing using clipping, using attribute values intra-predicted at each subnode within the same parent node as the node to be decoded.
[0010] A fourth feature of the present invention is summarized as a point group decoding device including a RAHT unit that applies a scaling factor to a predicted value of an AC coefficient or a predicted value of an attribute value in inter prediction of an AC coefficient of each node.
[0011] A fifth feature of the present invention is a point cloud decoding device comprising a RAHT unit that predicts a DC coefficient of each node, wherein the RAHT unit applies a scaling factor to a predicted value of the DC coefficient in inter prediction of the DC coefficient. Effect of the Invention
[0012] According to the present invention, it is possible to provide a point cloud decoding device, a point cloud decoding method, and a program that can improve the coding efficiency of attribute information coding. [Brief description of the drawings]
[0013] [Figure 1] FIG. 1 is a diagram showing an example of the configuration of a point cloud processing system 10 according to an embodiment. [Diagram 2] FIG. 2 is a diagram illustrating an example of functional blocks of a point group decoding device 200 according to an embodiment. [Diagram 3] FIG. 3 is a diagram showing an example of the configuration of encoded data (bit stream) received by the geometric information decoding unit 2010 of the point cloud decoding device 200 according to an embodiment. [Figure 4] FIG. 4 is a diagram showing an example of the syntax configuration of GPS2011. [Diagram 5] FIG. 5 is a flowchart showing an example of the operation of the tree synthesis unit 2020 of the point group decoding device 200 according to an embodiment. [Figure 6] FIG. 6 is a flowchart showing an example of the process of decoding the predictor information and the spherical coordinate residual in step S504. [Figure 7] FIG. 7 is a flowchart showing an example of the process of decoding the predictor information and the spherical coordinate residual in step S504. [Figure 8] FIG. 8 is a flowchart showing an example of the process of decoding the predictor information and the spherical coordinate residual in step S504. [Figure 9] FIG. 9 shows an example of the configuration of encoded data (bit stream) received by the attribute information decoding unit 2060 of the point group decoding device 200 according to an embodiment. [Figure 10] FIG. 10 is an example of the syntax configuration of the APS2611 shown in FIG. [Figure 11] FIG. 11 is a flowchart showing an example of processing by the RAHT unit 2080. [Figure 12] FIG. 12 is a flowchart showing an example of the process of step S28004. [Figure 13] FIG. 13 is a flowchart showing an example of the process of step S28104. [Figure 14] FIG. 14 is a flowchart showing an example of the intra prediction process in step S28112. [Figure 15] FIG. 15 is a diagram showing the relationship between a decoding target node and adjacent nodes in a higher layer. [Figure 16] FIG. 16 is a diagram showing the relationship between a decoding target node and adjacent nodes in the subnode hierarchy. [Figure 17] FIG. 17 is a flowchart showing an example of the intra prediction process in step S28112. [Figure 18] FIG. 18 is a flowchart showing an example of processing by the RAHT unit 2080. [Figure 19] FIG. 19 is a diagram showing an example of the inter prediction process in step S28111. [Figure 20] FIG. 20 is a diagram showing an example of functional blocks of the point group encoding device 100 according to this embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0014] Hereinafter, the embodiments of the present invention will be described with reference to the drawings. Note that the components in the following embodiments can be appropriately replaced with existing components, and various variations including combinations with other existing components are possible. Therefore, the description of the following embodiments does not limit the content of the invention described in the claims.
[0015] (First embodiment) A point cloud processing system 10 according to a first embodiment of the present invention will be described below with reference to Fig. 1 to Fig. 20. Fig. 1 is a diagram showing a point cloud processing system 10 according to the embodiment.
[0016] As shown in FIG. 1, a point cloud processing system 10 includes a point cloud encoding device 100 and a point cloud decoding device 200.
[0017] The point cloud encoding device 100 is configured to generate encoded data (bit stream) by encoding an input point cloud signal. The point cloud decoding device 200 is configured to generate an output point cloud signal by decoding the bit stream.
[0018] The input point cloud signal and the output point cloud signal are composed of position information and attribute information of each point in the point cloud, such as color information and reflectance of each point.
[0019] Here, such a bit stream may be transmitted from the point group encoding device 100 to the point group decoding device 200 via a transmission path. Also, the bit stream may be stored in a storage medium and then provided from the point group encoding device 100 to the point group decoding device 200.
[0020] (Point Cloud Decoding Device 200) Hereinafter, the point group decoding device 200 according to this embodiment will be described with reference to Fig. 2. Fig. 2 is a diagram showing an example of functional blocks of the point group decoding device 200 according to this embodiment.
[0021] As shown in FIG. 2, the point cloud decoding device 200 includes a geometric information decoding unit 2010, a tree synthesis unit 2020, an approximate surface synthesis unit 2030, a geometric information reconstruction unit 2040, an inverse coordinate transformation unit 2050, an attribute information decoding unit 2060, an inverse quantization unit 2070, a RAHT unit 2080, a LoD calculation unit 2090, an inverse lifting unit 2100, an inverse color transformation unit 2110, and a frame buffer 2120.
[0022] The geometric information decoding unit 2010 is configured to receive as input a bit stream relating to geometric information (geometric information bit stream) out of the bit streams output from the point group encoding device 100, and to decode the syntax.
[0023] The decoding process is, for example, a context-adaptive binary arithmetic decoding process. Here, for example, the syntax includes control data (flags and parameters) for controlling the decoding process of the position information.
[0024] The tree synthesis unit 2020 is configured to receive as input the control data decoded by the geometric information decoding unit 2010 and an occupancy code indicating which node in the tree described below the point group exists in, and generate tree information indicating in which area in the space to be decoded the point exists.
[0025] The tree synthesis unit 2020 may be configured to perform the decoding process of the occupancy code within itself.
[0026] This process divides the space to be decoded into rectangles, determines whether a point exists in each rectangle by referring to the occupancy code, divides the rectangle in which the point exists into multiple rectangles, and then generates tree information by recursively repeating the process of referring to the occupancy code.
[0027] Here, when decoding the occupancy code, inter prediction, which will be described later, may be used.
[0028] In this embodiment, a method called "Octree" can be used, which always treats the above-mentioned rectangular parallelepiped as a cube and performs octree division recursively, and a method called "QtBt" can be used, which performs quadtree division and binary tree division in addition to octree division. Whether or not to use "QtBt" is transmitted from the point cloud encoding device 100 as control data.
[0029] Alternatively, when the control data specifies that predictive geometry coding is to be used, the tree synthesis unit 2020 is configured to decode the coordinates of each point based on an arbitrary tree structure determined in the point cloud encoding device 100.
[0030] The approximate surface synthesis unit 2030 is configured to generate approximate surface information using the tree information generated by the tree synthesis unit 2020, and to decode the point cloud based on the approximate surface information.
[0031] Approximate surface information is used when, for example, decoding three-dimensional point cloud data of an object, in cases where the point cloud is densely distributed on the object surface, to approximate the area in which the point cloud exists using a small plane, rather than decoding each individual point cloud.
[0032] Specifically, the approximate surface synthesis unit 2030 can generate approximate surface information and decode the point cloud using a method called "Trisoup", for example. A specific processing example of "Trisoup" will be described later. In addition, when decoding a sparse point cloud acquired by Lidar or the like, this processing can be omitted.
[0033] The geometric information reconstruction unit 2040 is configured to reconstruct geometric information (position information in the coordinate system assumed by the decoding process) of each point of the point cloud data to be decoded, based on the tree information generated by the tree synthesis unit 2020 and the approximate surface information generated by the approximate surface synthesis unit 2030.
[0034] The inverse coordinate transformation unit 2050 is configured to receive the geometric information reconstructed by the geometric information reconstruction unit 2040 as input, transform the information from the coordinate system assumed by the decoding process to the coordinate system of the output point cloud signal, and output position information.
[0035] The frame buffer 2120 is configured to store, as an input, the geometric information reconstructed by the geometric information reconstruction unit 2040 as a reference frame. When the tree synthesis unit 2020 performs inter-prediction of temporally different frames, the stored reference frame is read out from the frame buffer 2130 and used as the reference frame.
[0036] Here, which reference frame at which time is to be used for each frame may be determined based on control data transmitted from the point group encoding device 100 as a bit stream, for example.
[0037] The attribute information decoding unit 2060 is configured to receive as input a bit stream relating to attribute information (attribute information bit stream) out of the bit streams output from the point group encoding device 100, and to decode the syntax.
[0038] The decoding process is, for example, a context-adaptive binary arithmetic decoding process. Here, for example, the syntax includes control data (flags and parameters) for controlling the decoding process of the attribute information.
[0039] Moreover, the attribute information decoding unit 2060 is configured to decode the quantized residual information from the decoded syntax.
[0040] The inverse quantization unit 2070 is configured to perform inverse quantization processing based on the quantized residual information decoded by the attribute information decoding unit 2060 and a quantization parameter, which is one of the control data decoded by the attribute information decoding unit 2060, to generate inverse quantized residual information.
[0041] The dequantized residual information is output to either the RAHT unit 2080 or the LoD calculation unit 2090 according to the features of the point group to be decoded. Control data decoded by the attribute information decoding unit 2060 specifies which unit the dequantized residual information is output to.
[0042] The RAHT unit 2080 is configured to receive the inverse quantized residual information generated by the inverse quantization unit 2070 and the geometric information generated by the geometric information reconstruction unit 2040 as input, and to decode the attribute information of each point using a type of Haar transform (inverse Haar transform in the decoding process) called RAHT (Region Adaptive Hierarchical Transform). As a specific process of RAHT, for example, the method described in Non-Patent Document 1 can be used.
[0043] The LoD calculation unit 2090 is configured to receive the geometric information generated by the geometric information reconstruction unit 2040 as an input and generate a Level of Detail (LoD).
[0044] LoD is information for defining a reference relationship (a referencing point and a referenced point) to realize predictive coding, such as predicting attribute information of a certain point from attribute information of another point and encoding or decoding the prediction residual.
[0045] In other words, LoD is information that defines a hierarchical structure in which each point contained in geometric information is classified into multiple levels, and attributes of points belonging to lower levels are encoded or decoded using attribute information of points belonging to higher levels.
[0046] As a specific method for determining the LoD, for example, the method described in Non-Patent Document 1 mentioned above may be used.
[0047] The inverse lifting unit 2100 is configured to decode attribute information of each point based on the hierarchical structure defined by the LoD, using the LoD generated by the LoD calculation unit 2090 and the inverse quantized residual information generated by the inverse quantization unit 2070. As a specific process of inverse lifting, for example, the method described in the above-mentioned Non-Patent Document 1 can be used.
[0048] The inverse color conversion unit 2110 is configured to perform inverse color conversion processing on the attribute information output from the RAHT unit 2080 or the inverse lifting unit 2100 when the attribute information to be decoded is color information and color conversion has been performed on the point cloud encoding device 100 side. Whether or not to perform such inverse color conversion processing is determined by the control data decoded by the attribute information decoding unit 2060.
[0049] The point group decoding device 200 is configured to decode and output the attribute information of each point in the point group through the above processing.
[0050] (Geometric Information Decoding Part 2010) Hereinafter, the control data decoded by the geometric information decoding unit 2010 will be described with reference to FIGS.
[0051] FIG. 3 shows an example of the structure of the coded data (bit stream) received by the geometric information decoding unit 2010. In FIG.
[0052] First, the bit stream may include a GPS 2011. The GPS 2011 is also called a geometry parameter set, and is a set of control data related to decoding of geometric information. A specific example will be described later. Each GPS 2011 includes at least GPS ID information for identifying each GPS 2011 when there are multiple GPS 2011s.
[0053] Secondly, the bit stream may include GSH2012A / 2012B. GSH2012A / 2012B is also called a geometry slice header or geometry data unit header, and is a set of control data corresponding to a slice described later. In the following description, the term "slice" is used, but slice can also be read as data unit. A specific example will be described later. GSH2012A / 2012B includes at least GPS id information for specifying GPS2011 corresponding to each GSH2012A / 2012B.
[0054] Thirdly, the bitstream may include slice data 2013A / 2013B following the GSH 2012A / 2012B. The slice data 2013A / 2013B includes data in which geometric information is encoded.
[0055] As described above, the bit stream is configured such that each slice data 2013A / 2013B corresponds to one GSH 2012A / 2012B and one GPS 2011.
[0056] As described above, since the GPS ID information is used to specify which GPS 2011 to refer to in the GSH 2012A / 2012B, a common GPS 2011 can be used for a plurality of slice data 2013A / 2013B.
[0057] In other words, it is not necessary to transmit the GPS 2011 for each slice. For example, as shown in Fig. 3, the bit stream may be configured such that the GPS 2011 is not coded immediately before the GSH 2012B and the slice data 2013B.
[0058] 3 is merely an example. As long as the slice data 2013A / 2013B corresponds to the GSH 2012A / 2012B and the GPS 2011, elements other than those described above may be added as components of the bit stream.
[0059] For example, as shown in Fig. 3, the bitstream may include a sequence parameter set (SPS) 2001. Similarly, when transmitted, the bitstream may be shaped into a configuration different from that shown in Fig. 3. Furthermore, the bitstream may be combined with a bitstream decoded by an attribute information decoding unit 2060 (described later) and transmitted as a single bitstream.
[0060] FIG. 4 is an example of the syntax configuration of GPS2011.
[0061] Note that the syntax names described below are merely examples. If the syntax functions described below are similar, the syntax names may be different.
[0062] The GPS 2011 may include GPS ID information (gps_geom_parameter_set_id) for identifying each GPS 2011.
[0063] The Descriptor column in Fig. 4 indicates how each syntax is coded. ue(v) indicates an unsigned zeroth-order exponential Golomb code, and u(1) indicates a 1-bit flag.
[0064] The GPS 2011 may include a flag (geom_tree_type) for controlling the tree type in the tree synthesis unit 2020 .
[0065] For example, if the value of geom_tree_type is "1", it may be defined that predictive geometry coding is used, and if the value of geom_tree_type is "0", it may be defined that Octree is used.
[0066] The GPS 2011 may include a flag (geom_angular_enabled) for controlling whether or not the tree synthesis unit 2020 performs processing in angular mode.
[0067] For example, when the value of geom_angular_enabled is “1”, it may be defined that predictive geometry coding processing is performed as angular mode, and when the value of geom_angular_enabled is “0”, it may be defined that predictive geometry coding processing is not performed as angular mode.
[0068] The GPS 2011 may include a value related to the number of points in the same laser (angularNumPhiPerTurn) according to the laser ID of the point cloud acquisition device in the angular mode in the tree synthesis unit 2020. The number of points in the same laser is the number of points acquired in the same laser.
[0069] In addition, the number of points in the same laser is a unique value for each laser, and exists for each laser ID. For example, if there are 64 laser IDs, there will also be 64 points in the same laser.
[0070] The GPS 2011 may include a flag (ptree_ang_azimuth_scaling_enabled) for controlling whether or not the adaptive azimuth angle quantization mode is in the angular mode in the tree synthesis unit 2020. The adaptive azimuth angle quantization mode is a mode in which adaptive quantization of the azimuth angle is performed according to the radius.
[0071] For example, if the value of ptree_ang_azimuth_scaling_enabled is "1", it is defined that adaptive quantization of the azimuth angle according to the radius is performed, and if the value of ptree_ang_azimuth_scaling_enabled is "0", it is defined that adaptive quantization of the azimuth angle according to the radius is not performed.
[0072] It may also be used as a flag to control whether to use the predictor list when calculating (selecting) a predictor in Angular mode.
[0073] For example, if the value of ptree_azimuth_scaling_enabled is “1”, it may be defined that a predictor list is used in the calculation of the predictor, and if the value of ptree_ang_azimuth_scaling_enabled is “0”, it may be defined that a predictor list is not used in the calculation of the predictor.
[0074] The GPS 2011 may include a value (ptree_ang_azimuth_step_minus1) related to the laser rotation rate for use in calculating the predicted azimuth angle in the tree synthesis unit 2020 in angular mode.
[0075] The GPS 2011 may include a threshold value (resR_context_qphi_threshold) for the number of azimuth angle steps used when decoding the radial residual in the tree synthesis unit 2020 in the angular mode.
[0076] The GPS 2011 may include a flag (resR_context_qphi_threshold_present_flag) for controlling whether or not to transmit a threshold value for the number of azimuth angle steps to the decoder in the tree synthesis unit 2020 in the angular mode.
[0077] For example, when the value of resR_context_qphi_threshold_present_flag is "1", it may be defined that the threshold is transmitted to the decoder, and when the value of resR_context_qphi_threshold_present_flag is "0", it may be defined that the threshold is not transmitted to the decoder.
[0078] (Tree Synthesis Department 2020) An example of the operation of the tree synthesis unit 2020 will be described below with reference to FIGS.
[0079] 5 is a flowchart showing an example of processing in the tree synthesis unit 2020. Note that, below, an example will be described in which trees are synthesized using "Predictive geometry coding".
[0080] Predictive geometry coding is also called predictive tree coding. Predictive geometry coding is a method for decoding the residual of position information predicted based on an arbitrary tree structure determined by the point cloud encoding device 100 and the position information of the point cloud data, and adding the two together to decode the position information of the point cloud data.
[0081] As shown in FIG. 5, in step S501, the tree synthesis unit 2020 determines whether or not decoding of position information of all point cloud data included in the slice has been completed.
[0082] This process, for example, transmits information indicating the number of point cloud data contained in the slice to the GSH, and by comparing this number of point cloud data with the number of data already processed, it can be determined whether processing of all points has been completed.
[0083] If the decoding of the position information of all point cloud data is completed, the operation proceeds to step S513 and ends the process. If the decoding of the position information of all point cloud data is not completed, the operation proceeds to step S502.
[0084] In step S502, the tree merge unit 2020 sets a parent node of a node to be decoded (node to be processed) of the point cloud data.
[0085] For example, the tree synthesis unit 2020 decodes the number of child nodes of each node to be decoded, and stores the indexes of the nodes to be decoded for the number of child nodes.
[0086] When the tree synthesis unit 2020 processes a node to be decoded after a certain node, it may refer to the array of indexes of the node, obtain one index stored at the end of the array, and set the node of the obtained index as the parent node of the node to be decoded.
[0087] After the parent node setting is completed, the operation proceeds to step S503.
[0088] In step S503, the tree synthesis unit 2020 determines whether to perform processing in the Angular mode.
[0089] For example, the tree synthesis unit 2020 can refer to the value of the above-mentioned geom_angular_enabled to determine whether to perform processing in Angular mode.
[0090] If the processing is to be performed in Angular mode, the operation proceeds to step S504, and if the processing is not to be performed in Angular mode, the operation proceeds to step S510.
[0091] In step S504, the tree synthesis unit 2020 decodes the predictor information and the spherical coordinate residual. Here, the spherical coordinate residual indicates the residual of the radius, the azimuth angle, and the laser ID. When the decoding is completed, the operation proceeds to step S505.
[0092] In step S505, the tree synthesis unit 2020 predicts the position information based on the predictor information decoded in step S504. Here, the predictor information is a predictor index or a prediction mode.
[0093] In this process, the tree synthesis unit 2020 first determines the type of predictor to be used for prediction.
[0094] For example, the tree synthesis unit 2020 may determine whether or not to perform processing in adaptive azimuth angle quantization mode based on the value of ptree_ang_azimuth_scaling_enabled, and may determine the type of predictor to be used based on the result of this determination.
[0095] For example, in the case of an adaptive azimuth angle quantization mode, the tree synthesis unit 2020 may select a predictor to be used based on the decoded prediction mode from among a plurality of predictors calculated using a tree structure.
[0096] Alternatively, when performing processing in the adaptive azimuth angle quantization mode, the tree synthesis unit 2020 may store position information of the decoded nodes in a list as a predictor, and refer to the list for a predictor assigned to a decoded predictor index to select the predictor type to be used.
[0097] Once the type of predictor is determined, the tree synthesis unit 2020 uses the predictor as the predicted value of the position information.
[0098] After the prediction of the position information is completed, the operation proceeds to step S506.
[0099] In step S506, the tree synthesis unit 2020 reconstructs the spherical coordinates. In this process, the tree synthesis unit 2020 reconstructs the spherical coordinates by adding the decoded spherical coordinate residual and the predictor.
[0100] After the reconfiguration is completed, the operation proceeds to step S507.
[0101] In step S507, the tree synthesis unit 2020 reconstructs the orthogonal integer coordinates. In this process, the tree synthesis unit 2020 can convert the spherical coordinates into orthogonal integer coordinates based on the reconstructed spherical coordinates. A specific method for this can be realized by, for example, the method described in Non-Patent Document 1.
[0102] After the reconstruction of the orthogonal integer coordinates is completed, the operation proceeds to step S508.
[0103] In step S508, the tree synthesis unit 2020 decodes the orthogonal integer coordinate residual.
[0104] After the decoding of the orthogonal integer coordinate residual is completed, the operation proceeds to step S509.
[0105] In step S509, the tree synthesis unit 2020 reconstructs the original coordinates. In this process, the tree synthesis unit 2020 reconstructs the original coordinates by adding the decoded orthogonal integer coordinate residual and the reconstructed orthogonal integer coordinates.
[0106] After the reconstruction of the original coordinates is completed, the operation returns to step S501.
[0107] In step S510, the tree synthesis unit 2020 predicts the position information. Specifically, the tree synthesis unit 2020 selects a predictor and sets the predictor as the predicted value of the position information.
[0108] For example, the tree synthesis unit 2020 may select a predictor based on the decoded predictor mode from among a plurality of predictors calculated based on a tree structure.
[0109] After the prediction of the position information is completed, the operation proceeds to step S511.
[0110] In step S511, the tree synthesis unit 2020 decodes the orthogonal integer coordinate residual.
[0111] After the decoding of the orthogonal integer coordinate residual is completed, the operation proceeds to step S512.
[0112] In step S512, the tree synthesis unit 2020 reconstructs the original coordinates. In this process, the tree synthesis unit 2020 reconstructs the original coordinates by adding the residual of the orthogonal integer coordinates decoded in step S511 and the position information predicted in step S510.
[0113] After the reconstruction of the original coordinates is completed, the operation returns to step S501.
[0114] FIG. 6 is a flowchart showing an example of the process of decoding the predictor information and the spherical coordinate residual in step S504.
[0115] As shown in FIG. 6, in step S601, the tree synthesis unit 2020 determines whether or not the adaptive azimuth angle quantization mode is selected based on the value of ptree_ang_azimuth_scaling_enabled.
[0116] If the mode is the adaptive azimuth angle quantization mode, the operation proceeds to step S602, whereas if the mode is not the adaptive azimuth angle quantization mode, the operation proceeds to step S603.
[0117] In step S602, the tree synthesis unit 2020 decodes the predictor index. After the decoding of the predictor index is completed, the operation proceeds to step S604.
[0118] In step S603, the tree synthesis unit 2020 decodes the prediction mode. After the prediction mode has been decoded, the operation proceeds to step S604.
[0119] In step S604, the tree synthesis unit 2020 decodes the azimuth angle step number. After the azimuth angle step number has been decoded, the operation proceeds to step S605.
[0120] In step S605, the tree synthesis unit 2020 decodes the spherical coordinate residual. The tree synthesis unit 2020 may perform such decoding using the method described in Non-Patent Document 2. After the decoding is completed, the operation proceeds to step S606, where the process ends.
[0121] In the above, an example has been shown in which the azimuth angle step number is decoded as is, but for example, the tree synthesis unit 2020 may correct the decoded azimuth angle step number based on the number of points acquired within the same laser.
[0122] For example, as shown in FIG. 7, in step S701, the tree synthesis unit 2020 may correct the decoded azimuth angle step number based on the interval of the point group.
[0123] Specifically, the tree synthesis unit 2020 may correct the number of azimuth angle steps based on angularNumPhiPerTurn.
[0124] First, the tree synthesis unit 2020 calculates the maximum ratio of the number of points in the same laser.
[0125] Here, the maximum value ratio of the number of points in the same laser is the value obtained by dividing the number of points in the same laser corresponding to the laser ID of the parent node of the node to be decoded by the maximum value of the number of points in the same laser. Also, the maximum value of the number of points in the same laser is the maximum value among the values of the number of points in the same laser that exist as many times as the number of laser IDs.
[0126] For example, if the maximum value of the number of points in the same laser is 4000 and the number of points in the same laser corresponding to the laser ID of the parent node of the node to be decoded is 800, the maximum value ratio of the number of points in the same laser is 5.
[0127] In addition, the tree synthesis unit 2020 may calculate the maximum value ratio of the number of points within the same laser in step S701.
[0128] Alternatively, after decoding angularNumPhiPerTurn, the tree synthesis unit 2020 may calculate the maximum value ratio of the number of points in the same laser corresponding to each laser ID before step S701, and in step S701, obtain the maximum value ratio of the number of points in the same laser corresponding to the laser ID of the parent node of the node to be decoded.
[0129] For example, the tree synthesis unit 2020 may perform the above correction by adding the maximum value ratio of the number of points in the same laser to the decoded number of azimuth angle steps.
[0130] Alternatively, the tree synthesis unit 2020 may perform the above correction by multiplying the decoded azimuth angle step number by the maximum value ratio of the number of points within the same laser.
[0131] As described above, the tree synthesis unit 2020 may be configured to correct the number of decoded azimuth angle steps based on the number of points acquired within the same laser.
[0132] With this configuration, it is possible to improve the coding efficiency of the number of azimuth angle steps.
[0133] Alternatively, the tree combiner 2020 may correct the rotational speed of the laser based on, for example, the number of points acquired within the same laser.
[0134] For example, as shown in FIG. 8, in step S801, the tree synthesis unit 2020 may calculate the maximum value ratio of the number of points in the same laser using the method described above, and correct ptree_ang_azimuth_step_minus1 by dividing it by the maximum value ratio of the number of points in the same laser.
[0135] In step S801, the tree synthesis unit 2020 may calculate the maximum ratio of the number of points in the same laser. Alternatively, after decoding angularNumPhiPerTurn, the tree synthesis unit 2020 may calculate the maximum value ratio of the number of points in the same laser corresponding to each laser ID before step S801, and in step S801, obtain the maximum value ratio of the number of points in the same laser corresponding to the laser ID of the parent node of the node to be decoded.
[0136] ptree_ang_azimuth_step_minus1 is used in decoding the spherical coordinate residual in step S505.
[0137] As described above, the tree synthesis unit 2020 may be configured to correct the rotation speed of the laser based on the number of points acquired within the same laser.
[0138] With this configuration, it is possible to improve the coding efficiency of the number of azimuth angle steps.
[0139] In step S605, when determining the context to be used for decoding the radial residual, the tree synthesis unit 2020 may use a threshold related to the number of azimuth angle steps, make a judgment based on the threshold and the number of decoded azimuth angle steps, and determine the context based on the result.
[0140] The context may be determined, for example, by using the decoded predictor index and the decoded azimuth angle step number, and selecting one context index that satisfies a condition from among four context indexes ctxIdx using a threshold value for the azimuth angle step number as follows, and then based on that context index.
[0141]
number
[0142] For example, a specific value such as 0 may be hard-coded as the threshold value x. Alternatively, resR_context_qphi_threshold may be referenced and that value may be used.
[0143] Alternatively, for example, when it is determined not to transmit the threshold to the decoder by referring to the value of resR_context_qphi_threshold_present_flag, the threshold x may be derived using a syntax held by GPS2011. For example, the threshold x may be derived based on the value of ptree_ang_azimuth_step_minus1, which is a value related to the rotation speed of the laser used to calculate the predicted value of the azimuth angle. For example, the threshold x may be calculated as follows.
[0144]
number
[0145] (Attribute information decoding unit 2060) The control data decoded by the attribute information decoding unit 2060 will be described below with reference to FIGS.
[0146] FIG. 9 shows an example of the structure of the encoded data (bit stream) received by the attribute information decoding unit 2060, and FIG. 10 shows an example of the syntax structure of the APS 2611 shown in FIG.
[0147] Note that the syntax names described below are merely examples. If the syntax functions described below are similar, the syntax names may be different.
[0148] The APS 2611 may include APS id information (aps_geom_parameter_set_id) for identifying each APS 2611.
[0149] The Descriptor column in Fig. 10 indicates how each syntax is coded. ue(v) indicates an unsigned zeroth-order exponential Golomb code, and u(1) indicates a 1-bit flag.
[0150] The APS 2611 may include a flag (attr_coding_type) for controlling whether the inverse quantization unit 2070 outputs the inverse quantized residual information to the RAHT unit 2080 or the LoD calculation unit 2090.
[0151] For example, when the value of attr_coding_type is “1”, it may be defined that the data is output to the LoD calculation unit 2090 , and when the value of attr_coding_type is “0”, it may be defined that the data is output to the RAHT unit 2080 .
[0152] The APS 2611 may include a flag (raht_prediction_enabled) for controlling whether or not to perform prediction of attribute information in the RAHT unit 2080.
[0153] For example, when the value of raht_prediction_enabled is "1", it may be defined that prediction of attribute information is performed, and when the value of raht_prediction_enabled is "0", it may be defined that prediction of attribute information is not performed.
[0154] The APS 2611 may include a flag (raht_subnode_prediction_enable_flag) that controls whether or not the RAHT unit 2080 uses subnodes to predict attribute information.
[0155] For example, when the value of raht_subnode_prediction_enable_flag is "1", it may be defined that subnodes are used to predict attribute information, and when the value of raht_subnode_prediction_enable_flag is "0", it may be defined that subnodes are not used to predict attribute information.
[0156] The APS 2611 may include weight parameters (raht_prediction_weights) used when the RAHT unit 2080 performs intra-prediction of attribute information.
[0157] For example, the value of raht_prediction_weights may be defined according to the manner in which the node to be decoded is adjacent to an adjacent node used for intra prediction.
[0158] The APS 2611 may include a flag (raht_smoothing_enable_flag) that controls whether or not to perform smoothing after the RAHT unit 2080 performs intra prediction of the attribute information.
[0159] For example, when the value of raht_smoothing_enable_flag is "1", it may be defined that smoothing is performed after predicting the attribute information, and when the value of raht_smoothing_enable_flag is "0", it may be defined that smoothing is not performed.
[0160] The APS 2611 may include a weighting parameter (raht_smoothing_weighted_average_weights) for performing smoothing by weighted averaging after the RAHT unit 2080 performs intra prediction of attribute information.
[0161] For example, a maximum of eight such weight parameters may be defined according to the manner in which the node to be decoded is adjacent to each subnode of the same parent node of the node to be decoded.
[0162] The APS 2611 may include a weighting parameter (raht_smoothing_clipping_weights) for performing smoothing by clipping after the RAHT unit 2080 performs intra prediction of attribute information.
[0163] For example, a maximum of eight such weight parameters may be defined according to the manner in which the node to be decoded is adjacent to each subnode of the same parent node of the node to be decoded.
[0164] The APS 2611 may include a threshold (raht_smoothing_clipping_threshold) for performing smoothing by clipping after the RAHT unit 2080 performs intra prediction of attribute information.
[0165] The APS 2611 may include a flag (raht_inter_prediction_enabled) for controlling whether or not to perform inter prediction of attribute information in the RAHT unit 2080.
[0166] For example, when the value of raht_inter_prediction_enabled is "1", it may be defined that prediction of attribute information is performed, and when the value of raht_inter_prediction_enabled is "0", it may be defined that prediction of attribute information is not performed.
[0167] The APS 2611 may include a value (raht_inter_prediction_depth_minus1) indicating a layer at which inter prediction of attribute information is enabled in the RAHT unit 2080.
[0168] For example, when raht_inter_prediction_depth_minus1 is "N-1", inter prediction may be enabled in up to the top N layers of the Octree structure.
[0169] (RAHT Division 2080) An example of the processing of the RAHT unit 2080 will be described with reference to FIGS.
[0170] FIG. 11 is a flowchart showing an example of processing by the RAHT unit 2080.
[0171] 11, in step S28001, the RAHT unit 2080 recursively divides the nodes into octrees until a predetermined size is reached, using a method called Octree. After the division is completed, the operation proceeds to step S28002.
[0172] In step S28002, the RAHT unit 2080 counts up the total number of points belonging to the lower layer of each node divided by the octree.
[0173] Specifically, the RAHT unit 2080 scans the nodes of a certain layer in order and records the number of points belonging to each node. Next, the RAHT unit 2080 adds up the numbers of points recorded in the child nodes of each node in the node one layer above to calculate the number of points belonging to each node.
[0174] The RAHT unit 2080 repeats the above scanning from the lowest layer to the highest layer. The total number of acquired points is used as a weight for the inverse transformation of the RAHT in step S28005 described later. After this calculation is completed, the operation proceeds to step S28003.
[0175] In step S28003, the RAHT unit 2080 decodes the DC coefficient of the node belonging to the highest layer of the Octree. Alternatively, the RAHT unit 2080 may calculate the DC coefficient by predicting the DC coefficient using intra prediction and decoding and adding up the prediction residual of the DC coefficient.
[0176] After completing the decoding of the DC coefficients, the RAHT unit 2080 calculates the total number of points belonging to the root node acquired in step S28002, w root and the decoded DC coefficient DC root Using this, the attribute value A of the root node root Calculate.
[0177]
number
[0178] In step S28004, the RAHT unit 2080 determines whether or not decoding of the attribute information of all nodes included in the layer has been completed.
[0179] If not completed, the operation proceeds to step S28005; if completed, the operation proceeds to step S28007.
[0180] In step S28005, the RAHT unit 2080 decodes the AC coefficients. The details will be described later. After the decoding is completed, the operation proceeds to step S28006.
[0181] In step S28006, the RAHT unit 2080 calculates attribute values using the inverse transform of the RAHT based on the total number of points belonging to the lower hierarchy of each node, the decoded AC coefficients, and the DC coefficients calculated from the nodes in the upper hierarchy using the method described below.
[0182] Here, the inverse transformation of RAHT is performed in units of 8 nodes (2 x 2 x 2) divided by the Octree.
[0183] Specifically, attribute values A1, A2, … A k is the DC coefficient DC of a node that holds k subnodes, and the AC coefficients AC1,AC2,…AC k-1 and the total number of points belonging to the lower hierarchy of each subnode, w=w1, w2, … w k Using this, it can be calculated using the following formula (1).
[0184]
number
[0185] It is assumed that the conversion process is performed repeatedly from the higher-level nodes to the lower-level nodes,
[0186]
number
[0187] In step S28007, the RAHT unit 2080 determines whether the decoding of nodes in all layers is complete.
[0188] If not completed, this operation moves the processing target hierarchical level to the next lower hierarchical level and proceeds to step S28004, If completed, this operation proceeds to step S28008 and ends the processing.
[0189] FIG. 12 is a flowchart showing an example of the process of step S28004.
[0190] 12, in step S28101, the RAHT unit 2080 determines whether to predict AC coefficients. When making such a determination, the RAHT unit 2080 may refer to raht_prediction_enabled and use the value thereof.
[0191] The RAHT unit 2080 may decode a flag indicating whether or not to perform prediction of an AC coefficient in the currently processed node, and use the value of the flag.
[0192] The flag may be decoded for each node or for each layer. The flag may be decoded only if the value of raht_prediction_enabled is "1", indicating that prediction is enabled. The flag may be included in the slice data.
[0193] If the result of the determination is that the AC coefficients are not to be predicted, the operation proceeds to step S28102, and if the result is that the AC coefficients are to be predicted, the operation proceeds to steps S28103 and S28104.
[0194] In step S28102, the RAHT unit 2080 decodes the AC coefficients. After the decoding is completed, the operation proceeds to step S28106, and the process ends.
[0195] In step S28103, the RAHT 2080 decodes the AC coefficient residuals. After the decoding is completed, the operation proceeds to step S28105.
[0196] In step S28104, the RAHT unit 2080 predicts AC coefficients. Inter prediction or intra prediction may be used for predicting the AC coefficients.
[0197] The RAHT unit 2080 may first predict the attribute values, and then calculate the predicted values of the AC coefficients by the RAHT. This will be described in detail later. After the prediction of the AC coefficients is completed, the operation proceeds to step S28105.
[0198] In step S28105, the RAHT unit 2080 adds the residual of the decoded AC coefficients to the predicted AC coefficients to reconstruct the AC coefficients. After the reconstruction is completed, the operation proceeds to step S28106, and the process ends.
[0199] FIG. 13 is a flowchart showing an example of the process of step S28104.
[0200] As shown in Fig. 13, in step S28107, the RAHT unit 2080 determines whether inter prediction is enabled. The RAHT unit 2080 may refer to raht_inter_prediction_enabled for the determination and use the value. If the result of the determination is that inter prediction is enabled, this operation proceeds to step S28109, and if inter prediction is disabled, this operation proceeds to step S28112.
[0201] In step S28109, the RAHT unit 2080 determines whether the depth of the layer including the node to be processed is equal to or less than a threshold. The RAHT unit 2080 may refer to raht_inter_prediction_depth_minus1 as the threshold and use that value.
[0202] If the result of the determination is that the depth is equal to or less than the threshold, the operation proceeds to step S28110, and if the depth is greater than the threshold, the operation proceeds to step S28112.
[0203] In step S28110, the RAHT unit 2080 determines whether or not to perform inter prediction on the AC coefficient of the node to be processed.
[0204] For this determination, the RAHT unit 2080 may check whether inter prediction is possible, and perform inter prediction if possible, and may not perform inter prediction if not possible.
[0205] The RAHT unit 2080 may decode a flag indicating whether or not to perform inter-prediction on the AC coefficients of the node to be processed, and may use the value of the flag for the determination. The flag may be decoded for each node, or may be decoded for each layer. The flag may be decoded and a determination may be made only when it is determined that inter-prediction is executable. The flag may be included in slice data.
[0206] In step S28111, the RAHT unit 2080 performs inter prediction of the AC coefficients of the node to be processed. This will be described in detail later.
[0207] In step S28112, the RAHT unit 2080 performs intra prediction of the AC coefficients of the node to be processed. This will be described in detail later.
[0208] In step S28113, the process of step S28104 ends. Note that the conditional branch in step S28109 may be omitted.
[0209] In the inter prediction process in step S28111, a process equivalent to the intra prediction process in step S28112 may also be performed, and prediction may be performed by combining the results of inter prediction and intra prediction.
[0210] FIG. 14 is a flowchart showing an example of the intra prediction process in step S28112.
[0211] 14, in step S28201, the RAHT unit 2080 determines whether or not to perform intra prediction using adjacent nodes in the subnode hierarchy. The RAHT unit 2080 may refer to raht_subnode_prediction_enable_flag and use the value thereof for the determination.
[0212] When the RAHT unit 2080 does not use adjacent nodes in the subnode hierarchy, it performs intra prediction using only adjacent nodes in a higher hierarchy.
[0213] In this case, the adjacent nodes in the higher hierarchy are the six nodes adjacent to the parent node of the node to be decoded on the face side, the 12 nodes adjacent to the edge side, and the parent node itself, which are a total of 19 nodes, three nodes adjacent to the face side of the node to be decoded on the face side, three nodes adjacent to the edge side, and the seven nodes of the parent node itself.
[0214] FIG. 15 is a diagram showing the relationship between a decoding target node and adjacent nodes in a higher layer.
[0215] When using adjacent nodes in the subnode hierarchy, the RAHT unit 2080 performs intra prediction using adjacent nodes in a higher hierarchy and adjacent nodes in the subnode hierarchy.
[0216] Here, an adjacent node in the subnode hierarchy is a subnode of an adjacent node in a higher hierarchy that has a face or edge adjacent to the decoding target node and has already been decoded.
[0217] FIG. 16 is a diagram showing the relationship between a decoding target node and adjacent nodes in the subnode hierarchy.
[0218] If the determination result is that intra prediction is to be performed without using adjacent nodes in the subnode hierarchy, this operation proceeds to step S28202, and if the determination result is that intra prediction is to be performed using adjacent nodes in the subnode hierarchy, this operation proceeds to step S28204.
[0219] In step S28202, the RAHT unit 2080 acquires the attribute value of the adjacent node in the upper layer. After acquiring the attribute value of the adjacent node in the upper layer, this operation proceeds to step S28203.
[0220] In step S28203, the RAHT unit 2080 predicts the attribute value of the node to be decoded.
[0221] The RAHT unit 2080 obtains the attribute values attr i and the weight w according to the type of adjacent node i i Using these, the attribute value attr may be predicted using the following formula:
[0222]
number
[0223] After the prediction of the attribute value is completed, the operation proceeds to step S28207.
[0224] In step S28204, the RAHT unit 2080 acquires the attribute value of the adjacent node in the upper layer.
[0225] Here, the targets for which attribute values are to be obtained are adjacent nodes in a higher hierarchy whose subnodes have not yet been decoded, or adjacent nodes in a higher hierarchy whose subnodes have been decoded but whose subnodes do not have any subnodes adjacent to the node to be decoded via faces or edges.
[0226] After the attribute value has been acquired, the operation proceeds to step S28205.
[0227] In step S28205, the RAHT unit 2080 acquires the attribute value of the adjacent node in the subnode hierarchy. After acquiring the attribute value of the adjacent node in the subnode hierarchy, this operation proceeds to step S28206.
[0228] In step S28206, the RAHT unit 2080 predicts the attribute value of the node to be decoded.
[0229] The RAHT unit 2080 obtains the attribute values attr i and the weight w according to the type of adjacent node i i Using these, the attribute value attr may be predicted using the following formula:
[0230]
number
[0231] After the attribute value prediction is completed, the operation proceeds to step S28207.
[0232] In step S28207, the RAHT unit 2080 converts the predicted attribute values into AC coefficients. The AC coefficients are generated by performing RAHT on the predicted attribute values. For example, the RAHT unit 2080 may use the method described in Non-Patent Document 1 as such a conversion method.
[0233] The above describes an example in which the RAHT unit 2080 uses the attribute values predicted in step S28206 directly to convert the AC coefficients in step S28207, but the RAHT unit 2080 may smooth the predicted attribute values and then convert the AC coefficients.
[0234] For example, as shown in FIG. 17, after predicting the attribute value, the RAHT unit 2080 may determine whether or not to perform smoothing in step S1301.
[0235] In making such a determination, the RAHT unit 2080 may refer to raht_smoothing_enable_flag and use the value thereof.
[0236] If smoothing is to be performed, the operation proceeds to step S1302. If smoothing is not to be performed, the operation proceeds to step S28207.
[0237] In step S1302, the RAHT unit 2080 may smooth the attribute values.
[0238] For example, the RAHT unit 2080 calculates the smoothed attribute value Attr smoothing The predicted attribute value Attr i and weight α i Alternatively, the weighted average may be calculated as follows:
[0239]
number
[0240] In addition, the RAHT unit 2080 uses the weight α i A hard-coded value may be used as the weights, or the value may be referenced and used as the weights.
[0241] In addition, the RAHT unit 2080 performs, for example, the smoothed attribute value Attr smoothing The predicted value Attr0 of the decoded node itself, the predicted attribute value Attr of the subnode i other than the decoded node among the subnodes in the same parent node as the decoded node, i , weight β i and a threshold value Thr, may be used to perform clipping as follows:
[0242]
number
[0243] The clipping function Clip3 is
number
[0244] Here, for the target subnode i, the RAHT unit 2080 may treat the decode target node as a face-adjacent node, a face-adjacent node and an edge-adjacent node, or all subnodes within the same parent node.
[0245] In addition, the RAHT unit 2080 uses the weight βi A hard-coded value may be used as the weights, or the value may be referenced and used as the raht_smoothing_clipping_weights.
[0246] Furthermore, the RAHT unit 2080 may use a hard-coded value as the threshold value Thr, or may refer to raht_smoothing_clipping_threshold and use that value.
[0247] In the above, an example has been described in which the RAHT unit 2080 decodes AC coefficients of both the color difference signal and the luminance signal, but the RAHT unit 2080 may skip decoding AC coefficients of the color difference signal only in the bottom layer of the octree.
[0248] For example, as shown in FIG. 18, in step S1401, the RAHT unit 2080 may determine whether or not to skip decoding of AC coefficients of chrominance signals only in the bottom layer of the octree.
[0249] If skipping, the operation proceeds to step S1402, otherwise the operation proceeds to step S28004.
[0250] In step S1402, the RAHT unit 2080 determines whether the node to be decoded is at the bottom layer of the octree.
[0251] If it is the bottom layer, the operation proceeds to step S1403. If it is not the bottom layer, the operation proceeds to step S28004.
[0252] In step S1403, the RAHT unit 2080 decodes the AC coefficients other than the color difference signals.
[0253] The RAHT unit 2080 performs the same process as in step S28004 for decoding AC coefficients other than those of the color difference signals, sets the AC coefficients of the color difference signals to 0, and calculates attribute values in the subsequent step S28005.
[0254] After the decoding of AC coefficients other than the color difference signal is completed, the operation proceeds to step S28006.
[0255] FIG. 19 is a diagram showing an example of the inter prediction process in step S28111.
[0256] The RAHT unit 2080 predicts the AC coefficients of the processing target node using information of a reference node, which is a corresponding node in a reference frame. Here, the information of the reference node may be its attribute value or AC coefficient. The reference frame may indicate another decoded frame, and the information may be included in the previous frame buffer 2120.
[0257] The RAHT unit 2080 may apply the same octree structure as the processing target frame to the reference frame. In such a case, a node may be set at a position where there is no point. Such a node is called an empty node. If the reference node is an empty node, the RAHT unit 2080 may disable inter prediction in step S28110.
[0258] The RAHT unit 2080 may apply an Octree to the reference frame independently of the processing target frame, and set an Octree structure different from that of the processing target frame. In such a case, a node may not necessarily exist at the same position as that of the processing target frame. If a reference node is not found at a position corresponding to the processing target node, the RAHT unit 2080 may disable inter prediction in step S28143.
[0259] If the reference node is an empty node, or if the reference node cannot be found, the RAHT unit 2080 may estimate and interpolate the information of the reference node using information of nodes in nearby positions within the reference frame.
[0260] For example, the RAHT unit 2080 may estimate and interpolate the average value of the attribute values or AC coefficients of adjacent nodes, nearest neighbor nodes, or k-nearest neighbor nodes with respect to the reference node position as the attribute value or AC coefficient of the reference node, respectively.
[0261] The RAHT unit 2080 may predict the AC coefficients of the processing target node from, for example, the attribute values of the reference node.
[0262] Specifically, the RAHT unit 2080 converts the decoded attribute value Attr inter The predicted value Attr of the attribute value of the node to be processed is calculated using pred The predicted value Attr pred By applying RAHT to the target node, the predicted value AC pred It is also possible to ask for
[0263] Attr pred =Attr inter AC pred =RAHT(Attr pred ) The RAHT unit 2080 may predict the AC coefficients of the processing target node directly from the AC coefficients of the reference node, for example.
[0264] Specifically, the RAHT unit 2080 calculates the AC coefficient values AC of the reference nodes using the RAHT in the reference frame. inter The calculated value is used as the predicted value AC pred It is also possible to use the following.
[0265] AC pred =AC inter The RAHT unit 2080 may obtain the AC coefficients of the reference node by recording the AC coefficients of each node of the reference frame in the frame buffer 2120 and referring to the values in the frame buffer 2120. In this case, if there are no AC coefficients of the reference node in the frame buffer 2120, the RAHT unit 2080 may determine in step S28110 that inter prediction is not executable.
[0266] In addition, RAHT section 2080 is Attr inter and A.C. intermay be multiplied by a scaling factor α.
[0267] Attr pred =αAttr inter or AC pred = αAC inter The coefficient α may be any real number. The coefficient α may be decoded for each node or for each layer. The coefficient α may be included in the slice data.
[0268] For example, the coefficient α may be defined using the hierarchical depth depth as follows, and α′ may be decoded instead of the coefficient α.
[0269] α=1+α'·2 -depth For example, the integer β may be defined as an integer ranging from integer a to integer b, and β may be decoded. The coefficient α may be calculated by adding an integer c to the decoded β and then dividing the result by the integer c, as follows:
[0270] α=(β+c) / c The integer β may be decoded using an exponential-Golomb code.
[0271] Alternatively, the coefficient α may be derived at the decoder.
[0272] For example, the AC coefficients AC parent and the inter prediction value AC when the parent node is decoded. parent_inter may be used to calculate as follows:
[0273] α=AC parent / AC parent_inter For example, the AC coefficients AC neighbor1、 AC neighbor2、…、 AC neighborN and the inter prediction value AC neighbor_inter1、 AC neighbor_inter2、…、 ACneighbor_interN may be used to calculate α so that the cost is minimized. The cost may be, for example, the sum of squared errors of the AC coefficients of each adjacent node and the predictor of the AC coefficients. The adjacent nodes may be, for example, only the nodes adjacent to the faces, or the nodes adjacent to the faces and the nodes adjacent to the edges.
[0274] The RAHT unit 2080 may perform a similar operation in the inter prediction of the DC coefficient in step S28003.
[0275] DC pred = αDC inter Here, we use the DC coefficient of the reference node as DC inter The predicted value of the DC coefficient of the root node is DC pred Let us assume that.
[0276] Furthermore, the RAHT unit 2080 may calculate predicted values of attribute values or AC coefficients by combining inter prediction and intra prediction.
[0277] For example, an example in which the RAHT unit 2080 requests a prediction of an attribute value will be shown below.
[0278] Attr pred =W inter Attr inter +W intra Attr intra Here, Attr inter and Attr intra are the inter prediction and intra prediction of the attribute value, respectively. inter and W intra are the weights for inter-prediction and intra-prediction, respectively. inter and W intra may be determined such that the deeper the layer, the more importance is placed on intra prediction, depending on the depth of the layer to be processed. For example, W inter =1-depth / N W intra =depth / N Let N be the maximum depth of the hierarchy for which inter prediction is valid. The combination of inter prediction and intra prediction may be valid only at a specific hierarchy. For example, the combination of inter prediction and intra prediction may be valid only when M < depth < N. M is an arbitrary real number less than N and may be decoded as header information such as APS.
[0279] (Point cloud encoding device 100) Hereinafter, with reference to FIG. 20, the point cloud encoding device 100 according to the present embodiment will be described. FIG. 20 is a diagram showing an example of the functional blocks of the point cloud encoding device 100 according to the present embodiment.
[0280] As shown in FIG. 20, the point cloud encoding device 100 includes a coordinate conversion unit 1010, a geometric information quantization unit 1020, a tree analysis unit 1030, an approximate surface analysis unit 1040, a geometric information encoding unit 1050, a geometric information reconstruction unit 1060, a color conversion unit 1070, an attribute transfer unit 1080, a RAHT unit 1090, a LoD calculation unit 1100, a lifting unit 1110, an attribute information quantization unit 1120, an attribute information encoding unit 1130, and a frame buffer 1140.
[0281] The coordinate conversion unit 1010 is configured to perform conversion processing from the three-dimensional coordinate system of the input point cloud to an arbitrary different coordinate system. The coordinate conversion may, for example, convert the x, y, and z coordinates of the input point cloud to arbitrary s, t, and u coordinates by rotating the input point cloud. Also, as one variation of the conversion, the coordinate system of the input point cloud may be used as it is.
[0282] The geometric information quantization unit 1020 is configured to perform quantization of the position information of the input point cloud after coordinate conversion and removal of points with overlapping coordinates. Note that when the quantization step size is 1, the position information of the input point cloud and the quantized position information coincide. That is, when the quantization step size is 1, it is equivalent to the case where quantization is not performed.
[0283] The tree analysis unit 1030 is configured to receive position information of the quantized point group as input, and to generate an occupancy code indicating at which node in the encoding target space a point exists, based on a tree structure described below.
[0284] In this process, the tree analysis unit 1030 is configured to recursively divide the encoding target space into rectangular parallelepipeds to generate a tree structure.
[0285] If a point exists within a certain rectangular parallelepiped, a tree structure can be generated by recursively dividing the rectangular parallelepiped into multiple rectangular parallelepipeds until the rectangular parallelepiped reaches a specified size. Each such rectangular parallelepiped is called a node. Each rectangular parallelepiped generated by dividing a node is called a child node, and the occupancy code is expressed as 0 or 1 to indicate whether or not a point is included in the child node.
[0286] As described above, the tree analysis unit 1030 is configured to generate occupancy codes while recursively dividing nodes until a predetermined size is reached.
[0287] In this embodiment, a method called "Octree" can be used, which recursively performs octree division on the above-mentioned rectangular parallelepiped, always treating it as a cube, and a method called "QtBt" can be used, which performs quadtree division and binary tree division in addition to octree division.
[0288] Here, whether or not to use “QtBt” is transmitted to the point cloud decoding device 200 as control data.
[0289] Alternatively, predictive geometry coding using an arbitrary tree structure may be specified to be used. In this case, the tree analysis unit 1030 determines the tree structure, and the determined tree structure is transmitted to the point cloud decoding device 200 as control data.
[0290] For example, the tree-structured control data may be configured so as to be decoded according to the procedures described with reference to FIGS.
[0291] The approximate surface analyzer 1040 is configured to generate approximate surface information using the tree information generated by the tree analyzer 1030 .
[0292] Approximate surface information is used when, for example, decoding three-dimensional point cloud data of an object, in cases where the point cloud is densely distributed on the object surface, to approximate the area in which the point cloud exists using a small plane, rather than decoding each individual point cloud.
[0293] Specifically, the approximate surface analysis unit 1040 may be configured to generate approximate surface information using, for example, a method called "Trisoup." In addition, when decoding a sparse point cloud acquired by Lidar or the like, this process can be omitted.
[0294] The geometric information encoding unit 1050 is configured to generate a bit stream (geometric information bit stream) by encoding syntax such as the occupancy code generated by the tree analysis unit 1030 and the approximate surface information generated by the approximate surface analysis unit 1040. Here, the bit stream may include, for example, the syntax described in FIG. 4.
[0295] The encoding process is, for example, a context-adaptive binary arithmetic encoding process. Here, for example, the syntax includes control data (flags and parameters) for controlling the decoding process of the position information.
[0296] The geometric information reconstruction unit 1060 is configured to reconstruct geometric information (the coordinate system assumed by the encoding process, i.e., the position information after coordinate transformation in the coordinate transformation unit 1010) of each point of the point cloud data to be encoded, based on the tree information generated by the tree analysis unit 1030 and the approximate surface information generated by the approximate surface analysis unit 1040.
[0297] The frame buffer 1140 is configured to receive the geometric information reconstructed by the geometric information reconstruction unit 1060 as an input and store it as a reference frame.
[0298] The stored reference frame is read out from the frame buffer 1140 and used as a reference frame when inter-prediction of a temporally different frame is performed in the tree analysis unit 1030.
[0299] Here, which reference frame to use for each frame may be determined based on, for example, the value of a cost function representing encoding efficiency, and information on the reference frame to be used may be transmitted to the point cloud decoding device 200 as control data.
[0300] The color conversion unit 1070 is configured to perform color conversion when the input attribute information is color information. The color conversion does not necessarily have to be performed, and the presence or absence of the color conversion process is coded as part of the control data and transmitted to the point cloud decoding device 200.
[0301] The attribute transfer unit 1080 is configured to correct the attribute values so as to minimize distortion of the attribute information, based on the position information of the input point cloud, the position information of the point cloud after reconstruction in the geometric information reconstruction unit 1060, and the attribute information after color change in the color conversion unit 1070. As a specific correction method, for example, the method described in Non-Patent Document 1 can be applied.
[0302] The RAHT unit 1090 is configured to receive as input the attribute information after transfer by the attribute transfer unit 1080 and the geometric information generated by the geometric information reconstruction unit 1060, and to generate residual information for each point using a type of Haar transform called RAHT (Region Adaptive Hierarchical Transform).
[0303] The information to be decoded is the direct current component (DC coefficient) and alternating current component (AC coefficient) of the attribute information generated by using RAHT in the encoding process, and in the decoding process, it is converted into attribute information by using the inverse transform of RAHT.
[0304] As a specific example of the RAHT process, the method described in the above-mentioned Non-Patent Document 1 can be used.
[0305] The LoD calculation unit 1100 is configured to receive the geometric information generated by the geometric information reconstruction unit 1060 as an input and generate a Level of Detail (LoD).
[0306] LoD is information for defining a reference relationship (a referencing point and a referenced point) to realize predictive coding, such as predicting attribute information of a certain point from attribute information of another point and encoding or decoding the prediction residual.
[0307] In other words, LoD is information that defines a hierarchical structure in which each point contained in geometric information is classified into multiple levels, and attributes of points belonging to lower levels are encoded or decoded using attribute information of points belonging to higher levels.
[0308] As a specific method for determining the LoD, for example, the method described in Non-Patent Document 1 mentioned above may be used.
[0309] The lifting unit 1110 is configured to generate residual information by a lifting process using the LoD generated by the LoD calculation unit 1100 and the attribute information after attribute transfer by the attribute transfer unit 1080.
[0310] As a specific example of the lifting process, the method described in the above-mentioned non-patent document 1 may be used.
[0311] The attribute information quantization unit 1120 is configured to quantize the residual information output from the RAHT unit 1090 or the lifting unit 1110. Here, a quantization step size of 1 is equivalent to no quantization being performed.
[0312] The attribute information encoding unit 1130 is configured to perform encoding processing using the quantized residual information and the like output from the attribute information quantization unit 1120 as syntax, and to generate a bit stream related to the attribute information (attribute information bit stream).
[0313] The encoding process is, for example, a context-adaptive binary arithmetic encoding process. Here, for example, the syntax includes control data (flags and parameters) for controlling the decoding process of the attribute information.
[0314] Through the above processing, the point group encoding device 100 is configured to perform encoding processing using position information and attribute information of each point in a point group as input, and to output a geometric information bit stream and an attribute information bit stream.
[0315] Furthermore, the above-mentioned point group encoding device 100 and point group decoding device 200 may be realized as a program that causes a computer to execute each function (each process).
[0316] In each of the above embodiments, the present invention has been described using the point cloud encoding device 100 and the point cloud decoding device 200 as examples, but the present invention is not limited to such examples and can be similarly applied to a point cloud encoding / decoding system having the functions of the point cloud encoding device 100 and the point cloud decoding device 200. [Industrial Applicability]
[0317] According to this embodiment, for example, it is possible to improve the overall service quality in video communication, which makes it possible to contribute to Goal 9 of the Sustainable Development Goals (SDGs) led by the United Nations, which is to "build resilient infrastructure, promote sustainable industrialization and foster innovation." [Explanation of symbols]
[0318] 10...Point cloud processing system 100...Point cloud encoding device 1010... Coordinate conversion section 1020...Geometric information quantization section 1030…Tree analysis section 1040…Approximate surface analysis section 1050...Geometric information encoding unit 1060...Geometric information reconstruction unit 1070…Color conversion section 1080…Attribute transfer section 1090…RAHT Section 1100…LoD calculation section 1110…Lifting section 1120...Attribute information quantization section 1130...Attribute information encoding unit 1140...Frame buffer 200...Point cloud decoding device 2010…Geometric Information Decoding Department 2020…Tree synthesis section 2030…Approximate surface synthesis part 2040...Geometric information reconstruction unit 2050…Inverse coordinate conversion section 2060…Attribute information decoding unit 2070...Inverse quantization section 2080…RAHT Division 2090…LoD calculation section 2100…Reverse lifting section 2110…Color inversion unit 2120...Frame buffer
Claims
1. A point cloud decoding device, characterized in that it comprises a RAHT unit that applies a scaling factor to a predicted value of the AC coefficient or a predicted value of an attribute value in the inter prediction of the AC coefficient of each node.
2. The point cloud decoding device according to claim 1, wherein the RAHT unit decodes the scaling factor for each layer.
3. The point cloud decoding device according to claim 2, wherein the scaling factor is derived as a value obtained by adding a second integer to a first integer decoded for each layer and then dividing by the second integer.
4. A point cloud decoding method, characterized in that a scaling factor is applied to a predicted value of the AC coefficient or a predicted value of an attribute value in the inter prediction of the AC coefficient of each node.
5. A program for causing a computer to function as a point cloud decoding device, wherein the point cloud decoding device comprises a RAHT unit that applies a scaling factor to a predicted value of the AC coefficient or a predicted value of an attribute value in the inter prediction of the AC coefficient of each node.