Point cloud decoding device, point cloud decoding method and program
Patent Information
- Application Number
- JP2023173790
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-05
- Publication Date
- 2025-06-13
AI Technical Summary
In the prior art, the primary scaling factor in the Octree hierarchy cannot be effectively optimized, resulting in low encoding efficiency.
In point cloud decoding devices, RAHT (Region Adaptive Hierarchical Transform) technology is used to independently scale each Octree level and frequency index, thereby improving coding efficiency.
Scaling each Octree level and frequency index through different scaling factors significantly improves encoding efficiency and improves the decoding performance of point cloud data.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a point group decoding device, a point group decoding method, and a program. [Background technology]
[0002] Conventionally, in decoding attribute information, a method is known in which the AC coefficient of the attribute value is inter-predicted, the inter-predicted value is scaled using a scaling factor, the residual of the scaled inter-predicted value and the decoded AC coefficient is added, the AC coefficient is reconstructed, and the attribute value is decoded by inverse RAHT. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] G-PCC codec description, ISO / IEC JTC1 / SC29 / WG7 N00271 [Non-Patent Document 2] G-PCC 2nd Edition codec description, ISO / IEC JTC1 / SC29 / WG7 N00506 Summary of the Invention [Problem to be solved by the invention]
[0004] However, in the conventional technology, there is one scaling factor for each Octree hierarchy, and therefore there is a problem in that the value of the scaling factor is not optimized.
[0005] Therefore, the present invention has been made in consideration of the above-mentioned problems, and aims to provide a point cloud decoding device, a point cloud decoding method, and a program that can improve the encoding efficiency of attribute information encoding. [Means for solving the problem]
[0006] A first feature of the present invention is summarized as a point group decoding device including a RAHT unit that performs scaling on intra-predicted values of AC coefficients using a different scaling factor for each Octree layer and for each frequency index of the AC coefficient.
[0007] A second feature of the present invention is summarized as a point group decoding method including a step of scaling an intra-predicted value of an AC coefficient by a different scaling factor for each Octree layer and for each frequency index of the AC coefficient.
[0008] A third feature of the present invention is a program for causing a computer to function as a point cloud decoding device, the point cloud decoding device comprising a RAHT unit that performs scaling on intra-predicted values of AC coefficients using different scaling factors for each Octree hierarchy and for each frequency index of the AC coefficients.
[0009] A fourth feature of the present invention is summarized as a point cloud decoding device including an attribute information decoding unit that derives the number of scaling factors in inter prediction based on a syntax that specifies a layer to which the decoded inter prediction is applied. Effect of the Invention
[0010] According to the present invention, it is possible to provide a point cloud decoding device, a point cloud decoding method, and a program that can improve the coding efficiency of attribute information coding. [Brief description of the drawings]
[0011] [Figure 1] FIG. 1 is a diagram showing an example of the configuration of a point cloud processing system 10 according to an embodiment. [Diagram 2] FIG. 2 is a diagram illustrating an example of functional blocks of a point group decoding device 200 according to an embodiment. [Diagram 3] FIG. 3 is a diagram showing an example of the configuration of encoded data (bit stream) received by the geometric information decoding unit 2010 of the point cloud decoding device 200 according to an embodiment. [Figure 4] FIG. 4 is a diagram showing an example of the syntax configuration of GPS2011. [Diagram 5] FIG. 5 shows an example of the configuration of encoded data (bit stream) received by the attribute information decoding unit 2060 of the point group decoding device 200 according to an embodiment. [Figure 6] FIG. 6 is an example of the syntax configuration of the APS2611 shown in FIG. [Figure 7] FIG. 7 is a flowchart showing an example of processing by the RAHT unit 2080. [Figure 8] FIG. 8 is a flowchart showing an example of the process of step S28004. [Figure 9] FIG. 9 is a flowchart showing an example of the process of step S28104. [Figure 10] FIG. 10 is a flowchart showing an example of the intra prediction process in step S28112. [Figure 11] FIG. 11 is a diagram showing the relationship between a decoding target node and adjacent nodes in a higher layer. [Figure 12] FIG. 12 is a diagram showing the relationship between a decoding target node and adjacent nodes in the subnode hierarchy. [Figure 13] FIG. 13 is a flowchart showing an example of the intra prediction process in step S28112. [Figure 14] FIG. 14 is a flowchart showing an example of processing by the RAHT unit 2080. [Figure 15] FIG. 15 is a diagram showing an example of the inter prediction process in step S28111. [Figure 16] FIG. 16 is a flowchart showing an example of the operation of the tree synthesis unit 2020 of the point group decoding device 200 according to an embodiment. [Figure 17] FIG. 17 is a flowchart showing an example of the process of decoding the predictor information and the spherical coordinate residual in step S1604. [Figure 18] FIG. 18 is a diagram showing an example of functional blocks of the point group encoding device 100 according to an embodiment. [Figure 19] FIG. 19 is a diagram for explaining the first modification. [Figure 20] FIG. 20 is a diagram for explaining the second modification. [Figure 21] FIG. 21 is a diagram for explaining the second modification. [Figure 22] FIG. 22 is a diagram for explaining the second modification. [Diagram 23] FIG. 23 is a diagram for explaining the second modification. [Figure 24] FIG. 24 is a diagram for explaining the second modification. [Diagram 25] FIG. 25 is a diagram for explaining the second modification. [Figure 26] FIG. 26 is a diagram for explaining the third modification. [Figure 27] FIG. 27 is a diagram for explaining the third modification. [Figure 28] FIG. 28 is a diagram for explaining the third modification. [Figure 29] FIG. 29 is a diagram for explaining the third modification. [Diagram 30] FIG. 30 is a diagram for explaining the third modification. [Diagram 31] FIG. 31 is a diagram for explaining the third modification. [Diagram 32] FIG. 32 is a diagram for explaining the third modification. [Diagram 33] FIG. 33 is a diagram for explaining the third modification. [Diagram 34] FIG. 34 is a diagram for explaining the third modification. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] Hereinafter, the embodiments of the present invention will be described with reference to the drawings. Note that the components in the following embodiments can be appropriately replaced with existing components, and various variations including combinations with other existing components are possible. Therefore, the description of the following embodiments does not limit the content of the invention described in the claims.
[0013] (First embodiment) A point cloud processing system 10 according to a first embodiment of the present invention will be described below with reference to Fig. 1 to Fig. 18. Fig. 1 is a diagram showing a point cloud processing system 10 according to the embodiment.
[0014] As shown in FIG. 1, a point cloud processing system 10 includes a point cloud encoding device 100 and a point cloud decoding device 200.
[0015] The point cloud encoding device 100 is configured to generate encoded data (bit stream) by encoding an input point cloud signal. The point cloud decoding device 200 is configured to generate an output point cloud signal by decoding the bit stream.
[0016] The input point cloud signal and the output point cloud signal are composed of position information and attribute information of each point in the point cloud, such as color information and reflectance of each point.
[0017] Here, such a bit stream may be transmitted from the point group encoding device 100 to the point group decoding device 200 via a transmission path. Also, the bit stream may be stored in a storage medium and then provided from the point group encoding device 100 to the point group decoding device 200.
[0018] (Point Cloud Decoding Device 200) Hereinafter, the point group decoding device 200 according to this embodiment will be described with reference to Fig. 2. Fig. 2 is a diagram showing an example of functional blocks of the point group decoding device 200 according to this embodiment.
[0019] As shown in FIG. 2, the point cloud decoding device 200 includes a geometric information decoding unit 2010, a tree synthesis unit 2020, an approximate surface synthesis unit 2030, a geometric information reconstruction unit 2040, an inverse coordinate transformation unit 2050, an attribute information decoding unit 2060, an inverse quantization unit 2070, a RAHT unit 2080, a LoD calculation unit 2090, an inverse lifting unit 2100, an inverse color transformation unit 2110, and a frame buffer 2120.
[0020] The geometric information decoding unit 2010 is configured to receive as input a bit stream relating to geometric information (geometric information bit stream) out of the bit streams output from the point group encoding device 100, and to decode the syntax.
[0021] The decoding process is, for example, a context-adaptive binary arithmetic decoding process. Here, for example, the syntax includes control data (flags and parameters) for controlling the decoding process of the position information.
[0022] The tree synthesis unit 2020 is configured to receive as input the control data decoded by the geometric information decoding unit 2010 and an occupancy code indicating which node in the tree described below the point group exists in, and generate tree information indicating in which area in the space to be decoded the point exists.
[0023] The tree synthesis unit 2020 may be configured to perform the decoding process of the occupancy code within itself.
[0024] This process divides the space to be decoded into rectangles, determines whether a point exists in each rectangle by referring to the occupancy code, divides the rectangle in which the point exists into multiple rectangles, and then generates tree information by recursively repeating the process of referring to the occupancy code.
[0025] Here, when decoding the occupancy code, inter prediction, which will be described later, may be used.
[0026] In this embodiment, a method called "Octree" can be used, which always treats the above-mentioned rectangular parallelepiped as a cube and performs octree division recursively, and a method called "QtBt" can be used, which performs quadtree division and binary tree division in addition to octree division. Whether or not to use "QtBt" is transmitted from the point cloud encoding device 100 as control data.
[0027] Alternatively, when the control data specifies that predictive geometry coding is to be used, the tree synthesis unit 2020 is configured to decode the coordinates of each point based on an arbitrary tree structure determined in the point cloud encoding device 100.
[0028] The approximate surface synthesis unit 2030 is configured to generate approximate surface information using the tree information generated by the tree synthesis unit 2020, and to decode the point cloud based on the approximate surface information.
[0029] Approximate surface information is used when, for example, decoding three-dimensional point cloud data of an object, in cases where the point cloud is densely distributed on the object surface, to approximate the area in which the point cloud exists using a small plane, rather than decoding each individual point cloud.
[0030] Specifically, the approximate surface synthesis unit 2030 can generate approximate surface information and decode the point cloud using a method called "Trisoup", for example. A specific processing example of "Trisoup" will be described later. In addition, when decoding a sparse point cloud acquired by Lidar or the like, this processing can be omitted.
[0031] The geometric information reconstruction unit 2040 is configured to reconstruct geometric information (position information in the coordinate system assumed by the decoding process) of each point of the point cloud data to be decoded, based on the tree information generated by the tree synthesis unit 2020 and the approximate surface information generated by the approximate surface synthesis unit 2030.
[0032] The inverse coordinate transformation unit 2050 is configured to receive the geometric information reconstructed by the geometric information reconstruction unit 2040 as input, transform the information from the coordinate system assumed by the decoding process to the coordinate system of the output point cloud signal, and output position information.
[0033] The frame buffer 2120 is configured to store, as an input, the geometric information reconstructed by the geometric information reconstruction unit 2040 as a reference frame. When the tree synthesis unit 2020 performs inter-prediction of temporally different frames, the stored reference frame is read out from the frame buffer 2130 and used as the reference frame.
[0034] Here, which reference frame at which time is to be used for each frame may be determined based on control data transmitted from the point group encoding device 100 as a bit stream, for example.
[0035] The attribute information decoding unit 2060 is configured to receive as input a bit stream relating to attribute information (attribute information bit stream) out of the bit streams output from the point group encoding device 100, and to decode the syntax.
[0036] The decoding process is, for example, a context-adaptive binary arithmetic decoding process. Here, for example, the syntax includes control data (flags and parameters) for controlling the decoding process of the attribute information.
[0037] Moreover, the attribute information decoding unit 2060 is configured to decode the quantized residual information from the decoded syntax.
[0038] The inverse quantization unit 2070 is configured to perform inverse quantization processing based on the quantized residual information decoded by the attribute information decoding unit 2060 and a quantization parameter, which is one of the control data decoded by the attribute information decoding unit 2060, to generate inverse quantized residual information.
[0039] The dequantized residual information is output to either the RAHT unit 2080 or the LoD calculation unit 2090 according to the features of the point group to be decoded. Control data decoded by the attribute information decoding unit 2060 specifies which unit the dequantized residual information is output to.
[0040] The RAHT unit 2080 is configured to receive the inverse quantized residual information generated by the inverse quantization unit 2070 and the geometric information generated by the geometric information reconstruction unit 2040 as input, and to decode the attribute information of each point using a type of Haar transform (inverse Haar transform in the decoding process) called RAHT (Region Adaptive Hierarchical Transform). The decoded information is the direct current component (DC coefficient) and alternating current component (AC coefficient) of the attribute information generated by using RAHT in the encoding process, and is converted into the attribute information by using the inverse transform of RAHT in the decoding process. As a specific process of RAHT, for example, the method described in Non-Patent Document 1 can be used.
[0041] The LoD calculation unit 2090 is configured to receive the geometric information generated by the geometric information reconstruction unit 2040 as an input and generate a Level of Detail (LoD).
[0042] LoD is information for defining a reference relationship (a referencing point and a referenced point) to realize predictive coding, such as predicting attribute information of a certain point from attribute information of another point and encoding or decoding the prediction residual.
[0043] In other words, LoD is information that defines a hierarchical structure in which each point contained in geometric information is classified into multiple levels, and attributes of points belonging to lower levels are encoded or decoded using attribute information of points belonging to higher levels.
[0044] As a specific method for determining the LoD, for example, the method described in Non-Patent Document 1 mentioned above may be used.
[0045] The inverse lifting unit 2100 is configured to decode attribute information of each point based on the hierarchical structure defined by the LoD, using the LoD generated by the LoD calculation unit 2090 and the inverse quantized residual information generated by the inverse quantization unit 2070. As a specific process of inverse lifting, for example, the method described in the above-mentioned Non-Patent Document 1 can be used.
[0046] The inverse color conversion unit 2110 is configured to perform inverse color conversion processing on the attribute information output from the RAHT unit 2080 or the inverse lifting unit 2100 when the attribute information to be decoded is color information and color conversion has been performed on the point cloud encoding device 100 side. Whether or not to perform such inverse color conversion processing is determined by the control data decoded by the attribute information decoding unit 2060.
[0047] The point group decoding device 200 is configured to decode and output the attribute information of each point in the point group through the above processing.
[0048] (Geometric Information Decoding Part 2010) The control data decoded by the geometric information decoding unit 2010 will be described below with reference to FIGS.
[0049] FIG. 3 shows an example of the structure of the coded data (bit stream) received by the geometric information decoding unit 2010. In FIG.
[0050] First, the bit stream may include a GPS 2011. The GPS 2011 is also called a geometry parameter set, and is a set of control data related to decoding of geometric information. A specific example will be described later. Each GPS 2011 includes at least GPS ID information for identifying each GPS 2011 when there are multiple GPS 2011s.
[0051] Secondly, the bit stream may include GSH2012A / 2012B. GSH2012A / 2012B is also called a geometry slice header or geometry data unit header, and is a set of control data corresponding to a slice described later. In the following description, the term "slice" is used, but slice can also be read as data unit. A specific example will be described later. GSH2012A / 2012B includes at least GPS id information for specifying GPS2011 corresponding to each GSH2012A / 2012B.
[0052] Thirdly, the bitstream may include slice data 2013A / 2013B following the GSH 2012A / 2012B. The slice data 2013A / 2013B includes data in which geometric information is encoded.
[0053] As described above, the bit stream is configured such that each slice data 2013A / 2013B corresponds to one GSH 2012A / 2012B and one GPS 2011.
[0054] As described above, since the GPS ID information is used to specify which GPS 2011 to refer to in the GSH 2012A / 2012B, a common GPS 2011 can be used for a plurality of slice data 2013A / 2013B.
[0055] In other words, it is not necessary to transmit the GPS 2011 for each slice. For example, as shown in Fig. 3, the bit stream may be configured such that the GPS 2011 is not coded immediately before the GSH 2012B and the slice data 2013B.
[0056] 3 is merely an example. As long as the slice data 2013A / 2013B corresponds to the GSH 2012A / 2012B and the GPS 2011, elements other than those described above may be added as components of the bit stream.
[0057] For example, as shown in Fig. 3, the bitstream may include a sequence parameter set (SPS) 2001. Similarly, when transmitted, the bitstream may be shaped into a configuration different from that shown in Fig. 3. Furthermore, the bitstream may be combined with a bitstream decoded by an attribute information decoding unit 2060 (described later) and transmitted as a single bitstream.
[0058] FIG. 4 is an example of the syntax configuration of GPS2011.
[0059] Note that the syntax names described below are merely examples. If the functions of the syntaxes described below are similar, the syntax names may be different.
[0060] The GPS 2011 may include GPS ID information (gps_geom_parameter_set_id) for identifying each GPS 2011.
[0061] The Descriptor column in Fig. 4 indicates how each syntax is coded. ue(v) indicates an unsigned zeroth-order exponential Golomb code, and u(1) indicates a 1-bit flag.
[0062] The GPS 2011 may include a flag (geom_tree_type) for controlling the tree type in the tree synthesis unit 2020 .
[0063] For example, if the value of geom_tree_type is "1", it may be defined that predictive geometry coding is used, and if the value of geom_tree_type is "0", it may be defined that Octree is used.
[0064] The GPS 2011 may include a flag (geom_angular_enabled) for controlling whether or not the tree synthesis unit 2020 performs processing in angular mode.
[0065] For example, when the value of geom_angular_enabled is “1”, it may be defined that predictive geometry coding processing is performed as angular mode, and when the value of geom_angular_enabled is “0”, it may be defined that predictive geometry coding processing is not performed as angular mode.
[0066] The GPS 2011 may include a flag (ptree_ang_azimuth_scaling_enabled) for controlling whether or not the adaptive azimuth angle quantization mode is in the angular mode in the tree synthesis unit 2020. The adaptive azimuth angle quantization mode is a mode in which adaptive quantization of the azimuth angle is performed according to the radius.
[0067] For example, when the value of ptree_ang_azimuth_scaling_enabled is "1", it may be defined that adaptive quantization of the azimuth angle according to the radius is performed, and when the value of ptree_ang_azimuth_scaling_enabled is "0", it may be defined that adaptive quantization of the azimuth angle according to the radius is not performed.
[0068] It may also be used as a flag to control whether to use the predictor list when calculating (selecting) a predictor in Angular mode.
[0069] For example, if the value of ptree_azimuth_scaling_enabled is “1”, it may be defined that a predictor list is used in the calculation of the predictor, and if the value of ptree_ang_azimuth_scaling_enabled is “0”, it may be defined that a predictor list is not used in the calculation of the predictor.
[0070] The GPS 2011 may include a value (ptree_ang_azimuth_step_minus1) related to the laser rotation rate for use in calculating the predicted azimuth angle in the tree synthesis unit 2020 in angular mode.
[0071] (Tree Synthesis Department 2020) An example of the operation of the tree synthesis unit 2020 will be described below with reference to FIGS.
[0072] 16 is a flowchart showing an example of processing in the tree synthesis unit 2020. Note that, below, an example will be described in which trees are synthesized using "Predictive geometry coding".
[0073] Predictive geometry coding is also called predictive tree coding. Predictive geometry coding is a method for decoding the residual of position information predicted based on an arbitrary tree structure determined by the point cloud encoding device 100 and the position information of the point cloud data, and adding the two together to decode the position information of the point cloud data.
[0074] As shown in FIG. 16, in step S1601, the tree synthesis unit 2020 determines whether or not decoding of position information of all point cloud data included in the slice has been completed.
[0075] This process, for example, transmits information indicating the number of point cloud data contained in the slice to the GSH, and by comparing this number of point cloud data with the number of data already processed, it can be determined whether processing of all points has been completed.
[0076] If the decoding of the position information of all point cloud data is completed, the operation proceeds to step S1613 and ends the process. If the decoding of the position information of all point cloud data is not completed, the operation proceeds to step S1602.
[0077] In step S1602, the tree merging unit 2020 sets a parent node of a node to be decoded (node to be processed) of the point cloud data.
[0078] For example, the tree synthesis unit 2020 decodes the number of child nodes of each node to be decoded, and stores the indexes of the nodes to be decoded for the number of child nodes.
[0079] When the tree synthesis unit 2020 processes a node to be decoded after a certain node, it may refer to the array of indexes of the node, obtain one index stored at the end of the array, and set the node of the obtained index as the parent node of the node to be decoded.
[0080] After the parent node setting is completed, the operation proceeds to step S1603.
[0081] In step S1603, the tree merge unit 2020 determines whether to perform processing in the Angular mode.
[0082] For example, the tree synthesis unit 2020 can refer to the value of the above-mentioned geom_angular_enabled to determine whether to perform processing in Angular mode.
[0083] If the processing is to be performed in Angular mode, the operation proceeds to step S1604; if the processing is not to be performed in Angular mode, the operation proceeds to step S1610.
[0084] In step S1604, the tree synthesis unit 2020 decodes the predictor information and spherical coordinate residual used in step S1605. Here, the spherical coordinate residual indicates the residual of the radius, azimuth angle, and laser ID. When this decoding is completed, the operation proceeds to step S1605.
[0085] In step S1605, the tree synthesis unit 2020 predicts the position information based on the predictor information decoded in step S1604. Here, the predictor information is a predictor index or a prediction mode.
[0086] In this process, the tree synthesis unit 2020 first determines the type of predictor to be used for prediction.
[0087] For example, the tree synthesis unit 2020 may determine whether or not to perform processing in adaptive azimuth angle quantization mode based on the value of ptree_ang_azimuth_scaling_enabled, and may determine the type of predictor to be used based on the result of this determination.
[0088] For example, in the case of an adaptive azimuth angle quantization mode, the tree synthesis unit 2020 may select a predictor to be used based on the decoded prediction mode from among a plurality of predictors calculated using a tree structure.
[0089] Alternatively, when performing processing in the adaptive azimuth angle quantization mode, the tree synthesis unit 2020 may store position information of the decoded nodes in a list as a predictor, and refer to the list for a predictor assigned to a decoded predictor index to select the predictor type to be used.
[0090] Once the type of predictor is determined, the tree synthesis unit 2020 uses the predictor as the predicted value of the position information.
[0091] After the prediction of the position information is completed, the operation proceeds to step S1606.
[0092] In step S1606, the tree synthesis unit 2020 reconstructs the spherical coordinates. In this process, the tree synthesis unit 2020 reconstructs the spherical coordinates by adding the decoded spherical coordinate residual and the predictor.
[0093] After the reconfiguration is completed, the operation proceeds to step S1607.
[0094] In step S1607, the tree synthesis unit 2020 reconstructs the orthogonal integer coordinates. In this process, the tree synthesis unit 2020 can convert the spherical coordinates into orthogonal integer coordinates based on the reconstructed spherical coordinates. A specific method for this can be realized by, for example, the method described in Non-Patent Document 1.
[0095] After the reconstruction of the orthogonal integer coordinates is completed, the operation proceeds to step S1608.
[0096] In step S1608, the tree synthesis unit 2020 decodes the orthogonal integer coordinate residual.
[0097] After the decoding of the orthogonal integer coordinate residual is completed, the operation proceeds to step S1609.
[0098] In step S1609, the tree synthesis unit 2020 reconstructs the original coordinates. In this process, the tree synthesis unit 2020 reconstructs the original coordinates by adding the decoded orthogonal integer coordinate residual and the reconstructed orthogonal integer coordinates.
[0099] After the reconstruction of the original coordinates is completed, the operation returns to step S1601.
[0100] In step S1610, the tree synthesis unit 2020 predicts the position information. Specifically, the tree synthesis unit 2020 selects a predictor and sets the predictor as the predicted value of the position information.
[0101] For example, the tree synthesis unit 2020 may select a predictor based on the decoded predictor mode from among a plurality of predictors calculated based on a tree structure.
[0102] After the prediction of the position information is completed, the operation proceeds to step S1611.
[0103] In step S1611, the tree synthesis unit 2020 decodes the orthogonal integer coordinate residual.
[0104] After the decoding of the orthogonal integer coordinate residual is completed, the operation proceeds to step S1612.
[0105] In step S1612, the tree synthesis unit 2020 reconstructs the original coordinates. In this process, the tree synthesis unit 2020 reconstructs the original coordinates by adding the residual of the orthogonal integer coordinates decoded in step S1611 and the position information predicted in step S1610.
[0106] After the reconstruction of the original coordinates is completed, the operation returns to step S1601.
[0107] FIG. 17 is a flowchart showing an example of the process of decoding the predictor information and the spherical coordinate residual in step S1604.
[0108] As shown in FIG. 17, in step S1701, the tree synthesis unit 2020 determines whether or not the adaptive azimuth angle quantization mode is selected based on the value of ptree_ang_azimuth_scaling_enabled.
[0109] If the mode is the adaptive azimuth angle quantization mode, the operation proceeds to step S1702. On the other hand, if the mode is not the adaptive azimuth angle quantization mode, the operation proceeds to step S1703.
[0110] In step S1702, the tree synthesis unit 2020 decodes the predictor index. After the decoding of the predictor index is completed, the operation proceeds to step S1704.
[0111] In step S1703, the tree synthesis unit 2020 decodes the prediction mode. After the decoding of the prediction mode is completed, the operation proceeds to step S1704.
[0112] In step S1704, the tree synthesis unit 2020 decodes the azimuth angle step number. After the azimuth angle step number has been decoded, the operation proceeds to step S1705.
[0113] In step S1705, the tree synthesis unit 2020 decodes the spherical coordinate residual. The tree synthesis unit 2020 may perform such decoding using the method described in Non-Patent Document 2. After the decoding is completed, the operation proceeds to step S1706, where the process ends.
[0114] (Attribute information decoding unit 2060) The control data decoded by the attribute information decoding unit 2060 will be described below with reference to FIGS.
[0115] FIG. 5 shows an example of the structure of the encoded data (bit stream) received by the attribute information decoding unit 2060, and FIG. 6 shows an example of the syntax structure of the APS 2611 shown in FIG.
[0116] Note that the syntax names described below are merely examples. If the functions of the syntaxes described below are similar, the syntax names may be different.
[0117] The APS 2611 may include APS id information (aps_geom_parameter_set_id) for identifying each APS 2611.
[0118] The Descriptor column in Fig. 4 indicates how each syntax is coded. ue(v) indicates an unsigned zeroth-order exponential Golomb code, and u(1) indicates a 1-bit flag.
[0119] The APS 2611 may include a flag (attr_coding_type) for controlling whether the inverse quantization unit 2070 outputs the inverse quantized residual information to the RAHT unit 2080 or the LoD calculation unit 2090.
[0120] For example, when the value of attr_coding_type is “1”, it may be defined that the data is output to the LoD calculation unit 2090 , and when the value of attr_coding_type is “0”, it may be defined that the data is output to the RAHT unit 2080 .
[0121] The APS 2611 may include a flag (raht_prediction_enabled) for controlling whether or not to perform prediction of attribute information in the RAHT unit 2080.
[0122] For example, when the value of raht_prediction_enabled is "1", it may be defined that prediction of attribute information is performed, and when the value of raht_prediction_enabled is "0", it may be defined that prediction of attribute information is not performed.
[0123] The APS 2611 may include a flag (raht_subnode_prediction_enable_flag) that controls whether or not the RAHT unit 2080 uses subnodes to predict attribute information.
[0124] For example, when the value of raht_subnode_prediction_enable_flag is "1", it may be defined that subnodes are used to predict attribute information, and when the value of raht_subnode_prediction_enable_flag is "0", it may be defined that subnodes are not used to predict attribute information.
[0125] The APS 2611 may include weight parameters (raht_prediction_weights) used when the RAHT unit 2080 performs intra-prediction of attribute information.
[0126] For example, the value of raht_prediction_weights may be defined according to the manner in which the node to be decoded is adjacent to an adjacent node used for intra prediction.
[0127] The APS 2611 may include a flag (raht_smoothing_enable_flag) that controls whether or not to perform smoothing after the RAHT unit 2080 performs intra prediction of the attribute information.
[0128] For example, when the value of raht_smoothing_enable_flag is "1", it may be defined that smoothing is performed after predicting the attribute information, and when the value of raht_smoothing_enable_flag is "0", it may be defined that smoothing is not performed.
[0129] The APS 2611 may include a weighting parameter (raht_smoothing_weighted_average_weights) for performing smoothing by weighted averaging after the RAHT unit 2080 performs intra prediction of attribute information.
[0130] For example, a maximum of eight such weight parameters may be defined according to the manner in which the node to be decoded is adjacent to each subnode of the same parent node of the node to be decoded.
[0131] The APS 2611 may include a weighting parameter (raht_smoothing_clipping_weights) for performing smoothing by clipping after the RAHT unit 2080 performs intra prediction of attribute information.
[0132] For example, a maximum of eight such weight parameters may be defined according to the manner in which the node to be decoded is adjacent to each subnode of the same parent node of the node to be decoded.
[0133] The APS 2611 may include a threshold (raht_smoothing_clipping_threshold) for performing smoothing by clipping after the RAHT unit 2080 performs intra prediction of attribute information.
[0134] The APS 2611 may include a flag (raht_inter_prediction_enabled) for controlling whether or not to perform inter prediction of attribute information in the RAHT unit 2080.
[0135] For example, when the value of raht_inter_prediction_enabled is "1", it may be defined that prediction of attribute information is performed, and when the value of raht_inter_prediction_enabled is "0", it may be defined that prediction of attribute information is not performed.
[0136] The APS 2611 may include a value (raht_inter_prediction_depth_minus1) indicating a layer at which inter prediction of attribute information is enabled in the RAHT unit 2080.
[0137] For example, when raht_inter_prediction_depth_minus1 is "N-1", inter prediction may be enabled in up to the top N layers of the Octree structure.
[0138] (RAHT Division 2080) An example of the processing of the RAHT unit 2080 will be described with reference to FIGS.
[0139] FIG. 7 is a flowchart showing an example of processing by the RAHT unit 2080.
[0140] 7, in step S28001, the RAHT unit 2080 recursively divides the nodes into octrees until a predetermined size is reached, using a method called Octree. After the division is completed, the operation proceeds to step S28002.
[0141] In step S28002, the RAHT unit 2080 counts up the total number of points belonging to the lower layer of each node divided by the octree.
[0142] Specifically, the RAHT unit 2080 scans the nodes of a certain layer in order and records the number of points belonging to each node. Next, the RAHT unit 2080 adds up the numbers of points recorded in the child nodes of each node in the node one layer above to calculate the number of points belonging to each node.
[0143] The RAHT unit 2080 repeats the above scanning from the lowest layer to the highest layer. The total number of acquired points is used as a weight for the inverse transformation of the RAHT in step S28005 described later. After this calculation is completed, the operation proceeds to step S28003.
[0144] In step S28003, the RAHT unit 2080 decodes the DC coefficient of the node belonging to the highest layer of the Octree. Alternatively, the RAHT unit 2080 may calculate the DC coefficient by predicting the DC coefficient using intra prediction and decoding and adding up the prediction residual of the DC coefficient.
[0145] After completing the decoding of the DC coefficient, the RAHT unit 2080 calculates the attribute value Aroot of the root node using the total number of points wroot belonging to the root node acquired in step S28002 and the decoded DC coefficient DCroot according to the following formula.
[0146]
number
[0147] In step S28004, the RAHT unit 2080 determines whether or not decoding of the attribute information of all nodes included in the layer has been completed.
[0148] If not completed, the operation proceeds to step S28005; if completed, the operation proceeds to step S28007.
[0149] In step S28005, the RAHT unit 2080 decodes the AC coefficients. The details will be described later. After the decoding is completed, the operation proceeds to step S28006.
[0150] In step S28006, the RAHT unit 2080 calculates attribute values using the inverse transform of the RAHT based on the total number of points belonging to the lower hierarchy of each node, the decoded AC coefficients, and the DC coefficients calculated from the nodes in the upper hierarchy using the method described below.
[0151] Here, the inverse transformation of RAHT is performed in units of 8 nodes (2 x 2 x 2) divided by the Octree.
[0152] Specifically, attribute values A1, A2, … A k is the DC coefficient DC of a node that holds k subnodes, and the AC coefficients AC1,AC2,…AC k-1 and the total number of points belonging to the lower hierarchy of each subnode, w=w1, w2, … w k Using this, it can be calculated using the following formula (1).
[0153]
number
[0154] It is assumed that the conversion process is performed repeatedly from the higher-level nodes to the lower-level nodes,
[0155]
number
[0156] In step S28007, the RAHT unit 2080 determines whether the decoding of nodes in all layers is complete.
[0157] If not completed, this operation moves the processing target hierarchical level to the next lower hierarchical level and proceeds to step S28004, If completed, this operation proceeds to step S28008 and ends the processing.
[0158] FIG. 8 is a flowchart showing an example of the process of step S28004.
[0159] 8, in step S28101, the RAHT unit 2080 determines whether to predict AC coefficients. When making such a determination, the RAHT unit 2080 may refer to raht_prediction_enabled and use the value thereof.
[0160] The RAHT unit 2080 may decode a flag indicating whether or not to perform prediction of an AC coefficient in the currently processed node, and use the value of the flag.
[0161] The flag may be decoded for each node or for each layer. The flag may be decoded only if the value of raht_prediction_enabled is "1", indicating that prediction is enabled. The flag may be included in the slice data.
[0162] If the result of the determination is that the AC coefficients are not to be predicted, the operation proceeds to step S28102, and if the result is that the AC coefficients are to be predicted, the operation proceeds to steps S28103 and S28104.
[0163] In step S28102, the RAHT unit 2080 decodes the AC coefficients. After the decoding is completed, the operation proceeds to step S28106, and the process ends.
[0164] In step S28103, the RAHT 2080 decodes the AC coefficient residuals. After the decoding is completed, the operation proceeds to step S28105.
[0165] In step S28104, the RAHT unit 2080 predicts AC coefficients. Inter prediction or intra prediction may be used for predicting the AC coefficients.
[0166] The RAHT unit 2080 may first predict the attribute values, and then calculate the predicted values of the AC coefficients by the RAHT. This will be described in detail later. After the prediction of the AC coefficients is completed, the operation proceeds to step S28105.
[0167] In step S28105, the RAHT unit 2080 adds the residual of the decoded AC coefficients to the predicted AC coefficients to reconstruct the AC coefficients. After the reconstruction is completed, the operation proceeds to step S28106, and the process ends.
[0168] FIG. 9 is a flowchart showing an example of the process of step S28104.
[0169] As shown in Fig. 9, in step S28107, the RAHT unit 2080 determines whether inter prediction is enabled. The RAHT unit 2080 may refer to raht_inter_prediction_enabled for the determination and use the value. If the result of the determination is that inter prediction is enabled, this operation proceeds to step S28109, and if inter prediction is disabled, this operation proceeds to step S28112.
[0170] In step S28109, the RAHT unit 2080 determines whether the depth of the layer including the node to be processed is equal to or less than a threshold. The RAHT unit 2080 may refer to raht_inter_prediction_depth_minus1 as the threshold and use that value.
[0171] If the result of the determination is that the depth is equal to or less than the threshold, the operation proceeds to step S28110, and if the depth is greater than the threshold, the operation proceeds to step S28112.
[0172] In step S28110, the RAHT unit 2080 determines whether or not to perform inter prediction on the AC coefficient of the node to be processed.
[0173] For this determination, the RAHT unit 2080 may check whether inter prediction is possible, and perform inter prediction if possible, and may not perform inter prediction if not possible.
[0174] The RAHT unit 2080 may decode a flag indicating whether or not to perform inter-prediction on the AC coefficients of the node to be processed, and may use the value of the flag for the determination. The flag may be decoded for each node, or may be decoded for each layer. The flag may be decoded and a determination may be made only when it is determined that inter-prediction is executable. The flag may be included in slice data.
[0175] In step S28111, the RAHT unit 2080 performs inter prediction of the AC coefficients of the node to be processed. This will be described in detail later.
[0176] In step S28112, the RAHT unit 2080 performs intra prediction of the AC coefficients of the node to be processed. This will be described in detail later.
[0177] In step S28113, the process of step S28104 ends. Note that the conditional branch in step S28109 may be omitted.
[0178] In the inter prediction process in step S28111, a process equivalent to the intra prediction process in step S28112 may also be performed, and prediction may be performed by combining the results of inter prediction and intra prediction.
[0179] FIG. 10 is a flowchart showing an example of the intra prediction process in step S28112.
[0180] 10, in step S28201, the RAHT unit 2080 determines whether or not to perform intra prediction using adjacent nodes in the subnode hierarchy. The RAHT unit 2080 may refer to raht_subnode_prediction_enable_flag and use the value thereof for the determination.
[0181] When the RAHT unit 2080 does not use adjacent nodes in the subnode hierarchy, it performs intra prediction using only adjacent nodes in a higher hierarchy.
[0182] In this case, the adjacent nodes in the higher hierarchy are the six nodes adjacent to the parent node of the node to be decoded on the face side, the 12 nodes adjacent to the edge side, and the parent node itself, which are a total of 19 nodes, three nodes adjacent to the face side of the node to be decoded on the face side, three nodes adjacent to the edge side, and the seven nodes of the parent node itself.
[0183] FIG. 11 is a diagram showing the relationship between a decoding target node and adjacent nodes in a higher layer.
[0184] When using adjacent nodes in the subnode hierarchy, the RAHT unit 2080 performs intra prediction using adjacent nodes in a higher hierarchy and adjacent nodes in the subnode hierarchy.
[0185] Here, an adjacent node in the subnode hierarchy is a subnode of an adjacent node in a higher hierarchy that has a face or edge adjacent to the decoding target node and has already been decoded.
[0186] FIG. 12 is a diagram showing the relationship between a decoding target node and adjacent nodes in the subnode hierarchy.
[0187] If the determination result is that intra prediction is to be performed without using adjacent nodes in the subnode hierarchy, this operation proceeds to step S28202, and if intra prediction is to be performed using adjacent nodes in the subnode hierarchy, this operation proceeds to step S28204.
[0188] In step S28202, the RAHT unit 2080 acquires the attribute value of the adjacent node in the upper layer. After acquiring the attribute value of the adjacent node in the upper layer, this operation proceeds to step S28203.
[0189] In step S28203, the RAHT unit 2080 predicts the attribute value of the node to be decoded.
[0190] The RAHT unit 2080 obtains the attribute values attr i and the weight w according to the type of adjacent node i i Using these, the attribute value attr may be predicted using the following formula:
[0191]
number
[0192] Here, the RAHT unit 2080 uses the weight w i Depending on whether the adjacent node i is a face adjacent node in a higher hierarchy, an edge adjacent node in a higher hierarchy, or a parent node, a hard-coded value may be used, or the weight w may be calculated by referring to raht_prediction_weights. i may be calculated.
[0193] After the prediction of the attribute value is completed, the operation proceeds to step S28207.
[0194] In step S28204, the RAHT unit 2080 acquires the attribute value of the adjacent node in the upper layer.
[0195] Here, the targets for which attribute values are to be obtained are adjacent nodes in a higher hierarchy whose subnodes have not yet been decoded, or adjacent nodes in a higher hierarchy whose subnodes have already been decoded but whose subnodes do not have any subnodes adjacent to the node to be decoded via faces or edges.
[0196] After the attribute value has been acquired, the operation proceeds to step S28205.
[0197] In step S28205, the RAHT unit 2080 acquires the attribute value of the adjacent node in the subnode hierarchy. After acquiring the attribute value of the adjacent node in the subnode hierarchy, this operation proceeds to step S28206.
[0198] In step S28206, the RAHT unit 2080 predicts the attribute value of the node to be decoded.
[0199] The RAHT unit 2080 obtains the attribute values attr i and the weight w according to the type of adjacent node i i Using these, the attribute value attr may be predicted using the following formula:
[0200]
number
[0201] Here, the RAHT unit 2080 uses the weight w i Depending on whether the adjacent node i is a face adjacent node in a higher hierarchy, an edge adjacent node in a higher hierarchy, a parent node, a face adjacent node in a subnode hierarchy, or an edge adjacent node in a subnode hierarchy, a hard-coded value may be used, or the weight w may be calculated by referring to raht_prediction_weights. i may be calculated.
[0202] After the attribute value prediction is completed, the operation proceeds to step S28207.
[0203] In step S28207, the RAHT unit 2080 converts the predicted attribute values into AC coefficients. The AC coefficients are generated by performing RAHT on the predicted attribute values. For example, the RAHT unit 2080 may use the method described in Non-Patent Document 1 as such a conversion method.
[0204] The above describes an example in which the RAHT unit 2080 uses the attribute values predicted in step S28206 directly to convert the AC coefficients in step S28207, but the RAHT unit 2080 may smooth the predicted attribute values and then convert the AC coefficients.
[0205] For example, as shown in FIG. 13, after predicting an attribute value, the RAHT unit 2080 may determine whether or not to perform smoothing in step S1301.
[0206] In making such a determination, the RAHT unit 2080 may refer to raht_smoothing_enable_flag and use the value thereof.
[0207] If smoothing is to be performed, the operation proceeds to step S1302. If smoothing is not to be performed, the operation proceeds to step S28207.
[0208] In step S1302, the RAHT unit 2080 may smooth the attribute values.
[0209] For example, the RAHT unit 2080 calculates the attribute value Attr i and weight α i Alternatively, the weighted average may be calculated as follows:
[0210]
number
[0211] Here, the RAHT unit 2080 may treat the decryption target node as a node adjacent to the decryption target node, or may treat all subnodes within the same parent node, for the target subnode i.
[0212] In addition, the RAHT unit 2080 uses the weight α i A hard-coded value may be used as the weights, or the value may be referenced and used as the weights.
[0213] In addition, the RAHT unit 2080 performs, for example, the smoothed attribute value Attr smoothing The predicted value Attr0 of the decoded node itself, the predicted attribute value Attr of the subnode i other than the decoded node among the subnodes in the same parent node as the decoded node, i , weight β i and threshold Th r It may be obtained by clipping using the following:
[0214]
number
[0215] Here, clipping is a process in which if the input value is greater than a predetermined maximum value, the maximum value is output, if the input value is less than a predetermined minimum value, the minimum value is output, and in all other cases, the input value is used as the output value as is.
[0216] The clipping function Clip3 is
[0217]
number
[0218] Here, for the target subnode i, the RAHT unit 2080 may treat the decode target node as a face-adjacent node, a face-adjacent node and an edge-adjacent node, or all subnodes within the same parent node.
[0219] In addition, the RAHT unit 2080 uses the weight β i A hard-coded value may be used as the weights, or the value may be referenced and used as the raht_smoothing_clipping_weights.
[0220] In addition, the RAHT unit 2080 sets the threshold value Th r You can use a hard-coded value for raht_smoothing_clipping_threshold, or you can refer to raht_smoothing_clipping_threshold and use that value.
[0221] In the above, an example has been described in which the RAHT unit 2080 decodes AC coefficients of both the color difference signal and the luminance signal, but the RAHT unit 2080 may skip decoding AC coefficients of the color difference signal only in the bottom layer of the octree.
[0222] For example, as shown in FIG. 14, the RAHT unit 2080 may determine in step S1401 whether or not to skip decoding of AC coefficients of chrominance signals only in the bottom layer of the octree.
[0223] If skipping, the operation proceeds to step S1402, otherwise the operation proceeds to step S28004.
[0224] In step S1402, the RAHT unit 2080 determines whether the node to be decoded is at the bottom layer of the octree.
[0225] If it is the bottom layer, the operation proceeds to step S1403. If it is not the bottom layer, the operation proceeds to step S28004.
[0226] In step S1403, the RAHT unit 2080 decodes the AC coefficients other than the color difference signals.
[0227] The RAHT unit 2080 performs the same process as in step S28004 for decoding AC coefficients other than those of the color difference signals, sets the AC coefficients of the color difference signals to 0, and calculates attribute values in the subsequent step S28005.
[0228] After the decoding of AC coefficients other than the color difference signal is completed, the operation proceeds to step S28006.
[0229] FIG. 15 is a diagram showing an example of the inter prediction process in step S28111.
[0230] The RAHT unit 2080 predicts the AC coefficients of the processing target node using information of a reference node, which is a corresponding node in a reference frame. Here, the information of the reference node may be its attribute value or AC coefficient. The reference frame may indicate another decoded frame, and the information may be included in the previous frame buffer 2120.
[0231] The RAHT unit 2080 may apply the same octree structure as the processing target frame to the reference frame. In such a case, a node may be set at a position where there is no point. Such a node is called an empty node. If the reference node is an empty node, the RAHT unit 2080 may disable inter prediction in step S28110.
[0232] The RAHT unit 2080 may apply an Octree to the reference frame independently of the processing target frame, and set an Octree structure different from that of the processing target frame. In such a case, a node may not necessarily exist at the same position as that of the processing target frame. If a reference node is not found at a position corresponding to the processing target node, the RAHT unit 2080 may disable inter prediction in step S28143.
[0233] If the reference node is an empty node, or if the reference node cannot be found, the RAHT unit 2080 may estimate and interpolate the information of the reference node using information of nodes in nearby positions within the reference frame.
[0234] For example, the RAHT unit 2080 may estimate and interpolate the average value of the attribute values or AC coefficients of adjacent nodes, nearest neighbor nodes, or k-nearest neighbor nodes with respect to the reference node position as the attribute value or AC coefficient of the reference node, respectively.
[0235] The RAHT unit 2080 may predict the AC coefficients of the processing target node from, for example, the attribute values of the reference node.
[0236] Specifically, the RAHT unit 2080 converts the decoded attribute value Attr inter The predicted value Attr of the attribute value of the node to be processed is calculated using pred The predicted value Attr pred By applying RAHT to the target node, the predicted value AC pred It is also possible to ask for
[0237] Attr pred =Attr inter AC pred =RAHT(Attr pred ) The RAHT unit 2080 may predict the AC coefficients of the processing target node directly from the AC coefficients of the reference node, for example.
[0238] Specifically, the RAHT unit 2080 calculates the AC coefficient values AC of the reference nodes using the RAHT in the reference frame. inter The calculated value is used as the predicted value AC pred It is also possible to use the following.
[0239] AC pred =AC inter The RAHT unit 2080 may obtain the AC coefficients of the reference node by recording the AC coefficients of each node of the reference frame in the frame buffer 2120 and referring to the values in the frame buffer 2120. In this case, if there are no AC coefficients of the reference node in the frame buffer 2120, the RAHT unit 2080 may determine in step S28110 that inter prediction is not executable.
[0240] In addition, RAHT section 2080 is Attr inter and A.C. inter may be multiplied by a scaling factor α.
[0241] Attr pred =αAttr inter or AC pred = αAC inter The coefficient α may be any real number. The coefficient α may be decoded for each node or for each layer. The coefficient α may be included in the slice data.
[0242] For example, the coefficient α may be defined using the hierarchical depth depth as follows, and α′ may be decoded instead of the coefficient α.
[0243] α=1+α'·2-depth For example, the integer β may be defined as an integer ranging from integer a to integer b, and β may be decoded. The coefficient α may be calculated by adding an integer c to the decoded β and then dividing the result by the integer c, as follows:
[0244] α=(β+c) / c The integer β may be decoded using an exponential-Golomb code.
[0245] Alternatively, the coefficient α may be derived at the decoder.
[0246] For example, the AC coefficients AC parentand the inter prediction value AC when the parent node is decoded. parent_inter may be used to calculate as follows:
[0247] α=AC parent / AC parent_inter For example, the RAHT unit 2080 calculates the AC coefficients AC neighbor1 , A.C. neighbor2 , … , A.C. neighborN and the inter prediction value AC neighbor_inter1 , A.C. neighbor_inter2 , … , A.C. neighbor_interN and α may be calculated so as to minimize the cost.
[0248] The cost may be, for example, the sum of the AC coefficients of each adjacent node and the squared error of the predictor of the AC coefficients. The adjacent nodes may be, for example, only the nodes adjacent to the faces, or the nodes adjacent to the faces and the nodes adjacent to the edges.
[0249] The RAHT unit 2080 may perform a similar operation in the inter prediction of the DC coefficient in step S28003.
[0250] DC pred = αDC inter Here, we use the DC coefficient of the reference node as DC inter The predicted value of the DC coefficient of the root node is DC pred Let us assume that.
[0251] Furthermore, the RAHT unit 2080 may calculate predicted values of attribute values or AC coefficients by combining inter prediction and intra prediction.
[0252] For example, the following is an example in which the RAHT unit 2080 requests a prediction of an attribute value.
[0253] Attr pred =W inter Attr inter +W intra Attrintra Here, Attr inter and Attr intra are respectively the inter-prediction and intra-prediction of the attribute values. Also, W inter and W intra are respectively the weights of the inter-prediction and intra-prediction. W inter and W intra may be determined such that the deeper the layer, the more the intra-prediction is emphasized according to the depth depth of the processing target layer. For example, W inter = 1 - depth / N W intra = depth / N N is the maximum value of the depth of the layer where the inter-prediction is effective. The combination of the inter-prediction and intra-prediction may be effective only at a specific layer. For example, the combination of the inter-prediction and intra-prediction may be effective only when M < depth < N. M is an arbitrary real number less than N and may be decoded as header information such as APS.
[0254] (Point cloud encoding device 100) Hereinafter, with reference to FIG. 18, the point cloud encoding device 100 according to the present embodiment will be described. FIG. 18 is a diagram showing an example of the functional blocks of the point cloud encoding device 100 according to the present embodiment.
[0255] As shown in FIG. 18, the point cloud encoding device 100 includes a coordinate conversion unit 1010, a geometric information quantization unit 1020, a tree analysis unit 1030, an approximate surface analysis unit 1040, a geometric information encoding unit 1050, a geometric information reconstruction unit 1060, a color conversion unit 1070, an attribute transfer unit 1080, a RAHT unit 1090, a LoD calculation unit 1100, a lifting unit 1110, an attribute information quantization unit 1120, an attribute information encoding unit 1130, and a frame buffer 1140.
[0256] The coordinate conversion unit 1010 is configured to convert the three-dimensional coordinate system of the input point cloud into any different coordinate system. The coordinate conversion may convert the x, y, and z coordinates of the input point cloud into any s, t, and u coordinates by rotating the input point cloud, for example. As a variation of the conversion, the coordinate system of the input point cloud may be used as it is.
[0257] The geometric information quantization unit 1020 is configured to quantize the position information of the input point group after the coordinate transformation and to remove points with overlapping coordinates. When the quantization step size is 1, the position information of the input point group and the position information after quantization match. In other words, when the quantization step size is 1, it is equivalent to the case where quantization is not performed.
[0258] The tree analysis unit 1030 is configured to receive position information of the quantized point group as input, and to generate an occupancy code indicating at which node in the encoding target space a point exists, based on a tree structure described below.
[0259] In this process, the tree analysis unit 1030 is configured to recursively divide the encoding target space into rectangular parallelepipeds to generate a tree structure.
[0260] If a point exists within a certain rectangular parallelepiped, a tree structure can be generated by recursively dividing the rectangular parallelepiped into multiple rectangular parallelepipeds until the rectangular parallelepiped reaches a specified size. Each such rectangular parallelepiped is called a node. Each rectangular parallelepiped generated by dividing a node is called a child node, and the occupancy code is expressed as 0 or 1 to indicate whether or not a point is included in the child node.
[0261] As described above, the tree analysis unit 1030 is configured to generate occupancy codes while recursively dividing nodes until a predetermined size is reached.
[0262] In this embodiment, a method called "Octree" can be used, which recursively performs octree division on the above-mentioned rectangular parallelepiped, always treating it as a cube, and a method called "QtBt" can be used, which performs quadtree division and binary tree division in addition to octree division.
[0263] Here, whether or not to use “QtBt” is transmitted to the point cloud decoding device 200 as control data.
[0264] Alternatively, predictive geometry coding using an arbitrary tree structure may be specified to be used. In this case, the tree analysis unit 1030 determines the tree structure, and the determined tree structure is transmitted to the point cloud decoding device 200 as control data.
[0265] For example, the tree-structured control data may be configured so as to be decoded in accordance with the procedures described with reference to FIGS.
[0266] The approximate surface analyzer 1040 is configured to generate approximate surface information using the tree information generated by the tree analyzer 1030 .
[0267] Approximate surface information is used when, for example, decoding three-dimensional point cloud data of an object, in cases where the point cloud is densely distributed on the object surface, to approximate the area in which the point cloud exists using a small plane, rather than decoding each individual point cloud.
[0268] Specifically, the approximate surface analysis unit 1040 may be configured to generate approximate surface information using, for example, a method called "Trisoup." In addition, when decoding a sparse point cloud acquired by Lidar or the like, this process can be omitted.
[0269] The geometric information encoding unit 1050 is configured to generate a bit stream (geometric information bit stream) by encoding syntax such as the occupancy code generated by the tree analysis unit 1030 and the approximate surface information generated by the approximate surface analysis unit 1040. Here, the bit stream may include, for example, the syntax described in FIG. 4.
[0270] The encoding process is, for example, a context-adaptive binary arithmetic encoding process. Here, for example, the syntax includes control data (flags and parameters) for controlling the decoding process of the position information.
[0271] The geometric information reconstruction unit 1060 is configured to reconstruct geometric information (the coordinate system assumed by the encoding process, i.e., the position information after coordinate transformation in the coordinate transformation unit 1010) of each point of the point cloud data to be encoded, based on the tree information generated by the tree analysis unit 1030 and the approximate surface information generated by the approximate surface analysis unit 1040.
[0272] The frame buffer 1140 is configured to receive the geometric information reconstructed by the geometric information reconstruction unit 1060 as an input and store it as a reference frame.
[0273] The stored reference frame is read out from the frame buffer 1140 and used as a reference frame when inter-prediction of a temporally different frame is performed in the tree analysis unit 1030.
[0274] Here, which reference frame to use for each frame may be determined based on, for example, the value of a cost function representing encoding efficiency, and information on the reference frame to be used may be transmitted to the point cloud decoding device 200 as control data.
[0275] The color conversion unit 1070 is configured to perform color conversion when the input attribute information is color information. The color conversion does not necessarily have to be performed, and the presence or absence of the color conversion process is coded as part of the control data and transmitted to the point cloud decoding device 200.
[0276] The attribute transfer unit 1080 is configured to correct the attribute values so as to minimize distortion of the attribute information, based on the position information of the input point cloud, the position information of the point cloud after reconstruction in the geometric information reconstruction unit 1060, and the attribute information after color change in the color conversion unit 1070. As a specific correction method, for example, the method described in Non-Patent Document 1 can be applied.
[0277] The RAHT unit 1090 is configured to receive as input the attribute information after transfer by the attribute transfer unit 1080 and the geometric information generated by the geometric information reconstruction unit 1060, and to generate residual information for each point using a type of Haar transform called RAHT (Region Adaptive Hierarchical Transform).
[0278] The information to be decoded is the direct current component (DC coefficient) and alternating current component (AC coefficient) of the attribute information generated by using RAHT in the encoding process, and in the decoding process, it is converted into attribute information by using the inverse transform of RAHT.
[0279] As a specific example of the RAHT process, the method described in the above-mentioned Non-Patent Document 1 can be used.
[0280] The LoD calculation unit 1100 is configured to receive the geometric information generated by the geometric information reconstruction unit 1060 as an input and generate a Level of Detail (LoD).
[0281] LoD is information for defining a reference relationship (a referencing point and a referenced point) to realize predictive coding, such as predicting attribute information of a certain point from attribute information of another point and encoding or decoding the prediction residual.
[0282] In other words, LoD is information that defines a hierarchical structure in which each point contained in geometric information is classified into multiple levels, and attributes of points belonging to lower levels are encoded or decoded using attribute information of points belonging to higher levels.
[0283] As a specific method for determining the LoD, for example, the method described in Non-Patent Document 1 mentioned above may be used.
[0284] The lifting unit 1110 is configured to generate residual information by a lifting process using the LoD generated by the LoD calculation unit 1100 and the attribute information after attribute transfer by the attribute transfer unit 1080.
[0285] As a specific example of the lifting process, the method described in the above-mentioned non-patent document 1 may be used.
[0286] The attribute information quantization unit 1120 is configured to quantize the residual information output from the RAHT unit 1090 or the lifting unit 1110. Here, a quantization step size of 1 is equivalent to no quantization being performed.
[0287] The attribute information encoding unit 1130 is configured to perform encoding processing using the quantized residual information and the like output from the attribute information quantization unit 1120 as syntax, and to generate a bit stream related to the attribute information (attribute information bit stream).
[0288] The encoding process is, for example, a context-adaptive binary arithmetic encoding process. Here, for example, the syntax includes control data (flags and parameters) for controlling the decoding process of the attribute information.
[0289] Through the above processing, the point group encoding device 100 is configured to perform encoding processing using position information and attribute information of each point in a point group as input, and to output a geometric information bit stream and an attribute information bit stream.
[0290] (Change example 1) Hereinafter, a first modification of the first embodiment will be described with reference to FIG. 19, focusing on differences from the first embodiment.
[0291] In the first embodiment described above, in step S28111, the RAHT unit 2080 calculates the inter-predicted value AC of the AC coefficient. inter This example illustrates a case where the above is scaled by different scaling factors for each Octree layer.
[0292] Here, scaling means that Attr inter and A.C. inter Using the scaling factor α, Attr pred =αAttr inter Or, AC pred = αAC inter and multiply it by α.
[0293] In contrast to this, modification 1 illustrates a case in which the RAHT unit 2080 performs scaling using different scaling factors for each Octree layer and for each frequency index of the AC coefficients.
[0294] Here, the frequency index of an AC coefficient is a number assigned to an AC coefficient generated using RAHT.
[0295] The scaling factor may also be decoded as a syntax included in APS2611 or ASH2612.
[0296] Figures 19 and 33 show examples of the syntax configuration of APS2611 and ASH2612 when a scaling factor is transmitted by ASH2612. Note that, hereinafter, only the difference between the syntax configuration shown in Figure 33 and the syntax configuration explained in Figure 6 will be explained.
[0297] The APS 2611 may include a value (raht_send_inter_filters) indicating whether or not to transmit a scaling factor in inter prediction of the attribute information.
[0298] For example, when raht_send_inter_filters is "1", it may be defined that a scaling factor in inter prediction of the attribute information is transmitted, and when raht_send_inter_filters is "0", it may be defined that a scaling factor in inter prediction of the attribute information is not transmitted.
[0299] The APS2611 may include a value (raht_inter_skip_layers) indicating how many layers above the root node of the Octree are to be exempt from inter prediction scaling application for inter prediction in the attribute information.
[0300] For example, when raht_inter_skip_layers is "3", it may be defined that inter prediction is not applied to the first layer to the third layer.
[0301] The APS 2611 may include a value (raht_enable_code_layer) indicating whether or not to transmit an applicability mode of inter prediction for each layer. Alternatively, the APS 2611 may include raht_enable_code_layer when raht_prediction_enabled is "1" and raht_inter_prediction_enabled is "1".
[0302] For example, when raht_enable_code_layer is '1', it may be defined that the applicability mode of inter prediction for each layer is transmitted, and when raht_enable_code_layer is '0', it may be defined that the applicability mode of inter prediction for each layer is not transmitted.
[0303] ASH2612 may include a value (raht_attr_layer_depth_num) indicating the number of layers of the frame if either raht_enable_code_layer or raht_send_inter_filters is "1".
[0304] Or, for example, ASH2612 may include raht_attr_layer_depth_num when only raht_send_inter_filters is "1".
[0305] Alternatively, raht_attr_layer_depth_num may be defined as the number of layers of the frame minus 1, or may be used by adding 1 after decoding.
[0306] Alternatively, if raht_attr_layer_depth_num is "0", raht_attr_layer_depth_num may be used as "0", and if it is other than "0", 1 may be subtracted from it after decoding.
[0307] When raht_enable_code_layer is "1", ASH2612 may include an inter prediction applicability mode (raht_attr_layer_code_mode) for each layer, the number of which is equal to raht_attr_layer_depth_num.
[0308] For example, in each layer, when inter prediction is applied, it may be defined as "1", and when inter prediction is not applied, it may be defined as "0".
[0309] In the ASH2612, when raht_send_inter_filters is “1”, the number of scaling factor values (raht_filter_taps) in the inter prediction (num_filter_taps) may be included. The num_filter_taps may be derived based on a syntax that specifies a decoded layer to which the inter prediction is applied.
[0310] For simplicity, an example of a method for deriving num_filter_taps will be described below using an example in which there is one type of scaling factor for each layer.
[0311] For example, num_filter_taps may be derived based on a value (raht_inter_skip_layers) indicating how many top layers should not be subject to inter prediction scaling, a value (raht_inter_prediction_depth_minus1) indicating the number of effective layers for inter prediction, and the number of layers for the frame (raht_attr_layer_depth_num).
[0312] Here, the effective number of layers of inter prediction is a numerical value indicating a threshold value of a layer to which inter prediction is applied. For example, the effective number of layers of inter prediction may be a value obtained by adding 1 to raht_inter_prediction_depth_minus1, and when raht_inter_prediction_depth_minus1 is "N-1", the effective number of layers of inter prediction may be defined as "N".
[0313] Specifically, for example, when the number of hierarchical layers of the frame is greater than the number of effective hierarchical layers for inter prediction, num_filter_taps may be obtained by subtracting a value indicating how many top layers are not to be applied to inter prediction scaling from the number of effective hierarchical layers for inter prediction; when the number of hierarchical layers of the frame is smaller than the number of effective hierarchical layers for inter prediction, num_filter_taps may be obtained by subtracting a value indicating how many top layers are not to be applied to inter prediction scaling from the number of hierarchical layers of the frame.
[0314] That is, num_filter_taps may be derived as follows:
[0315]
number
[0316] Alternatively, for example, num_filter_taps may be derived based on a value (raht_inter_skip_layers) indicating how many top layers inter prediction scaling should not be applied to, a value (raht_inter_prediction_depth_minus1) indicating the number of valid layers for inter prediction, a value (raht_attr_layer_depth_num) indicating the number of layers for the frame, and the applicability mode of inter prediction for each layer (raht_attr_layer_code_mode).
[0317] Figure 34 is a flowchart showing an example of a process for deriving num_filter_taps based on a value indicating how many top layers should not be subject to inter-prediction scaling, a value indicating the number of effective layers for inter-prediction, the number of layers for the frame, and the applicability mode of inter-prediction for each layer.
[0318] In step S3401, the attribute information decoding unit 2060 determines whether or not to transmit the applicability mode of inter prediction for each layer. The determination may be made using the raht_enable_code_layer.
[0319] If it is determined that the applicability mode of inter prediction for each layer is to be transmitted, this operation proceeds to step S3405, and if it is determined that the applicability mode of inter prediction for each layer is not to be transmitted, this operation proceeds to step S3402.
[0320] In step S3402, the attribute information decoding unit 2060 determines whether or not the number of layers of the frame is greater than the number of effective layers of inter prediction. For example, the determination may be made using the above-mentioned raht_attr_layer_depth_num and raht_inter_prediction_depth_minus1.
[0321] If it is determined that the number of hierarchical layers of the frame is smaller than the number of effective hierarchical layers for inter prediction, the operation proceeds to step S3403; if it is determined that the number of hierarchical layers of the frame is greater than the number of effective hierarchical layers for inter prediction, the operation proceeds to step S3404.
[0322] In step S3403, the attribute information decoding unit 2060 derives the number of scaling factors using the number of layers of the frame. Specifically, the attribute information decoding unit 2060 may obtain the number of scaling factors by subtracting a value indicating up to which upper layers the inter prediction scaling should not be applied from the number of layers of the frame. After deriving the number of scaling factors, the operation proceeds to step S3406, and the process ends.
[0323] In step S3404, the attribute information decoding unit 2060 derives the number of scaling factors using the number of effective layers of inter prediction. Specifically, the attribute information decoding unit 2060 may obtain the number of scaling factors by subtracting a value indicating up to which upper layers the inter prediction scaling should not be applied from the number of effective layers of inter prediction. After deriving the number of scaling factors, this operation proceeds to step S3406, and the process ends.
[0324] That is, the number of scaling factors in steps S3402 to S3404 can be derived by the following formula.
[0325]
number
[0326] In step S3405, the attribute information decoding unit 2060 derives the number of scaling factors using the applicability mode of inter prediction for each layer. Specifically, the attribute information decoding unit 2060 may count the layers to which inter prediction is applied based on the applicability mode of inter prediction for each layer. However, the attribute information decoding unit 2060 may exclude layers to which inter prediction scaling is not applicable from the count based on a value indicating up to which upper layers inter prediction scaling is not applicable. After deriving the number of scaling factors, this operation proceeds to step S3406 and ends the process.
[0327] In the above, an example has been described in which the above-mentioned information is decoded by the APS 2611, but the information may be included in the ASH 2612 or in the SPS 2601. In other words, the information may be included in any of the headers.
[0328] For example, when the above-mentioned decoded raht_filter_taps is X, the RAHT unit 2080 may subtract X from 128, shift the result to the right by 7 bits, and use the result as the scaling factor α for inter prediction.
[0329] For example, when the value of raht_filter_taps is "0", the value of the scaling factor α in inter prediction in the attribute information may be defined as "1" obtained by subtracting 0 from 128 and shifting the result 7 bits to the right.
[0330] The RAHT unit 2080 may determine whether to apply inter prediction based on a syntax that specifies a layer to which inter prediction is applied, and may scale the inter prediction value using the decoded raht_filter_taps if it is determined that inter prediction is applied to the layer. If it is determined that inter prediction is not applied to the layer, it is not necessary to scale the inter prediction.
[0331] Specifically, the RAHT unit 2080 may determine to scale the inter-prediction value if the depth of the hierarchy in which the node to be processed is included is equal to or less than the number of effective hierarchical layers for inter-prediction, and if the depth of the hierarchy in which the node to be processed is included is equal to or greater than a value indicating how many upper layers are exempt from inter-prediction scaling.
[0332] Here, the RAHT unit 2080 may refer to raht_inter_prediction_depth_minus1 and use the value thereof for the number of effective layers for inter prediction.
[0333] Furthermore, the RAHT unit 2080 may refer to raht_inter_skip_layers and use the value indicating up to which upper layers the scaling of inter prediction is not applied.
[0334] Alternatively, for example, when the depth of the hierarchical layer including the node to be processed is equal to or less than the number of effective hierarchical layers for inter prediction, and the depth of the hierarchical layer including the node to be processed is equal to or greater than a value indicating how many upper layers are exempt from inter prediction scaling, and it is determined that inter prediction is applied in the hierarchical layer including the node to be processed, the RAHT unit 2080 may determine to scale the inter prediction value.
[0335] Here, the RAHT unit 2080 may refer to raht_attr_layer_code_mode and determine whether or not inter prediction is applied in the layer that includes the node to be processed, based on the value thereof.
[0336] In the above, the case where there is one scaling factor for each layer has been described, but even when a scaling factor is transmitted for each frequency index idx of the AC coefficient, the number of scaling factors to be decoded can be derived by multiplying the number of scaling factors calculated above by the number of scaling factors for each layer. The number of scaling factors may be, for example, 7.
[0337] Also, for example, a scaling factor α_(depth_idx) that differs for each layer and for each frequency index idx of the AC coefficient may refer to raht_filter_taps and use the value thereof.
[0338] Alternatively, when it is determined with reference to raht_send_inter_filters that raht_filter_taps is not to be transmitted, the scaling factor α_(depth_idx) may be set to an arbitrary hard-coded value.
[0339] Furthermore, the RAHT unit 2080 may group frequency indexes of AC coefficients, assign a group to each frequency index, and use the scaling factor of the corresponding group as the scaling factor of each frequency index.
[0340] Furthermore, the RAHT unit 2080 may be configured to derive other scaling factors based on some of the decoded scaling factors.
[0341] For example, the scaling factor may be derived based on a scaling factor of another layer that has already been decoded.
[0342] That is, the RAHT unit 2080 may determine the scaling factor α_(d2_idx) for each frequency index idx in the layer d2 based on the scaling factor α_(d1_idx) for each frequency index idx in the decoded layer d1.
[0343] Alternatively, for example, the scaling factor may be derived based on a scaling factor of another frequency index that has already been decoded.
[0344] That is, the RAHT unit 2080 may determine the scaling factor α_(depth_i2) for the frequency index i2 of each layer based on the scaling factor α_(depth_i1) for the decoded frequency index i1 of each layer.
[0345] Alternatively, for example, the scaling factor may be derived based on a scaling factor of another layer that has already been decoded and another frequency index that has already been decoded.
[0346] That is, the RAHT unit 2080 may determine the scaling factor α_(d2_i2) for frequency index i2 in layer d2 based on the scaling factor α_(d1_i1) for frequency index i1 in the decoded layer d1.
[0347] Furthermore, after decoding the scaling factors, the RAHT unit 2080 may rearrange the order of the frequency indexes in any order.
[0348] For example, the RAHT unit 2080 may rearrange the decoded scaling factors in the order of idx=3, 1, 5, 2, 6, 4, 7 in the hierarchy depth.
[0349] (Change example 2) Hereinafter, a second modification of the first embodiment will be described with reference to FIGS. 20 to 25, focusing on the differences from the first embodiment.
[0350] Fig. 20 shows an example of the syntax configuration of the APS2611 in this modified example. Note that only the differences from the syntax configuration explained in Fig. 6 will be explained.
[0351] The APS 2611 may include values (raht_prediction_threshould0, raht_prediction_threshould1) indicating thresholds for the number of adjacent nodes of the grandparent node and parent node of the node to be processed, as conditions for determining whether or not to predict attribute information in the RAHT unit 2080.
[0352] For example, if the number of adjacent nodes of the grandparent node of the node to be processed is smaller than raht_prediction_threshould0, or if the number of adjacent nodes of the parent node of the node to be processed is smaller than raht_prediction_threshould1, intra prediction of attribute information may not be performed, and in other cases, intra prediction of attribute information may be performed.
[0353] Fig. 21 is a flowchart showing an example of the process of step S28104. Note that, hereinafter, only the differences from the flowchart explained using Fig. 9 will be explained, and the same reference numerals will be used for the parts that have not been changed from Fig. 9, and the explanation will be omitted.
[0354] As shown in FIG. 21, in step S28114, the RAHT unit 2080 determines whether or not to intra-predict the AC coefficients of the node to be processed.
[0355] In making this determination, the RAHT unit 2080 may check whether intra prediction is feasible, and if intra prediction is feasible, perform intra prediction, and if intra prediction is not feasible, may not perform intra prediction.
[0356] For example, the RAHT unit 2080 may determine that intra prediction is executable when the number of adjacent nodes of the grandparent node and the parent node of the processing target node are equal to or greater than raht_prediction_threshould0 and raht_prediction_threshould1, respectively.
[0357] Alternatively, the RAHT unit 2080 may determine that intra prediction is executable when the number of adjacent nodes of the parent node of the processing target node is equal to or greater than raht_prediction_threshould1.
[0358] In this determination, the RAHT unit 2080 may decode a flag indicating whether or not to intra-predict the AC coefficient of the node to be processed, and use the value.
[0359] The flag may be decoded for each node or for each layer. The flag may be decoded and the above-mentioned determination may be made only if it is determined that intra prediction is possible by the above-mentioned method. The flag may be included in slice data.
[0360] If intra prediction is possible, the operation proceeds to step S28112, and if intra prediction is not possible, the operation proceeds to step S28115.
[0361] In step S28115, the RAHT unit 2080 skips prediction of the AC coefficient of the node to be processed. For example, the RAHT unit 2080 may not perform prediction, and may input a predicted value of 0 to the subsequent process.
[0362] When step S28115 ends, the operation proceeds to step S28113, and the processing of step S28104 ends.
[0363] Note that the determination in step S28109 may be replaced with a determination based on syntax (flags, etc.) included in the bitstream. Such syntax may be decoded for each layer of the RAHT. Such syntax may be included in the APS2611, ASH2612A / 2612B, or slice data 2613A / 2613B.
[0364] Figures 22 and 23 are flowcharts showing an example of the process of step S28104. Note that, hereinafter, only the differences from the flowchart explained using Figure 21 will be explained, and the same reference numerals will be used for the parts that have not been changed from Figure 21, and the explanation will be omitted.
[0365] As shown in FIG. 22, if the RAHT unit 2080 determines in step S28109 that the depth of the layer including the processing target node is greater than the threshold value, the operation proceeds to the intra-priority flow in FIG.
[0366] In the intra priority flow of FIG. 23, operation proceeds to step S28116.
[0367] In step S28116, the RAHT unit 2080 determines whether or not intra prediction is executable, using a method similar to that in step S28114.
[0368] If intra prediction is feasible, the operation proceeds to step S28112, and if intra prediction is not feasible, the operation proceeds to step S28117.
[0369] In step S28117, the RAHT unit 2080 determines whether inter prediction is executable. Here, the RAHT unit 2080 may decode a flag indicating the presence or absence of a reference node or whether inter prediction is to be executed, and make a determination based on the value, as in step S28110. Such a flag may be decoded for each node at a certain layer or higher, and the same flag as that of the parent node of the processing node may be used in layers lower than the certain layer. The "certain layer" may be specified by a syntax included in a header such as APS2611 or ASH2612A / 2612B, or may be specified by a fixed value set in advance.
[0370] If inter prediction is possible, the operation proceeds to step S28118, and if inter prediction is not possible, the operation proceeds to step S28115.
[0371] In step S28118, the RAHT unit 2080 performs inter prediction of the AC coefficients of the node to be processed.
[0372] Here, the RAHT unit 2080 may perform inter-prediction of the AC coefficient of the node to be processed using the AC coefficients or attribute values of the reference node or the adjacent nodes of the reference node, in the same manner as in the method described with reference to FIG.
[0373] Furthermore, when the RAHT unit 2080 uses the attribute values of the reference node or the adjacent nodes of the reference node, the RAHT unit 2080 may predict the AC coefficients of the node to be processed by applying RAHT using the weights of the frame to be processed to the attribute values.
[0374] Furthermore, the RAHT unit 2080 may perform inter prediction using AC coefficients or attribute values of both the reference node and the neighboring nodes of the reference node.
[0375] For example, the RAHT unit 2080 may take the weighted average of the AC coefficients of the reference node and the adjacent nodes of the reference node as the predicted value of the AC coefficient of the processing target node.
[0376] Furthermore, the RAHT unit 2080 may use the AC coefficient obtained by applying RAHT using the weight of the processing target frame to the weighted average of the attribute values of the reference node and the adjacent nodes of the reference node as the predicted value of the AC coefficient of the processing target node.
[0377] Here, the RAHT unit 2080 may set the weights of the weighted average so that the weight of the reference node is large and the weight of the adjacent node of the reference node is small.
[0378] Furthermore, the RAHT unit 2080 may set the weights of the weighted average to be large for face adjacent nodes and small for edge adjacent nodes among the adjacent nodes of the reference node.
[0379] When an empty node is included in the reference node or an adjacent node of the reference node, the RAHT unit 2080 may calculate the weighted average using values of nodes other than the empty node.
[0380] When the inter prediction is completed, the operation proceeds to step S28113, and the processing of step S28104 ends.
[0381] Fig. 24 is a flowchart showing an example of the process of step S28104. Note that, hereinafter, only the differences from the flowchart explained using Fig. 22 will be explained, and the same reference numerals will be used for the parts that have not been changed from Fig. 22, and the explanation will be omitted.
[0382] As shown in FIG. 24, if it is determined in step S28114 that intra prediction is not executable, the operation proceeds to step S28119.
[0383] In step S28119, the RAHT unit 2080 determines whether or not to predict the AC coefficients of the node to be processed by the enhanced intra prediction. The enhanced intra prediction will be described later.
[0384] The RAHT unit 2080 may determine that the enhanced intra prediction is executable, for example, when the number of adjacent nodes of the grandparent node and the parent node of the processing target node are equal to or greater than the respective thresholds. Here, the thresholds may be decoded as header information such as the APS.
[0385] The RAHT unit 2080 may decode and determine a flag indicating whether or not the AC coefficient of the processing target node is predicted by the extended intra prediction. Such a flag may be included in slice data. Such a flag may be decoded for each node, or may be decoded for each layer. Such a flag may be decoded and determined only when it is determined that the extended intra prediction is executable by the above-mentioned method. Such a flag may be decoded for each node at a certain layer or higher, and the same flag as that of the parent node of the processing node may be used at a layer lower than the certain layer. The "certain layer" may be specified by a syntax included in a header such as APS2611 or ASH2612A / 2612B, or may be specified by a fixed value set in advance.
[0386] If it is determined that the enhanced intra prediction is executable, the operation proceeds to step S28120, and if it is determined that the enhanced intra prediction is not executable, the operation proceeds to step S28115.
[0387] In step S28120, the RAHT unit 2080 predicts the AC coefficients of the node to be processed by extended intra prediction.
[0388] Here, the enhanced intra prediction predicts the attribute value of the node to be decoded by referring to the values of neighboring nodes that are not directly adjacent to the node to be processed as well as the values of neighboring nodes.
[0389] As shown in FIG. 25, the neighboring nodes used in the scaled intra prediction may be nodes that are adjacent to the parent node of the decoding target node and that are not directly adjacent to the decoding target node.
[0390] The prediction of the attribute value of the node to be decoded is calculated as a weighted average of the attribute values of adjacent nodes in the upper layer of the node to be processed, adjacent nodes in the subnode layer, and nearby nodes, as in steps S28203 and S28206.
[0391] The predicted values of the attribute values are converted into predicted values of the AC coefficients in the same manner as in step S28207.
[0392] (Change example 3) Hereinafter, a second modification of the first embodiment will be described with reference to FIGS. 26 to 34, focusing on differences from the first embodiment.
[0393] Figure 26 is a diagram explaining the processing flow of RAHT in this modified example 3. Note that, below, only the differences from the points explained using Figure 7 will be explained, and the points that have not changed from Figure 7 will be assigned the same reference numerals and explanations will be omitted.
[0394] In this third modification, as shown in step S30001 in FIG. 7, the RAHT unit 2080 performs region division processing.
[0395] An example of such area division processing will be described below with reference to FIG.
[0396] FIG. 27(a) is an example of an octree constructed in step S28001.
[0397] In the flow of FIG. 7, the RAHT unit 2080 performs subsequent processing on the octree of FIG. 27(a).
[0398] On the other hand, in this third modified example, as shown in FIG. 27(b), the RAHT unit 2080 divides the above-mentioned Octree into a predetermined hierarchy (a hierarchy in which the node size is 2N in the example of FIG. 27(b)).
[0399] Hereafter, a group of nodes connected to a single root node will be referred to as a "tree."
[0400] The process of step S30001 can be said to be a process of dividing one tree as shown in Fig. 27(a) into multiple trees. Note that dividing a tree for each node of a predetermined size can also be said to be a process of dividing the space to be decoded into areas corresponding to each root node.
[0401] In this third modified example, the RAHT unit 2080 performs RAHT processing on the multiple trees divided as described above.
[0402] Here, in the example of Figure 26, the RAHT unit 2080 first performs the Octree processing of step S28001 and then performs the region segmentation processing of step S30001, but this order may be reversed to first perform the region segmentation processing of step S30001 and then perform the Octree processing of step S28001 for each region.
[0403] When the region division process in step S30001 is executed first, this can be realized by classifying each point belonging to the same root node based on the coordinate information of each point to be decoded and a predetermined node size.
[0404] Specifically, for example, the RAHT unit 2080 may first sort the points to be decoded in the order of Morton codes, and then divide them into regions for each root node.
[0405] When Octree processing is performed in Morton code order, if the points to be decoded are sorted in Morton code order, points belonging to the same root node will appear consecutively, making region division easier.
[0406] In addition, the specified node size may be decoded as syntax included in APS2611 or ASH2612.
[0407] FIG. 28 shows an example of a syntax table when the relevant information is transmitted by the APS2611.
[0408] The attribute information decoding unit 2060 may decode a flag (raht_split_enabled) that controls whether or not to execute the above-mentioned region splitting process. If the value of this flag is "1", the RAHT unit 2080 may be specified to execute the splitting process and perform the process in the flow of Fig. 26. On the other hand, if the value of this flag is "0", the RAHT unit 2080 may be specified to not perform the splitting process and to perform the process in the flow of Fig. 7, for example.
[0409] In addition, when the flag that controls whether or not to perform region division processing indicates "perform region division," the attribute information decoding unit 2060 may decode a syntax (raht_split_nodesize_log2) that specifies the root node size at the time of region division.
[0410] When the root node size can only take values that are powers of 2, the attribute information decoding unit 2060 may decode, as such syntax, a value obtained by converting the root node size into a logarithm with base 2.
[0411] Furthermore, when a minimum value of the root node size is specified, the attribute information decoding unit 2060 may decode, as the syntax, a value obtained by subtracting the minimum value in advance.
[0412] For example, if the minimum value of the root node size is 4 (=22), the attribute information decoding unit 2060 may first add 2 to the value decoded as raht_split_nodesize_log2, and then convert it to a power of 2 to calculate the final root node size.
[0413] Furthermore, the attribute information decoding unit 2060 may decode, instead of the root node size, information on at which level of the octree the region should be divided.
[0414] For example, the attribute information decoding unit 2060 may decode a value indicating at what level the tree is divided when the densest hierarchy of the octree is defined as level 0, and each level is successively sparser as level 1, level 2, and so on.
[0415] In the above, an example has been described in which the above-mentioned information is decoded by the APS 2611, but the information may be included in the ASH 2612 or in the SPS 2601. In other words, the information may be included in any of the headers.
[0416] After the above-mentioned region division process is performed, the operation proceeds to step S30002.
[0417] In step S30002, the RAHT unit 2080 checks whether the processing has been completed for all areas (trees). If the processing has been completed, the operation proceeds to step S28008 and ends. On the other hand, if the processing has not been completed, the operation proceeds to step S28002.
[0418] Next, with reference to FIG. 29, the difference between the decoding process of the DC coefficient in step S30003 and step S28003 will be described.
[0419] In the example of step S28003, there is only one root node, but in this modification example 3, as described above, there are multiple root nodes and their corresponding DC coefficients.
[0420] Figure 29(a) shows the four root nodes (R a , R b , R c , R d ) is present. Here, the decoding order is R a Tree → R b Tree → R c Tree → R d The order of the tree is as follows:
[0421] For example, the DC coefficients of the root nodes may be decoded independently of each other. With this configuration, the DC coefficients of the root nodes can be decoded in parallel, thereby reducing the processing time.
[0422] Also, for example, the RAHT unit 2080 may generate a predicted value of the DC coefficient of the root node of the tree to be decoded using the DC coefficient of the root node that has already been decoded, and may decode the DC coefficient of the root node of the tree to be decoded by adding the predicted value to the decoded DC coefficient value (residual of the DC coefficient).
[0423] Specifically, for example, the RAHT unit 2080 may use the DC coefficient decoded immediately before in decoding order as the above-mentioned predicted value.
[0424] In the example of FIG. 29(a), the RAHT unit 2080 b The prediction of the DC coefficient of the root node of the tree is done using R a Use the DC coefficient of the root node of the tree, R c The prediction of the DC coefficient of the root node of the tree is done using R b The DC coefficient of the root node of the tree is used.
[0425] Also, for example, the RAHT unit 2080 may perform intra prediction as described with reference to FIG. 10 using the DC coefficient of an already decoded root node.
[0426] For example, as shown in FIG. 29(b), the root node (R a , Rb , R c , R d ) may be spatially adjacent, so in this case, intra prediction as described in FIG. 10 can be performed.
[0427] Since there are no nodes at a higher level than the root node, the process of generating a predicted value in step S28206 is executed based only on the attribute values of the subnode hierarchy in step S28205 in FIG.
[0428] With this configuration, the amount of code for the DC coefficients can be reduced.
[0429] Furthermore, the RAHT unit 2080 may perform inter prediction using DC coefficients of other frames that have already been decoded.
[0430] Next, the decoding process of AC coefficients in step S30005 will be described.
[0431] The process of step S30005 is basically the same as the process described in step S28005 and Fig. 10. The difference is the search range when acquiring the attribute value of an adjacent node or a subnode described in step S28202, step S28204, and step S28205 in Fig. 10.
[0432] For example, when acquiring the attribute values of adjacent nodes or subnodes described in steps S28202, S28204, and S28205 of FIG. 10, the RAHT unit 2080 may set the search range to only nodes belonging to the same root node.
[0433] For example, when performing intra prediction of the AC coefficients of node C6 in FIG. 30(a), the RAHT unit 2080 may set the area surrounded by the dotted line as the search range.
[0434] Specifically, the RAHT unit 2080 may set the search range to nodes C1 to C3 in steps S28202 and S28204, and may set the search range to nodes C4 and C5 in step S28205.
[0435] With this configuration, it becomes possible to decode AC coefficients independently for each tree, and therefore parallel processing can be performed on a tree-by-tree basis.
[0436] Also, for example, when obtaining attribute values of adjacent nodes or subnodes described in steps S28202, S28204, and S28205 of FIG. 10, the RAHT unit 2080 may include all decoded nodes, including nodes belonging to other root nodes, in the search range.
[0437] For example, when performing intra prediction of the AC coefficients of node C6 in FIG. 30(b), the RAHT unit 2080 may set the area surrounded by the dotted line as the search range.
[0438] Specifically, the RAHT unit 2080 may set the search range to nodes A1 to A3, B1, B2, and C1 to C3 in steps S28202 and S28204, and may set the search range to nodes A4 to A10, B3 to B6, C4, and C5 in step S28205.
[0439] With this configuration, a larger number of nodes can be used for prediction compared to the configuration in FIG. 30(a), so that the prediction accuracy is improved and the amount of code for AC coefficients can be reduced.
[0440] Next, in relation to both steps S30003 and S30005, construction of a context when decoding coefficients will be described with reference to FIG.
[0441] The context (probability distribution used for arithmetic decoding) when decoding coefficients may be initialized for each tree, as shown in FIG. 31(a).
[0442] Here, the arrows in the figure indicate the order in which the context used for coefficient decoding is updated.
[0443] The RAHT unit 2080 first decodes the DC coefficient and then the AC coefficient in each root node, and then one-dimensionally decodes the AC coefficients and updates the context in the order of breadth-first search.
[0444] With the configuration of FIG. 31(a), the context is independent for each tree, so that the coefficient decoding process can be executed in parallel for each tree.
[0445] The context (probability distribution used for arithmetic decoding) when decoding coefficients may be initialized only once before decoding the first root node, as shown in Figure 31(b), and then updated sequentially in the order of the decoding process.
[0446] The context (probability distribution used for arithmetic decoding) used when decoding coefficients may be initialized and updated at the same level between different trees, as shown in FIG. 32(a).
[0447] When the distribution of absolute values of coefficients differs for each level, as shown in Figure 32(a), by updating the context (probability distribution) for each level, it becomes easier to converge to a probability distribution appropriate for each level, thereby improving coding efficiency.
[0448] The context (probability distribution used for arithmetic decoding) used when decoding coefficients may be initialized and updated separately for the root node and the other nodes, as shown in FIG. 32(b).
[0449] Furthermore, the RAHT unit 2080 may initialize and update the context by dividing it into the DC coefficient of the root node and other coefficients (the AC coefficient of the root node and the AC coefficients of other nodes).
[0450] Since the absolute value of a DC coefficient tends to be much larger than that of an AC coefficient, by separating the contexts (probability distributions) of the DC and AC coefficients, they can be converged to appropriate probability distributions, improving coding efficiency.
[0451] The first embodiment and modified examples 1 to 3 described above can be arbitrarily combined with each other.
[0452] Furthermore, the above-mentioned point group encoding device 100 and point group decoding device 200 may be realized as a program that causes a computer to execute each function (each process).
[0453] In each of the above embodiments, the present invention has been described using the point cloud encoding device 100 and the point cloud decoding device 200 as examples, but the present invention is not limited to such examples and can be similarly applied to a point cloud encoding / decoding system having the functions of the point cloud encoding device 100 and the point cloud decoding device 200. [Industrial Applicability]
[0454] According to this embodiment, for example, it is possible to improve the overall service quality in video communication, which makes it possible to contribute to Goal 9 of the Sustainable Development Goals (SDGs) led by the United Nations, which is to "build resilient infrastructure, promote sustainable industrialization and foster innovation." [Explanation of symbols]
[0455] 10...Point cloud processing system 100...Point cloud encoding device 1010... Coordinate conversion section 1020...Geometric information quantization section 1030…Tree analysis section 1040…Approximate surface analysis section 1050...Geometric information encoding unit 1060...Geometric information reconstruction unit 1070…Color conversion section 1080…Attribute transfer section 1090…RAHT Section 1100…LoD calculation section 1110…Lifting section 1120...Attribute information quantization section 1130...Attribute information encoding unit 200...Point cloud decoding device 2010…Geometric Information Decoding Department 2020…Tree synthesis section 2030…Approximate surface synthesis part 2040...Geometric information reconstruction unit 2050…Inverse coordinate conversion section 2060…Attribute information decoding unit 2070...Inverse quantization section 2080…RAHT Division 2090…LoD calculation section 2100…Reverse lifting section 2110…Color inversion unit
Claims
1. A point cloud decoding device, comprising: in the inter prediction of the AC coefficient of RAH T, a RAH T unit that applies a scaling factor to the predicted value of the AC coefficient or the predicted value of the attribute value; an attribute information decoding unit that derives the number of scaling factors used in the inter prediction and decodes the scaling factors by the number, characterized in that the point cloud decoding device comprises the attribute information decoding unit.
2. The point cloud decoding device according to claim 1, wherein the attribute information decoding unit derives the number of scaling factors using the number of effective layers of the inter prediction.
3. The point cloud decoding device according to claim 2, wherein the attribute information decoding unit determines whether to transmit the applicability mode of the inter prediction for each layer, and controls the method for deriving the number of effective layers of the inter prediction based on the determination result.
4. The point cloud decoding device according to claim 3, wherein when the determination result of whether to transmit the applicability mode of the inter prediction for each layer is "transmit the applicability mode of the inter prediction for each layer", the attribute information decoding unit decodes the applicability mode of the inter prediction for each layer, and derives the number of effective layers of the inter prediction based on the applicability mode of the inter prediction for each layer.
5. The point cloud decoding device according to any one of claims 2 to 4, wherein the attribute information decoding unit derives the number of scaling factors by subtracting a value indicating up to which upper layer the scaling application of the inter prediction is excluded from the number of effective layers of the inter prediction.
6. A point cloud decoding method, comprising: in the inter prediction of the AC coefficient of RAH T, a step of applying a scaling factor to the predicted value of the AC coefficient or the predicted value of the attribute value; a step of deriving the number of scaling factors used in the inter prediction and decoding the scaling factors by the number, characterized in that the point cloud decoding method comprises the step.
7. A program for causing a computer to function as a point cloud decoding device, wherein the point cloud decoding device in the inter prediction of the AC coefficient of RAH T, a RAH T unit that applies a scaling factor to the predicted value of the AC coefficient or the predicted value of the attribute value; A program comprising: an attribute information decoding unit that derives the number of scaling factors used in the inter prediction and decodes the scaling factors by the number.