Point cloud decoding device, point cloud decoding method, and program

By searching for adjacent nodes in the upper layer and setting a search range based on past results, the decoding processing amount of attribute information is reduced, addressing inefficiencies in conventional Morton code-based methods.

WO2025150309A1PCT designated stage expired Publication Date: 2025-07-17KDDI CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/042926
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-12
Filing Date
2024-12-04
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

The conventional method of searching for adjacent nodes in intra prediction using Morton code results in increased processing amounts due to varying spatial Morton codes, leading to inefficient decoding of attribute information in point cloud processing.

Method used

The proposed solution involves searching for adjacent nodes in the upper layer of intra prediction and setting a search range based on past search results, utilizing a RAHT unit to reduce decoding processing amounts.

Benefits of technology

This approach effectively reduces the decoding processing amount of attribute information by optimizing the search range and improving efficiency in point cloud decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024042926_17072025_PF_FP_ABST
    Figure JP2024042926_17072025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention reduces the decoding processing amount of attribute information. A point cloud decoding device 200 according to the present invention includes a RAHT unit 2080 that searches for adjacent nodes in the upper hierarchy in intra-prediction and sets the search range of the adjacent nodes in the upper hierarchy on the basis of the past search result.
Need to check novelty before this filing date? Find Prior Art

Description

Point group decoding device, point group decoding method and program

[0001] The present invention relates to a point group decoding device, a point group decoding method, and a program.

[0002] Conventionally, when searching for adjacent nodes in RAHT intra prediction, the maximum range in which adjacent nodes can exist is searched in the order of the motion code.

[0003] G-PCC codec description, ISO / IEC JTC 1 / SC 29 / WG 7 N 00271G-PCC 2nd edition codec description, ISO / IEC JTC 1 / SC 29 / WG 7 N00506

[0004] However, in the motion code order, the motion code values ​​of spatially adjacent nodes may differ greatly, which increases the search range and the amount of processing required for the search.

[0005] Therefore, the present invention has been made in consideration of the above-mentioned problems, and aims to provide a point cloud decoding device, a point cloud decoding method, and a program that can reduce the amount of decoding processing required for attribute information.

[0006] A first feature of the present invention is a point cloud decoding device that includes an RAHT unit that searches for adjacent nodes in a higher layer in intra prediction and sets a search range for the adjacent nodes in the higher layer based on past search results.

[0007] A second feature of the present invention is summarized as a point cloud decoding method including a step of searching for adjacent nodes in a higher layer in intra prediction and setting a search range for the adjacent nodes in the higher layer based on past search results.

[0008] A third feature of the present invention is a program that causes a computer to function as a point cloud decoding device, wherein the point cloud decoding device includes an RAHT unit that searches for adjacent nodes in a higher hierarchy in intra prediction and sets a search range for the adjacent nodes in the higher hierarchy based on past search results.

[0009] According to the present invention, it is possible to provide a point group decoding device, a point group decoding method, and a program that can reduce the amount of decoding processing for attribute information.

[0010] FIG. 1 is a diagram showing an example of the configuration of a point cloud processing system 10 according to an embodiment. FIG. 2 is a diagram showing an example of functional blocks of a point cloud decoding device 200 according to an embodiment. FIG. 3 is a diagram showing an example of the configuration of coded data (bit stream) received by a geometric information decoding unit 2010 of the point cloud decoding device 200 according to an embodiment. FIG. 4 is a diagram showing an example of the syntax configuration of GPS2011. FIG. 5 is an example of the configuration of coded data (bit stream) received by an attribute information decoding unit 2060 of the point cloud decoding device 200 according to an embodiment. FIG. 6 is an example of the syntax configuration of APS2611 shown in FIG. 5. FIG. 7 is a flowchart showing an example of the processing of the RAHT unit 2080. FIG. 8 is a flowchart showing an example of the processing of step S28004. FIG. 9 is a flowchart showing an example of the processing of step S28104. FIG. 10 is a flowchart showing an example of the intra prediction processing of step S28112. FIG. 11 is a diagram showing the relationship between a node to be decoded and adjacent nodes in a higher layer. FIG. 12 is a diagram showing the relationship between a node to be decoded and adjacent nodes in the subnode hierarchy. FIG. 13 is a flowchart showing an example of the intra prediction process of step S28112. FIG. 14 is a flowchart showing an example of the process of the RAHT unit 2080. FIG. 15 is a diagram showing an example of the inter prediction process of step S28111. FIG. 16 is a flowchart showing an example of the operation of the tree synthesis unit 2020 of the point cloud decoding device 200 according to an embodiment. FIG. 17 is a flowchart showing an example of the process of decoding predictor information and spherical coordinate residuals in step S1604. FIG. 18 is a diagram showing an example of functional blocks of the point cloud encoding device 100 according to an embodiment. FIG. 19 is a flowchart showing an example of the process of searching for adjacent nodes in an upper hierarchy of the node to be decoded. FIG. 20 is a flowchart showing an example of the process of searching for adjacent nodes in an upper hierarchy of the node to be decoded. FIG. 21 is a flowchart showing an example of the process of searching for adjacent nodes in an upper hierarchy of the node to be decoded. FIG. 22 is a flowchart showing an example of the process of step S28103. FIG. 23 is a flowchart showing an example of the intra prediction process of step S28112. FIG. 24 shows an example of the syntax configuration of the APS 2611 shown in FIG.

[0011] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that the components in the following embodiments can be appropriately replaced with existing components, etc., and various variations, including combinations with other existing components, are possible. Therefore, the description of the following embodiments does not limit the content of the invention described in the claims.

[0012] First Embodiment Hereinafter, a point cloud processing system 10 according to a first embodiment of the present invention will be described with reference to Figures 1 to 24. Figure 1 is a diagram showing a point cloud processing system 10 according to this embodiment.

[0013] As shown in FIG. 1 , the point cloud processing system 10 includes a point cloud encoding device 100 and a point cloud decoding device 200 .

[0014] The point cloud encoding device 100 is configured to generate encoded data (bitstream) by encoding an input point cloud signal, and the point cloud decoding device 200 is configured to generate an output point cloud signal by decoding the bitstream.

[0015] The input point cloud signal and the output point cloud signal are composed of position information and attribute information of each point in the point cloud, such as color information and reflectance of each point.

[0016] Here, the bit stream may be transmitted from the point group encoding device 100 to the point group decoding device 200 via a transmission path. Alternatively, the bit stream may be stored in a storage medium and then provided from the point group encoding device 100 to the point group decoding device 200.

[0017] (Point Group Decoding Device 200) The point group decoding device 200 according to this embodiment will be described below with reference to Fig. 2. Fig. 2 is a diagram showing an example of functional blocks of the point group decoding device 200 according to this embodiment.

[0018] As shown in Figure 2, the point cloud decoding device 200 has a geometric information decoding unit 2010, a tree synthesis unit 2020, an approximate surface synthesis unit 2030, a geometric information reconstruction unit 2040, an inverse coordinate transformation unit 2050, an attribute information decoding unit 2060, an inverse quantization unit 2070, an RAHT unit 2080, an LoD calculation unit 2090, an inverse lifting unit 2100, an inverse color transformation unit 2110, and a frame buffer 2120.

[0019] The geometric information decoding unit 2010 is configured to receive as input a bit stream relating to geometric information (geometric information bit stream) from among the bit streams output from the point group encoding device 100, and to decode the syntax.

[0020] The decoding process is, for example, a context-adaptive binary arithmetic decoding process. Here, for example, the syntax includes control data (flags and parameters) for controlling the decoding process of the position information.

[0021] The tree synthesis unit 2020 is configured to receive as input the control data decoded by the geometric information decoding unit 2010 and an occupancy code indicating at which node in the tree (described later) the point group exists, and to generate tree information indicating in which area in the decoding target space the point exists.

[0022] The decoding process of the occupancy code may be performed within the tree synthesis unit 2020.

[0023] This process divides the decoding target space into rectangular parallelepipeds, determines whether a point exists in each rectangular parallelepiped by referring to the occupancy code, divides the rectangular parallelepiped in which the point exists into multiple rectangular parallelepipeds, and refers to the occupancy code. By recursively repeating this process, tree information can be generated.

[0024] Here, when decoding such an occupancy code, inter prediction, which will be described later, may be used.

[0025] In this embodiment, a method called "Octree" can be used, which recursively performs octree division on the above-mentioned rectangular parallelepiped as a cube, and a method called "QtBt" can be used, which performs quadtree division and binary tree division in addition to octree division. Whether or not to use "QtBt" is transmitted from the point cloud encoding device 100 as control data.

[0026] Alternatively, when the use of predictive geometry coding is specified by the control data, the tree synthesis unit 2020 is configured to decode the coordinates of each point based on an arbitrary tree configuration determined by the point cloud encoding device 100.

[0027] The approximate surface synthesis unit 2030 is configured to generate approximate surface information using the tree information generated by the tree synthesis unit 2020, and to decode the point cloud based on the approximate surface information.

[0028] For example, when decoding three-dimensional point cloud data of an object, if the point cloud is densely distributed on the surface of the object, approximate surface information is used to represent the area where the point cloud exists by approximating it with a small plane, rather than decoding each individual point cloud.

[0029] Specifically, the approximate surface synthesis unit 2030 can generate approximate surface information and decode the point cloud using a technique called "Trisoup," for example. A specific example of the "Trisoup" process will be described later. Furthermore, when decoding a sparse point cloud acquired by Lidar or the like, this process can be omitted.

[0030] The geometric information reconstruction unit 2040 is configured to reconstruct the geometric information (position information in the coordinate system assumed by the decoding process) of each point of the point cloud data to be decoded based on the tree information generated by the tree synthesis unit 2020 and the approximate surface information generated by the approximate surface synthesis unit 2030.

[0031] The inverse coordinate transformation unit 2050 is configured to receive the geometric information reconstructed by the geometric information reconstruction unit 2040 as input, transform it from the coordinate system assumed by the decoding process to the coordinate system of the output point cloud signal, and output position information.

[0032] The frame buffer 2120 is configured to store, as a reference frame, the geometric information reconstructed by the geometric information reconstruction unit 2040 as an input. The stored reference frame is read from the frame buffer 2130 and used as the reference frame when the tree synthesis unit 2020 performs inter-prediction of a temporally different frame.

[0033] Here, which reference frame at which time point is to be used for each frame may be determined based on control data transmitted as a bit stream from the point cloud encoding device 100, for example.

[0034] The attribute information decoding unit 2060 is configured to receive as input a bit stream relating to attribute information (attribute information bit stream) from among the bit streams output from the point group encoding device 100, and to decode the syntax.

[0035] The decoding process is, for example, a context-adaptive binary arithmetic decoding process, where the syntax includes, for example, control data (flags and parameters) for controlling the decoding process of the attribute information.

[0036] Moreover, the attribute information decoding unit 2060 is configured to decode the quantized residual information from the decoded syntax.

[0037] The inverse quantization unit 2070 is configured to perform inverse quantization processing based on the quantized residual information decoded by the attribute information decoding unit 2060 and the quantization parameter, which is one of the control data decoded by the attribute information decoding unit 2060, to generate inverse quantized residual information.

[0038] The dequantized residual information is output to either the RAHT unit 2080 or the LoD calculation unit 2090, depending on the characteristics of the point group to be decoded. Which unit the information is output to is specified by control data decoded by the attribute information decoding unit 2060.

[0039] The RAHT unit 2080 is configured to receive as input the inverse-quantized residual information generated by the inverse quantization unit 2070 and the geometric information generated by the geometric information reconstruction unit 2040, and to decode the attribute information of each point using a type of Haar transform (inverse Haar transform in the decoding process) called RAHT (Region Adaptive Hierarchical Transform). The decoded information is the direct current component (DC coefficient) and alternating current component (AC coefficient) of the attribute information generated by using the RAHT in the encoding process, and is converted into attribute information by using the inverse transform of the RAHT in the decoding process. As a specific process of the RAHT, for example, the method described in Non-Patent Document 1 can be used.

[0040] The LoD calculation unit 2090 is configured to receive the geometric information generated by the geometric information reconstruction unit 2040 as an input and generate an LoD (Level of Detail).

[0041] LoD is information for defining a reference relationship (a point to be referenced and a point to be referenced) to realize predictive coding, such as predicting attribute information of another point from attribute information of another point and encoding or decoding the prediction residual.

[0042] In other words, LoD is information that defines a hierarchical structure in which each point contained in geometric information is classified into multiple levels, and the attributes of points belonging to lower levels are encoded or decoded using the attribute information of points belonging to higher levels.

[0043] As a specific method for determining the LoD, for example, the method described in Non-Patent Document 1 mentioned above may be used.

[0044] The inverse lifting unit 2100 is configured to decode attribute information of each point based on the hierarchical structure defined by the LoD, using the LoD generated by the LoD calculation unit 2090 and the inverse-quantized residual information generated by the inverse quantization unit 2070. As a specific process of inverse lifting, for example, the method described in the above-mentioned Non-Patent Document 1 can be used.

[0045] The inverse color conversion unit 2110 is configured to perform inverse color conversion processing on the attribute information output from the RAHT unit 2080 or the inverse lifting unit 2100 when the attribute information to be decoded is color information and color conversion has been performed on the point cloud encoding device 100 side. Whether or not to perform such inverse color conversion processing is determined by the control data decoded by the attribute information decoding unit 2060.

[0046] The point cloud decoding device 200 is configured to decode and output attribute information of each point in the point cloud through the above processing.

[0047] (Geometric Information Decoding Unit 2010) Hereinafter, the control data decoded by the geometric information decoding unit 2010 will be described with reference to FIGS.

[0048] FIG. 3 shows an example of the structure of coded data (bit stream) received by the geometric information decoding unit 2010.

[0049] First, the bitstream may include a GPS 2011. The GPS 2011 is also called a geometry parameter set and is a set of control data related to decoding of geometric information. A specific example will be described later. Each GPS 2011 includes at least GPS ID information for identifying each GPS 2011 when multiple GPS 2011 exist.

[0050] Second, the bitstream may include GSH2012A / 2012B. GSH2012A / 2012B is also called a geometry slice header or geometry data unit header, and is a collection of control data corresponding to a slice, which will be described later. Hereinafter, the term "slice" will be used, but "slice" can also be interpreted as "data unit." Specific examples will be described later. GSH2012A / 2012B includes at least GPS ID information for specifying the GPS2011 corresponding to each GSH2012A / 2012B.

[0051] Third, the bitstream may include slice data 2013A / 2013B following the GSH 2012A / 2012B. The slice data 2013A / 2013B includes data that encodes geometric information.

[0052] As described above, the bit stream is configured such that one GSH 2012A / 2012B and one GPS 2011 correspond to each slice data 2013A / 2013B.

[0053] As described above, the GPS id information is used to specify which GPS 2011 to refer to in the GSH 2012A / 2012B, so a common GPS 2011 can be used for multiple slice data 2013A / 2013B.

[0054] In other words, it is not necessary to transmit the GPS 2011 for each slice. For example, as shown in Fig. 3, the bit stream may be configured such that the GPS 2011 is not coded immediately before the GSH 2012B and slice data 2013B.

[0055] 3 is merely an example. As long as the GSH 2012A / 2012B and the GPS 2011 correspond to each slice data 2013A / 2013B, elements other than those described above may be added as components of the bit stream.

[0056] For example, as shown in Fig. 3, the bitstream may include a sequence parameter set (SPS) 2001. Similarly, when transmitted, the bitstream may be shaped into a configuration different from that shown in Fig. 3. Furthermore, the bitstream may be combined with a bitstream decoded by an attribute information decoding unit 2060 (described later) and transmitted as a single bitstream.

[0057] FIG. 4 shows an example of the syntax configuration of GPS2011.

[0058] Note that the syntax names explained below are merely examples. If the syntax functions explained below are similar, the syntax names may be different.

[0059] The GPS 2011 may include GPS id information (gps_geom_parameter_set_id) for identifying each GPS 2011.

[0060] 4 indicates how each syntax element is coded. ue(v) indicates an unsigned zeroth-order exponential-Golomb code, and u(1) indicates a 1-bit flag.

[0061] The GPS 2011 may include a flag (geom_tree_type) for controlling the tree type in the tree synthesis unit 2020 .

[0062] For example, when the value of geom_tree_type is "1", it may be defined that predictive geometry coding is used, and when the value of geom_tree_type is "0", it may be defined that Octree is used.

[0063] The GPS 2011 may include a flag (geom_angular_enabled) for controlling whether or not the tree synthesis unit 2020 performs processing in angular mode.

[0064] For example, when the value of geom_angular_enabled is "1", it may be defined that predictive geometry coding processing is performed as angular mode, and when the value of geom_angular_enabled is "0", it may be defined that predictive geometry coding processing is not performed as angular mode.

[0065] The GPS 2011 may include a flag (ptree_ang_azimuth_scaling_enabled) for controlling whether or not the adaptive azimuth angle quantization mode is in the angular mode in the tree synthesis unit 2020. The adaptive azimuth angle quantization mode is a mode in which adaptive azimuth angle quantization is performed according to the radius.

[0066] For example, when the value of ptree_ang_azimuth_scaling_enabled is "1", it may be defined that adaptive quantization of the azimuth angle according to the radius is performed, and when the value of ptree_ang_azimuth_scaling_enabled is "0", it may be defined that adaptive quantization of the azimuth angle according to the radius is not performed.

[0067] It may also be used as a flag to control whether or not to use the predictor list in predictor calculation (selection) in angular mode.

[0068] For example, if the value of ptree_azimuth_scaling_enabled is "1", it may be defined that a predictor list is used in the calculation of such a predictor, and if the value of ptree_ang_azimuth_scaling_enabled is "0", it may be defined that a predictor list is not used in the calculation of such a predictor.

[0069] The GPS 2011 may include a value (ptree_ang_azimuth_step_minus1) related to the laser rotation speed for use in calculating the predicted value of the azimuth angle in the tree synthesis unit 2020 in angular mode.

[0070] (Tree Merging Unit 2020) An example of the operation of the tree merging unit 2020 will now be described with reference to FIGS.

[0071] 16 is a flowchart showing an example of processing in the tree merging unit 2020. Note that the following describes an example in which trees are merged using "Predictive geometry coding."

[0072] Predictive geometry coding is also called predictive tree coding. Predictive geometry coding is a method for decoding the position information of the point cloud data by decoding the residual of position information predicted based on an arbitrary tree structure determined by the point cloud encoding device 100 and the position information of the point cloud data, and adding the two together.

[0073] As shown in FIG. 16, in step S1601, the tree synthesis unit 2020 determines whether decoding of position information of all point cloud data included in the slice has been completed.

[0074] This process can be performed, for example, by transmitting information indicating the number of point cloud data contained in the slice to the GSH, and comparing this number of point cloud data with the number of data already processed to determine whether processing of all points has been completed.

[0075] If the decoding of the position information of all point cloud data has been completed, the operation proceeds to step S1613, where the processing ends. If the decoding of the position information of all point cloud data has not been completed, the operation proceeds to step S1602.

[0076] In step S1602, the tree merging unit 2020 sets the parent node of the node to be decoded (node ​​to be processed) of the point cloud data.

[0077] For example, the tree synthesis unit 2020 decodes the number of child nodes of each node to be decoded, and stores the indexes of the nodes to be decoded for the number of child nodes.

[0078] When the tree synthesis unit 2020 processes a node to be decoded after a certain node, it may refer to the array of indexes of the node, obtain one index stored at the end of the array, and set the node of the obtained index as the parent node of the node to be decoded.

[0079] After the setting of the parent node is completed, the operation proceeds to step S1603.

[0080] In step S1603, the tree merging unit 2020 determines whether to perform processing in angular mode.

[0081] For example, the tree synthesis unit 2020 can refer to the value of the above-mentioned geom_angular_enabled to determine whether to perform processing in angular mode.

[0082] If processing is to be performed in angular mode, the operation proceeds to step S1604, and if processing is not to be performed in angular mode, the operation proceeds to step S1610.

[0083] In step S1604, the tree synthesis unit 2020 decodes the predictor information and spherical coordinate residuals to be used in step S1605. Here, the spherical coordinate residuals indicate the residuals of the radius, azimuth angle, and laser ID. Once this decoding is complete, the operation proceeds to step S1605.

[0084] In step S1605, the tree synthesis unit 2020 predicts the position information based on the predictor information decoded in step S1604. Here, the predictor information is a predictor index or a prediction mode.

[0085] In this process, the tree synthesis unit 2020 first determines the type of predictor to be used for prediction.

[0086] For example, the tree synthesis unit 2020 may determine whether or not to perform processing in adaptive azimuth angle quantization mode based on the value of ptree_ang_azimuth_scaling_enabled, and may determine the type of predictor to use based on the result of this determination.

[0087] For example, in the case of adaptive azimuth angle quantization mode, the tree synthesis unit 2020 may select a predictor to use from among multiple predictors calculated using a tree structure based on the decoded prediction mode.

[0088] Alternatively, when performing processing in adaptive azimuth angle quantization mode, the tree synthesis unit 2020 may store the position information of the decoded node as a predictor in a list, refer to the list for the predictor assigned to the decoded predictor index, and select the predictor type to be used.

[0089] Once the type of predictor is determined, the tree synthesis unit 2020 uses the determined predictor as the predicted value of the position information.

[0090] After the prediction of the position information is completed, the operation proceeds to step S1606.

[0091] In step S1606, the tree synthesis unit 2020 reconstructs the spherical coordinates by adding the decoded spherical coordinate residual and the predictor.

[0092] After the reconfiguration is completed, the operation proceeds to step S1607.

[0093] In step S1607, the tree synthesis unit 2020 reconstructs orthogonal integer coordinates. In this process, the tree synthesis unit 2020 can convert spherical coordinates into orthogonal integer coordinates based on the reconstructed spherical coordinates. A specific method for this can be realized, for example, by the technique described in Non-Patent Document 1.

[0094] After the reconstruction of the orthogonal integer coordinates is completed, the operation proceeds to step S1608.

[0095] In step S1608, the tree synthesis unit 2020 decodes the orthogonal integer coordinate residual.

[0096] After the decoding of the orthogonal integer coordinate residual is completed, the operation proceeds to step S1609.

[0097] In step S1609, the tree synthesis unit 2020 reconstructs the original coordinates by adding the decoded orthogonal integer coordinate residual and the reconstructed orthogonal integer coordinates.

[0098] After the reconstruction of the original coordinates is completed, the operation returns to step S1601.

[0099] In step S1610, the tree synthesis unit 2020 predicts the position information. Specifically, the tree synthesis unit 2020 selects a predictor and sets the selected predictor as the predicted value of the position information.

[0100] For example, the tree synthesis unit 2020 may select a predictor based on the decoded predictor mode from among a plurality of predictors calculated based on a tree structure.

[0101] After the prediction of the position information is completed, the operation proceeds to step S1611.

[0102] In step S1611, the tree synthesis unit 2020 decodes the orthogonal integer coordinate residual.

[0103] After the decoding of the orthogonal integer coordinate residual is completed, the operation proceeds to step S1612.

[0104] In step S1612, the tree synthesis unit 2020 reconstructs the original coordinates by adding the residual of the orthogonal integer coordinates decoded in step S1611 to the position information predicted in step S1610.

[0105] After the reconstruction of the original coordinates is completed, the operation returns to step S1601.

[0106] FIG. 17 is a flowchart showing an example of the process of decoding the predictor information and the spherical coordinate residual in step S1604.

[0107] As shown in FIG. 17, in step S1701, the tree synthesis unit 2020 determines whether or not the adaptive azimuth angle quantization mode is selected based on the value of ptree_ang_azimuth_scaling_enabled.

[0108] If the adaptive azimuth angle quantization mode is selected, the operation proceeds to step S1702. On the other hand, if the adaptive azimuth angle quantization mode is not selected, the operation proceeds to step S1703.

[0109] In step S1702, the tree synthesis unit 2020 decodes the predictor index. After the decoding of the predictor index is completed, the operation proceeds to step S1704.

[0110] In step S1703, the tree synthesis unit 2020 decodes the prediction mode. After the prediction mode has been decoded, the operation proceeds to step S1704.

[0111] In step S1704, the tree synthesis unit 2020 decodes the number of azimuth angle steps. After the decoding of the number of azimuth angle steps is completed, the operation proceeds to step S1705.

[0112] In step S1705, the tree synthesis unit 2020 decodes the spherical coordinate residual. The tree synthesis unit 2020 may perform this decoding using the method described in Non-Patent Document 2. After the decoding is complete, the operation proceeds to step S1706, where the process ends.

[0113] (Attribute Information Decoding Unit 2060) Hereinafter, the control data decoded by the attribute information decoding unit 2060 will be described with reference to FIGS. 5 to 6 and 24. FIG.

[0114] FIG. 5 shows an example of the structure of coded data (bit stream) received by the attribute information decoding unit 2060, and FIGS. 6 and 24 show examples of the syntax structure of the APS 2611 shown in FIG.

[0115] Note that the syntax names explained below are merely examples. If the syntax functions explained below are similar, the syntax names may be different.

[0116] The APS 2611 may include APS id information (aps_geom_parameter_set_id) for identifying each APS 2611.

[0117] Note that the Descriptor column in FIG. 6 indicates how each syntax is coded, where se(v) indicates a signed zeroth-order exponential-Golomb code, ue(v) indicates an unsigned zeroth-order exponential-Golomb code, and u(1) indicates a 1-bit flag.

[0118] The APS 2611 may include a flag (attr_coding_type) for controlling whether the inverse quantization unit 2070 outputs the inverse quantized residual information to the RAHT unit 2080 or the LoD calculation unit 2090 .

[0119] For example, when the value of attr_coding_type is “1”, it may be defined to be output to the LoD calculation unit 2090 , and when the value of attr_coding_type is “0”, it may be defined to be output to the RAHT unit 2080 .

[0120] The APS 2611 may include a flag (raht_prediction_enabled) for controlling whether or not the RAHT unit 2080 predicts attribute information.

[0121] For example, when the value of raht_prediction_enabled is "1", it may be defined that attribute information is predicted, and when the value of raht_prediction_enabled is "0", it may be defined that attribute information is not predicted.

[0122] The APS 2611 may include a flag (raht_subnode_prediction_enable_flag) that controls whether or not the RAHT unit 2080 uses subnodes to predict attribute information.

[0123] For example, if the value of raht_subnode_prediction_enable_flag is "1", it may be defined that subnodes are used to predict attribute information, and if the value of raht_subnode_prediction_enable_flag is "0", it may be defined that subnodes are not used to predict attribute information.

[0124] The APS 2611 may include weight parameters (raht_prediction_weights) used when the RAHT unit 2080 performs intra prediction of attribute information.

[0125] For example, the value of intra_prediction_weights may be defined according to the manner in which the node to be decoded is adjacent to the adjacent node used for intra prediction.

[0126] The APS 2611 may include a flag (raht_smoothing_enable_flag) that controls whether or not smoothing is performed after the RAHT unit 2080 performs intra prediction of attribute information.

[0127] For example, when the value of raht_smoothing_enable_flag is "1", it may be defined that smoothing is performed after predicting the attribute information, and when the value of raht_smoothing_enable_flag is "0", it may be defined that smoothing is not performed.

[0128] The APS 2611 may include a weighting parameter (raht_smoothing_weighted_average_weights) for performing smoothing by weighted averaging after the RAHT unit 2080 performs intra prediction of attribute information.

[0129] For example, up to eight such weight parameters may be defined depending on the manner in which the node to be decoded is adjacent to each subnode of the same parent node of the node to be decoded.

[0130] The APS 2611 may include weight parameters (raht_smoothing_clipping_weights) for performing smoothing by clipping after the RAHT unit 2080 performs intra prediction of attribute information.

[0131] For example, up to eight such weight parameters may be defined depending on the manner in which the node to be decoded is adjacent to each subnode of the same parent node of the node to be decoded.

[0132] The APS 2611 may include a threshold (raht_smoothing_clipping_threshold) for performing smoothing by clipping after the RAHT unit 2080 performs intra prediction of attribute information.

[0133] The APS 2611 may include a flag (raht_inter_prediction_enabled) for controlling whether or not the RAHT unit 2080 performs inter prediction of attribute information.

[0134] For example, when the value of raht_inter_prediction_enabled is "1", it may be defined that prediction of attribute information is performed, and when the value of raht_inter_prediction_enabled is "0", it may be defined that prediction of attribute information is not performed.

[0135] The APS 2611 may include a value (raht_inter_prediction_depth_minus1) indicating a layer for which inter prediction of attribute information is enabled in the RAHT unit 2080.

[0136] For example, if raht_inter_prediction_depth_minus1 is "N-1", inter prediction may be enabled in up to the top N layers of the Octree structure.

[0137] When predicting attribute information (for example, when the value of raht_prediction_enabled is "1"), the APS 2611 may additionally include the following syntax as shown in FIG.

[0138] The APS 2611 may include a syntax (raht_prediction_search_range) that indicates the maximum search range when searching for higher-level nodes in the RAHT.

[0139] The APS 2611 may include a flag (raht_last_comp_pred_enabled) indicating whether to predict the Cr signal from the Cb signal. For example, when the value of this flag is “1,” it may be defined that the Cr signal is predicted from the Cb signal, and when the value of this flag is “0,” it may be defined that the Cr signal is not predicted from the Cb signal.

[0140] When predicting the Cr signal from the Cb signal, the APS 2611 may additionally include syntax (raht_last_comp_pred_coeff_diff[dpth]) indicating a prediction coefficient value.

[0141] The attribute information decoding unit 2060 may decode the prediction coefficients as different values ​​for each layer of the RAHT. Furthermore, the attribute information decoding unit 2060 may decode the values ​​of the prediction coefficients as differential values ​​with respect to (already decoded) coefficient values ​​in a higher layer, rather than decoding the values ​​of the prediction coefficients as they are.

[0142] The APS2611 may include a flag (rht_inter_comp_pred_enabled) indicating whether or not to predict color difference signals (Cb and Cr signals) from the luminance signal. For example, it may be defined that when the value of this flag is “1,” the color difference signals (Cb and Cr signals) are predicted from the luminance signal, and when the value of this flag is “0,” the color difference signals (Cb and Cr signals) are not predicted from the luminance signal.

[0143] When predicting a color difference signal (Cb signal and Cr signal) from a luminance signal, the APS2611 may additionally include a syntax (rht_inter_comp_pred_coeff_diff[dpth]) indicating a prediction coefficient value.

[0144] The attribute information decoding unit 2060 may decode the prediction coefficients into different values ​​for each layer of the RAHT. Furthermore, the attribute information decoding unit 2060 may decode the prediction coefficients into different values ​​for each component to be predicted (Cb signal and Cr signal). Furthermore, the attribute information decoding unit 2060 may decode the values ​​of the prediction coefficients as differential values ​​with respect to (already decoded) coefficient values ​​in a higher layer, rather than decoding the values ​​directly.

[0145] When predicting a Cr signal from a Cb signal, or predicting a color difference signal (Cb signal and Cr signal) from a luminance signal, APS2611 may additionally include syntax (raht_coeff_calc_range) indicating a reference range when calculating prediction coefficients on the point cloud decoding device 200 side.

[0146] (RAHT Unit 2080) An example of the processing of the RAHT unit 2080 will be described with reference to FIGS.

[0147] FIG. 7 is a flowchart showing an example of processing by the RAHT unit 2080.

[0148] 7, in step S28001, the RAHT unit 2080 recursively divides the nodes into octrees until they reach a predetermined size, using a technique called Octree. After the division is complete, the operation proceeds to step S28002.

[0149] In step S28002, the RAHT unit 2080 counts the total number of points belonging to the layer below the node for each node divided by the Octree.

[0150] Specifically, the RAHT unit 2080 sequentially scans the nodes in a certain layer and records the number of points belonging to each node. Next, the RAHT unit 2080 adds up the numbers of points recorded in the child nodes of each node in the node one layer above to calculate the number of points belonging to each node.

[0151] The RAHT unit 2080 repeats the above scanning from the bottom layer to the top layer. The total number of acquired points is used as a weight for the inverse transformation of the RAHT in step S28005, which will be described later. After this calculation is completed, the operation proceeds to step S28003.

[0152] In step S28003, the RAHT unit 2080 decodes the DC coefficients of the nodes belonging to the highest layer of the Octree. Alternatively, the RAHT unit 2080 may calculate the DC coefficients by predicting the DC coefficients using intra prediction and decoding and adding up the prediction residuals of the DC coefficients.

[0153] After completing the decoding of the DC coefficient, the RAHT unit 2080 calculates the attribute value Aroot of the root node using the total number of points belonging to the root node acquired in step S28002, wroot, and the decoded DC coefficient DCroot, using the following formula.

[0154] After this calculation is completed, the operation proceeds to step S28004.

[0155] In step S28004, the RAHT unit 2080 determines whether or not the decoding of the attribute information of all nodes included in the layer has been completed.

[0156] If not completed, the operation proceeds to step S28005; if completed, the operation proceeds to step S28007.

[0157] In step S28005, the RAHT unit 2080 decodes the AC coefficients. Details will be described later. After the decoding is completed, the operation proceeds to step S28006.

[0158] In step S28006, the RAHT unit 2080 calculates attribute values ​​using the inverse transform of the RAHT based on the total number of points belonging to the lower layer of each aggregated node, the decoded AC coefficients, and the DC coefficients calculated from the nodes in the upper layer using the method described below.

[0159] Here, the inverse transformation of the RAHT is performed in units of 8 nodes (2×2×2) divided by the Octree.

[0160] Specifically, attribute value A 1 , A 2 , ...A k is the DC coefficient DC of a node that holds k subnodes, and the AC coefficient AC 1 , A.C. 2 , ...AC k-1 and the total number of points belonging to the lower layer of each subnode, w = w 1 , w 2 ,...w k is used to obtain the following equation (1).

[0161] Here, T(w) -1 is a matrix used for the inverse transformation of the RAHT, and can be generated by the method described in Non-Patent Document 1, for example.

[0162] It is assumed that this conversion process is performed repeatedly in the order from the higher-level nodes to the lower-level nodes,

[0163] is used as a DC coefficient in the inverse transform of the RAHT of each sub-node. After this transform process is completed, the operation proceeds to step S28004.

[0164] In step S28007, the RAHT unit 2080 determines whether the decoding of nodes in all layers has been completed.

[0165] If not, the operation moves to the next lower hierarchical level and proceeds to step S28004. If completed, the operation proceeds to step S28008 and ends the process.

[0166] FIG. 8 is a flowchart showing an example of the process in step S28004.

[0167] 8, in step S28101, the RAHT unit 2080 determines whether to predict AC coefficients. When making this determination, the RAHT unit 2080 may refer to raht_prediction_enabled and use the value thereof.

[0168] The RAHT unit 2080 may decode a flag indicating whether or not AC coefficient prediction is to be performed in the current node to be processed, and use the value of the flag.

[0169] The flag may be decoded for each node or for each layer. The flag may be decoded only if the value of raht_prediction_enabled is "1", which indicates that prediction is enabled. The flag may be included in slice data.

[0170] If the result of the determination is that AC coefficients are not to be predicted, the operation proceeds to step S28102, and if AC coefficients are to be predicted, the operation proceeds to steps S28103 and S28104.

[0171] In step S28102, the RAHT unit 2080 decodes the AC coefficients. After the decoding is completed, the operation proceeds to step S28106, where the process ends.

[0172] In step S28103, the RAHT unit 2080 decodes the AC coefficient residuals.

[0173] For example, the RAHT unit 2080 may perform inverse quantization processing on the quantized AC coefficient residuals decoded from the bitstream based on the quantization parameter decoded from the bitstream to calculate the inverse quantized AC coefficient residuals. In this case, the inverse quantized AC coefficient residuals correspond to the output in step S28103 (decoded AC coefficient residuals).

[0174] Alternatively, for example, the RAHT unit 2080 may output the quantized AC coefficient residuals decoded from the bitstream as they are in step S28103 (decoded AC coefficient residuals).

[0175] The RAHT unit 2080 may also decode the AC coefficient residuals using the method shown in FIG.

[0176] Fig. 22 is a flowchart showing an example of the process of step S28103. Note that Fig. 22 is a flowchart assuming a case where the attribute signal to be decoded has multiple components (for example, a luminance signal (Y signal) and color difference signals (Cb signal and Cr signal)).

[0177] As shown in FIG. 22, in step 2201, the RAHT unit 2080 determines whether or not the decoding of AC coefficient residuals of all components in the AC coefficient has been completed.

[0178] If the decoding of the AC coefficient residuals of all components is completed, the operation proceeds to step S2207, where the processing ends. On the other hand, if the decoding of the AC coefficient residuals of all components is not completed, the operation proceeds to step S2202.

[0179] In step S2202, the RAHT unit 2080 decodes the quantized AC coefficient residual of the component of the AC coefficient from the bitstream.

[0180] Note that the RAHT unit 2080 may decode the quantized AC coefficient residuals from the bitstream and store them in memory prior to the processing of step S2202, and then read out the AC coefficient residuals stored in memory in step S2202. After decoding the quantized AC coefficient residuals as described above, the operation proceeds to step S2203.

[0181] In step S2203, the RAHT unit 2080 performs inverse quantization processing on the quantized AC coefficient residuals decoded in step S2202 based on the quantization parameter decoded from the bitstream, and calculates the inverse quantized AC coefficient residuals. After the above processing is completed, the operation proceeds to step S2204.

[0182] In step S2204, the RAHT unit 2080 checks whether the component is a target of inter-component prediction.

[0183] The RAHT unit 2080 may determine whether or not a component is a target of inter-component prediction based on the value of a syntax element included in a header such as SPS, APS, or ASH.

[0184] For example, the RAHT unit 2080 may predict the color difference signal from the luminance signal. In this case, the target of inter-component prediction is the color difference signal (Cb signal and Cr signal).

[0185] Furthermore, the RAHT unit 2080 may perform inter-component prediction between chrominance signals. For example, the RAHT unit 2080 may predict a Cr signal from a Cb signal. In this case, the Cr signal is the target of inter-component prediction.

[0186] Note that the following description is based on the assumption that decoding is performed in the order of the Y signal, the Cb signal, and the Cr signal. When performing inter-component prediction, the RAHT unit 2080 can use only signals that have been decoded prior to the signal to be predicted.

[0187] If the component is not a target of inter-component prediction, the operation returns to step S2201 to process the next component. On the other hand, if the component is a target of inter-component prediction, the operation proceeds to step S2205.

[0188] In step S2205, the RAHT unit 2080 performs inter-component prediction of the AC coefficient residuals.

[0189] In the following, an example will be described in which inter-component prediction is performed from a luminance signal to a Cb signal, but similarly, inter-component prediction is also possible from a luminance signal to a Cr signal, and from a Cb signal to a Cr signal.

[0190] Here, the AC coefficient residual of the luminance signal after inverse quantization calculated in step S2203 is set to Y'. A predicted value Crp' of the AC coefficient residual of the Cb signal can be calculated, for example, as shown in the following equation.

[0191] Crp'=a×Y'+b where a and b are prediction coefficients.

[0192] The RAHT unit 2080 may decode prediction coefficients from headers such as SPS, APS, and ASH.

[0193] Alternatively, the RAHT unit 2080 may calculate prediction coefficients from other AC coefficient residuals that have already been decoded. Specifically, the RAHT unit 2080 may calculate and use AC coefficient residuals that minimize the sum of squared errors when predicting the AC coefficient residuals of the Cb signal after inverse quantization using the above-mentioned equation, for example, from the AC coefficient residuals of the luminance signal after inverse quantization of other AC coefficient residuals that have already been decoded.

[0194] The above process can be analytically calculated using the least squares method. At this time, the RAHT unit 2080 may use all decoded AC coefficient residuals to calculate the prediction coefficients.

[0195] Furthermore, the RAHT unit 2080 may use only the most recently decoded N AC coefficient residuals based on the AC coefficient residuals to calculate prediction coefficients. Here, the RAHT unit 2080 may determine the number of N before decoding, or may decode the value of N from a header such as SPS, APS, or ASH. For example, the RAHT unit 2080 may use the value of raht_coeff_calc_range as N.

[0196] Furthermore, the RAHT unit 2080 may use a weighted least squares method instead of a simple least squares method. Specifically, the RAHT unit 2080 may perform the weighting described above based on the absolute value of the difference between the inverse-quantized AC coefficient residual value Y′ of the Y signal of the AC coefficient residual and the inverse-quantized AC coefficient residual of the Y signal of the AC coefficient residual that has already been decoded, or the square value of such difference. In this case, the RAHT unit 2080 may define the weight so that the smaller the difference, the heavier the weight (the larger the absolute value of the weighting coefficient).

[0197] After the predicted value Crp' of the AC coefficient residual is calculated as described above, the operation proceeds to the next step S2206.

[0198] In step S2206, the RAHT unit 2080 updates the AC coefficient residuals. Specifically, the RAHT unit 2080 uses the predicted value Crp′ to update the AC coefficient residual Cb′ after inverse quantization of the Cb signal decoded in step S2203, as follows:

[0199] Cb'=Cb'+Crp' After this update is complete, the operation proceeds to step S2201 to process the next component.

[0200] As described above, after such decoding is complete, operation proceeds to step S28105.

[0201] In step S28104, the RAHT unit 2080 predicts AC coefficients. For predicting the AC coefficients, inter prediction or intra prediction may be used.

[0202] The RAHT unit 2080 may first predict the attribute values, and then calculate the predicted values ​​of the AC coefficients by RAHT. This will be described in detail later. After the prediction of the AC coefficients is completed, the operation proceeds to step S28105.

[0203] In step S28105, the RAHT unit 2080 adds the residuals of the decoded AC coefficients to the predicted AC coefficients to reconstruct the AC coefficients. After the reconstruction is completed, the operation proceeds to step S28106, where the process ends.

[0204] FIG. 9 is a flowchart showing an example of the process in step S28104.

[0205] As shown in Fig. 9 , in step S28107, the RAHT unit 2080 determines whether inter prediction is enabled. For this determination, the RAHT unit 2080 may refer to raht_inter_prediction_enabled and use the value thereof. If the determination result shows that inter prediction is enabled, this operation proceeds to step S28109, and if inter prediction is disabled, this operation proceeds to step S28112.

[0206] In step S28109, the RAHT unit 2080 determines whether the depth of the layer including the node to be processed is equal to or less than a threshold. The RAHT unit 2080 may refer to raht_inter_prediction_depth_minus1 and use the value as the threshold.

[0207] If the result of the determination is that the depth is equal to or less than the threshold, the operation proceeds to step S28110, and if the depth is greater than the threshold, the operation proceeds to step S28112.

[0208] In step S28110, the RAHT unit 2080 determines whether or not to perform inter prediction on the AC coefficients of the node to be processed.

[0209] For this determination, the RAHT unit 2080 may check whether inter prediction is possible, and perform inter prediction if possible, or may not perform inter prediction if possible. Specific details will be described later.

[0210] The RAHT unit 2080 may decode a flag indicating whether or not to perform inter prediction on the AC coefficients of the node to be processed, and use the value of the flag to make the determination. The flag may be decoded for each node or for each layer. The flag may be decoded and the determination may be made only when it is determined that inter prediction is feasible. The flag may be included in the slice data.

[0211] In step S28111, the RAHT unit 2080 performs inter prediction of the AC coefficients of the node to be processed. Specific details will be described later.

[0212] In step S28112, the RAHT unit 2080 performs intra prediction of the AC coefficients of the node to be processed. Specific details will be described later.

[0213] In step S28113, the process of step S28104 ends. Note that the conditional branch in step S28109 may be omitted.

[0214] In the inter prediction process of step S28111, a process equivalent to the intra prediction process of step S28112 may also be performed, and prediction may be performed by combining the results of the inter prediction and the intra prediction. Specific details will be described later.

[0215] FIG. 10 is a flowchart showing an example of the intra prediction process in step S28112.

[0216] 10 , in step S28201, the RAHT unit 2080 determines whether to perform intra prediction using adjacent nodes in the subnode hierarchy. The RAHT unit 2080 may refer to raht_subnode_prediction_enable_flag and use the value thereof for the determination.

[0217] If the RAHT unit 2080 does not use adjacent nodes in the subnode hierarchy, it performs intra prediction using only adjacent nodes in the upper hierarchy.

[0218] Here, the adjacent nodes in the higher hierarchy are the six nodes adjacent to the parent node of the node to be decoded on the face side, the 12 nodes adjacent to the edge side and the parent node itself, a total of 19 nodes, three nodes adjacent to the node to be decoded on the face side, three nodes adjacent to the edge side, and seven nodes of the parent node itself.

[0219] FIG. 11 is a diagram showing the relationship between a decoding target node and adjacent nodes in a higher layer.

[0220] When using adjacent nodes in the subnode hierarchy, the RAHT unit 2080 performs intra prediction using adjacent nodes in the upper hierarchy and adjacent nodes in the subnode hierarchy.

[0221] Here, an adjacent node in the subnode hierarchy is a subnode of an adjacent node in a higher hierarchy, which has a face or edge adjacent to the decoding target node and has already been decoded.

[0222] FIG. 12 is a diagram showing the relationship between a node to be decoded and adjacent nodes in the subnode hierarchy.

[0223] If the determination result is that intra prediction is to be performed without using adjacent nodes in the subnode hierarchy, this operation proceeds to step S28202, and if intra prediction is to be performed using adjacent nodes in the subnode hierarchy, this operation proceeds to step S28204.

[0224] In step S28202, the RAHT unit 2080 acquires the attribute value of the adjacent node in the upper layer.

[0225] The attribute value acquisition process can be realized in two stages: a search process for adjacent nodes, and a process for acquiring attribute values ​​from the adjacent nodes identified by the search process. An example of the search process for adjacent nodes will be described below with reference to FIGS.

[0226] FIG. 19 is a flowchart showing an example of a process for searching for an adjacent node in an upper layer of a decoding target node.

[0227] In FIG. 19, similarly to FIG. 11, a node in a layer above the node to be decoded is specially called a parent node.

[0228] In the following, in Figures 19 to 21, it is assumed that the Morton codes corresponding to the positions of each node in the upper layer are stored in ascending order of the Morton codes in a one-dimensional array (hereinafter referred to as the upper layer node array).

[0229] Furthermore, it is assumed that the index (hereinafter referred to as the parent node index) at which the Morton code of the parent node is stored in the one-dimensional array is known in advance.

[0230] As shown in FIG. 19, in step S1901, the RAHT unit 2080 checks whether or not the search for all adjacent nodes in the upper layer has been completed.

[0231] If the search for all adjacent nodes has been completed, the operation proceeds to step S1905 and ends the process. On the other hand, if the search for all adjacent nodes has not been completed, the operation proceeds to step S1902.

[0232] In step S1902, the RAHT unit 2080 calculates the Morton code corresponding to the position of the adjacent node to be searched.

[0233] Specifically, the RAHT unit 2080 may, for example, first convert the Morton code corresponding to the position of the parent node into a value of Cartesian coordinates (x, y, z), second calculate the coordinates (x', y', z') of the adjacent node in the Cartesian coordinate space, and third convert the Cartesian coordinates of the adjacent node back into Morton code, thereby calculating the Morton code corresponding to the position of the adjacent node to be searched for.

[0234] After the calculation of the Morton code is completed, the operation proceeds to step S1903.

[0235] In step S1903, the RAHT unit 2080 sets a search range, which is the range of indices to be searched in the upper layer node array, with the parent node index as the reference.

[0236] If the Morton code of the adjacent node is smaller than the Morton code of the parent node, the start and end points of the search range can be set as follows:

[0237] Search range = min (Morton code of parent node - Morton code of adjacent node, maximum search range) Start point of search range = max (parent node index - search range, 0) End point of search range = parent node index - 1 Here, min(a, b) is a function that returns the smaller of the two arguments a and b. Also, max(a, b) is a function that returns the larger of the two arguments a and b. Also, the maximum search range is the maximum value of the search range determined in advance.

[0238] The value of the maximum search range may be decoded from a header such as SPS, APS, ASH, etc. For example, the RAHT unit 2080 may use the value of raht_prediction_search_range as the value of such maximum search range.

[0239] On the other hand, if the Morton code of the adjacent node is greater than the Morton code of the parent node, the start point and end point of the search range can be set as follows:

[0240] Search range = min (Morton code of adjacent node - Morton code of parent node, maximum search range) Start point of search range = parent node index + 1 End point of search range = min (Morton code of parent node + search range, number of nodes in upper hierarchy - 1) After determining the start point and end point of the search as described above, this operation proceeds to step S1906.

[0241] In step S1906, the RAHT unit 2080 determines whether or not the adjacent node to be searched for is not stored or is likely to be stored between the start point and end point of the search set in step S1903.

[0242] If it is determined that the node is not stored, the RAHT unit 2080 determines that there is no adjacent node to be searched for, and proceeds to step S1901 to search for the next adjacent node.

[0243] On the other hand, if the RAHT unit 2080 determines that there is a possibility that the file is stored, the process proceeds to the next step S1904, where a search is performed.

[0244] Here, such a determination can be made, for example, as follows.

[0245] If the Morton code of the adjacent node is smaller than the Morton code of the parent node, the RAHT unit 2080 checks the Morton code of the node stored in the index corresponding to the start point of the search range.

[0246] If the Morton code of the node stored in the index corresponding to the start point of the search range is greater than the Morton code of the adjacent node, the RAHT unit 2080 determines that the adjacent node to be searched is not stored between the start point and the end point of the search.

[0247] If this is not the case (if the Morton code of the node stored in the index corresponding to the start point of the search range is equal to or less than the Morton code of the adjacent node), the RAHT unit 2080 determines that the adjacent node to be searched for may be stored between the start point and the end point of the search.

[0248] Here, if the index corresponding to the start point of the search range is 0, the RAHT unit 2080 may determine that there is a possibility that the adjacent node to be searched for is stored between the start point and the end point of the search.

[0249] Furthermore, if the "parent node index - index corresponding to the starting point" is equal to or less than the maximum search range, the RAHT unit 2080 may determine that the adjacent node to be searched for may be stored between the starting point and the end point of the search.

[0250] On the other hand, if the Morton code of the adjacent node is greater than the Morton code of the parent node, the RAHT unit 2080 checks the Morton code of the node stored in the index corresponding to the end point of the search range.

[0251] If the Morton code of the node stored in the index corresponding to the end point of the search range is smaller than the Morton code of the adjacent node, the RAHT unit 2080 determines that the adjacent node to be searched is not stored between the start point and the end point of the search.

[0252] If this is not the case (if the Morton code of the node stored in the index corresponding to the end point of the search range is equal to or greater than the Morton code of the adjacent node), the RAHT unit 2080 determines that the adjacent node to be searched for may be stored between the start point and the end point of the search.

[0253] Here, the RAHT unit 2080 may determine that if the index corresponding to the end point of the search range is "the number of nodes in the upper hierarchy - 1", there is a possibility that the adjacent node to be searched is stored between the start point and the end point of the search.

[0254] In addition, if "index corresponding to the end point - parent node index" is less than the maximum search range, the RAHT unit 2080 may determine that there is a possibility that the adjacent node to be searched for is stored between the start point and the end point of the search.

[0255] In other words, the above process is a process of determining whether or not there is an adjacent node to be searched for within the search range, based on the Morton code of the node stored at the start or end point of the search range.

[0256] In this way, prior to the search process, it is determined whether or not there is a possibility that the node to be searched for exists within the search range, and if it does not exist, the search process is omitted, thereby reducing unnecessary processing. If implemented in software, this can reduce execution time, and if implemented in hardware, this can reduce power consumption.

[0257] In step S1904, the RAHT unit 2080 searches for an element that stores the same Morton code as the adjacent node, from within the range of the upper layer node array specified by the start and end points of the search described above.

[0258] Here, if the RAHT unit 2080 finds an element that stores the same Morton code as an adjacent node, it returns the index of that element.

[0259] On the other hand, if the RAHT unit 2080 does not find an element that stores the same Morton code as an adjacent node, it returns a value (for example, -1) indicating that the adjacent node could not be found. In this case, because the Morton codes are stored in ascending order in the upper layer node array, the RAHT unit 2080 can perform a search with fewer searches than a full search by using a binary search or the like.

[0260] After the above process is completed, the operation proceeds to step S1901 to search for the next adjacent node.

[0261] 20 is a flowchart showing an example of a process for searching for an adjacent node in an upper layer of a decoding target node. An example of the process for searching for an adjacent node in an upper layer will be described below with reference to FIG.

[0262] As shown in FIG. 20, in step 2001, the RAHT unit 2080 calculates the Morton codes of all adjacent nodes.

[0263] Here, the method for calculating the Morton code is the same as the method described in step S1902. The difference from step S1902 is that the Morton codes of all adjacent nodes (e.g., 19 nodes) are calculated. After this calculation process is completed, the operation proceeds to step S2002.

[0264] In step S2002, the RAHT unit 2080 sorts the adjacent nodes, for example, in ascending order based on the Morton code calculated in step S2001. Thereafter, the RAHT unit 2080 searches for adjacent nodes in ascending order of Morton code value.

[0265] In step S2003, the RAHT unit 2080 determines whether or not the search for all adjacent nodes having a Morton code smaller than that of the parent node has been completed.

[0266] If the search is completed, the operation proceeds to step S2006. On the other hand, if the search is not completed, the operation proceeds to step S2004.

[0267] In step S2004, the RAHT unit 2080 sets a search range. For example, the RAHT unit 2080 sets the search range in the following manner.

[0268] Search range = min (min (parent node index - index of other adjacent node most recently discovered, Morton code of parent node - Morton code of adjacent node), maximum search range) Start point of search range = max (parent node index - search range, 0) End point of search range = parent node index - 1 Here, the difference from step S1903 is that the search range is set using the index of other adjacent node most recently discovered. Because adjacent nodes are searched for in ascending order of Morton code, it is guaranteed that the Morton code of the adjacent node currently being searched is a larger value than the Morton code of other adjacent node most recently discovered. Similarly, it is guaranteed that the index of the element in the hierarchical node array in which the Morton code of the adjacent node currently being searched is stored is larger than the index of the element in which the Morton code of other adjacent node discovered is stored. Therefore, by using the index of the other adjacent node that was discovered immediately before, the search range can be reduced, and the number of search processes and the time required for the search can be reduced.

[0269] The RAHT unit 2080 may initialize the index value of the adjacent node discovered just before to 0, and thereafter update the index value each time an adjacent node is discovered. After the above-mentioned processing is completed, the operation proceeds to step S2005.

[0270] In step S2005, the RAHT unit 2080 performs the same search process as in step S1904.

[0271] In step S2006, the RAHT unit 2080 changes the search order. Specifically, the RAHT unit 2080 changes the search order from ascending order of the Morton codes of adjacent nodes in the previous processing to descending order of the Morton codes. After changing the search order, the operation proceeds to step S2007.

[0272] In step S2007, the RAHT unit 2080 determines whether or not the search for all adjacent nodes having a Morton code greater than that of the parent node has been completed.

[0273] If all such searches have been completed, the operation proceeds to step S2009, where the process ends. If all such searches have not been completed, the operation proceeds to step S2008.

[0274] In step S2008, the RAHT unit 2080 sets the search range as follows: Here, the RAHT unit 2080 uses the index of the other adjacent node discovered immediately before, as in step S2004.

[0275] Search range = min (min (index of other adjacent node discovered most recently - parent node index, Morton code of adjacent node - Morton code of parent node), maximum search range) Start point of search range = parent node index + 1 End point of search range = min (Morton code of parent node + search range, number of nodes in the upper hierarchy) Here, the Morton codes of adjacent nodes are searched in descending order, so it is guaranteed that the Morton code currently being searched is smaller than the Morton code of other adjacent nodes discovered most recently.

[0276] The index of the other adjacent node discovered immediately before may be initialized to the number of nodes in the upper layer (= the number of elements in the upper layer node array) at the timing of step S2006, and thereafter may be updated with this index value each time an adjacent node is discovered.

[0277] After the search range is set as described above, the operation proceeds to step S2009.

[0278] In step S2009, the RAHT unit 2080 performs the same search process as in step S2005.

[0279] Fig. 21 is a flowchart showing an example of a process for searching for an adjacent node in an upper layer of a decoding target node. An example of the process for searching for an adjacent node in an upper layer will be described below with reference to Fig. 21. Note that the same processes as those in Fig. 20 are denoted by the same reference numerals as in Fig. 20, and description thereof will be omitted.

[0280] As shown in FIG. 21, in step S2101, the RAHT unit 2080 restores the search results recorded in step S2102, which will be described later.

[0281] Specifically, the RAHT unit 2080 refers to information stored in an array or the like to identify the index value of the adjacent node.

[0282] In step S2102, if an adjacent node having a larger Morton code than the parent node is found, the RAHT unit 2080 records the index of the parent node as an adjacent node when the adjacent node becomes the parent node.

[0283] When viewed from the adjacent node, the parent node is also the adjacent node (even if the adjacent node becomes the parent node). In other words, there is a symmetrical relationship.

[0284] In this embodiment, the RAHT unit 2080 processes parent nodes in ascending order of Morton code.

[0285] Therefore, the RAHT unit 2080 can reduce the search process by storing the index of the parent node as an adjacent node for the adjacent node of the parent node that has a larger Morton code than the parent node (in case the adjacent node becomes the parent node).

[0286] That is, as explained in step S2101, adjacent nodes with a smaller Morton code than the parent node have already been searched when the adjacent node was the parent node, so the RAHT unit 2080 can avoid performing the search process again by storing the results.

[0287] This allows the search process to be reduced to about half compared to when such storage is not performed.

[0288] After the attribute values ​​of the adjacent nodes in the upper layer are acquired as described above, the operation proceeds to step S28203.

[0289] In step S28203, the RAHT unit 2080 predicts the attribute value of the node to be decoded.

[0290] The RAHT unit 2080 calculates the attribute values ​​attr of the k adjacent nodes in the higher hierarchy. i and the weight w according to the type of adjacent node i i The attribute value attr may be predicted using the following formula:

[0291] Here, the RAHT unit 2080 uses the weight w i Depending on whether the adjacent node i is a face adjacent node in a higher layer, an edge adjacent node in a higher layer, or a parent node, a hard-coded value may be used, or the weight w may be calculated from the value by referring to the weight_prediction_weights. i may be calculated.

[0292] After the prediction of the attribute value is completed, the operation proceeds to step S28207.

[0293] In step S28204, the RAHT unit 2080 acquires the attribute value of the adjacent node in the upper layer.

[0294] Here, the targets for acquiring attribute values ​​are adjacent nodes in a higher hierarchy whose subnodes have not yet been decoded, or adjacent nodes in a higher hierarchy whose subnodes have been decoded but whose faces or edges are not adjacent to the node to be decoded.

[0295] After the attribute value has been acquired, the operation proceeds to step S28205.

[0296] In step S28205, the RAHT unit 2080 acquires the attribute value of the adjacent node in the subnode hierarchy. After acquiring the attribute value of the adjacent node in the subnode hierarchy, the operation proceeds to step S28206.

[0297] In step S28206, the RAHT unit 2080 predicts the attribute value of the node to be decoded.

[0298] The RAHT unit 2080 calculates the attribute values ​​attr of the acquired k adjacent nodes in the upper layer and the adjacent nodes in the subnode layer. i and the weight w according to the type i of the adjacent node i The attribute value attr may be predicted using the following formula:

[0299] Here, the RAHT unit 2080 uses the weight w i Depending on whether the adjacent node i is a face adjacent node in a higher layer, an edge adjacent node in a higher layer, a parent node, a face adjacent node in a subnode layer, or an edge adjacent node in a subnode layer, a hard-coded value may be used, or the weight w may be calculated by referring to the weights in the root_prediction_weights. i may be calculated.

[0300] After the attribute value prediction is completed, the operation proceeds to step S28207.

[0301] In step S28207, the RAHT unit 2080 converts the predicted attribute values ​​into AC coefficients. The AC coefficients are generated by performing RAHT on the predicted attribute values. For example, the RAHT unit 2080 may use the method described in Non-Patent Document 1 as the conversion method.

[0302] In step S28207, instead of converting the predicted attribute values ​​into AC coefficients, the RAHT unit 2080 may convert the residuals of the decoded AC coefficients into residuals of the attribute values, and add the predicted values ​​and residuals in the attribute value domain.

[0303] Fig. 23 is a flowchart showing an example of the intra prediction process of step S28112. An example of the intra prediction process will be described below with reference to Fig. 23. Note that the same processes as those in Fig. 10 are denoted by the same reference numerals as in Fig. 10, and description thereof will be omitted.

[0304] In addition, in FIG. 23, the attribute information of the decoded point group data has a plurality of components, and the processing in FIG. 23 will be described as processing for one of those components.

[0305] By applying the process of FIG. 23 to each component for each node to be decoded, it is possible to decode signals of all components included in the node to be decoded.

[0306] In step S2201, the RAHT unit 2080 determines whether or not the component is a target of inter-component prediction.

[0307] Here, the RAHT unit 2080 may determine whether or not the component is a target of inter-component prediction based on the value of a syntax element included in a header such as SPS, APS, or ASH.

[0308] For example, the RAHT unit 2080 may predict the chrominance signal from the luminance signal when the value of raht_inter_comp_pred_enabled is 1. In this case, the chrominance signals (Cb and Cr signals) are the targets of inter-component prediction.

[0309] Furthermore, the RAHT unit 2080 may perform inter-component prediction between chrominance signals. For example, when the value of raht_last_comp_pred_enabled is "1," the RAHT unit 2080 may predict the Cr signal from the Cb signal. In this case, the Cr signal is the target of inter-component prediction.

[0310] Note that the following description is based on the assumption that the Y signal, Cb signal, and Cr signal are decoded in this order. When performing inter-component prediction, only signals that have been decoded prior to the signal to be predicted can be used.

[0311] If the component is a target of inter-component prediction, the operation proceeds to step S2302; if not, the operation proceeds to step S28201.

[0312] In step S2302, the RAHT unit 2080 performs inter-component prediction.

[0313] In the following, an example of predicting the Cb signal from the luminance signal will be described, but it is also possible to predict the Cr signal from the luminance signal and the Cr signal from the Cb signal. Here, the AC coefficient of the luminance signal calculated in step S28207 is set to Y.

[0314] The predicted value Crp of the AC coefficient of the Cb signal can be calculated, for example, by the following equation.

[0315] Crp=a×Y+b where a and b are prediction coefficients. The RAHT unit 2080 may decode the prediction coefficients from a header such as SPS, APS, or ASH.

[0316] For example, the RAHT unit 2080 may decode such AC coefficients from the values ​​of raht_last_comp_pred_coeff_diff[dpth] and raht_inter_comp_pred_coeff_diff[dpth].

[0317] Alternatively, the RAHT unit 2080 may calculate such prediction coefficients from other AC coefficient values ​​that have already been decoded.

[0318] Specifically, the RAHT unit 2080 may calculate and use AC coefficients that minimize the sum of squared errors when predicting the AC coefficients of the Cb signal using the above-mentioned formula, for example, from the AC coefficient values ​​of the luminance signal of other AC coefficients that have already been decoded.

[0319] The above process can be analytically calculated using the least squares method. At this time, the RAHT unit 2080 may use all of the decoded AC coefficients to calculate the AC coefficients.

[0320] Alternatively, the RAHT unit 2080 may use only the most recently decoded N AC coefficients based on the AC coefficient in calculating the AC coefficient. The RAHT unit 2080 may determine the number N before decoding, or may decode the value of N from a header such as SPS, APS, or ASH.

[0321] For example, the RAHT unit 2080 may use the value of raht_coeff_calc_range as N.

[0322] Furthermore, the RAHT unit 2080 may calculate the AC coefficients using the weighted least squares method instead of the simple least squares method.

[0323] Specifically, the RAHT unit 2080 may perform weighting based on the absolute value of the difference between the value Y of the Y signal of the AC coefficient and the Y signal of an AC coefficient that has already been decoded, or the square value of the difference. In this case, the weight may be defined so that the smaller the difference, the heavier the weight.

[0324] Although the above description has been given taking as an example a case where the RAHT unit 2080 performs the above prediction in the AC coefficient domain, the RAHT unit 2080 may perform the above prediction in the attribute signal domain.

[0325] That is, the RAHT unit 2080 may predict the attribute signal value of the Cb signal from the attribute signal value of the Y signal.

[0326] The above describes an example in which the RAHT unit 2080 uses the attribute values ​​predicted in step S28206 as they are to convert the AC coefficients in step S28207, but the RAHT unit 2080 may also smooth the predicted attribute values ​​before converting the AC coefficients.

[0327] For example, as shown in FIG. 13, the RAHT unit 2080 may predict the attribute value, and then determine whether or not to perform smoothing in step S1301.

[0328] In making such a determination, the RAHT unit 2080 may refer to the raht_smoothing_enable_flag and use the value thereof.

[0329] If smoothing is to be performed, the operation proceeds to step S1302. If smoothing is not to be performed, the operation proceeds to step S28207.

[0330] In step S1302, the RAHT unit 2080 may smooth the attribute values.

[0331] For example, the RAHT unit 2080 calculates the attribute value Attr predicted at the subnode i in the same parent node as the node to be decoded for the attribute value Attr smoothing after smoothing of the node to be decoded. i and weight αi Alternatively, the weighted average may be calculated as follows:

[0332] Here, the RAHT unit 2080 may select, for the target subnode i, the node to be decoded as a node adjacent to the target node, or may select all subnodes within the same parent node.

[0333] In addition, the RAHT unit 2080 uses a weight α i A hard-coded value may be used as the weight, or the value may be referenced and used as the weight.

[0334] Furthermore, the RAHT unit 2080 calculates, for example, the smoothed attribute value Attr of the node to be decoded. smoothing For the node to be decoded, the predicted value Attr 0 , the attribute value Attr predicted at a subnode i other than the node to be decoded among the subnodes in the same parent node as the node to be decoded i , weight β i and threshold Th r It may be obtained by clipping using the following:

[0335] Here, clipping is a process in which if the input value is greater than a predetermined maximum value, the maximum value is output, if the input value is less than a predetermined minimum value, the minimum value is output, and in all other cases the input value is used as the output value as is.

[0336] The clipping function Clip3 is

[0337] is.

[0338] Here, for the target subnode i, the RAHT unit 2080 may consider the node to be decoded as a face-adjacent node, a face-adjacent node and an edge-adjacent node, or all subnodes within the same parent node.

[0339] In addition, the RAHT unit 2080 uses the weight β iA hard-coded value may be used as the weight, or the value may be referenced and used as the weight.

[0340] In addition, the RAHT unit 2080 sets a threshold value Th r A hard-coded value may be used as the threshold, or the value may be referenced and used.

[0341] In the above, an example has been described in which the RAHT unit 2080 decodes AC coefficients of both the color difference signal and the luminance signal, but the RAHT unit 2080 may skip decoding AC coefficients of the color difference signal only in the lowest layer of the Octree.

[0342] For example, as shown in FIG. 14, the RAHT unit 2080 may determine in step S1401 whether to skip decoding of AC coefficients of color difference signals only in the bottom layer of the Octree.

[0343] If skipping is to be performed, the operation proceeds to step S1402. If not skipping is to be performed, the operation proceeds to step S28004.

[0344] In step S1402, the RAHT unit 2080 determines whether the node to be decoded is in the lowest layer of the Octree.

[0345] If it is the bottom layer, the operation proceeds to step S1403. If it is not the bottom layer, the operation proceeds to step S28004.

[0346] In step S1403, the RAHT unit 2080 decodes AC coefficients other than the color difference signals.

[0347] The RAHT unit 2080 performs the same process as in step S28004 for decoding AC coefficients other than the color difference signals, sets the AC coefficients of the color difference signals to 0, and calculates the attribute values ​​in the subsequent step S28005.

[0348] After the decoding of AC coefficients other than the color difference signal is completed, the operation proceeds to step S28006.

[0349] FIG. 15 is a diagram showing an example of the inter prediction process in step S28111.

[0350] The RAHT unit 2080 predicts the AC coefficients of the target node using information about a reference node, which is the corresponding node in a reference frame. Here, the information about the reference node may be its attribute value or AC coefficient. The reference frame may also refer to another decoded frame, and the information about the reference frame may be included in the previous frame buffer 2120.

[0351] The RAHT unit 2080 may apply the same Octree structure as the frame to be processed to the reference frame. In such a case, a node may be set at a position where there is no point. Such a node is called an empty node. If the reference node is an empty node, the RAHT unit 2080 may disable inter prediction in step S28110.

[0352] The RAHT unit 2080 may apply an Octree to the reference frame independently of the current frame and set an Octree structure different from that of the current frame. In such a case, a node may not necessarily exist at the same position as in the current frame. If a reference node is not found at a position corresponding to the current node, the RAHT unit 2080 may disable inter prediction in step S28143.

[0353] If the reference node is an empty node, or if the reference node cannot be found, the RAHT unit 2080 may estimate and interpolate the information of the reference node using information of nodes in nearby positions within the reference frame.

[0354] For example, the RAHT unit 2080 may estimate and interpolate the average value of the attribute values ​​or AC coefficients of the adjacent nodes, the nearest nodes, or the k nearest nodes relative to the reference node position as the attribute value or AC coefficient of the reference node, respectively.

[0355] The RAHT unit 2080 may predict the AC coefficients of the node to be processed from, for example, the attribute values ​​of the reference node.

[0356] Specifically, the RAHT unit 2080 calculates the value Attr of the decoded attribute value of the reference node.inter The predicted value Attr of the attribute value of the node to be processed is calculated using pred and the predicted value Attr pred By applying RAHT to the AC coefficients of the node to be processed, the predicted value AC pred It may also be possible to ask for

[0357] Attr pred = Attr inter AC pred =RAHT(Attr pred The RAHT unit 2080 may predict the AC coefficients of the node to be processed directly from the AC coefficients of the reference node, for example.

[0358] Specifically, the RAHT unit 2080 calculates the AC coefficient values ​​AC of the reference node using the RAHT in the reference frame. inter is calculated, and the calculated value is used as the predicted value AC of the AC coefficient of the node to be processed. pred It may also be possible to use the following.

[0359] AC pred =AC inter The RAHT unit 2080 may obtain the AC coefficients of the reference node by recording the AC coefficients of each node of the reference frame in the frame buffer 2120 and referring to the values ​​in the frame buffer 2120. In this case, if there are no AC coefficients of the reference node in the frame buffer 2120, the RAHT unit 2080 may determine in step S28110 that inter prediction is not executable.

[0360] The RAHT unit 2080 inter and A.C. inter may be multiplied by a scaling factor α.

[0361] Attr pred = αAttr inter Or AC pred = αAC inter The coefficient α may be any real number. The coefficient α may be decoded for each node or for each layer. The coefficient α may be included in the slice data.

[0362] For example, the coefficient α may be defined using the depth of the hierarchy as follows, and α′ may be decoded instead of the coefficient α.

[0363] For example, the integer β may be defined as an integer ranging from integer a to integer b, and β may be decoded. The coefficient α may be calculated by adding an integer c to the decoded β and then dividing the result by the integer c, as follows:

[0364] α=(β+c) / c The integer β may be decoded using Exponential-Golomb coding.

[0365] Alternatively, the coefficient α may be derived at the decoder.

[0366] For example, the AC coefficients AC parent and the inter predicted value AC when the parent node is decoded. parent_inter may be calculated as follows using

[0367] α = AC parent / AC parent_inter For example, the RAHT unit 2080 calculates the AC coefficients AC of N neighboring nodes of the node to be decoded. neighbor1 , A.C. neighbor2 , ..., AC neighborN and the inter-predicted value AC when each adjacent node is decoded. neighbor_inter1 , A.C. neighbor_inter2 , ..., AC neighbor_interN and may be used to calculate α so as to minimize the cost.

[0368] The cost may be, for example, the sum of the AC coefficients of each adjacent node and the squared error of the predictor of the AC coefficients. The adjacent nodes may be, for example, only nodes adjacent to a face, or nodes adjacent to a face and nodes adjacent to an edge.

[0369] The RAHT unit 2080 may perform a similar operation in the inter prediction of DC coefficients in step S28003.

[0370] DC pred = αDC inter Here, the DC coefficient of the reference node is DC inter and the predicted value of the DC coefficient of the root node ispred Let's say.

[0371] Furthermore, the RAHT unit 2080 may calculate predicted values ​​of attribute values ​​or AC coefficients by combining inter prediction and intra prediction.

[0372] For example, an example in which the RAHT unit 2080 requests a prediction of an attribute value will be shown below.

[0373] Attr pred =W inter ・Attr inter +W intra ・Attr intra Here, Attr inter and Attr intra are the inter prediction and intra prediction of the attribute value, respectively. inter and W intra are the weights of inter prediction and intra prediction, respectively. inter and W intra may be determined depending on the depth of the layer to be processed so that the deeper the layer, the more importance is placed on intra prediction. For example, W inter =1-depth / N W intra =depth / N, where N is the maximum value of the depth of a layer for which inter prediction is enabled. The combination of inter prediction and intra prediction may be enabled only in a specific layer. For example, the combination of inter prediction and intra prediction may be enabled only when M<depth<N. M may be any real number less than N, and may be decoded as header information such as APS.

[0374] (Point Cloud Encoding Device 100) The point cloud encoding device 100 according to this embodiment will be described below with reference to Fig. 18. Fig. 18 is a diagram showing an example of functional blocks of the point cloud encoding device 100 according to this embodiment.

[0375] As shown in FIG. 18, the point cloud encoding device 100 includes a coordinate transformation unit 1010, a geometric information quantization unit 1020, a tree analysis unit 1030, an approximate surface analysis unit 1040, a geometric information encoding unit 1050, a geometric information reconstruction unit 1060, a color conversion unit 1070, an attribute transfer unit 1080, an RAHT unit 1090, an LoD calculation unit 1100, a lifting unit 1110, an attribute information quantization unit 1120, an attribute information encoding unit 1130, and a frame buffer 1140.

[0376] The coordinate transformation unit 1010 is configured to perform transformation processing from the three-dimensional coordinate system of the input point cloud to any different coordinate system. For example, the coordinate transformation may involve rotating the input point cloud to transform the x, y, and z coordinates of the input point cloud into any s, t, and u coordinates. Alternatively, as a variation of the transformation, the coordinate system of the input point cloud may be used as is.

[0377] The geometric information quantization unit 1020 is configured to quantize the position information of the input point group after coordinate transformation and remove points with overlapping coordinates. Note that when the quantization step size is 1, the position information of the input point group and the position information after quantization match. In other words, when the quantization step size is 1, it is equivalent to not performing quantization.

[0378] The tree analysis unit 1030 is configured to receive position information of the quantized point group as input, and to generate an occupancy code indicating at which node in the encoding target space a point exists, based on a tree structure described below.

[0379] In this process, the tree analysis unit 1030 is configured to recursively divide the encoding target space into rectangular parallelepipeds to generate a tree structure.

[0380] If a point exists within a rectangular parallelepiped, a tree structure can be generated by recursively dividing the rectangular parallelepiped into multiple rectangular parallelepipeds until the rectangular parallelepiped reaches a predetermined size. Each such rectangular parallelepiped is called a node. Each rectangular parallelepiped generated by dividing a node is called a child node, and the occupancy code is a value of 0 or 1 indicating whether or not the point is contained within the child node.

[0381] As described above, the tree analysis unit 1030 is configured to generate an occupancy code while recursively dividing a node until it reaches a predetermined size.

[0382] In this embodiment, a method called "Octree" can be used, which recursively performs octree division on the above-mentioned rectangular parallelepiped, always treating it as a cube, and a method called "QtBt" can be used, which performs quadtree division and binary tree division in addition to octree division.

[0383] Here, whether or not to use “QtBt” is transmitted to the point cloud decoding device 200 as control data.

[0384] Alternatively, predictive geometry coding using an arbitrary tree structure may be specified. In this case, the tree analysis unit 1030 determines the tree structure, and the determined tree structure is transmitted to the point cloud decoding device 200 as control data.

[0385] For example, the tree-structured control data may be configured so that it can be decoded according to the procedures described with reference to FIGS.

[0386] The approximate surface analysis unit 1040 is configured to generate approximate surface information using the tree information generated by the tree analysis unit 1030 .

[0387] For example, when decoding three-dimensional point cloud data of an object, if the point cloud is densely distributed on the surface of the object, approximate surface information is used to represent the area where the point cloud exists by approximating it with a small plane, rather than decoding each individual point cloud.

[0388] Specifically, the approximate surface analysis unit 1040 may be configured to generate approximate surface information using, for example, a method called "Trisoup." Furthermore, when decoding a sparse point cloud acquired by Lidar or the like, this process can be omitted.

[0389] The geometric information encoding unit 1050 is configured to generate a bitstream (geometric information bitstream) by encoding syntax such as the occupancy code generated by the tree analysis unit 1030 and the approximate surface information generated by the approximate surface analysis unit 1040. Here, the bitstream may include, for example, the syntax described in FIG. 4 .

[0390] The encoding process is, for example, a context-adaptive binary arithmetic coding process. Here, for example, the syntax includes control data (flags and parameters) for controlling the decoding process of the position information.

[0391] The geometric information reconstruction unit 1060 is configured to reconstruct the geometric information of each point of the point cloud data to be encoded (the coordinate system assumed by the encoding process, i.e., the position information after coordinate transformation in the coordinate transformation unit 1010) based on the tree information generated by the tree analysis unit 1030 and the approximate surface information generated by the approximate surface analysis unit 1040.

[0392] The frame buffer 1140 is configured to receive the geometric information reconstructed by the geometric information reconstruction unit 1060 as an input and store it as a reference frame.

[0393] The stored reference frame is read from the frame buffer 1140 and used as a reference frame when the tree analysis unit 1030 performs inter-prediction of a temporally different frame.

[0394] Here, which reference frame to use for each frame may be determined based on, for example, the value of a cost function representing encoding efficiency, and information on the reference frame to be used may be transmitted to the point cloud decoding device 200 as control data.

[0395] The color conversion unit 1070 is configured to perform color conversion when the input attribute information is color information. The color conversion does not necessarily have to be performed, and whether or not the color conversion process is to be performed is coded as part of the control data and transmitted to the point cloud decoding device 200.

[0396] The attribute transfer unit 1080 is configured to correct the attribute values ​​so as to minimize distortion of the attribute information, based on the position information of the input point cloud, the position information of the point cloud after reconstruction by the geometric information reconstruction unit 1060, and the attribute information after color change by the color conversion unit 1070. As a specific correction method, for example, the method described in Non-Patent Document 1 can be applied.

[0397] The RAHT unit 1090 is configured to receive as input the attribute information transferred by the attribute transfer unit 1080 and the geometric information generated by the geometric information reconstruction unit 1060, and to generate residual information for each point using a type of Haar transform called RAHT (Region Adaptive Hierarchical Transform).

[0398] The information to be decoded is the direct current component (DC coefficient) and alternating current component (AC coefficient) of the attribute information generated by using RAHT in the encoding process, and in the decoding process, it is converted into attribute information by using the inverse transform of RAHT.

[0399] As a specific example of the RAHT process, the method described in Non-Patent Document 1 can be used.

[0400] The LoD calculation unit 1100 is configured to receive the geometric information generated by the geometric information reconstruction unit 1060 as an input and generate an LoD (Level of Detail).

[0401] LoD is information for defining a reference relationship (a point to be referenced and a point to be referenced) to realize predictive coding, such as predicting attribute information of another point from attribute information of another point and encoding or decoding the prediction residual.

[0402] In other words, LoD is information that defines a hierarchical structure in which each point contained in geometric information is classified into multiple levels, and the attributes of points belonging to lower levels are encoded or decoded using the attribute information of points belonging to higher levels.

[0403] As a specific method for determining the LoD, for example, the method described in Non-Patent Document 1 mentioned above may be used.

[0404] The lifting unit 1110 is configured to generate residual information by a lifting process using the LoD generated by the LoD calculation unit 1100 and the attribute information after attribute transfer by the attribute transfer unit 1080 .

[0405] As a specific example of the lifting process, the method described in Non-Patent Document 1 above may be used.

[0406] The attribute information quantization unit 1120 is configured to quantize the residual information output from the RAHT unit 1090 or the lifting unit 1110. Here, a quantization step size of 1 is equivalent to no quantization being performed.

[0407] The attribute information encoding unit 1130 is configured to perform encoding processing using the quantized residual information, etc. output from the attribute information quantization unit 1120 as syntax, and to generate a bit stream related to the attribute information (attribute information bit stream).

[0408] The encoding process is, for example, a context-adaptive binary arithmetic coding process, where the syntax includes, for example, control data (flags and parameters) for controlling the decoding process of the attribute information.

[0409] Through the above processing, the point cloud encoding device 100 is configured to perform encoding processing using the position information and attribute information of each point in a point cloud as input, and to output a geometry information bit stream and an attribute information bit stream.

[0410] Furthermore, the above-described point group encoding device 100 and point group decoding device 200 may be realized as a program that causes a computer to execute each function (each process).

[0411] In each of the above embodiments, the present invention has been described using the example of applying it to the point cloud encoding device 100 and the point cloud decoding device 200, but the present invention is not limited to such an example and can be similarly applied to a point cloud encoding / decoding system having the functions of the point cloud encoding device 100 and the point cloud decoding device 200.

[0412] According to this embodiment, for example, it is possible to improve the overall service quality in video communication, which makes it possible to contribute to Goal 9 of the Sustainable Development Goals (SDGs) led by the United Nations, which is to "Develop resilient infrastructure, promote sustainable industrialization and foster innovation."

[0413] 10... Point cloud processing system 100... Point cloud encoding device 1010... Coordinate transformation unit 1020... Geometric information quantization unit 1030... Tree analysis unit 1040... Approximate surface analysis unit 1050... Geometric information encoding unit 1060... Geometric information reconstruction unit 1070... Color conversion unit 1080... Attribute transfer unit 1090... RAHT unit 1100... LoD calculation unit 1110... Lifting unit 1120... Attribute information quantization unit 1130... Attribute information encoding unit 200... Point cloud decoding device 2010... Geometric information decoding unit 2020... Tree synthesis unit 2030... Approximate surface synthesis unit 2040... Geometric information reconstruction unit 2050... Inverse coordinate transformation unit 2060... Attribute information decoding unit 2070... Inverse quantization unit 2080... RAHT unit 2090...LoD calculation unit 2100...inverse lifting unit 2110...inverse color conversion unit

Claims

1. A point cloud decoding apparatus, comprising: a RAHT unit that searches for adjacent nodes in a higher layer in intra prediction and sets a search range for the adjacent nodes in the higher layer based on past search results.

2. The point cloud decoding apparatus according to claim 1, wherein the RAHT unit sets a search range for an adjacent node to be searched using information on other searched adjacent nodes at a parent node.

3. The point cloud decoding apparatus according to claim 2, wherein the RAHT unit determines whether there is no adjacent node to be searched within the search range or there is a possibility of existence based on the Morton code of a node stored at a start point or an end point of the search range.

4. The point cloud decoding apparatus according to claim 3, wherein the RAHT unit stores the nodes in the higher layer on a one-dimensional array, uses an index corresponding to the parent node in the array as a starting point of search, sorts the adjacent nodes at the parent node in ascending order of Morton code, searches for the adjacent nodes in the higher layer in ascending order of Morton code, and sets a search range for the adjacent node to be searched using an index of the other searched adjacent nodes at the parent node in the array during the search.

5. The point cloud decoding apparatus according to claim 4, wherein the RAHT unit records a search result of an adjacent node in the higher layer and omits the search if the recorded search result exists.

6. The point cloud decoding apparatus according to claim 5, wherein the RAHT unit sequentially searches for adjacent nodes starting from a node with a smaller Morton code among the nodes in the higher layer, records the search result when the Morton code of the adjacent node to be searched is larger than the Morton code of the parent node, and omits the search by referring to the recorded search result when the Morton code of the adjacent node to be searched is smaller than the Morton code of the parent node.

7. A point cloud decoding method, comprising: a step of searching for adjacent nodes in a higher layer in intra prediction and setting a search range for the adjacent nodes in the higher layer based on past search results.

8. A program that causes a computer to function as a point cloud decoding device, wherein the point cloud decoding device includes a RAHT unit that searches for adjacent nodes in a higher layer in intra prediction and sets a search range for the adjacent nodes in the higher layer based on past search results.