Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device
By encoding normal vectors as attribute information within three-dimensional data, the method addresses inefficiencies in processing, enhancing the efficiency of encoding and decoding three-dimensional data.
Patent Information
- Application Number
- JP2025030850
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-05-28
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2040-05-27
AI Technical Summary
Existing methods for encoding and decoding three-dimensional data, particularly point clouds, are inefficient in terms of processing requirements, lacking effective techniques to reduce the amount of processing needed for compression and transmission.
A method that encodes three-dimensional data by dividing it into split data and generating a bit stream with control information that includes attribute type information, specifically encoding normal vectors as attribute information, allowing for reduced processing by treating them similarly to other attributes.
This approach reduces the processing load by encoding normal vectors as attribute information, enabling more efficient encoding and decoding of three-dimensional data, thereby optimizing data transmission and storage.
Smart Images

Figure 2025078684000001_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, and a three-dimensional data decoding device. [Background technology]
[0002] In the future, devices and services that utilize 3D data are expected to become widespread in a wide range of fields, such as computer vision for autonomous operation of automobiles or robots, map information, monitoring, infrastructure inspection, video distribution, etc. 3D data is acquired in various ways, such as distance sensors such as range finders, stereo cameras, or a combination of multiple monocular cameras.
[0003] One method of expressing three-dimensional data is a method called a point cloud, which represents the shape of a three-dimensional structure using a group of points in a three-dimensional space. In a point cloud, the positions and colors of the points are stored. Point clouds are expected to become the mainstream method of expressing three-dimensional data, but point clouds have a very large amount of data. Therefore, when storing or transmitting three-dimensional data, it is essential to compress the amount of data by encoding, just like two-dimensional video images (examples include MPEG-4 AVC or HEVC standardized by MPEG).
[0004] In addition, compression of point clouds is partially supported by public libraries (Point Cloud Library) that perform point cloud-related processing.
[0005] Furthermore, a technique is known that uses three-dimensional map data to search for and display facilities located around a vehicle (see, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0006] [Patent Document 1] International Publication No. 2014 / 020663 Summary of the Invention [Problem to be solved by the invention]
[0007] In the encoding process and the decoding process of three-dimensional data, it is desirable to reduce the amount of processing.
[0008] An object of the present disclosure is to provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, and a three-dimensional data decoding device that can reduce the amount of processing. [Means for solving the problem]
[0009] A three-dimensional data encoding method according to one embodiment of the present disclosure includes: encoding, by a processor, three-dimensional data including position information and one or more attribute information; generating, by the processor, control information common to the position information and the one or more attribute information; and generating, by the processor, a bit stream including the encoded three-dimensional data and the control information, wherein the control information includes one or more attribute type information corresponding to the one or more attribute information, and when first attribute type information included in the one or more attribute type information includes first information, the first attribute information included in the one or more attribute information relates to a normal vector, and the first attribute information corresponds to the first attribute type information.
[0010] A three-dimensional data encoding method according to one aspect of the present disclosure divides point cloud data into multiple split data and generates a bit stream by encoding the multiple split data, wherein the bit stream includes information indicating normal vectors of each of the multiple split data.
[0011] A three-dimensional data decoding method according to one embodiment of the present disclosure includes: acquiring, by a processor, a bit stream including encoded three-dimensional data including encoding position information and one or more encoded attribute information; acquiring, by the processor, control information common to the encoding position information and the one or more encoded attribute information from the bit stream; and decoding, by the processor, the one or more encoded attribute information, wherein the control information includes one or more attribute type information corresponding to the one or more encoded attribute information, and when first attribute type information included in the one or more attribute type information includes first information, the first encoded attribute information included in the one or more encoded attribute information relates to a normal vector, and the first encoded attribute information corresponds to the first attribute type information.
[0012] A three-dimensional data decoding method according to one aspect of the present disclosure obtains a bit stream generated by encoding multiple split data generated by dividing point cloud data, and obtains information indicating the normal vectors of each of the multiple split data from the bit stream. Effect of the Invention
[0013] The present disclosure can provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device that can reduce the amount of processing. [Brief description of the drawings]
[0014] [Figure 1] FIG. 1 is a diagram showing a configuration of a three-dimensional data encoding / decoding system according to the first embodiment. [Diagram 2] FIG. 2 is a diagram illustrating an example of a configuration of point cloud data according to the first embodiment. [Diagram 3] FIG. 3 is a diagram showing an example of a configuration of a data file in which point cloud data information according to the first embodiment is described. [Figure 4] FIG. 4 is a diagram showing types of point cloud data according to the first embodiment. [Diagram 5] FIG. 5 is a diagram showing a configuration of a first encoding unit according to the first embodiment. [Figure 6] FIG. 6 is a block diagram of a first encoding unit according to the first embodiment. [Figure 7] FIG. 7 is a diagram showing a configuration of a first decoding unit according to the first embodiment. [Figure 8] FIG. 8 is a block diagram of a first decoding unit according to the first embodiment. [Figure 9] FIG. 9 is a block diagram of a three-dimensional data encoding device according to the first embodiment. [Figure 10] FIG. 10 is a diagram showing an example of location information according to the first embodiment. [Figure 11] FIG. 11 is a diagram showing an example of an octree representation of position information according to the first embodiment. [Figure 12] FIG. 12 is a block diagram of the three-dimensional data decoding device according to the first embodiment. [Figure 13] FIG. 13 is a block diagram of the attribute information encoding unit according to the first embodiment. [Figure 14] FIG. 14 is a block diagram of the attribute information decoding unit according to the first embodiment. [Figure 15] FIG. 15 is a block diagram showing a configuration of an attribute information encoding unit according to the first embodiment. As shown in FIG. [Figure 16] FIG. 16 is a block diagram of the attribute information encoding unit according to the first embodiment. [Figure 17] FIG. 17 is a block diagram showing a configuration of an attribute information decoding unit according to the first embodiment. As shown in FIG. [Figure 18] FIG. 18 is a block diagram of the attribute information decoding unit according to the first embodiment. As shown in FIG. [Figure 19] FIG. 19 is a diagram showing a configuration of a second encoding unit according to the first embodiment. As shown in FIG. [Figure 20] FIG. 20 is a block diagram of a second encoding unit according to the first embodiment. [Figure 21] FIG. 21 is a diagram illustrating a configuration of a second decoding unit according to the first embodiment. As shown in FIG. [Figure 22] FIG. 22 is a block diagram of a second decoding unit according to the first embodiment. [Diagram 23]FIG. 23 is a diagram showing a protocol stack related to PCC encoded data according to the first embodiment. [Figure 24] FIG. 24 is a diagram illustrating configurations of an encoding unit and a multiplexing unit according to the second embodiment. In FIG. [Diagram 25] FIG. 25 is a diagram illustrating an example of a structure of encoded data according to the second embodiment. In FIG. [Figure 26] FIG. 26 is a diagram showing an example of the structure of coded data and an NAL unit according to the second embodiment. [Figure 27] FIG. 27 is a diagram illustrating an example of the semantics of pcc_nal_unit_type according to the second embodiment. [Figure 28] FIG. 28 is a block diagram of a first encoding unit according to the third embodiment. [Figure 29] FIG. 29 is a block diagram of a first decoding unit according to the third embodiment. [Diagram 30] FIG. 30 is a block diagram of a division unit according to the third embodiment. [Diagram 31] FIG. 31 is a diagram showing an example of division into slices and tiles according to the third embodiment. [Diagram 32] FIG. 32 is a diagram showing an example of a division pattern of slices and tiles according to the third embodiment. [Diagram 33] FIG. 33 is a diagram illustrating an example of a dependency relationship according to the third embodiment. [Diagram 34] FIG. 34 is a diagram showing an example of a data decoding order according to the third embodiment. [Diagram 35] FIG. 35 is a flowchart of the encoding process according to the third embodiment. [Diagram 36] FIG. 36 is a block diagram of a coupling unit according to the third embodiment. [Figure 37] FIG. 37 is a diagram showing an example of the structure of coded data and an NAL unit according to the third embodiment. [Figure 38] FIG. 38 is a flowchart of the encoding process according to the third embodiment. [Figure 39] FIG. 39 is a flowchart of the decoding process according to the third embodiment. [Diagram 40] FIG. 40 is a diagram illustrating an example of the syntax of the tile additional information according to the fourth embodiment. [Diagram 41] FIG. 41 is a block diagram of a coding / decoding system according to the fourth embodiment. [Diagram 42] FIG. 42 is a diagram showing an example of the syntax of slice additional information according to the fourth embodiment. [Diagram 43] FIG. 43 is a flowchart of the encoding process according to the fourth embodiment. [Diagram 44] FIG. 44 is a flowchart of the decoding process according to the fourth embodiment. [Diagram 45] FIG. 45 is a diagram showing an example of a division method according to the fifth embodiment. [Figure 46] FIG. 46 is a diagram illustrating an example of division of point cloud data according to the fifth embodiment. [Figure 47] FIG. 47 is a diagram illustrating an example of the syntax of the tile additional information according to the fifth embodiment. [Figure 48] FIG. 48 is a diagram showing an example of index information according to the fifth embodiment. [Figure 49] FIG. 49 is a diagram illustrating an example of a dependency relationship according to the fifth embodiment. [Figure 50] FIG. 50 is a diagram showing an example of transmission data according to the fifth embodiment. In FIG. [Figure 51] FIG. 51 is a diagram showing an example of the structure of a NAL unit according to the fifth embodiment. [Figure 52] FIG. 52 is a diagram illustrating an example of a dependency relationship according to the fifth embodiment. In FIG. [Diagram 53] FIG. 53 is a diagram showing an example of a data decoding order according to the fifth embodiment. [Figure 54] FIG. 54 is a diagram showing an example of a dependency relationship according to the fifth embodiment. In FIG. [Figure 55] FIG. 55 is a diagram showing an example of a data decoding order according to the fifth embodiment. [Figure 56] FIG. 56 is a flowchart of the encoding process according to the fifth embodiment. [Figure 57]FIG. 57 is a flowchart of the decoding process according to the fifth embodiment. [Figure 58] FIG. 58 is a flowchart of the encoding process according to the fifth embodiment. [Figure 59] FIG. 59 is a flowchart of the encoding process according to the fifth embodiment. [Figure 60] FIG. 60 is a diagram showing an example of transmission data and reception data according to the fifth embodiment. In FIG. [Figure 61] FIG. 61 is a flowchart of the decoding process according to the fifth embodiment. [Figure 62] FIG. 62 is a diagram showing an example of transmission data and reception data according to the fifth embodiment. In FIG. [Figure 63] FIG. 63 is a flowchart of the decoding process according to the fifth embodiment. [Figure 64] FIG. 64 is a flowchart of the encoding process according to the fifth embodiment. [Figure 65] FIG. 65 is a diagram showing an example of index information according to the fifth embodiment. [Figure 66] FIG. 66 is a diagram showing an example of a dependency relationship according to the fifth embodiment. In FIG. [Figure 67] FIG. 67 is a diagram showing an example of transmission data according to the fifth embodiment. In FIG. [Figure 68] FIG. 68 is a diagram showing an example of transmission data and reception data according to the fifth embodiment. In FIG. [Figure 69] FIG. 69 is a flowchart of the decoding process according to the fifth embodiment. [Figure 70] FIG. 70 is a diagram illustrating a quantization process for each tile according to the sixth embodiment. [Figure 71] FIG. 71 is a diagram illustrating an example of syntax of a GPS according to the sixth embodiment. [Figure 72] FIG. 72 is a diagram illustrating an example of syntax of the tile information according to the sixth embodiment. [Figure 73] FIG. 73 is a diagram illustrating an example of the syntax of the node information according to the sixth embodiment. [Figure 74]FIG. 74 is a flowchart of three-dimensional data encoding processing according to the sixth embodiment. [Figure 75] FIG. 75 is a flowchart of three-dimensional data encoding processing according to the sixth embodiment. [Figure 76] FIG. 76 is a flowchart of three-dimensional data decoding processing according to the sixth embodiment. [Figure 77] FIG. 77 is a diagram showing an example of tile division according to the sixth embodiment. [Figure 78] FIG. 78 is a diagram showing an example of tile division according to the sixth embodiment. [Figure 79] FIG. 79 is a flowchart of three-dimensional data encoding processing according to the sixth embodiment. [Figure 80] FIG. 80 is a block diagram of a three-dimensional data encoding device according to the sixth embodiment. [Figure 81] FIG. 81 is a diagram illustrating an example of syntax of a GPS according to the sixth embodiment. [Figure 82] FIG. 82 is a flowchart of three-dimensional data decoding processing according to the sixth embodiment. [Figure 83] FIG. 83 is a diagram showing an example of an application according to the sixth embodiment. In FIG. [Figure 84] FIG. 84 is a diagram showing an example of tile division and slice division according to the sixth embodiment. In FIG. [Figure 85] FIG. 85 is a flowchart of processing in the system according to the sixth embodiment. [Figure 86] FIG. 86 is a flowchart of processing in the system according to the sixth embodiment. [Figure 87] FIG. 87 is a block diagram of a three-dimensional data encoding device according to the seventh embodiment. [Figure 88] FIG. 88 is a block diagram of a three-dimensional data decoding device according to the seventh embodiment. [Figure 89] FIG. 89 is a block diagram of a three-dimensional data encoding device according to the seventh embodiment. [Figure 90]FIG. 90 is a block diagram showing a configuration of a three-dimensional data decoding device according to the seventh embodiment. [Figure 91] FIG. 91 is a diagram illustrating an example of point cloud data according to the seventh embodiment. As shown in FIG. [Figure 92] FIG. 92 is a diagram showing an example of normal vectors for each point according to the seventh embodiment. [Figure 93] FIG. 93 is a diagram illustrating an example of the syntax of a normal vector according to the seventh embodiment. [Figure 94] FIG. 94 is a flowchart of three-dimensional data encoding processing according to the seventh embodiment. [Figure 95] FIG. 95 is a flowchart of three-dimensional data decoding processing according to the seventh embodiment. [Figure 96] FIG. 96 is a diagram showing an example of a configuration of a bit stream according to the seventh embodiment. [Figure 97] FIG. 97 is a diagram showing an example of point cloud information according to the seventh embodiment. In FIG. [Figure 98] FIG. 98 is a flowchart of three-dimensional data encoding processing according to the seventh embodiment. [Figure 99] FIG. 99 is a flowchart of three-dimensional data decoding processing according to the seventh embodiment. [Figure 100] FIG. 100 is a diagram showing an example of division of a normal vector according to the seventh embodiment. [Figure 101] FIG. 101 is a diagram showing an example of division of a normal vector according to the seventh embodiment. [Figure 102] FIG. 102 is a diagram showing an example of point cloud data according to the seventh embodiment. In FIG. [Figure 103] FIG. 103 is a diagram showing an example of a normal vector according to the seventh embodiment. In FIG. [Figure 104] FIG. 104 is a diagram showing an example of information on normal vectors according to the seventh embodiment. In FIG. [Figure 105] FIG. 105 is a diagram showing an example of a cube according to the seventh embodiment. In FIG. [Fig. 106] FIG. 106 is a diagram showing an example of a face of a cube according to the seventh embodiment. In FIG. [Figure 107] FIG. 107 is a diagram showing an example of a face of a cube according to the seventh embodiment. In FIG. [Figure 108] FIG. 108 is a diagram showing an example of a face of a cube according to the seventh embodiment. In FIG. [Fig. 109] FIG. 109 is a diagram showing an example of visibility of slices according to the seventh embodiment. [Figure 110] FIG. 110 is a diagram showing an example of a configuration of a bit stream according to the seventh embodiment. [Figure 111] FIG. 111 is a diagram illustrating an example of the syntax of a slice header of position information according to the seventh embodiment. [Figure 112] FIG. 112 is a diagram illustrating an example of the syntax of a slice header of position information according to the seventh embodiment. [Figure 113] FIG. 113 is a flowchart of three-dimensional data encoding processing according to the seventh embodiment. [Fig. 114] FIG. 114 is a flowchart of three-dimensional data decoding processing according to the seventh embodiment. [Fig. 115] FIG. 115 is a flowchart of three-dimensional data decoding processing according to the seventh embodiment. [Fig. 116] FIG. 116 is a diagram showing an example of a configuration of a bit stream according to the seventh embodiment. [Fig. 117] FIG. 117 is a diagram showing an example of the syntax of slice information according to the seventh embodiment. [Fig. 118] FIG. 118 is a diagram illustrating an example of the syntax of slice information according to the seventh embodiment. [Figure 119] FIG. 119 is a flowchart of three-dimensional data decoding processing according to the seventh embodiment. [Figure 120] FIG. 120 is a diagram illustrating an example of partial decoding processing according to the seventh embodiment. [Figure 121] FIG. 121 is a diagram illustrating an example of the configuration of a three-dimensional data decoding device according to the seventh embodiment. [Figure 122] FIG. 122 is a diagram illustrating an example of processing by a random access control unit according to the seventh embodiment. In FIG. [Figure 123] FIG. 123 is a diagram illustrating an example of processing by a random access control unit according to the seventh embodiment. In FIG. [Figure 124] FIG. 124 is a diagram showing an example of the relationship between distance and resolution according to the seventh embodiment. In FIG. [Fig. 125] FIG. 125 is a diagram showing an example of bricks and normal vectors according to the seventh embodiment. [Fig. 126] FIG. 126 is a diagram showing an example of levels according to the seventh embodiment. [Figure 127] FIG. 127 is a diagram showing an example of an octree structure according to the seventh embodiment. [Figure 128] FIG. 128 is a flowchart of three-dimensional data decoding processing according to the seventh embodiment. [Figure 129] FIG. 129 is a flowchart of three-dimensional data decoding processing according to the seventh embodiment. [Fig. 130] FIG. 130 is a diagram showing an example of a brick to be decoded according to the seventh embodiment. [Fig. 131] FIG. 131 is a diagram showing an example of levels of decoding targets according to the seventh embodiment. In FIG. [Fig. 132] FIG. 132 is a diagram illustrating an example of the syntax of a slice header of position information according to the seventh embodiment. [Fig. 133] FIG. 133 is a flowchart of three-dimensional data encoding processing according to the seventh embodiment. [Fig. 134] FIG. 134 is a flowchart of three-dimensional data decoding processing according to the seventh embodiment. [Fig. 135] FIG. 135 is a diagram showing an example of point cloud data according to the seventh embodiment. As shown in FIG. [Fig. 136] FIG. 136 is a diagram showing an example of point cloud data according to the seventh embodiment. As shown in FIG. [Fig. 137] FIG. 137 is a diagram showing an example of the configuration of a system according to the seventh embodiment. In FIG. [Fig. 138] FIG. 138 is a diagram illustrating a configuration example of a system according to the seventh embodiment. In FIG. [Fig. 139]FIG. 139 is a diagram showing an example of the configuration of a system according to the seventh embodiment. In FIG. [Fig. 140] FIG. 140 is a diagram illustrating a configuration example of a system according to the seventh embodiment. In FIG. [Fig. 141] FIG. 141 is a diagram showing an example of a configuration of a bit stream according to the seventh embodiment. [Fig. 142] FIG. 142 is a diagram illustrating an example of the configuration of a three-dimensional data encoding device according to the seventh embodiment. [Fig. 143] FIG. 143 is a diagram showing an example of the configuration of a three-dimensional data decoding device according to the seventh embodiment. [Fig. 144] FIG. 144 is a diagram showing the basic structure of an ISOBMFF according to the seventh embodiment. [Fig. 145] Figure 145 is a protocol stack diagram when a NAL unit common to the PCC codec in embodiment 7 is stored in ISOBMFF. [Fig. 146] FIG. 146 is a diagram showing an example of converting a bitstream according to the seventh embodiment into a file format. [Fig. 147] FIG. 147 is a diagram illustrating an example of the syntax of slice information according to the seventh embodiment. [Fig. 148] FIG. 148 is a diagram showing an example of syntax of a PCC random access table relating to embodiment 7. [Figure 149] FIG. 149 is a diagram showing an example of syntax of a PCC random access table relating to embodiment 7. [Fig. 150] FIG. 150 is a diagram showing an example of syntax of a PCC random access table according to the seventh embodiment. [Fig. 151] FIG. 151 is a flowchart of three-dimensional data encoding processing according to the seventh embodiment. [Fig. 152] FIG. 152 is a flowchart of three-dimensional data decoding processing according to the seventh embodiment. [Fig. 153] FIG. 153 is a flowchart of three-dimensional data encoding processing according to the seventh embodiment. [Fig. 154]FIG. 154 is a flowchart of three-dimensional data decoding processing according to the seventh embodiment. [Fig. 155] FIG. 155 is a block diagram of a three-dimensional data creation device according to the eighth embodiment. [Fig. 156] FIG. 156 is a flowchart of a three-dimensional data creating method according to the eighth embodiment. [Fig. 157] FIG. 157 is a diagram showing a configuration of a system according to the eighth embodiment. In FIG. [Fig. 158] FIG. 158 is a block diagram of a client device according to the eighth embodiment. [Fig. 159] FIG. 159 is a block diagram of a server according to the eighth embodiment. [Fig. 160] FIG. 160 is a flowchart of three-dimensional data creation processing by a client device according to the eighth embodiment. [Fig. 161] FIG. 161 is a flowchart of a sensor information transmission process by a client device according to the eighth embodiment. [Fig. 162] FIG. 162 is a flowchart of three-dimensional data creation processing by the server according to the eighth embodiment. [Fig. 163] FIG. 163 is a flowchart of a 3D map transmission process by the server according to the eighth embodiment. [Fig. 164] FIG. 164 is a diagram showing a configuration of a modified example of the system according to the eighth embodiment. In FIG. [Fig. 165] FIG. 165 is a diagram showing the configurations of a server and a client device according to the eighth embodiment. In FIG. [Fig. 166] FIG. 166 is a diagram showing the configurations of a server and a client device according to the eighth embodiment. In FIG. [Fig. 167] FIG. 167 is a flowchart of processing by a client device according to the eighth embodiment. [Fig. 168] FIG. 168 is a diagram illustrating a configuration of a sensor information collecting system according to the eighth embodiment. As shown in FIG. [Fig. 169] FIG. 169 is a diagram illustrating an example of a system according to the eighth embodiment. [Fig. 170] FIG. 170 is a diagram showing a modification of the system according to the eighth embodiment. [Fig. 171] FIG. 171 is a flowchart showing an example of application processing according to the eighth embodiment. [Fig. 172] FIG. 172 is a diagram showing the sensor ranges of various sensors according to the eighth embodiment. [Fig. 173] FIG. 173 is a diagram illustrating a configuration example of an autonomous driving system according to the eighth embodiment. [Fig. 174] FIG. 174 is a diagram showing an example of a bitstream configuration according to the eighth embodiment. In FIG. [Fig. 175] FIG. 175 is a flowchart of the point group selection process according to the eighth embodiment. [Fig. 176] FIG. 176 is a diagram showing an example of a screen for the point group selection process according to the eighth embodiment. [Fig. 177] FIG. 177 is a diagram showing an example of a screen for the point group selection process according to the eighth embodiment. [Fig. 178] FIG. 178 is a diagram showing an example of a screen for the point group selection process according to the eighth embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0015] A three-dimensional data encoding method according to one embodiment of the present disclosure generates a bit stream by encoding position information and one or more attribute information of each of a plurality of three-dimensional points included in point cloud data, and in the encoding, a normal vector of each of the plurality of three-dimensional points is encoded as a piece of attribute information included in the one or more attribute information.
[0016] According to this, the three-dimensional data encoding method can process the normal vector in the same way as other attribute information by encoding the normal vector as attribute information, thereby reducing the amount of processing.
[0017] For example, in the encoding, the normal vector expressed by a floating point may be converted into an integer and then encoded.
[0018] According to this, the three-dimensional data encoding method can process normal vectors in the same way as other attribute information expressed by integers.
[0019] For example, the bit stream may include control information common to the position information and the one or more attribute information, and the control information may include at least one of information indicating that one of the attribute information included in the one or more attribute information indicates the normal vector, or information indicating that the normal vector is data having three elements for each point.
[0020] A three-dimensional data encoding method according to one aspect of the present disclosure divides point cloud data into multiple split data and generates a bit stream by encoding the multiple split data, wherein the bit stream includes information indicating normal vectors of each of the multiple split data.
[0021] According to this, the three-dimensional data encoding method can reduce the amount of processing and the amount of code by encoding the normal vector for each divided data, compared to encoding the normal vector for each point.
[0022] For example, each of the plurality of divided data may be a random access unit.
[0023] A three-dimensional data decoding method according to one embodiment of the present disclosure obtains a bit stream generated by encoding position information and one or more attribute information of each of a plurality of three-dimensional points included in point cloud data, in which normal vectors of each of the plurality of three-dimensional points are encoded as a piece of attribute information included in the one or more attribute information, and obtains the normal vectors from the bit stream by decoding the piece of attribute information.
[0024] According to this, the three-dimensional data decoding method decodes the normal vector as attribute information, and can process the normal vector in the same manner as other attribute information, thereby reducing the amount of processing.
[0025] For example, in obtaining the normal vector, the normal vector indicated by an integer may be obtained.
[0026] According to this, the three-dimensional data decoding method can process normal vectors in the same way as other attribute information expressed as integers.
[0027] For example, the bit stream may include control information common to the position information and the one or more attribute information, and the control information may include at least one of information indicating that first attribute information included in the one or more attribute information indicates the normal vector, or information indicating that the normal vector is data having three elements for each point.
[0028] A three-dimensional data decoding method according to one aspect of the present disclosure obtains a bit stream generated by encoding multiple split data generated by dividing point cloud data, and obtains information indicating the normal vectors of each of the multiple split data from the bit stream.
[0029] According to this, the three-dimensional data decoding method can reduce the amount of processing by decoding the normal vector for each divided data, compared to decoding the normal vector for each point.
[0030] For example, each of the plurality of divided data may be a random access unit.
[0031] For example, the three-dimensional data decoding method may further include determining a piece of divided data to be decoded from the plurality of pieces of divided data based on the normal vector, and decoding the piece of divided data to be decoded.
[0032] For example, the three-dimensional data decoding method may further include determining a decoding order of the plurality of divided data based on the normal vector, and decoding the plurality of divided data in the determined decoding order.
[0033] In addition, a three-dimensional data encoding device according to one embodiment of the present disclosure includes a processor and a memory, and the processor uses the memory to generate a bit stream by encoding position information and one or more attribute information of each of a plurality of three-dimensional points included in the point cloud data, and in the encoding, encodes a normal vector of each of the plurality of three-dimensional points as a piece of attribute information included in the one or more attribute information.
[0034] According to this, the three-dimensional data encoding device can process the normal vector in the same way as other attribute information by encoding the normal vector as attribute information, thereby reducing the amount of processing.
[0035] In addition, a three-dimensional data encoding device according to one aspect of the present disclosure includes a processor and a memory, and the processor uses the memory to divide point cloud data into a plurality of split data and encode the plurality of split data to generate a bit stream, the bit stream including information indicating a normal vector of each of the plurality of split data.
[0036] According to this, the three-dimensional data encoding device can reduce the amount of processing and the amount of code by encoding the normal vector for each divided data, compared to encoding the normal vector for each point.
[0037] In addition, a three-dimensional data decoding device according to one embodiment of the present disclosure includes a processor and a memory, and the processor uses the memory to obtain a bit stream generated by encoding position information and one or more attribute information of each of a plurality of three-dimensional points included in point cloud data, in which normal vectors of each of the plurality of three-dimensional points are encoded as a piece of attribute information included in the one or more attribute information, and obtains the normal vectors from the bit stream by decoding the piece of attribute information.
[0038] According to this, the three-dimensional data decoding device decodes the normal vector as attribute information, and can process the normal vector in the same way as other attribute information, thereby reducing the amount of processing.
[0039] In addition, a three-dimensional data decoding device according to one aspect of the present disclosure includes a processor and a memory, and the processor uses the memory to obtain a bit stream generated by encoding multiple split data generated by dividing point cloud data, and obtains information indicating the normal vector of each of the multiple split data from the bit stream.
[0040] This allows the three-dimensional data decoding device to reduce the amount of processing by decoding the normal vector for each divided data, compared to decoding the normal vector for each point.
[0041] These comprehensive or specific aspects may be realized as a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.
[0042] Hereinafter, the embodiments will be described in detail with reference to the drawings. Note that each of the embodiments described below shows a specific example of the present disclosure. The numerical values, shapes, materials, components, arrangement and connection forms of the components, steps, and order of steps shown in the following embodiments are merely examples and are not intended to limit the present disclosure. In addition, among the components in the following embodiments, components that are not described in the independent claims will be described as optional components.
[0043] (Embodiment 1) When using the encoded data of a point cloud in an actual device or service, it is desirable to transmit and receive the necessary information according to the purpose in order to reduce the network bandwidth. However, until now, such a function has not existed in the encoding structure of three-dimensional data, and there has been no encoding method for this purpose.
[0044] In this embodiment, we describe a three-dimensional data encoding method and a three-dimensional data encoding device for providing the function of transmitting and receiving necessary information depending on the application in encoded data of a three-dimensional point cloud, as well as a three-dimensional data decoding method and a three-dimensional data decoding device for decoding the encoded data, a three-dimensional data multiplexing method for multiplexing the encoded data, and a three-dimensional data transmission method for transmitting the encoded data.
[0045] In particular, the first and second encoding methods are currently being considered as methods (encoding schemes) for encoding point cloud data; however, the structure of the encoded data and the method for storing the encoded data in a system format have not been defined, and as things stand, there is a problem that MUX processing (multiplexing) in the encoding unit, or transmission or storage, is not possible.
[0046] Also, there has been no method to date that supports a format in which two codecs, the first encoding method and the second encoding method, are mixed, such as PCC (Point Cloud Compression).
[0047] In this embodiment, a structure of PCC encoded data in which two codecs, a first encoding method and a second encoding method, are mixed, and a method of storing the encoded data in a system format will be described.
[0048] First, the configuration of a three-dimensional data (point cloud data) encoding / decoding system according to this embodiment will be described. Fig. 1 is a diagram showing an example of the configuration of a three-dimensional data encoding / decoding system according to this embodiment. As shown in Fig. 1, the three-dimensional data encoding / decoding system includes a three-dimensional data encoding system 4601, a three-dimensional data decoding system 4602, a sensor terminal 4603, and an external connection unit 4604.
[0049] The three-dimensional data encoding system 4601 generates encoded data or multiplexed data by encoding point cloud data, which is three-dimensional data. The three-dimensional data encoding system 4601 may be a three-dimensional data encoding device realized by a single device, or may be a system realized by multiple devices. The three-dimensional data encoding device may include some of the multiple processing units included in the three-dimensional data encoding system 4601.
[0050] The three-dimensional data encoding system 4601 includes a point cloud data generation system 4611, a presentation unit 4612, an encoding unit 4613, a multiplexing unit 4614, an input / output unit 4615, and a control unit 4616. The point cloud data generation system 4611 includes a sensor information acquisition unit 4617 and a point cloud data generation unit 4618.
[0051] The sensor information acquisition unit 4617 acquires sensor information from the sensor terminal 4603, and outputs the sensor information to the point cloud data generation unit 4618. The point cloud data generation unit 4618 generates point cloud data from the sensor information, and outputs the point cloud data to the encoding unit 4613.
[0052] The presentation unit 4612 presents the sensor information or the point cloud data to the user. For example, the presentation unit 4612 displays information or an image based on the sensor information or the point cloud data.
[0053] The encoding unit 4613 encodes (compresses) the point cloud data, and outputs the obtained encoded data, control information obtained in the encoding process, and other additional information to the multiplexing unit 4614. The additional information includes, for example, sensor information.
[0054] The multiplexing unit 4614 generates multiplexed data by multiplexing the coded data, control information, and additional information input from the coding unit 4613. The format of the multiplexed data is, for example, a file format for storage, or a packet format for transmission.
[0055] The input / output unit 4615 (e.g., a communication unit or an interface) outputs the multiplexed data to the outside. Alternatively, the multiplexed data is stored in a storage unit such as an internal memory. The control unit 4616 (or an application execution unit) controls each processing unit. In other words, the control unit 4616 controls encoding, multiplexing, etc.
[0056] The sensor information may be input to the encoding unit 4613 or the multiplexing unit 4614. Furthermore, the input / output unit 4615 may output the point cloud data or the encoded data directly to the outside.
[0057] The transmission signal (multiplexed data) output from the three-dimensional data encoding system 4601 is input to the three-dimensional data decoding system 4602 via the external connection unit 4604.
[0058] The three-dimensional data decoding system 4602 generates point cloud data, which is three-dimensional data, by decoding the encoded data or multiplexed data. The three-dimensional data decoding system 4602 may be a three-dimensional data decoding device realized by a single device, or may be a system realized by multiple devices. The three-dimensional data decoding device may also include some of the multiple processing units included in the three-dimensional data decoding system 4602.
[0059] The three-dimensional data decoding system 4602 includes a sensor information acquisition unit 4621 , an input / output unit 4622 , a demultiplexing unit 4623 , a decoding unit 4624 , a presentation unit 4625 , a user interface 4626 , and a control unit 4627 .
[0060] The sensor information acquisition unit 4621 acquires sensor information from the sensor terminal 4603 .
[0061] The input / output unit 4622 acquires a transmission signal, decodes multiplexed data (file format or packets) from the transmission signal, and outputs the multiplexed data to the demultiplexer 4623.
[0062] The demultiplexer 4623 obtains the coded data, control information, and additional information from the multiplexed data, and outputs the coded data, control information, and additional information to the decoder 4624.
[0063] The decoding unit 4624 reconstructs the point cloud data by decoding the encoded data.
[0064] The presentation unit 4625 presents the point cloud data to the user. For example, the presentation unit 4625 displays information or an image based on the point cloud data. The user interface 4626 acquires an instruction based on a user's operation. The control unit 4627 (or the application execution unit) controls each processing unit. That is, the control unit 4627 controls demultiplexing, decoding, presentation, etc.
[0065] The input / output unit 4622 may obtain the point cloud data or the encoded data directly from the outside. The presentation unit 4625 may obtain additional information such as sensor information and present information based on the additional information. The presentation unit 4625 may present information based on a user instruction obtained by the user interface 4626.
[0066] The sensor terminal 4603 generates sensor information obtained by a sensor. The sensor terminal 4603 is a terminal equipped with a sensor or a camera, and examples of the sensor terminal include a moving body such as an automobile, a flying object such as an airplane, a mobile terminal, and a camera.
[0067] The sensor information that can be acquired by the sensor terminal 4603 is, for example, (1) the distance between the sensor terminal 4603 and an object, or the reflectance of the object, obtained from a LIDAR, millimeter wave radar, or an infrared sensor, and (2) the distance between a camera and an object, or the reflectance of the object, obtained from a plurality of monocular camera images or stereo camera images. The sensor information may also include the attitude, direction, gyro (angular velocity), position (GPS information or altitude), speed, acceleration, etc. of the sensor. The sensor information may also include temperature, air pressure, humidity, magnetism, etc.
[0068] The external connection unit 4604 is realized by an integrated circuit (LSI or IC), an external storage unit, communication with a cloud server via the Internet, broadcasting, or the like.
[0069] Next, the point cloud data will be described. Fig. 2 is a diagram showing the configuration of the point cloud data. Fig. 3 is a diagram showing an example of the configuration of a data file in which information on the point cloud data is written.
[0070] Point cloud data includes data on multiple points. The data on each point includes position information (three-dimensional coordinates) and attribute information for that position information. A collection of multiple points is called a point cloud. For example, a point cloud may represent the three-dimensional shape of an object.
[0071] Position information such as three-dimensional coordinates is sometimes called geometry. Data for each point may include attribute information of multiple attribute types. The attribute types may be, for example, color or reflectance.
[0072] One piece of attribute information may be associated with one piece of location information, or multiple pieces of attribute information having different attribute types may be associated with one piece of location information, or multiple pieces of attribute information of the same attribute type may be associated with one piece of location information.
[0073] The configuration example of the data file shown in FIG. 3 is an example in which position information and attribute information correspond one-to-one, and shows the position information and attribute information of N points that make up the point cloud data.
[0074] The position information is, for example, information on three axes, x, y, and z. The attribute information is, for example, RGB color information. A typical data file is a ply file.
[0075] Next, the types of point cloud data will be described. Fig. 4 is a diagram showing the types of point cloud data. As shown in Fig. 4, point cloud data includes static objects and dynamic objects.
[0076] A static object is 3D point cloud data at any time (a certain time). A dynamic object is 3D point cloud data that changes over time. Hereinafter, 3D point cloud data at a certain time will be referred to as a PCC frame, or a frame.
[0077] The object may be a point cloud whose area is restricted to a certain extent, such as ordinary video data, or a large-scale point cloud whose area is not restricted, such as map information.
[0078] Also, there may be point cloud data of various densities, and there may be sparse point cloud data and dense point cloud data.
[0079] The details of each processing unit will be described below. The sensor information is acquired by various methods such as a distance sensor such as a LIDAR or a range finder, a stereo camera, or a combination of multiple monocular cameras. The point cloud data generation unit 4618 generates point cloud data based on the sensor information acquired by the sensor information acquisition unit 4617. The point cloud data generation unit 4618 generates position information as point cloud data, and adds attribute information for the position information to the position information.
[0080] The point cloud data generating unit 4618 may process the point cloud data when generating the position information or adding the attribute information. For example, the point cloud data generating unit 4618 may reduce the amount of data by deleting point clouds with overlapping positions. In addition, the point cloud data generating unit 4618 may convert the position information (position shift, rotation, normalization, etc.) or render the attribute information.
[0081] In FIG. 1, the point cloud data generation system 4611 is included in the three-dimensional data encoding system 4601, but it may be provided independently outside the three-dimensional data encoding system 4601.
[0082] The encoding unit 4613 generates encoded data by encoding the point cloud data based on a predefined encoding method. There are two main types of encoding methods: the first is an encoding method using position information, and this encoding method will be referred to as the first encoding method hereinafter; the second is an encoding method using a video codec, and this encoding method will be referred to as the second encoding method hereinafter.
[0083] The decoding unit 4624 decodes the encoded data based on a predefined encoding method to decode the point cloud data.
[0084] The multiplexing unit 4614 generates multiplexed data by multiplexing the encoded data using an existing multiplexing method. The generated multiplexed data is transmitted or stored. In addition to the PCC encoded data, the multiplexing unit 4614 multiplexes other media such as video, audio, subtitles, applications, and files, or reference time information. The multiplexing unit 4614 may further multiplex attribute information related to sensor information or point cloud data.
[0085] Multiplexing methods or file formats include ISOBMFF, MPEG-DASH, which is an ISOBMFF-based transmission method, MMT, MPEG-2 TS Systems, and RMP.
[0086] The demultiplexer 4623 extracts the PCC encoded data, other media, time information, and the like from the multiplexed data.
[0087] The input / output unit 4615 transmits the multiplexed data using a method suited to a transmission medium or a storage medium, such as broadcasting or communication. The input / output unit 4615 may communicate with other devices via the Internet, or may communicate with a storage unit such as a cloud server.
[0088] The communication protocol used may be http, ftp, TCP, UDP, etc. A PULL type communication method or a PUSH type communication method may be used.
[0089] Either wired transmission or wireless transmission may be used. For wired transmission, Ethernet (registered trademark), USB, RS-232C, HDMI (registered trademark), coaxial cable, etc. are used. For wireless transmission, wireless LAN, Wi-Fi (registered trademark), Bluetooth (registered trademark), millimeter wave, etc. are used.
[0090] As a broadcasting system, for example, DVB-T2, DVB-S2, DVB-C2, ATSC3.0, or ISDB-S3 is used.
[0091] Fig. 5 is a diagram showing a configuration of a first encoding unit 4630 which is an example of the encoding unit 4613 that performs encoding using the first encoding method. Fig. 6 is a block diagram of the first encoding unit 4630. The first encoding unit 4630 generates encoded data (encoded stream) by encoding point cloud data using the first encoding method. The first encoding unit 4630 includes a position information encoding unit 4631, an attribute information encoding unit 4632, an additional information encoding unit 4633, and a multiplexing unit 4634.
[0092] The first encoding unit 4630 has a feature of performing encoding in consideration of a three-dimensional structure. Also, the first encoding unit 4630 has a feature that the attribute information encoding unit 4632 performs encoding using information obtained from the position information encoding unit 4631. The first encoding method is also called GPCC (Geometry based PCC).
[0093] The point cloud data is PCC point cloud data such as a PLY file, or PCC point cloud data generated from sensor information, and includes position information (Position), attribute information (Attribute), and other additional information (MetaData). The position information is input to a position information encoding unit 4631, the attribute information is input to an attribute information encoding unit 4632, and the additional information is input to an additional information encoding unit 4633.
[0094] The position information encoding unit 4631 generates encoded position information (Compressed Geometry) that is encoded data by encoding the position information. For example, the position information encoding unit 4631 encodes the position information using an N-ary tree structure such as an octet tree. Specifically, in the octet tree, the target space is divided into eight nodes (subspaces), and 8-bit information (occupancy code) indicating whether or not a point cloud is included in each node is generated. In addition, the node that includes the point cloud is further divided into eight nodes, and 8-bit information indicating whether or not a point cloud is included in each of the eight nodes is generated. This process is repeated until the number of point clouds included in a predetermined hierarchy or node is equal to or less than a threshold value.
[0095] The attribute information encoding unit 4632 generates encoded attribute information (Compressed Attribute) that is encoded data by encoding using the configuration information generated by the position information encoding unit 4631. For example, the attribute information encoding unit 4632 determines a reference point (reference node) to be referenced in encoding a target point (target node) to be processed based on the octree structure generated by the position information encoding unit 4631. For example, the attribute information encoding unit 4632 refers to a node, among peripheral nodes or adjacent nodes, whose parent node in the octree is the same as that of the target node. Note that the method of determining the reference relationship is not limited to this.
[0096] Furthermore, the encoding process of the attribute information may include at least one of a quantization process, a prediction process, and an arithmetic coding process. In this case, the reference means using a reference node to calculate a predicted value of the attribute information, or using a state of the reference node (e.g., occupancy information indicating whether or not a point group is included in the reference node) to determine a parameter of the encoding. For example, the parameter of the encoding is a quantization parameter in the quantization process, or a context in the arithmetic coding.
[0097] The additional information encoding unit 4633 generates encoded additional information (Compressed MetaData) that is encoded data by encoding compressible data from the additional information.
[0098] The multiplexing unit 4634 multiplexes the encoding position information, the encoding attribute information, the encoding additional information, and other additional information to generate an encoded stream (Compressed Stream) that is encoded data. The generated encoded stream is output to a processing unit of a system layer (not shown).
[0099] Next, a first decoding unit 4640 which is an example of the decoding unit 4624 that performs decoding of the first encoding method will be described. FIG. 7 is a diagram showing a configuration of the first decoding unit 4640. FIG. 8 is a block diagram of the first decoding unit 4640. The first decoding unit 4640 generates point group data by decoding, by the first encoding method, encoded data (encoded stream) encoded by the first encoding method. The first decoding unit 4640 includes a demultiplexing unit 4641, a position information decoding unit 4642, an attribute information decoding unit 4643, and an additional information decoding unit 4644.
[0100] A coded stream (compressed stream), which is coded data, is input to the first decoding unit 4640 from a processing unit in a system layer (not shown).
[0101] The demultiplexer 4641 separates the encoded position information (Compressed Geometry), the encoded attribute information (Compressed Attribute), the encoded additional information (Compressed MetaData), and other additional information from the encoded data.
[0102] The position information decoding unit 4642 generates position information by decoding the encoded position information. For example, the position information decoding unit 4642 restores the position information of a point group represented by three-dimensional coordinates from the encoded position information represented by an N-ary tree structure such as an octree.
[0103] The attribute information decoding unit 4643 decodes the encoded attribute information based on the configuration information generated by the position information decoding unit 4642. For example, the attribute information decoding unit 4643 determines a reference point (reference node) to be referenced in decoding a target point (target node) to be processed based on the octree structure obtained by the position information decoding unit 4642. For example, the attribute information decoding unit 4643 refers to a node, among peripheral nodes or adjacent nodes, whose parent node in the octree is the same as that of the target node. Note that the method of determining the reference relationship is not limited to this.
[0104] In addition, the decoding process of the attribute information may include at least one of an inverse quantization process, a prediction process, and an arithmetic decoding process. In this case, the reference means using a reference node to calculate a predicted value of the attribute information, or using a state of the reference node (e.g., occupancy information indicating whether or not a point group is included in the reference node) to determine a parameter of the decoding. For example, the parameter of the decoding is a quantization parameter in the inverse quantization process, or a context in the arithmetic decoding.
[0105] The additional information decoding unit 4644 generates additional information by decoding the encoded additional information. The first decoding unit 4640 uses the additional information required for decoding the position information and attribute information during decoding, and outputs the additional information required for an application to the outside.
[0106] Next, a configuration example of the position information encoding unit will be described. Fig. 9 is a block diagram of the position information encoding unit 2700 according to this embodiment. The position information encoding unit 2700 includes an octree generating unit 2701, a geometric information calculating unit 2702, a coding table selecting unit 2703, and an entropy encoding unit 2704.
[0107] The octree generating unit 2701 generates, for example, an octree from the input position information, and generates an occupancy code of each node of the octree. The geometric information calculating unit 2702 obtains information indicating whether an adjacent node of the target node is an occupied node. For example, the geometric information calculating unit 2702 calculates occupancy information of the adjacent node (information indicating whether the adjacent node is an occupied node) from the occupancy code of the parent node to which the target node belongs. The geometric information calculating unit 2702 may also store encoded nodes in a list and search for adjacent nodes from the list. The geometric information calculating unit 2702 may also switch adjacent nodes depending on the position of the target node in the parent node.
[0108] The coding table selection unit 2703 selects a coding table to be used for entropy coding of the target node, using the occupancy information of the adjacent nodes calculated by the geometric information calculation unit 2702. For example, the coding table selection unit 2703 may generate a bit string using the occupancy information of the adjacent nodes, and select a coding table for an index number generated from the bit string.
[0109] The entropy coding unit 2704 generates the coding position information and metadata by performing entropy coding on the occupancy code of the target node using the coding table of the selected index number. The entropy coding unit 2704 may add information indicating the selected coding table to the coding position information.
[0110] The octree representation and the scanning order of position information will be described below. Position information (position data) is converted (octreeized) into an octree structure and then encoded. The octree structure is composed of nodes and leaves. Each node has eight nodes or leaves, and each leaf has voxel (VXL) information. Fig. 10 is a diagram showing an example of the structure of position information including multiple voxels. Fig. 11 is a diagram showing an example of the position information shown in Fig. 10 converted into an octree structure. Here, among the leaves shown in Fig. 11, leaves 1, 2, and 3 respectively represent voxels VXL1, VXL2, and VXL3 shown in Fig. 10, and represent a VXL including a point cloud (hereinafter, effective VXL).
[0111] Specifically, node 1 corresponds to the whole space including the position information in FIG. 10. The whole space corresponding to node 1 is divided into eight nodes, and among the eight nodes, the node including the valid VXL is further divided into eight nodes or leaves, and this process is repeated for the number of levels of the tree structure. Here, each node corresponds to a subspace, and has information (occupancy code) indicating at what position the next node or leaf will be located after division as node information. Also, the block at the bottom level is set as a leaf, and the number of point clouds included in the leaf, etc., is held as leaf information.
[0112] Next, a configuration example of the position information decoding unit will be described. Fig. 12 is a block diagram of the position information decoding unit 2710 according to this embodiment. The position information decoding unit 2710 includes an octree generating unit 2711, a geometric information calculating unit 2712, a coding table selecting unit 2713, and an entropy decoding unit 2714.
[0113] The octree generating unit 2711 generates an octree of a certain space (node) using header information or metadata of a bit stream. For example, the octree generating unit 2711 generates a large space (root node) using the size of the x-axis, y-axis, and z-axis directions of a certain space added to the header information, and generates an octree by dividing the space into two in the x-axis, y-axis, and z-axis directions to generate eight small spaces A (nodes A0 to A7). In addition, nodes A0 to A7 are set in order as target nodes.
[0114] The geometric information calculation unit 2712 obtains occupancy information indicating whether an adjacent node of the target node is an occupancy node. For example, the geometric information calculation unit 2712 calculates the occupancy information of the adjacent node from the occupancy code of the parent node to which the target node belongs. The geometric information calculation unit 2712 may also store decoded nodes in a list and search for adjacent nodes from the list. The geometric information calculation unit 2712 may also switch adjacent nodes depending on the position of the target node in the parent node.
[0115] The coding table selection unit 2713 selects a coding table (decoding table) to be used for entropy decoding of the target node using the occupancy information of the adjacent node calculated by the geometric information calculation unit 2712. For example, the coding table selection unit 2713 may generate a bit string using the occupancy information of the adjacent node, and select a coding table for an index number generated from the bit string.
[0116] The entropy decoding unit 2714 generates position information by entropy decoding the occupancy code of the target node using the selected coding table. Note that the entropy decoding unit 2714 may obtain information on the selected coding table by decoding from the bit stream, and entropy decode the occupancy code of the target node using the coding table indicated by the information.
[0117] The configurations of the attribute information encoding unit and the attribute information decoding unit will be described below. Fig. 13 is a block diagram showing an example of the configuration of the attribute information encoding unit A100. The attribute information encoding unit may include a plurality of encoding units that execute different encoding methods. For example, the attribute information encoding unit may switch between the following two methods depending on the use case.
[0118] The attribute information encoding unit A100 includes an LoD attribute information encoding unit A101 and a conversion attribute information encoding unit A102. The LoD attribute information encoding unit A101 classifies each 3D point into a plurality of layers using the position information of the 3D point, predicts the attribute information of the 3D point belonging to each layer, and encodes the prediction residual. Here, each classified layer is called LoD (Level of Detail).
[0119] The transformed attribute information encoding unit A102 encodes the attribute information using RAHT (Region Adaptive Hierarchical Transform). Specifically, the transformed attribute information encoding unit A102 applies RAHT or Haar transform to each piece of attribute information based on the position information of the three-dimensional point to generate high-frequency components and low-frequency components of each layer, and encodes the values using quantization, entropy coding, or the like.
[0120] 14 is a block diagram showing a configuration example of the attribute information decoding unit A110. The attribute information decoding unit may include multiple decoding units that execute different decoding methods. For example, the attribute information decoding unit may switch between the following two methods for decoding based on information included in the header or metadata.
[0121] The attribute information decoding unit A110 includes a LoD attribute information decoding unit A111 and a converted attribute information decoding unit A112. The LoD attribute information decoding unit A111 classifies each 3D point into a plurality of hierarchical levels using the position information of the 3D point, and decodes the attribute value while predicting the attribute information of the 3D point belonging to each hierarchical level.
[0122] The transformed attribute information decoding unit A112 decodes the attribute information using RAHT (Region Adaptive Hierarchical Transform). Specifically, the transformed attribute information decoding unit A112 decodes the attribute value by applying inverse RAHT or inverse Haar transform to high-frequency components and low-frequency components of each attribute value based on the position information of the three-dimensional point.
[0123] FIG. 15 is a block diagram showing a configuration of an attribute information encoding unit 3140 which is an example of the LoD attribute information encoding unit A101.
[0124] The attribute information encoding unit 3140 includes a LoD generation unit 3141, a surrounding search unit 3142, a prediction unit 3143, a prediction residual calculation unit 3144, a quantization unit 3145, an arithmetic encoding unit 3146, an inverse quantization unit 3147, a decoded value generation unit 3148, and a memory 3149.
[0125] The LoD generation unit 3141 generates LoD using the position information of the three-dimensional points.
[0126] The surrounding search unit 3142 searches for nearby 3D points adjacent to each 3D point, using the LoD generation result by the LoD generation unit 3141 and distance information indicating the distance between each 3D point.
[0127] The prediction unit 3143 generates a predicted value of the attribute information of the target 3D point to be encoded.
[0128] The prediction residual calculation unit 3144 calculates (generates) a prediction residual of the predicted value of the attribute information generated by the prediction unit 3143.
[0129] The quantization unit 3145 quantizes the prediction residual of the attribute information calculated by the prediction residual calculation unit 3144 .
[0130] The arithmetic coding unit 3146 arithmetically codes the prediction residuals after being quantized by the quantization unit 3145. The arithmetic coding unit 3146 outputs a bit stream including the arithmetically coded prediction residuals to, for example, a three-dimensional data decoding device.
[0131] Note that the prediction residual may be binarized by, for example, the quantization unit 3145 before being arithmetically coded by the arithmetic coding unit 3146.
[0132] Also, for example, the arithmetic coding unit 3146 may initialize a coding table used for arithmetic coding before arithmetic coding. The arithmetic coding unit 3146 may initialize a coding table used for arithmetic coding for each layer. Also, the arithmetic coding unit 3146 may output information indicating the position of the layer for which the coding table has been initialized, by including it in the bitstream.
[0133] The inverse quantization unit 3147 inverse quantizes the prediction residual after being quantized by the quantization unit 3145 .
[0134] The decoded value generation unit 3148 generates a decoded value by adding the predicted value of the attribute information generated by the prediction unit 3143 and the prediction residual after inverse quantization by the inverse quantization unit 3147.
[0135] The memory 3149 is a memory that stores the decoded values of the attribute information of each 3D point decoded by the decoded value generation unit 3148. For example, when generating a predicted value of a 3D point that has not yet been encoded, the prediction unit 3143 generates the predicted value by using the decoded value of the attribute information of each 3D point stored in the memory 3149.
[0136] 16 is a block diagram of an attribute information encoding unit 6600 which is an example of the transformed attribute information encoding unit A102. The attribute information encoding unit 6600 includes a sorting unit 6601, a Haar transform unit 6602, a quantization unit 6603, an inverse quantization unit 6604, an inverse Haar transform unit 6605, a memory 6606, and an arithmetic encoding unit 6607.
[0137] The sorting unit 6601 generates a Morton code using the position information of the 3D points, and sorts the multiple 3D points in Morton code order. The Haar transform unit 6602 generates coding coefficients by applying a Haar transform to the attribute information. The quantization unit 6603 quantizes the coding coefficients of the attribute information.
[0138] The inverse quantization unit 6604 inversely quantizes the quantized coding coefficients. The inverse Haar transform unit 6605 applies an inverse Haar transform to the coding coefficients. The memory 6606 stores values of attribute information of multiple decoded 3D points. For example, the attribute information of the decoded 3D points stored in the memory 6606 may be used for predicting uncoded 3D points.
[0139] The arithmetic coding unit 6607 calculates ZeroCnt from the coding coefficients after quantization, and arithmetically codes the ZeroCnt. Furthermore, the arithmetic coding unit 6607 arithmetically codes the non-zero coding coefficients after quantization. The arithmetic coding unit 6607 may binarize the coding coefficients before arithmetic coding. Furthermore, the arithmetic coding unit 6607 may generate and code various header information.
[0140] FIG. 17 is a block diagram showing a configuration of an attribute information decoding unit 3150 which is an example of the LoD attribute information decoding unit A111.
[0141] The attribute information decoding unit 3150 includes an LoD generation unit 3151 , a surrounding search unit 3152 , a prediction unit 3153 , an arithmetic decoding unit 3154 , an inverse quantization unit 3155 , a decoded value generation unit 3156 , and a memory 3157 .
[0142] The LoD generation unit 3151 generates LoD using position information of the 3D points decoded by a position information decoding unit (not shown in FIG. 17).
[0143] The surrounding search unit 3152 searches for nearby 3D points adjacent to each 3D point, using the LoD generation result by the LoD generation unit 3151 and distance information indicating the distance between each 3D point.
[0144] The prediction unit 3153 generates a predicted value of the attribute information of the target 3D point to be decoded.
[0145] The arithmetic decoding unit 3154 arithmetically decodes the prediction residual in the bitstream acquired from the attribute information encoding unit 3140 shown in FIG. 15. The arithmetic decoding unit 3154 may initialize a decoding table used for arithmetic decoding. The arithmetic decoding unit 3154 initializes a decoding table used for arithmetic decoding for a layer on which the arithmetic encoding unit 3146 shown in FIG. 15 has performed an encoding process. The arithmetic decoding unit 3154 may initialize a decoding table used for arithmetic decoding for each layer. The arithmetic decoding unit 3154 may initialize the decoding table based on information included in the bitstream and indicating the position of the layer for which the encoding table has been initialized.
[0146] The inverse quantization unit 3155 inverse quantizes the prediction residual arithmetically decoded by the arithmetic decoding unit 3154 .
[0147] The decoded value generation unit 3156 generates a decoded value by adding the predicted value generated by the prediction unit 3153 and the prediction residual after inverse quantization by the inverse quantization unit 3155. The decoded value generation unit 3156 outputs the decoded attribute information data to another device.
[0148] The memory 3157 is a memory that stores the decoded values of the attribute information of each 3D point decoded by the decoded value generation unit 3156. For example, when generating a predicted value of a 3D point that has not yet been decoded, the prediction unit 3153 generates the predicted value by using the decoded value of the attribute information of each 3D point stored in the memory 3157.
[0149] 18 is a block diagram of an attribute information decoding unit 6610 which is an example of the transformed attribute information decoding unit A112. The attribute information decoding unit 6610 includes an arithmetic decoding unit 6611, an inverse quantization unit 6612, an inverse Haar transform unit 6613, and a memory 6614.
[0150] The arithmetic decoding unit 6611 arithmetically decodes the ZeroCnt and the coding coefficients included in the bit stream. Note that the arithmetic decoding unit 6611 may also decode various types of header information.
[0151] The inverse quantization unit 6612 inverse quantizes the arithmetically decoded coding coefficients. The inverse Haar transform unit 6613 applies an inverse Haar transform to the coding coefficients after the inverse quantization. The memory 6614 stores values of attribute information of multiple decoded 3D points. For example, the attribute information of the decoded 3D points stored in the memory 6614 may be used to predict undecoded 3D points.
[0152] Next, a second encoding unit 4650, which is an example of the encoding unit 4613 that performs encoding by the second encoding method, will be described. Fig. 19 is a diagram showing a configuration of the second encoding unit 4650. Fig. 20 is a block diagram of the second encoding unit 4650.
[0153] The second encoding unit 4650 generates encoded data (encoded stream) by encoding the point cloud data by a second encoding method. The second encoding unit 4650 includes an additional information generating unit 4651, a position image generating unit 4652, an attribute image generating unit 4653, a video encoding unit 4654, an additional information encoding unit 4655, and a multiplexing unit 4656.
[0154] The second encoding unit 4650 has a feature of generating a position image and an attribute image by projecting a three-dimensional structure onto a two-dimensional image, and encoding the generated position image and attribute image using an existing video encoding method. The second encoding method is also called VPCC (Video based PCC).
[0155] The point cloud data is PCC point cloud data such as a PLY file, or PCC point cloud data generated from sensor information, and includes position information (Position), attribute information (Attribute), and other additional information (MetaData).
[0156] The additional information generating unit 4651 generates map information of a plurality of two-dimensional images by projecting a three-dimensional structure onto the two-dimensional images.
[0157] The position image generating unit 4652 generates a position image (Geometry Image) based on the position information and the map information generated by the additional information generating unit 4651. This position image is, for example, a distance image in which distance (Depth) is indicated as a pixel value. Note that this distance image may be an image in which a plurality of point clouds are viewed from one viewpoint (an image in which a plurality of point clouds are projected onto one two-dimensional plane), a plurality of images in which a plurality of point clouds are viewed from multiple viewpoints, or a single image in which these multiple images are integrated.
[0158] The attribute image generating unit 4653 generates an attribute image based on the attribute information and the map information generated by the additional information generating unit 4651. This attribute image is, for example, an image in which the attribute information (for example, color (RGB)) is indicated as pixel values. Note that this image may be an image in which multiple point clouds are viewed from one viewpoint (an image in which multiple point clouds are projected onto one two-dimensional plane), or multiple images in which multiple point clouds are viewed from multiple viewpoints, or a single image in which these multiple images are integrated.
[0159] The video encoding unit 4654 generates an encoded position image (Compressed Geometry Image) and an encoded attribute image (Compressed Attribute Image), which are encoded data, by encoding the position image and the attribute image using a video encoding method. Note that any known encoding method may be used as the video encoding method. For example, the video encoding method is AVC, HEVC, or the like.
[0160] The additional information encoding unit 4655 generates encoded additional information (Compressed MetaData) by encoding the additional information and map information included in the point cloud data.
[0161] The multiplexing unit 4656 generates a compressed stream, which is encoded data, by multiplexing the encoded position image, the encoded attribute image, the encoded additional information, and other additional information. The generated compressed stream is output to a processing unit of a system layer (not shown).
[0162] Next, a second decoding unit 4660 which is an example of the decoding unit 4624 that performs decoding of the second encoding method will be described. FIG. 21 is a diagram showing a configuration of the second decoding unit 4660. FIG. 22 is a block diagram of the second decoding unit 4660. The second decoding unit 4660 generates point group data by decoding, by the second encoding method, encoded data (encoded stream) encoded by the second encoding method. The second decoding unit 4660 includes a demultiplexing unit 4661, a video decoding unit 4662, an additional information decoding unit 4663, a position information generating unit 4664, and an attribute information generating unit 4665.
[0163] A compressed stream, which is encoded data, is input to the second decoding unit 4660 from a processing unit in a system layer (not shown).
[0164] The demultiplexer 4661 separates the encoded position image (Compressed Geometry Image), the encoded attribute image (Compressed Attribute Image), the encoded additional information (Compressed MetaData), and other additional information from the encoded data.
[0165] The video decoding unit 4662 generates a position image and an attribute image by decoding the encoded position image and the encoded attribute image using a video encoding method. Note that any known encoding method may be used as the video encoding method. For example, the video encoding method is AVC or HEVC.
[0166] The additional information decoding unit 4663 decodes the encoded additional information to generate additional information including map information and the like.
[0167] The position information generating unit 4664 generates position information using the position image and map information. The attribute information generating unit 4665 generates attribute information using the attribute image and map information.
[0168] The second decoding unit 4660 uses the additional information necessary for decoding during decoding, and outputs the additional information necessary for an application to the outside.
[0169] Problems with the PCC encoding method will be described below. Fig. 23 is a diagram showing a protocol stack related to PCC encoded data. Fig. 23 shows an example in which other media data such as video (e.g., HEVC) or audio is multiplexed with PCC encoded data and transmitted or stored.
[0170] The multiplexing method and file format have the function of multiplexing various encoded data and transmitting or storing them. In order to transmit or store the encoded data, the encoded data must be converted into the format of the multiplexing method. For example, HEVC specifies a technology for storing encoded data in a data structure called a NAL unit and storing the NAL unit in ISOBMFF.
[0171] Meanwhile, a first encoding method (Codec1) and a second encoding method (Codec2) are currently being considered as methods for encoding point cloud data. However, the structure of the encoded data and the method for storing the encoded data in a system format have not been defined, which poses the problem that MUX processing (multiplexing) in the encoding unit, transmission, and storage cannot be performed as is.
[0172] In the following description, unless a specific encoding method is specified, it refers to either the first encoding method or the second encoding method.
[0173] (Embodiment 2) In this embodiment, the types of encoded data (position information (Geometry), attribute information (Attribute), additional information (Metadata)) generated by the above-mentioned first encoding unit 4630 or second encoding unit 4650, a generation method of the additional information (metadata), and multiplexing processing in the multiplexing unit will be described. Note that the additional information (metadata) may also be referred to as a parameter set or control information.
[0174] In this embodiment, the dynamic object (three-dimensional point cloud data that changes over time) described in Figure 4 will be used as an example, but a similar method may also be used in the case of a static object (three-dimensional point cloud data at any time).
[0175] 24 is a diagram showing configurations of an encoding unit 4801 and a multiplexing unit 4802 included in the three-dimensional data encoding device according to this embodiment. The encoding unit 4801 corresponds to, for example, the first encoding unit 4630 or the second encoding unit 4650 described above. The multiplexing unit 4802 corresponds to the multiplexing unit 4634 or 4656 described above.
[0176] The encoding unit 4801 encodes point cloud data of multiple PCC (Point Cloud Compression) frames, and generates encoded data (Multiple Compressed Data) of multiple pieces of position information, attribute information, and additional information.
[0177] The multiplexing unit 4802 converts data of multiple data types (position information, attribute information, and additional information) into NAL units, thereby converting the data into a data structure that takes into consideration data access in the decoding device.
[0178] 25 is a diagram showing an example of the structure of coded data generated by the coding unit 4801. The arrows in the diagram indicate dependencies related to the decoding of coded data, with the source of the arrow depending on the data at the tip of the arrow. In other words, the decoding device decodes the data at the tip of the arrow, and uses the decoded data to decode the data at the tip of the arrow. In other words, dependency means that the data at the dependency destination is referenced (used) in the processing (encoding, decoding, etc.) of the data at the dependency destination.
[0179] First, the process of generating encoded data of position information will be described. The encoding unit 4801 generates encoded position data (compressed geometry data) for each frame by encoding the position information of each frame. The encoded position data is represented as G(i), where i indicates the frame number, the time of the frame, etc.
[0180] The encoding unit 4801 also generates a position parameter set (GPS(i)) corresponding to each frame. The position parameter set includes parameters that can be used to decode the encoded position data. The encoded position data for each frame depends on the corresponding position parameter set.
[0181] Moreover, encoded position data consisting of multiple frames is defined as a position sequence (Geometry Sequence). The encoding unit 4801 generates a position sequence parameter set (Geometry Sequence PS: also written as position SPS) that stores parameters commonly used in decoding processes for multiple frames in the position sequence. The position sequence depends on the position SPS.
[0182] Next, a process for generating encoded data of attribute information will be described. Encoding unit 4801 generates encoded attribute data (Compressed Attribute Data) for each frame by encoding the attribute information of each frame. The encoded attribute data is represented as A(i). FIG. 25 shows an example in which attribute X and attribute Y exist, and the encoded attribute data of attribute X is represented as AX(i) and the encoded attribute data of attribute Y is represented as AY(i).
[0183] The encoding unit 4801 also generates an attribute parameter set (APS(i)) corresponding to each frame. The attribute parameter set for attribute X is represented as AXPS(i), and the attribute parameter set for attribute Y is represented as AYPS(i). The attribute parameter set includes parameters that can be used to decode the encoded attribute information. The encoded attribute data depends on the corresponding attribute parameter set.
[0184] Moreover, encoded attribute data consisting of multiple frames is defined as an attribute sequence. The encoding unit 4801 generates an attribute sequence parameter set (Attribute Sequence PS: also referred to as attribute SPS) that stores parameters commonly used in decoding processes for multiple frames in the attribute sequence. The attribute sequence depends on the attribute SPS.
[0185] Furthermore, in the first encoding method, the encoded attribute data depends on the encoded position data.
[0186] 25 shows an example in which two types of attribute information (attribute X and attribute Y) exist. When there are two types of attribute information, for example, two encoding units generate respective data and metadata. Also, for example, an attribute sequence is defined for each type of attribute information, and an attribute SPS is generated for each type of attribute information.
[0187] In addition, while FIG. 25 shows an example in which there is one type of position information and two types of attribute information, the present invention is not limited to this, and there may be one type of attribute information, or three or more types. In this case, the encoded data can be generated in a similar manner. In addition, in the case of point cloud data that does not have attribute information, the attribute information may not be necessary. In this case, the encoding unit 4801 does not need to generate a parameter set related to the attribute information.
[0188] Next, a process of generating additional information (metadata) will be described. The encoding unit 4801 generates a PCC stream PS (also written as stream PS), which is a parameter set for the entire PCC stream. The encoding unit 4801 stores parameters that can be commonly used in decoding processes for one or more position sequences and one or more attribute sequences in the stream PS. For example, the stream PS includes identification information indicating the codec of the point cloud data, information indicating the algorithm used for encoding, and the like. The position sequence and the attribute sequence depend on the stream PS.
[0189] Next, an access unit and a GOF will be described. In this embodiment, the concept of an access unit (AU) and a group of frames (GOF) are newly introduced.
[0190] An access unit is a basic unit for accessing data during decoding, and is composed of one or more pieces of data and one or more pieces of metadata. For example, an access unit is composed of position information at the same time and one or more pieces of attribute information. A GOF is a random access unit, and is composed of one or more access units.
[0191] The encoding unit 4801 generates an access unit header (AU Header) as identification information indicating the beginning of an access unit. The encoding unit 4801 stores parameters related to the access unit in the access unit header. For example, the access unit header includes the configuration or information of the encoded data included in the access unit. The access unit header also includes parameters commonly used for data included in the access unit, such as parameters related to the decoding of the encoded data.
[0192] Alternatively, instead of an access unit header, the encoding unit 4801 may generate an access unit delimiter that does not include parameters related to the access unit. This access unit delimiter is used as identification information indicating the beginning of the access unit. The decoding device identifies the beginning of the access unit by detecting the access unit header or the access unit delimiter.
[0193] Next, generation of identification information for the start of a GOF will be described. The encoding unit 4801 generates a GOF header as identification information indicating the start of a GOF. The encoding unit 4801 stores parameters related to the GOF in the GOF header. For example, the GOF header includes the configuration or information of the encoded data included in the GOF. The GOF header also includes parameters commonly used for data included in the GOF, such as parameters related to the decoding of the encoded data.
[0194] Instead of a GOF header, the encoding unit 4801 may generate a GOF delimiter that does not include parameters related to the GOF. This GOF delimiter is used as identification information indicating the beginning of the GOF. The decoding device identifies the beginning of the GOF by detecting the GOF header or the GOF delimiter.
[0195] In PCC encoded data, for example, an access unit is defined as a PCC frame unit, and a decoding device accesses a PCC frame based on identification information at the beginning of the access unit.
[0196] Also, for example, GOF is defined as one random access unit. The decoding device accesses the random access unit based on the identification information at the beginning of the GOF. For example, if PCC frames are not dependent on each other and can be decoded independently, the PCC frames may be defined as the random access unit.
[0197] Note that two or more PCC frames may be assigned to one access unit, and multiple random access units may be assigned to one GOF.
[0198] Furthermore, the encoding unit 4801 may define and generate parameter sets or metadata other than those described above. For example, the encoding unit 4801 may generate SEI (Supplemental Enhancement Information) that stores parameters (optional parameters) that may not necessarily be used during decoding.
[0199] Next, the structure of the coded data and the method of storing the coded data in the NAL unit will be described.
[0200] For example, a data format is defined for each type of encoded data. Figure 26 shows an example of encoded data and NAL units.
[0201] For example, as shown in Fig. 26, the encoded data includes a header and a payload. The encoded data may include length information indicating the length (amount of data) of the encoded data, the header, or the payload. The encoded data may not include a header.
[0202] The header includes, for example, identification information for identifying the data, which indicates, for example, the data type or the frame number.
[0203] The header includes, for example, identification information indicating a reference relationship. This identification information is stored in the header when, for example, there is a dependency between data, and is information for referring to the reference destination from the reference source. For example, the header of the reference destination includes identification information for identifying the data. The header of the reference source includes identification information indicating the reference destination.
[0204] In addition, when the reference destination or the reference source can be identified or derived from other information, the identification information for specifying the data or the identification information indicating the reference relationship may be omitted.
[0205] The multiplexing unit 4802 stores the encoded data in the payload of the NAL unit. The NAL unit header includes pcc_nal_unit_type, which is identification information of the encoded data. Figure 27 shows an example of the semantics of pcc_nal_unit_type.
[0206] 27, when pcc_codec_type is codec 1 (Codec1: first encoding method), values 0 to 10 of pcc_nal_unit_type are assigned to the encoded position data (Geometry), encoded attribute X data (AttributeX), encoded attribute Y data (AttributeY), position PS (Geom.PS), attribute XPS (AttrX.PS), attribute YPS (AttrX.PS), position SPS (Geometry Sequence PS), attribute XSPS (AttributeX Sequence PS), attribute YSPS (AttributeY Sequence PS), AU header (AU Header), and GOF header (GOF Header) in codec 1. Values 11 and above are assigned as spares for codec 1.
[0207] When pcc_codec_type is Codec2 (Codec2: second encoding method), values 0 to 2 of pcc_nal_unit_type are assigned to codec data A (DataA), metadata A (MetaDataA), and metadata B (MetaDataB). Values 3 and above are assigned as spares for Codec2.
[0208] (Embodiment 3) HEVC encoding has data division tools such as slicing or tiling to enable parallel processing in the decoding device, but PCC (Point Cloud Compression) encoding does not yet have such tools.
[0209] In PCC, various data division methods can be considered depending on parallel processing, compression efficiency, and compression algorithms. This section describes the definitions of slices and tiles, the data structure, and the transmission and reception methods.
[0210] 28 is a block diagram showing a configuration of a first encoding unit 4910 included in the three-dimensional data encoding device according to this embodiment. The first encoding unit 4910 generates encoded data (encoded stream) by encoding point cloud data using a first encoding method (GPCC (Geometry based PCC)). The first encoding unit 4910 includes a division unit 4911, a plurality of position information encoding units 4912, a plurality of attribute information encoding units 4913, an additional information encoding unit 4914, and a multiplexing unit 4915.
[0211] The division unit 4911 generates a plurality of divided data by dividing the point cloud data. Specifically, the division unit 4911 generates a plurality of divided data by dividing the space of the point cloud data into a plurality of subspaces. Here, the subspace is one of a tile and a slice, or a combination of a tile and a slice. More specifically, the point cloud data includes position information, attribute information, and additional information. The division unit 4911 divides the position information into a plurality of divided position information, and divides the attribute information into a plurality of divided attribute information. In addition, the division unit 4911 generates additional information related to the division.
[0212] The position information encoding units 4912 generate a plurality of pieces of encoded position information by encoding the plurality of pieces of divided position information. For example, the position information encoding units 4912 process the plurality of pieces of divided position information in parallel.
[0213] The attribute information encoding units 4913 generate a plurality of pieces of encoded attribute information by encoding the plurality of pieces of divided attribute information. For example, the attribute information encoding units 4913 process the plurality of pieces of divided attribute information in parallel.
[0214] The additional information encoding unit 4914 generates encoded additional information by encoding the additional information included in the point cloud data and the additional information related to the data division generated by the division unit 4911 at the time of division.
[0215] The multiplexing unit 4915 multiplexes a plurality of pieces of encoding position information, a plurality of pieces of encoding attribute information, and encoding additional information to generate encoded data (encoded stream), and transmits the generated encoded data. In addition, the encoded additional information is used during decoding.
[0216] 28 shows an example in which there are two position information encoders 4912 and two attribute information encoders 4913, but the number of position information encoders 4912 and two attribute information encoders 4913 may be one, or three or more. Furthermore, multiple pieces of divided data may be processed in parallel within the same chip, such as multiple cores in a CPU, or may be processed in parallel by cores on multiple chips, or may be processed in parallel by multiple cores on multiple chips.
[0217] 29 is a block diagram showing a configuration of the first decoding unit 4920. The first decoding unit 4920 restores the point cloud data by decoding the encoded data (encoded stream) generated by encoding the point cloud data by the first encoding method (GPCC). The first decoding unit 4920 includes a demultiplexing unit 4921, a plurality of position information decoding units 4922, a plurality of attribute information decoding units 4923, an additional information decoding unit 4924, and a combining unit 4925.
[0218] The demultiplexer 4921 demultiplexes the coded data (coded stream) to generate a plurality of pieces of coding position information, a plurality of pieces of coding attribute information, and coded additional information.
[0219] The position information decoders 4922 generate a plurality of pieces of divided position information by decoding the plurality of pieces of encoded position information. For example, the position information decoders 4922 process the plurality of pieces of encoded position information in parallel.
[0220] The multiple attribute information decoding unit 4923 generates multiple pieces of divided attribute information by decoding the multiple pieces of encoded attribute information. For example, the multiple attribute information decoding unit 4923 processes the multiple pieces of encoded attribute information in parallel.
[0221] The multiple additional information decoders 4924 generate additional information by decoding the encoded additional information.
[0222] The combining unit 4925 generates position information by combining a plurality of pieces of divided position information using the additional information The combining unit 4925 generates attribute information by combining a plurality of pieces of divided attribute information using the additional information.
[0223] 29 shows an example in which the number of each of the position information decoding units 4922 and the attribute information decoding units 4923 is two, but the number of each of the position information decoding units 4922 and the attribute information decoding units 4923 may be one, or three or more. Furthermore, multiple pieces of divided data may be processed in parallel within the same chip, such as multiple cores within a CPU, or may be processed in parallel by cores on multiple chips, or may be processed in parallel by multiple cores on multiple chips.
[0224] Next, it is described how the dividing unit 4911 has a configuration. Fig. 30 is a block diagram of the dividing unit 4911. The dividing unit 4911 includes a slice dividing unit 4931, a position information tile dividing unit 4932, and an attribute information tile dividing unit 4933.
[0225] The slice division unit 4931 generates a plurality of slice position information by dividing position information (Position (Geometry)) into slices. The slice division unit 4931 also generates a plurality of slice attribute information by dividing attribute information (Attribute) into slices. The slice division unit 4931 also outputs slice additional information (SliceMetaData) including information related to slice division and information generated in the slice division.
[0226] The position information tile division unit 4932 divides a plurality of slice position information pieces into tiles to generate a plurality of pieces of divided position information pieces (a plurality of pieces of tile position information pieces). In addition, the position information tile division unit 4932 outputs position tile additional information (Geometry Tile MetaData) including information related to the tile division of the position information pieces and information generated in the tile division of the position information pieces.
[0227] The attribute information tile division unit 4933 divides a plurality of slice attribute information pieces into tiles to generate a plurality of pieces of divided attribute information (a plurality of pieces of tile attribute information). In addition, the attribute information tile division unit 4933 outputs attribute tile additional information (Attribute Tile MetaData) including information related to the tile division of the attribute information and information generated in the tile division of the attribute information.
[0228] The number of slices or tiles to be divided is equal to or greater than 1. In other words, it is not necessary to divide the slices or tiles.
[0229] Although an example in which tile division is performed after slice division has been shown here, slice division may be performed after tile division. Furthermore, a new division type may be defined in addition to slices and tiles, and division may be performed using three or more division types.
[0230] A method for dividing point cloud data will be described below. Fig. 31 is a diagram showing an example of division into slices and tiles.
[0231] First, a method of dividing into slices will be described. The dividing unit 4911 divides the three-dimensional point cloud data into arbitrary point clouds in slice units. In the slice division, the dividing unit 4911 does not divide the position information and attribute information constituting a point, but divides the position information and attribute information together. That is, the dividing unit 4911 divides into slices so that the position information and attribute information at an arbitrary point belong to the same slice. In addition, according to these, the number of divisions and the division method may be any method. In addition, the minimum unit of division is a point. For example, the number of divisions of the position information and the attribute information are the same. For example, the three-dimensional point corresponding to the position information after the slice division and the three-dimensional point corresponding to the attribute information are included in the same slice.
[0232] Furthermore, the division unit 4911 generates slice additional information, which is additional information related to the number of divisions and the division method when dividing a slice. The slice additional information is the same as the position information and the attribute information. For example, the slice additional information includes information indicating the reference coordinate position, size, or side length of the bounding box after division. Furthermore, the slice additional information includes information indicating the number of divisions, the division type, etc.
[0233] Next, a tile division method will be described. The division unit 4911 divides the data divided into slices into slice position information (G slices) and slice attribute information (A slices), and divides each of the slice position information and the slice attribute information into tiles.
[0234] Although FIG. 31 shows an example of division using an octree structure, the number of divisions and the division method may be any method.
[0235] Furthermore, the division unit 4911 may divide the position information and the attribute information using different division methods or the same division method. Furthermore, the division unit 4911 may divide a plurality of slices into tiles using different division methods or the same division method.
[0236] Furthermore, the division unit 4911 generates tile additional information related to the number of divisions and the division method when dividing tiles. The tile additional information (position tile additional information and attribute tile additional information) is independent of position information and attribute information. For example, the tile additional information includes information indicating the reference coordinate position, size, or side length of the bounding box after division. The tile additional information also includes information indicating the number of divisions, the division type, etc.
[0237] Next, an example of a method for dividing the point cloud data into slices or tiles will be described. The dividing unit 4911 may use a predetermined method as a method for dividing the point cloud data into slices or tiles, or may adaptively switch the method to be used depending on the point cloud data.
[0238] When dividing into slices, the dividing unit 4911 divides the three-dimensional space collectively based on the position information and the attribute information. For example, the dividing unit 4911 determines the shape of an object and divides the three-dimensional space into slices according to the shape of the object. For example, the dividing unit 4911 extracts objects such as trees or buildings, and divides the space on an object-by-object basis. For example, the dividing unit 4911 divides the space into slices such that one or more objects are entirely included in one slice. Alternatively, the dividing unit 4911 divides one object into multiple slices.
[0239] In this case, the encoding device may change the encoding method for each slice, for example. For example, the encoding device may use a high-quality compression method for a specific object or a specific part of an object. In this case, the encoding device may store information indicating the encoding method for each slice in additional information (metadata).
[0240] Furthermore, the dividing unit 4911 may divide the image into slices based on map information or location information so that each slice corresponds to a predetermined coordinate space.
[0241] When dividing into tiles, the dividing unit 4911 divides the position information and the attribute information independently. For example, the dividing unit 4911 divides a slice into tiles according to the amount of data or the amount of processing. For example, the dividing unit 4911 determines whether the amount of data of a slice (for example, the number of three-dimensional points included in a slice) is greater than a predetermined threshold. If the amount of data of a slice is greater than the threshold, the dividing unit 4911 divides the slice into tiles. If the amount of data of a slice is less than the threshold, the dividing unit 4911 does not divide the slice into tiles.
[0242] For example, the division unit 4911 divides the slice into tiles so that the processing amount or processing time in the decoding device is within a certain range (a predetermined value or less). This makes the processing amount per tile in the decoding device constant, facilitating distributed processing in the decoding device.
[0243] Furthermore, when the processing amount differs between the position information and the attribute information, for example when the processing amount of the position information is greater than the processing amount of the attribute information, dividing section 4911 divides the position information into a greater number of parts than the attribute information.
[0244] Also, for example, in the case where the decoding device may decode and display the position information quickly and the attribute information may be decoded and displayed slowly later depending on the content, the division unit 4911 may divide the position information into a larger number of parts than the attribute information. This allows the decoding device to process the position information in parallel in a larger number, so that the position information can be processed faster than the attribute information.
[0245] Note that the decoding device does not necessarily need to process sliced or tiled data in parallel, and may determine whether to process the data in parallel depending on the number or capabilities of the decoding processing units.
[0246] By dividing the image in the above manner, it is possible to realize adaptive encoding according to the content or object. In addition, it is possible to realize parallel processing in the decoding process. This improves the flexibility of the point cloud encoding system or the point cloud decoding system.
[0247] Fig. 32 is a diagram showing an example of a division pattern of slices and tiles. DU in the diagram is a data unit (DataUnit) and indicates data of a tile or slice. Each DU includes a slice index (SliceIndex) and a tile index (TileIndex). The numerical value in the upper right of the DU in the diagram indicates the slice index, and the numerical value in the lower left of the DU indicates the tile index.
[0248] In pattern 1, the number of divisions and the division method are the same for G slices and A slices in slice division. In tile division, the number of divisions and the division method for G slices are different from the number of divisions and the division method for A slices. Furthermore, the same number of divisions and the division method are used among multiple G slices. The same number of divisions and the division method are used among multiple A slices.
[0249] In pattern 2, the number of divisions and the division method are the same for G slices and A slices in slice division. In tile division, the number of divisions and the division method for G slices are different from the number of divisions and the division method for A slices. In addition, the number of divisions and the division method are different between multiple G slices. The number of divisions and the division method are different between multiple A slices.
[0250] Next, a method for encoding divided data will be described. A three-dimensional data encoding device (first encoding unit 4910) encodes each of the divided data. When encoding attribute information, the three-dimensional data encoding device generates dependency information indicating which configuration information (position information, additional information, or other attribute information) was used for encoding as additional information. That is, the dependency information indicates, for example, the configuration information of the reference destination (dependency destination). In this case, the three-dimensional data encoding device generates dependency information based on configuration information corresponding to the division shape of the attribute information. Note that the three-dimensional data encoding device may generate dependency information based on configuration information corresponding to a plurality of division shapes.
[0251] The dependency information may be generated by the three-dimensional data encoding device, and the generated dependency information may be sent to the three-dimensional data decoding device. Alternatively, the three-dimensional data decoding device may generate the dependency information, and the three-dimensional data encoding device may not send the dependency information. Also, the dependency relationships used by the three-dimensional data encoding device may be determined in advance, and the three-dimensional data encoding device may not send the dependency information.
[0252] Fig. 33 is a diagram showing an example of the dependency relationship of each data. The tip of the arrow in the diagram indicates the dependency destination, and the start of the arrow indicates the dependency source. The three-dimensional data decoding device decodes data in the order from the dependency destination to the dependency source. Also, data shown by solid lines in the diagram is data that is actually sent, and data shown by dotted lines is data that is not sent.
[0253] In the figure, G indicates location information, and A indicates attribute information. s1 indicates the position information of slice number 1, and G s2 indicates the position information of slice number 2. s1t1 indicates the position information of slice number 1 and tile number 1, and G s1t2 indicates the position information of slice number 1 and tile number 2, and G s2t1 indicates the position information of slice number 2 and tile number 1, and G s2t2 indicates the position information of slice number 2 and tile number 2. Similarly, A s1 indicates the attribute information of slice number 1, and A s2 indicates the attribute information of slice number 2. s1t1 indicates the attribute information of slice number 1 and tile number 1, and A s1t2 indicates the attribute information of slice number 1 and tile number 2, and A s2t1 indicates the attribute information of slice number 2 and tile number 1, and A s2t2 indicates attribute information of slice number 2 and tile number 2.
[0254] Mslice indicates slice additional information, MGtile indicates position tile additional information, and MAtile indicates attribute tile additional information. s1t1 is attribute information As1t1 D s2t1 is attribute information A s2t1 This shows dependency information for
[0255] Furthermore, the three-dimensional data encoding device may rearrange the data in the decoding order so that there is no need to rearrange the data in the three-dimensional data decoding device. Note that the data may be rearranged in the three-dimensional data decoding device, or the data may be rearranged in both the three-dimensional data encoding device and the three-dimensional data decoding device.
[0256] Fig. 34 is a diagram showing an example of the data decoding order. In the example of Fig. 34, decoding is performed in order from the left data. When data has a dependency relationship, the three-dimensional data decoding device decodes the dependent data first. For example, the three-dimensional data encoding device rearranges the data in advance so as to achieve this order, and transmits the data. Any order may be used as long as the dependent data comes first. The three-dimensional data encoding device may also transmit additional information and dependency information before the data.
[0257] Fig. 35 is a flowchart showing the flow of processing by the three-dimensional data encoding device. First, the three-dimensional data encoding device encodes data of a plurality of slices or tiles as described above (S4901). Next, the three-dimensional data encoding device rearranges the data so that the dependent data comes first, as shown in Fig. 34 (S4902). Next, the three-dimensional data encoding device multiplexes (NAL unitizes) the rearranged data (S4903).
[0258] Next, a description will be given of the configuration of the combining unit 4925 included in the first decoding unit 4920. Fig. 36 is a block diagram showing the configuration of the combining unit 4925. The combining unit 4925 includes a position information tile combining unit 4941 (geometry tile combiner), an attribute information tile combining unit 4942 (attribute tile combiner), and a slice combining unit (slice combiner).
[0259] The position information tile combining unit 4941 generates a plurality of slice position information pieces by combining a plurality of pieces of divided position information pieces using the position tile additional information. The attribute information tile combining unit 4942 generates a plurality of slice attribute information pieces by combining a plurality of pieces of divided attribute information pieces using the attribute tile additional information.
[0260] The slice combining unit 4943 generates position information by combining a plurality of slice position information pieces using the slice additional information. Also, the slice combining unit 4943 generates attribute information by combining a plurality of slice attribute information pieces using the slice additional information.
[0261] The number of slices or tiles to be divided is equal to or greater than 1. In other words, the division into slices or tiles does not necessarily have to be performed.
[0262] Although an example in which tile division is performed after slice division has been shown here, slice division may be performed after tile division. Furthermore, a new division type may be defined in addition to slices and tiles, and division may be performed using three or more division types.
[0263] Next, a structure of coded data divided into slices or tiles and a method of storing the coded data in NAL units (multiplexing method) will be described. Fig. 37 is a diagram showing the structure of coded data and a method of storing the coded data in NAL units.
[0264] The encoded data (partition position information and partition attribute information) is stored in the payload of the NAL unit.
[0265] The encoded data includes a header and a payload. The header includes identification information for identifying data included in the payload. This identification information includes, for example, a type of slice division or tile division (slice_type, tile_type), index information for identifying a slice or tile (slice_idx, tile_idx), position information of data (slice or tile), or an address of data (address). The index information for identifying a slice is also referred to as a slice index (SliceIndex). The index information for identifying a tile is also referred to as a tile index (TileIndex). The type of division is, for example, a method based on an object shape as described above, a method based on map information or position information, or a method based on a data amount or a processing amount.
[0266] In addition, all or part of the above information may be stored in one of the headers of the divided position information and the divided attribute information, and may not be stored in the other. For example, when the same division method is used for the position information and the attribute information, the division type (slice_type, tile_type) and index information (slice_idx, tile_idx) are the same for the position information and the attribute information. Therefore, these information may be included in the header of one of the position information and the attribute information. For example, when the attribute information depends on the position information, the position information is processed first. Therefore, these information may be included in the header of the position information, and may not be included in the header of the attribute information. In this case, the three-dimensional data decoding device determines that the dependent attribute information belongs to the same slice or tile as the slice or tile of the dependent position information, for example.
[0267] In addition, additional information related to slice division or tile division (slice additional information, position tile additional information, or attribute tile additional information), and dependency information indicating dependencies, etc. may be stored in and transmitted from existing parameter sets (such as GPS, APS, position SPS, or attribute SPS). When the division method changes for each frame, information indicating the division method may be stored in the parameter set for each frame (such as GPS or APS). When the division method does not change within a sequence, information indicating the division method may be stored in the parameter set for each sequence (position SPS or attribute SPS). Furthermore, when the same division method is used for position information and attribute information, information indicating the division method may be stored in the parameter set of the PCC stream (stream PS).
[0268] In addition, the above information may be stored in any of the above parameter sets, or may be stored in multiple parameter sets. Also, a parameter set for tile division or slice division may be defined, and the above information may be stored in the parameter set. Further, these information may be stored in the header of the encoded data.
[0269] In addition, the header of the encoded data includes identification information indicating dependencies. That is, when there is a dependency between data, the header includes identification information for referring from the dependent source to the dependent destination. For example, the header of the dependent destination data includes identification information for specifying the data. The header of the dependent source data includes identification information indicating the dependent destination. Note that when the identification information for specifying data, the additional information related to slice division or tile division, and the identification information indicating dependencies can be identified or derived from other information, these information may be omitted.
[0270] Next, the flow of the encoding process and decoding process of the point cloud data according to the present embodiment will be described. FIG. 38 is a flowchart of the encoding process of the point cloud data according to the present embodiment.
[0271] First, the three-dimensional data encoding device determines the division method to be used (S4911). The division method includes whether or not to perform slice division and whether or not to perform tile division. The division method may also include the number of divisions when performing slice division or tile division, and the type of division. The type of division may be a method based on the object shape as described above, a method based on map information or position information, or a method based on the amount of data or the amount of processing. The division method may be determined in advance.
[0272] When slice division is performed (Yes in S4912), the three-dimensional data encoding device generates a plurality of slice position information and a plurality of slice attribute information by dividing the position information and the attribute information together (S4913). In addition, the three-dimensional data encoding device generates slice additional information related to the slice division. Note that the three-dimensional data encoding device may divide the position information and the attribute information independently.
[0273] If tile division is performed (Yes in S4914), the three-dimensional data encoding device generates a plurality of division position information and a plurality of division attribute information by independently dividing a plurality of slice position information and a plurality of slice attribute information (or position information and attribute information) (S4915). In addition, the three-dimensional data encoding device generates position tile additional information and attribute tile additional information related to the tile division. Note that the three-dimensional data encoding device may divide the slice position information and slice attribute information together.
[0274] Next, the three-dimensional data encoding device generates a plurality of pieces of encoding position information and a plurality of pieces of encoding attribute information by encoding each of the plurality of pieces of division position information and the plurality of pieces of division attribute information (S4916). Also, the three-dimensional data encoding device generates dependency relationship information.
[0275] Next, the three-dimensional data encoding device generates encoded data (encoded stream) by forming (multiplexing) the plurality of pieces of encoding position information, the plurality of pieces of encoding attribute information, and the additional information into NAL units (S4917). In addition, the three-dimensional data encoding device transmits the generated encoded data.
[0276] 39 is a flowchart of a decoding process of point cloud data according to this embodiment. First, the three-dimensional data decoding device analyzes additional information (slice additional information, position tile additional information, and attribute tile additional information) related to the division method included in the encoded data (encoded stream) to determine the division method (S4921). This division method includes whether or not to perform slice division and whether or not to perform tile division. In addition, the division method may include the number of divisions when performing slice division or tile division, the type of division, and the like.
[0277] Next, the three-dimensional data decoding device generates split position information and split attribute information by decoding the multiple pieces of encoded position information and multiple pieces of encoded attribute information contained in the encoded data using dependency information contained in the encoded data (S4922).
[0278] When the additional information indicates that tile division has been performed (Yes in S4923), the three-dimensional data decoding device generates a plurality of slice position information and a plurality of slice attribute information by combining a plurality of pieces of division position information and a plurality of pieces of division attribute information by respective methods based on the position tile additional information and the attribute tile additional information (S4924). Note that the three-dimensional data decoding device may combine a plurality of pieces of division position information and a plurality of pieces of division attribute information by the same method.
[0279] When the additional information indicates that slice division has been performed (Yes in S4925), the three-dimensional data decoding device generates position information and attribute information by combining multiple slice position information and multiple slice attribute information (multiple division position information and multiple division attribute information) in the same manner based on the slice additional information (S4926). Note that the three-dimensional data decoding device may combine multiple slice position information and multiple slice attribute information in different manners.
[0280] In addition, attribute information of tiles or slices (identifiers, area information, address information, position information, etc.) may be stored in other control information, not limited to SEI. For example, the attribute information may be stored in control information indicating the configuration of the entire PCC data, or may be stored in control information for each tile or slice.
[0281] Furthermore, when transmitting PCC data to another device, the three-dimensional data encoding device (three-dimensional data transmission device) may convert control information such as SEI into control information specific to the protocol of that system and display it.
[0282] For example, when the three-dimensional data encoding device converts PCC data including attribute information into ISOBMFF (ISO Base Media File Format), the SEI may be stored in an "mdat box" together with the PCC data, or in a "track box" that describes control information related to the stream. In other words, the three-dimensional data encoding device may store the control information in a table for random access. Also, when the three-dimensional data encoding device packets the PCC data for transmission, the SEI may be stored in a packet header. In this way, by making it possible to acquire attribute information at a system layer, access to the attribute information and the tile data or slice data becomes easier, and the access speed can be improved.
[0283] In addition, in the configuration of the three-dimensional data decoding device, the memory management unit may determine in advance whether or not the information necessary for the decoding process is in the memory, and if the information necessary for the decoding process is not present, the information may be obtained from storage or a network.
[0284] When the three-dimensional data decoding device acquires PCC data from a storage or a network using Pull in a protocol such as MPEG-DASH, the memory management unit may identify attribute information of data necessary for decoding processing based on information from a localization unit or the like, request tiles or slices including the identified attribute information, and acquire the necessary data (PCC stream). Identification of tiles or slices including attribute information may be performed on the storage or network side, or may be performed by the memory management unit. For example, the memory management unit may acquire SEI of all PCC data in advance, and identify tiles or slices based on that information.
[0285] When all PCC data is transmitted from the storage or network using Push in a UDP protocol or the like, the memory management unit may identify attribute information of the data required for the decoding process and tiles or slices based on information from the localization unit or the like, and obtain the desired data by filtering the desired tiles or slices from the transmitted PCC data.
[0286] Furthermore, when acquiring data, the three-dimensional data encoding device may determine whether or not desired data is available, whether or not real-time processing is possible based on the data size, etc., or the communication state, etc. If the three-dimensional data encoding device determines that data acquisition is difficult based on the result of this determination, it may select and acquire another slice or tile with a different priority or amount of data.
[0287] In addition, the three-dimensional data decoding device may transmit information from a localization unit or the like to a cloud server, and the cloud server may determine the necessary information based on that information.
[0288] (Embodiment 4) Next, the tile additional information will be described. The three-dimensional data encoding device generates tile additional information, which is metadata related to a method of dividing tiles, and transmits the generated tile additional information to the three-dimensional data decoding device.
[0289] Fig. 40 is a diagram showing an example of the syntax of the tile additional information (TileMetaData). As shown in Fig. 40, for example, the tile additional information includes division method information (type_of_divide), shape information (topview_shape), overlap flag (tile_overlap_flag), overlap information (type_of_overlap), height information (tile_height), number of tiles (tile_number), and tile position information (global_position, relative_position).
[0290] The division method information (type_of_divide) indicates a division method of tiles. For example, the division method information indicates whether the division method of tiles is based on map information, that is, division based on a top view (top_view), or other (other).
[0291] The shape information (topview_shape) is included in the tile additional information, for example, when the tile division method is division based on a top view. The shape information indicates the shape of the tile when viewed from above. For example, this shape includes a square and a circle. Note that this shape may include an ellipse, a rectangle, or a polygon other than a quadrangle, or may include other shapes. Note that the shape information is not limited to the shape of the tile when viewed from above, and may also indicate the three-dimensional shape of the tile (for example, a cube, a cylinder, etc.).
[0292] The overlap flag (tile_overlap_flag) indicates whether tiles overlap. For example, the overlap flag is included in the tile additional information when the tile division method is division based on a top view. In this case, the overlap flag indicates whether tiles overlap in a top view. Note that the overlap flag may also indicate whether tiles overlap in a three-dimensional space.
[0293] The overlap information (type_of_overlap) is included in the tile additional information when tiles overlap, for example. The overlap information indicates how the tiles overlap, etc. For example, the overlap information indicates the size of the overlapping area, etc.
[0294] Height information (tile_height) indicates the height of a tile. The height information may include information indicating the shape of the tile. For example, if the shape of the tile in top view is rectangular, the information may indicate the lengths of the sides of the rectangle (the vertical length and the horizontal length). Furthermore, if the shape of the tile in top view is circular, the information may indicate the diameter or radius of the circle.
[0295] The height information may indicate the height of each tile, or may indicate a common height for multiple tiles. Multiple height types, such as roads and intersections, may be set in advance, and the height information may indicate the height of each height type and the height type of each tile. Alternatively, the height of each height type may be defined in advance, and the height information may indicate the height type of each tile. In other words, the height of each height type does not have to be indicated by the height information.
[0296] The tile number (tile_number) indicates the number of tiles. The tile additional information may include information indicating the spacing between tiles.
[0297] The tile position information (global_position, relative_position) is information for identifying the position of each tile. For example, the tile position information indicates the absolute coordinates or relative coordinates of each tile.
[0298] Note that some or all of the above information may be provided for each tile, or for each set of tiles (for example, for each frame or for each set of frames).
[0299] The three-dimensional data encoding device may include the tile additional information in SEI (Supplemental Enhancement Information) and transmit it. Alternatively, the three-dimensional data encoding device may store the tile additional information in an existing parameter set (such as PPS, GPS, or APS) and transmit it.
[0300] For example, if the tile additional information changes for each frame, the tile additional information may be stored in a parameter set for each frame (GPS or APS, etc.). If the tile additional information does not change within a sequence, the tile additional information may be stored in a parameter set for each sequence (position SPS or attribute SPS). Furthermore, if the same tile division information is used for position information and attribute information, the tile additional information may be stored in a parameter set of a PCC stream (stream PS).
[0301] Furthermore, the tile additional information may be stored in any one of the above parameter sets, or in multiple parameter sets. Furthermore, the tile additional information may be stored in a header of the encoded data. Furthermore, the tile additional information may be stored in a header of a NAL unit.
[0302] In addition, all or part of the tile additional information may be stored in one of the headers of the divided position information and the divided attribute information, and not in the other. For example, when the same tile additional information is used in the position information and the attribute information, the tile additional information may be included in the header of either the position information or the attribute information. For example, when the attribute information depends on the position information, the position information is processed first. Therefore, the header of the position information may include these tile additional information, and the header of the attribute information may not include the tile additional information. In this case, the three-dimensional data decoding device determines, for example, that the dependent attribute information belongs to the same tile as the tile of the dependent position information.
[0303] The three-dimensional data decoding device reconstructs the point cloud data divided into tiles based on the tile additional information. If there is overlapping point cloud data, the three-dimensional data decoding device identifies the overlapping multiple point cloud data and selects one of them or merges the multiple point cloud data.
[0304] The three-dimensional data decoding device may also perform decoding using tile additional information. For example, when multiple tiles overlap, the three-dimensional data decoding device may perform decoding for each tile, and perform processing (e.g., smoothing or filtering) using the multiple decoded data to generate point cloud data. This may enable highly accurate decoding.
[0305] 41 is a diagram showing an example of the configuration of a system including a three-dimensional data encoding device and a three-dimensional data decoding device. A tile dividing unit 5051 divides point cloud data including position information and attribute information into a first tile and a second tile. The tile dividing unit 5051 also sends tile additional information related to the tile division to a decoding unit 5053 and a tile combining unit 5054.
[0306] The encoding unit 5052 generates encoded data by encoding the first tile and the second tile.
[0307] The decoding unit 5053 restores the first tile and the second tile by decoding the coded data generated by the coding unit 5052. The tile combining unit 5054 restores the point cloud data (position information and attribute information) by combining the first tile and the second tile using the tile additional information.
[0308] Next, slice additional information will be described. The three-dimensional data encoding device generates slice additional information, which is metadata related to a method of dividing a slice, and transmits the generated slice additional information to the three-dimensional data decoding device.
[0309] Fig. 42 is a diagram showing an example of the syntax of slice additional information (SliceMetaData). As shown in Fig. 42, for example, the slice additional information includes division method information (type_of_divide), an overlap flag (slice_overlap_flag), overlap information (type_of_overlap), the number of slices (slice_number), slice position information (global_position, relative_position), and slice size information (slice_bounding_box_size).
[0310] The division method information (type_of_divide) indicates a method of dividing a slice. For example, the division method information indicates whether the division method of a slice is division based on object information (object). The slice additional information may include information indicating a method of dividing an object. For example, this information indicates whether one object is divided into multiple slices or assigned to one slice. This information may also indicate the number of divisions when one object is divided into multiple slices.
[0311] The overlap flag (slice_overlap_flag) indicates whether the slices overlap. The overlap information (type_of_overlap) is included in the slice additional information when the slices overlap, for example. The overlap information indicates how the slices overlap, etc. For example, the overlap information indicates the size of the overlapping area, etc.
[0312] The slice number (slice_number) indicates the number of slices.
[0313] Slice position information (global_position, relative_position) and slice size information (slice_bounding_box_size) are information related to the area of a slice. Slice position information is information for identifying the position of each slice. For example, slice position information indicates absolute coordinates or relative coordinates of each slice. Slice size information (slice_bounding_box_size) indicates the size of each slice. For example, slice size information indicates the size of the bounding box of each slice.
[0314] The three-dimensional data encoding device may include the slice additional information in the SEI and transmit it. Alternatively, the three-dimensional data encoding device may store the slice additional information in an existing parameter set (such as PPS, GPS, or APS) and transmit it.
[0315] For example, when slice additional information changes for each frame, the slice additional information may be stored in a parameter set for each frame (GPS or APS, etc.). When slice additional information does not change within a sequence, the slice additional information may be stored in a parameter set for each sequence (position SPS or attribute SPS). Furthermore, when the same slice division information is used for position information and attribute information, the slice additional information may be stored in a parameter set of a PCC stream (stream PS).
[0316] Furthermore, the slice additional information may be stored in any one of the above parameter sets, or in a plurality of parameter sets. Furthermore, the slice additional information may be stored in a header of the encoded data. Furthermore, the slice additional information may be stored in a header of the NAL unit.
[0317] In addition, all or a part of the slice additional information may be stored in one of the headers of the division position information and the header of the division attribute information, and may not be stored in the other. For example, when the same slice additional information is used in the position information and the attribute information, the slice additional information may be included in the header of one of the position information and the attribute information. For example, when the attribute information depends on the position information, the position information is processed first. Therefore, the header of the position information may include these slice additional information, and the header of the attribute information may not include the slice additional information. In this case, the three-dimensional data decoding device determines that the dependent attribute information belongs to the same slice as the slice of the dependent position information, for example.
[0318] The three-dimensional data decoding device reconstructs the point cloud data divided into slices based on the slice additional information. If there is overlapping point cloud data, the three-dimensional data decoding device identifies the overlapping multiple point cloud data and selects one of them or merges the multiple point cloud data.
[0319] The three-dimensional data decoding device may also perform decoding using slice additional information. For example, when multiple slices overlap, the three-dimensional data decoding device may perform decoding for each slice, and perform processing (e.g., smoothing or filtering) using the multiple decoded data to generate point cloud data. This may enable highly accurate decoding.
[0320] FIG. 43 is a flowchart of three-dimensional data encoding processing, including processing for generating tile additional information, performed by the three-dimensional data encoding device according to this embodiment.
[0321] First, the three-dimensional data encoding device determines a tile division method (S5031). Specifically, the three-dimensional data encoding device determines whether to use a division method based on a top view (top_view) or another method (other) as a tile division method. The three-dimensional data encoding device also determines the shape of the tile when using a division method based on a top view. The three-dimensional data encoding device also determines whether the tile overlaps with other tiles.
[0322] If the tile division method determined in step S5031 is a division method based on a top view (Yes in S5032), the three-dimensional data encoding device describes in the tile additional information that the tile division method is a division method based on a top view (top_view) (S5033).
[0323] On the other hand, if the tile division method determined in step S5031 is other than a division method based on a top view (No in S5032), the three-dimensional data encoding device describes in the tile additional information that the tile division method is other than a division method based on a top view (top_view) (S5034).
[0324] Furthermore, if the shape of the tile viewed from above determined in step S5031 is a square (square in S5035), the three-dimensional data encoding device records in the tile additional information that the shape of the tile viewed from above is a square (S5036).On the other hand, if the shape of the tile viewed from above determined in step S5031 is a circle (circle in S5035), the three-dimensional data encoding device records in the tile additional information that the shape of the tile viewed from above is a circle (S5037).
[0325] Next, the three-dimensional data encoding device determines whether the tile overlaps with other tiles (S5038). If the tile overlaps with other tiles (Yes in S5038), the three-dimensional data encoding device records in the tile additional information that the tile overlaps (S5039). On the other hand, if the tile does not overlap with other tiles (No in S5038), the three-dimensional data encoding device records in the tile additional information that the tile does not overlap (S5040).
[0326] Next, the three-dimensional data encoding device divides the tiles based on the tile division method determined in step S5031, encodes each tile, and transmits the generated encoded data and tile additional information (S5041).
[0327] FIG. 44 is a flowchart of three-dimensional data decoding processing using tile additional information by the three-dimensional data decoding device according to this embodiment.
[0328] First, the 3D data decoding device analyzes the tile additional information included in the bitstream (S5051).
[0329] If the tile additional information indicates that the tile does not overlap with other tiles (No in S5052), the 3D data decoding device generates point cloud data for each tile by decoding each tile (S5053). Next, the 3D data decoding device reconstructs point cloud data from the point cloud data for each tile based on the tile division method and tile shape indicated in the tile additional information (S5054).
[0330] On the other hand, if the tile additional information indicates that the tile overlaps with other tiles (Yes in S5052), the three-dimensional data decoding device generates point cloud data for each tile by decoding each tile. The three-dimensional data decoding device also identifies overlapping parts of the tiles based on the tile additional information (S5055). Note that the three-dimensional data decoding device may perform decoding processing for the overlapping parts using multiple pieces of overlapping information. Next, the three-dimensional data decoding device reconstructs point cloud data from the point cloud data of each tile based on the tile division method, tile shape, and overlap information indicated in the tile additional information (S5056).
[0331] Modifications and the like regarding slices will be described below. The three-dimensional data encoding device may transmit information indicating the type (road, building, tree, etc.) or attribute (dynamic information, static information, etc.) of an object as additional information. Alternatively, encoding parameters may be predefined according to the object, and the three-dimensional data encoding device may notify the three-dimensional data decoding device of the encoding parameters by sending the type or attribute of the object.
[0332] The following methods may be used for the coding order and transmission order of slice data. For example, the three-dimensional data encoding device may encode slice data in the order of data that is easiest to recognize or cluster. Alternatively, the three-dimensional data encoding device may encode slice data in the order of slice data that has been clustered first. Furthermore, the three-dimensional data encoding device may transmit the encoded slice data in the order of decreasing priority for decoding in an application. For example, when the priority of decoding dynamic information is high, the three-dimensional data encoding device may transmit the slice data in the order of decreasing priority for decoding dynamic information.
[0333] Furthermore, when the order of the encoded data differs from the order of the decoding priority, the three-dimensional data encoding device may rearrange the encoded data before transmitting it. Furthermore, when storing the encoded data, the three-dimensional data encoding device may store the encoded data after rearranging it.
[0334] An application (a three-dimensional data decoding device) requests a server (a three-dimensional data encoding device) to transmit slices containing desired data. The server transmits slice data required by the application, but does not need to transmit unnecessary slice data.
[0335] An application requests the server to send tiles containing desired data. The server sends the tile data required by the application, but does not need to send unnecessary tile data.
[0336] (Embodiment 5) In this embodiment, processing of division units (for example, tiles or slices) that do not include points will be described. First, a method of dividing point cloud data will be described.
[0337] In video coding standards such as HEVC, data exists for every pixel in a two-dimensional image, so even if a two-dimensional space is divided into multiple data regions, data exists in every data region. On the other hand, in coding of three-dimensional point cloud data, the points that are elements of the point cloud data are themselves data, and there is a possibility that data does not exist in some regions.
[0338] There are various methods for spatially dividing point cloud data, but the division methods can be classified according to whether a division unit (for example, a tile or slice), which is a divided data unit, always contains one or more point data.
[0339] A division method in which all of the division units contain one or more point data is called a first division method. For example, the first division method is a method in which the point cloud data is divided while considering the encoding processing time or the size of the encoded data. In this case, the number of points in each division unit is roughly equal.
[0340] Fig. 45 is a diagram showing an example of a division method. For example, as a first division method, a method of dividing points belonging to the same space into two identical spaces as shown in (a) of Fig. 45 may be used. Also, as shown in (b) of Fig. 45, a space may be divided into a plurality of subspaces (division units) such that each division unit includes a point.
[0341] These methods are point-aware divisions, so every division unit always contains at least one point.
[0342] A division method in which multiple division units may include one or more division units that do not contain point data is called a second division method. For example, as the second division method, a method of equally dividing a space can be used, as shown in (c) of FIG. 45. In this case, a point does not necessarily exist in a division unit. In other words, there are cases in which a point does not exist in a division unit.
[0343] When a three-dimensional data encoding device divides point cloud data, it may indicate in division additional information (metadata) related to the division (e.g., tile additional information or slice additional information) whether (1) a division method was used in which all of the multiple division units include one or more point data, (2) a division method was used in which the multiple division units have one or more division units that do not include point data, or (3) a division method was used in which the multiple division units may have one or more division units that do not include point data, and send out the division additional information.
[0344] The three-dimensional data encoding device may indicate the above information as the type of division method. Also, the three-dimensional data encoding device may perform division using a predetermined division method and not transmit the division additional information. In that case, the three-dimensional data encoding device indicates in advance whether the division method is the first division method or the second division method.
[0345] The second division method and an example of generating and transmitting encoded data will be described below. Note that, although tile division will be described as an example of a method of dividing a three-dimensional space, the following technique can also be applied to a division method using a division unit other than tiles. For example, tile division may be read as slice division.
[0346] Fig. 46 is a diagram showing an example of dividing point cloud data into six tiles. Fig. 46 shows an example in which the minimum unit is a point, and shows an example in which position information (Geometry) and attribute information (Attribute) are divided together. Note that the same applies when the position information and attribute information are divided by different division methods or division numbers, when there is no attribute information, and when there are multiple pieces of attribute information.
[0347] In the example shown in Figure 46, after tile division, there are tiles (#1, #2, #4, #6) that contain points and tiles (#3, #5) that do not contain points. Tiles that do not contain points are called null tiles.
[0348] Note that the division is not limited to six tiles, and any division method may be used. For example, the division unit may be a cube, or may be a non-cubic shape such as a rectangular parallelepiped or a cylinder. The multiple division units may have the same shape, or may include different shapes. In addition, a predetermined method may be used as the division method, or a different method may be used for each predetermined unit (e.g., PCC frame).
[0349] In this division method, when point cloud data is divided into tiles, if there is no data in a tile, a bitstream is generated that includes information indicating that the tile is a null tile.
[0350] Hereinafter, a method for transmitting null tiles and a method for signaling null tiles will be described. The three-dimensional data encoding device may generate, for example, the following information as additional information (metadata) related to data division, and transmit the generated information. Fig. 47 is a diagram showing an example of the syntax of tile additional information (TileMetaData). The tile additional information includes division method information (type_of_divide), division method null information (type_of_divide_null), number of tile divisions (number_of_tiles), and tile null flag (tile_null_flag).
[0351] The division method information (type_of_divide) is information related to the division method or division type. For example, the division method information indicates one or more division methods or division types. For example, the division method includes top view division and equal division. Note that when there is one division method definition, the division method information does not need to be included in the tile additional information.
[0352] The division method null information (type_of_divide_null) is information indicating whether the division method used is the first division method or the second division method described below. Here, the first division method is a division method in which all of the multiple division units always contain one or more point data. The second division method is a division method in which the multiple division units include one or more division units that do not contain point data, or in which the multiple division units may include one or more division units that do not contain point data.
[0353] The tile additional information may include, as division information for the entire tile, at least one of: (1) information indicating the number of divisions of the tile (number_of_tiles) or information for specifying the number of divisions of the tile, (2) information indicating the number of null tiles or information for specifying the number of null tiles, and (3) information indicating the number of tiles other than null tiles or information for specifying the number of tiles other than null tiles. The tile additional information may include, as division information for the entire tile, information indicating the shape of the tile or whether the tiles overlap.
[0354] Furthermore, the tile additional information indicates the division information for each tile in order. For example, the order of tiles is predetermined for each division method and is known in the three-dimensional data encoding device and the three-dimensional data decoding device. Note that, if the order of tiles is not predetermined, the three-dimensional data encoding device may send information indicating the order to the three-dimensional data decoding device.
[0355] The division information for each tile includes a tile null flag (tile_null_flag) that is a flag indicating whether or not data (points) exist in the tile. Note that if there is no data in the tile, the tile null flag may be included as the tile division information.
[0356] Furthermore, if the tile is not a null tile, the tile additional information includes division information for each tile (position information (e.g., coordinates of the origin (origin_x, origin_y, origin_z)) and height information of the tile, etc.). If the tile is a null tile, the tile additional information does not include division information for each tile.
[0357] For example, when slice division information for each tile is stored in the division information for each tile, the three-dimensional data encoding device does not need to store slice division information for null tiles in the additional information.
[0358] In this example, the number of tile divisions (number_of_tiles) indicates the number of tiles including null tiles. Fig. 48 is a diagram showing an example of tile index information (idx). In the example shown in Fig. 48, index information is also assigned to null tiles.
[0359] Next, a data structure and a transmission method of coded data including null tiles will be described. Figures 49 to 51 are diagrams showing data structures in the case where position information and attribute information are divided into six tiles and data does not exist in the third and fifth tiles.
[0360] FIG. 49 is a diagram showing an example of the dependency relationship of each data. The tip of the arrow in the figure indicates the dependency destination, and the base of the arrow indicates the dependency source. tn (n is 1 to 6) indicates the position information of tile number n, and A tn indicates the attribute information of tile number n. tile indicates tile additional information.
[0361] Fig. 50 is a diagram showing an example of the structure of transmission data, which is encoded data transmitted from a three-dimensional data encoding device. Fig. 51 is a diagram showing the structure of encoded data and a method of storing the encoded data in an NAL unit.
[0362] As shown in FIG. 51, the headers of the data of the position information (division position information) and the attribute information (division attribute information) each include tile index information (tile_idx).
[0363] Also, as shown in Structure 1 of Fig. 50, the three-dimensional data encoding device may not transmit position information or attribute information constituting a null tile. Alternatively, as shown in Structure 2 of Fig. 50, the three-dimensional data encoding device may transmit information indicating that the tile is a null tile as data of the null tile. For example, the three-dimensional data encoding device may indicate that the type of the data is a null tile in the tile_type stored in the header of the NAL unit or in the header in the payload (nal_unit_payload) of the NAL unit, and transmit the header. Note that the following description will be given assuming Structure 1.
[0364] In structure 1, when a null tile exists, the value of the tile index information (tile_idx) included in the header of the position information data or attribute information data in the transmission data is not consecutive and has gaps.
[0365] Furthermore, when there is a dependency between data, the three-dimensional data encoding device transmits the data so that the referenced data can be decoded before the referenced data. Note that the attribute information tiles have a dependency on the position information tiles. The attribute information and position information that have a dependency are assigned the same tile index number.
[0366] In addition, the tile additional information related to the tile division may be stored in both the parameter set of the position information (GPS) and the parameter set of the attribute information (APS), or in either one of them. When the tile additional information is stored in one of the GPS and the APS, reference information indicating the referenced GPS or APS may be stored in the other of the GPS and the APS. Furthermore, when the tile division method differs between the position information and the attribute information, different tile additional information is stored in the GPS and the APS, respectively. Furthermore, when the tile division method is the same for a sequence (multiple PCC frames), the tile additional information may be stored in the GPS, the APS, or the SPS (sequence parameter set).
[0367] For example, when tile additional information is stored in both the GPS and the APS, the tile additional information of position information is stored in the GPS, and the tile additional information of attribute information is stored in the APS. When tile additional information is stored in common information such as an SPS, tile additional information commonly used in the position information and the attribute information may be stored, or the tile additional information of position information and the tile additional information of attribute information may be stored separately.
[0368] A combination of tile division and slice division will be described below. First, a data structure and data transmission in the case where tile division is performed after slice division will be described.
[0369] Fig. 52 is a diagram showing an example of the dependency relationship of each piece of data when dividing into tiles after dividing into slices. The tip of the arrow in the diagram indicates the dependency destination, and the base of the arrow indicates the dependency source. In addition, data shown with a solid line in the diagram is data that is actually transmitted, and data shown with a dotted line is data that is not transmitted.
[0370] In the figure, G indicates location information, and A indicates attribute information. s1 indicates the position information of slice number 1, and G s2 indicates the position information of slice number 2. s1t1 indicates the position information of slice number 1 and tile number 1, and G s2t2indicates the position information of slice number 2 and tile number 2. Similarly, A s1 indicates the attribute information of slice number 1, and A s2 indicates the attribute information of slice number 2. s1t1 indicates the attribute information of slice number 1 and tile number 1, and A s2t1 indicates attribute information of slice number 2 and tile number 1.
[0371] Mslice indicates slice additional information, MGtile indicates position tile additional information, and MAtile indicates attribute tile additional information. s1t1 is attribute information A s1t1 D s2t1 is attribute information A s2t1 This shows dependency information for
[0372] The three-dimensional data encoding device does not need to generate and transmit position information and attribute information related to null tiles.
[0373] Even if the number of tile divisions is the same for all slices, the number of tiles generated and transmitted between slices may differ. For example, if the number of tile divisions differs between the position information and the attribute information, null tiles may exist in either the position information or the attribute information, but not in the other. In the example shown in FIG. 52, the position information (G s1 ) is G s1t1 and G s1t2 It is divided into two tiles, G s1t2 is a null tile. On the other hand, the attribute information of slice 1 (A s1 ) is not divided and is one A s1t1 exists and there is no null tile.
[0374] Furthermore, the three-dimensional data encoding device generates and transmits dependency information of the attribute information when data exists at least in the tile of the attribute information, regardless of whether or not a null tile is included in the slice of the position information. For example, when the three-dimensional data encoding device stores information on slice division for each tile in division information for each slice included in slice additional information related to slice division, the three-dimensional data encoding device stores information on whether the tile is a null tile in this information.
[0375] Fig. 53 is a diagram showing an example of the data decoding order. In the example of Fig. 53, decoding is performed in order from the left data. When data has a dependency relationship, the three-dimensional data decoding device decodes the dependent data first. For example, the three-dimensional data encoding device rearranges the data in advance so as to achieve this order, and transmits the data. Any order may be used as long as the dependent data comes first. The three-dimensional data encoding device may also transmit additional information and dependency information before the data.
[0376] Next, a data structure and data transmission in the case where slice division is performed after tile division will be described.
[0377] Fig. 54 is a diagram showing an example of the dependency relationship of each piece of data when dividing into slices after dividing into tiles. The tip of the arrow in the diagram indicates the dependency destination, and the base of the arrow indicates the dependency source. In addition, data shown with a solid line in the diagram is data that is actually sent, and data shown with a dotted line is data that is not sent.
[0378] In the figure, G indicates location information, and A indicates attribute information. t1 indicates the position information of tile number 1. t1s1 indicates the position information of tile number 1 and slice number 1, and G t1s2 indicates the position information of tile number 1 and slice number 2. Similarly, A t1 indicates the attribute information of tile number 1, and A t1s1 indicates attribute information of tile number 1 and slice number 1.
[0379] Mtile indicates tile additional information, MGslice indicates position slice additional information, and MAslice indicates attribute slice additional information. t1s1 is attribute information A t1s1 D t2s1 is attribute information A t2s1 This shows dependency information for
[0380] The three-dimensional data encoding device does not divide null tiles into slices, and does not need to generate and transmit position information, attribute information, and dependency information of attribute information related to null tiles.
[0381] FIG. 55 is a diagram showing an example of the data decoding order. In the example of FIG. 55, decoding is performed in order from the left data. When data has a dependency relationship, the three-dimensional data decoding device decodes the dependent data first. For example, the three-dimensional data encoding device rearranges the data in advance so as to achieve this order, and transmits the data. Any order may be used as long as the dependent data comes first. Furthermore, the three-dimensional data encoding device may transmit the additional information and dependency information before the data.
[0382] Next, a flow of a process of dividing and combining point cloud data will be described. Note that although an example of dividing into tiles and slices will be described here, a similar method can be applied to dividing other spaces.
[0383] 56 is a flowchart of a three-dimensional data encoding process including a data division process by a three-dimensional data encoding device. First, the three-dimensional data encoding device determines the division method to be used (S5101). Specifically, the three-dimensional data encoding device determines whether to use the first division method or the second division method. For example, the three-dimensional data encoding device may determine the division method based on a designation from a user or an external device (e.g., a three-dimensional data decoding device), or may determine the division method according to input point cloud data. In addition, the division method to be used may be predetermined.
[0384] Here, the first division method is a division method in which all of the multiple division units (tiles or slices) always contain one or more point data, and the second division method is a division method in which the multiple division units include one or more division units that do not include point data, or there is a possibility that the multiple division units include one or more division units that do not include point data.
[0385] If the determined division method is the first division method (first division method in S5102), the three-dimensional data encoding device describes that the division method used is the first division method in the division additional information (e.g., tile additional information or slice additional information), which is metadata related to data division (S5103).Then, the three-dimensional data encoding device encodes all division units (S5104).
[0386] On the other hand, if the determined division method is the second division method (second division method in S5102), the three-dimensional data encoding device describes in the division additional information that the division method used is the second division method (S5105).Then, the three-dimensional data encoding device encodes the division units, excluding the division units that do not include point data (e.g., null tiles), among the multiple division units (S5106).
[0387] 57 is a flowchart of a three-dimensional data decoding process including a data splicing process by a three-dimensional data decoding device. First, the three-dimensional data decoding device refers to the partitioning additional information included in the bitstream, and determines whether the partitioning method used is the first partitioning method or the second partitioning method (S5111).
[0388] If the division method used is the first division method (first division method in S5112), the three-dimensional data decoding device receives the encoded data of all division units and generates decoded data of all division units by decoding the received encoded data (S5113). Next, the three-dimensional data decoding device reconstructs a three-dimensional point cloud using the decoded data of all division units (S5114). For example, the three-dimensional data decoding device reconstructs a three-dimensional point cloud by combining multiple division units.
[0389] On the other hand, if the division method used is the second division method (second division method in S5112), the three-dimensional data decoding device receives the coded data of the division units including point data and the coded data of the division units not including point data, and generates decoded data by decoding the coded data of the received division units (S5115). Note that, if a division unit not including point data is not sent, the three-dimensional data decoding device does not need to receive and decode a division unit not including point cloud data. Next, the three-dimensional data decoding device reconstructs a three-dimensional point cloud using the decoded data of the division units including point data (S5116). For example, the three-dimensional data decoding device reconstructs a three-dimensional point cloud by combining a plurality of division units.
[0390] Other methods of dividing point cloud data will be described below. When dividing a space evenly as shown in (c) of Fig. 45, there are cases where no points exist in the divided spaces. In this case, the three-dimensional data encoding device combines the space where no points exist with another space where points exist. In this way, the three-dimensional data encoding device can form multiple division units so that all division units include one or more points.
[0391] 58 is a flowchart of data division in this case. First, the three-dimensional data encoding device divides the data in a specific method (S5121). For example, the specific method is the second division method described above.
[0392] Next, the three-dimensional data encoding device determines whether or not a point is included in the target division unit, which is the division unit to be processed (S5122). If a point is included in the target division unit (Yes in S5122), the three-dimensional data encoding device encodes the target division unit (S5123). On the other hand, if a point is not included in the target division unit (No in S5122), the three-dimensional data encoding device combines the target division unit with other division units that include points, and encodes the combined division unit (S5124). In other words, the three-dimensional data encoding device encodes the target division unit together with other division units that include points.
[0393] Although the example of performing the determination and combination for each division unit has been described here, the processing method is not limited to this. For example, the three-dimensional data encoding device may determine whether or not each of the multiple division units includes a point, combine the multiple division units so that there is no division unit that does not include a point, and encode each of the multiple division units after the combination.
[0394] Next, a method for transmitting data including null tiles will be described. If a target tile to be processed is a null tile, the three-dimensional data encoding device does not transmit data of the target tile. Figure 59 is a flowchart of the data transmission process.
[0395] First, the three-dimensional data encoding device determines a tile division method, and divides the point cloud data into tiles using the determined division method (S5131).
[0396] Next, the three-dimensional data encoding device determines whether or not the target tile is a null tile (S5132), that is, whether or not there is no data in the target tile.
[0397] If the target tile is a null tile (Yes in S5132), the three-dimensional data encoding device indicates that the target tile is a null tile in the tile additional information, and does not indicate information about the target tile (such as the tile's position and size) (S5133).In addition, the three-dimensional data encoding device does not send out the target tile (S5134).
[0398] On the other hand, if the target tile is not a null tile (No in S5132), the three-dimensional data encoding device indicates in the tile additional information that the target tile is not a null tile and indicates information for each tile (S5135).The three-dimensional data encoding device also outputs the target tile (S5136).
[0399] In this way, by not including information about null tiles in the tile additional information, the amount of information in the tile additional information can be reduced.
[0400] A method for decoding coded data including null tiles will be described below. First, a process for when there is no packet loss will be described.
[0401] 60 is a diagram showing an example of transmission data, which is encoded data transmitted from a three-dimensional data encoding device, and reception data input to a three-dimensional data decoding device. Note that, here, a system environment without packet loss is assumed, and the reception data is the same as the transmission data.
[0402] In the case of a system environment without packet loss, the three-dimensional data decoding device receives all of the transmitted data. Fig. 61 is a flowchart of the process performed by the three-dimensional data decoding device.
[0403] First, the three-dimensional data decoding device refers to the tile additional information (S5141), and determines whether or not each tile is a null tile (S5142).
[0404] If the tile additional information indicates that the target tile is not a null tile (No in S5142), the 3D data decoding device determines that the target tile is not a null tile and decodes the target tile (S5143). Next, the 3D data decoding device obtains tile information (tile position information (origin coordinates, etc.) and size, etc.) from the tile additional information, and reconstructs 3D data by combining multiple tiles using the obtained information (S5144).
[0405] On the other hand, if the tile additional information indicates that the target tile is not a null tile (Yes in S5142), the 3D data decoding device determines that the target tile is a null tile and does not decode the target tile (S5145).
[0406] The three-dimensional data decoding device may determine that missing data is a null tile by sequentially analyzing index information indicated in the header of the encoded data. The three-dimensional data decoding device may also combine a determination method using tile additional information and a determination method using index information.
[0407] Next, the process when packet loss occurs will be described. Fig. 62 is a diagram showing an example of data sent from a three-dimensional data encoding device and data received by a three-dimensional data decoding device. Here, a system environment with packet loss is assumed.
[0408] In a system environment with packet loss, the 3D data decoding device may not be able to receive all of the transmitted data. t2 and A t2 packets are being lost.
[0409] 63 is a flowchart of the process of the three-dimensional data decoding device in this case. First, the three-dimensional data decoding device analyzes the continuity of the index information indicated in the header of the encoded data (S5151), and determines whether the index number of the target tile exists (S5152).
[0410] If the index number of the target tile exists (Yes in S5152), the 3D data decoding device determines that the target tile is not a null tile, and performs a decoding process for the target tile (S5153). Next, the 3D data decoding device obtains tile information (tile position information (origin coordinates, etc.) and size, etc.) from the tile additional information, and reconstructs 3D data by combining multiple tiles using the obtained information (S5154).
[0411] On the other hand, if index information for the target tile does not exist (No in S5152), the 3D data decoding device determines whether or not the target tile is a null tile by referring to the tile additional information (S5155).
[0412] If the target tile is not a null tile (No in S5156), the three-dimensional data decoding device determines that the target tile has been lost (packet loss) and performs error decoding processing (S5157). The error decoding processing is, for example, processing that attempts to decode the original data assuming that the data exists. In this case, the three-dimensional data decoding device may reproduce the three-dimensional data and perform reconstruction of the three-dimensional data (S5154).
[0413] On the other hand, if the target tile is a null tile (Yes in S5156), the 3D data decoding device assumes that the target tile is a null tile and does not perform the decoding process or reconstruct the 3D data (S5158).
[0414] Next, a coding method in which null tiles are not explicitly indicated will be described. The three-dimensional data coding device may generate coded data and additional information in the following manner.
[0415] The three-dimensional data encoder does not indicate information about null tiles in the tile additional information. The three-dimensional data encoder adds index numbers of tiles other than null tiles to the data header. The three-dimensional data encoder does not send null tiles.
[0416] In this case, the number of tile divisions (number_of_tiles) indicates the number of divisions excluding null tiles. The three-dimensional data encoding device may store information indicating the number of null tiles separately in the bit stream. The three-dimensional data encoding device may also indicate information about null tiles in the additional information, or may indicate part of the information about null tiles.
[0417] 64 is a flowchart of the three-dimensional data encoding process by the three-dimensional data encoding device in this case. First, the three-dimensional data encoding device determines a tile division method, and divides the point cloud data into tiles using the determined division method (S5161).
[0418] Next, the three-dimensional data encoding device determines whether the target tile is a null tile (S5162), that is, whether there is no data in the target tile.
[0419] If the target tile is not a null tile (No in S5162), the three-dimensional data encoding device adds index information of tiles other than null tiles to the data header (S5163).Then, the three-dimensional data encoding device outputs the target tile (S5164).
[0420] On the other hand, if the target tile is a null tile (Yes in S5162), the three-dimensional data encoding device does not add index information of the target tile to the data header, and does not transmit the target tile.
[0421] Fig. 65 is a diagram showing an example of index information (idx) added to a data header. As shown in Fig. 65, index information for null tiles is not added, and consecutive numbers are added to tiles other than null tiles.
[0422] FIG. 66 is a diagram showing an example of the dependency relationship of each data. The tip of the arrow in the figure indicates the dependency destination, and the base of the arrow indicates the dependency source. tn (n is 1 to 4) indicates the position information of tile number n, and A tn indicates the attribute information of tile number n. tile indicates tile additional information.
[0423] FIG. 67 is a diagram showing an example of the structure of transmission data, which is encoded data transmitted from the three-dimensional data encoding device.
[0424] The following describes a decoding method when null tiles are not explicitly indicated. Figure 68 is a diagram showing an example of data sent from a three-dimensional data encoding device and data received by a three-dimensional data decoding device. Here, a system environment with packet loss is assumed.
[0425] 69 is a flowchart of the process of the three-dimensional data decoding device in this case. First, the three-dimensional data decoding device analyzes the tile index information indicated in the header of the encoded data, and determines whether the index number of the target tile exists. In addition, the three-dimensional data decoding device obtains the number of divisions of the tile from the tile additional information (S5171).
[0426] If the index number of the target tile exists (Yes in S5172), the 3D data decoding device performs a decoding process for the target tile (S5173). Next, the 3D data decoding device obtains tile information (tile position information (origin coordinates, etc.) and size, etc.) from the tile additional information, and reconstructs 3D data by combining multiple tiles using the obtained information (S5175).
[0427] On the other hand, if the index number of the target tile does not exist (No in S5172), the three-dimensional data decoding device determines that the target tile is a packet loss, and performs error decoding processing (S5174). In addition, the three-dimensional data decoding device determines that a space that does not exist in the data is a null tile, and reconstructs the three-dimensional data.
[0428] In addition, by explicitly indicating null tiles, the three-dimensional data encoding device can properly determine that no points exist within a tile, and that this is not due to measurement error or data loss due to data processing, etc., or packet loss.
[0429] The three-dimensional data encoding device may use both a method of explicitly indicating null packets and a method of not explicitly indicating null packets. In this case, the three-dimensional data encoding device may indicate information indicating whether or not null packets are explicitly indicated in the tile additional information. Also, whether or not null packets are explicitly indicated may be determined in advance according to the type of division method, and the three-dimensional data encoding device may indicate whether or not null packets are explicitly indicated by indicating the type of division method.
[0430] In addition, in Figure 47 etc., an example is shown in which the tile additional information indicates information related to all tiles, but the tile additional information may indicate information on some of the multiple tiles, or may indicate null tile information on some of the multiple tiles.
[0431] Also, an example has been described in which information about split data, such as information about whether or not split data (tiles) exists, is stored in the tile additional information, but some or all of this information may be stored in a parameter set or stored as data. When this information is stored as data, for example, nal_unit_type, which means information indicating whether or not split data exists, may be defined, and this information may be stored in the NAL unit. Also, this information may be stored in both the additional information and the data.
[0432] (Embodiment 6) The process of performing quantization for each tile will be described below. Figure 70 is a diagram showing a schematic diagram of the quantization process for each tile.
[0433] The three-dimensional data encoding device divides the point cloud data into a plurality of data units, such as tiles, that can be encoded and decoded independently, and quantizes each of the divided data obtained by the division.
[0434] By using the following method, the coding efficiency of the quantized data can be improved by quantizing and merging overlapping points when dividing into tiles.
[0435] When quantizing each divided data, the three-dimensional data encoding device quantizes the position information of points belonging to each tile and merges overlapping points into one point. After that, the three-dimensional data encoding device converts the position information of the point cloud data into an occupancy code for each tile, and arithmetically encodes the occupancy code.
[0436] For example, the three-dimensional data encoding device may merge point A and point B into point C, which has the same three-dimensional position information. The three-dimensional data encoding device may assign the average value of attribute information, such as color or reflectance, of points A and B to point C. The three-dimensional data encoding device may also merge point B into point A, or merge point A into point B.
[0437] When merging, the three-dimensional data encoding device sets MergeDuplicatedPointFlag to 1. This MergeDuplicatedPointFlag indicates that duplicated points in the tile have been merged and that there are no duplicated points in the tile. The three-dimensional data encoding device stores MergeDuplicatedPointFlag in the parameter set as metadata (additional information).
[0438] When MergeDuplicatedPointFlag is 1, each leaf node in the occupancy code for each tile includes one point. Therefore, the three-dimensional data encoding device does not need to encode information indicating the number of points included in the leaf node as the leaf node information. In addition, the three-dimensional data encoding device may encode the three-dimensional position information of one point and attribute information such as color and reflectance.
[0439] When MergeDuplicatedPointFlag is 1, the three-dimensional data encoding device may merge M duplicated points into N (M>N) points. At this time, the three-dimensional data encoding device may add information indicating the value N to a header or the like. Alternatively, the value N may be specified by a standard or the like. This eliminates the need for the three-dimensional data encoding device to add information indicating the value N to each leaf node, thereby reducing the amount of encoding generated.
[0440] When the three-dimensional data encoding device does not quantize the point cloud data divided into tiles, it sets MergeDuplicatedPointFlag to 0. When MergeDuplicatedPointFlag is 0, the three-dimensional data encoding device encodes information about duplicated points included in leaf nodes in the tile as leaf node information. For example, each leaf node may include one or more duplicated points. Therefore, the three-dimensional data encoding device may encode information indicating how many points a leaf node includes. In addition, the three-dimensional data encoding device may encode each attribute information of the duplicated points.
[0441] As described above, the three-dimensional data encoding device may change the encoded data structure based on MergeDuplicatedPointFlag.
[0442] The three-dimensional data encoding device may write MergeDuplicatedPointFlag in GPS, which is a parameter set for each frame. In this case, for example, a flag is used that indicates whether or not at least one of all tiles except null tiles includes a duplicated point. The flag may be added to each frame, or the same flag may be applied to all or multiple frames. When the same flag is applied to all or multiple frames, the three-dimensional data encoding device may write the flag in SPS, but not in GPS. This can reduce the amount of transmission data. Here, SPS is a parameter set for each sequence (multiple frames).
[0443] The three-dimensional data encoding device may change the quantization parameter for each tile, or may determine whether or not to quantize depending on the tile.
[0444] Furthermore, when determining whether or not to quantize or merge for each tile, the three-dimensional data encoding device stores a flag for each tile in the GPS (e.g., MergeDuplicatedPointFlag). Alternatively, the three-dimensional data encoding device stores MergeDuplicatedPointFlag in the header of data for each tile, and uses MergeDuplicatedPointFlag as a flag indicating whether or not there is a duplicated point in the tile. In this case, the three-dimensional data encoding device writes a flag indicating that the flag is stored in the header of the data in the GPS. Note that the three-dimensional data encoding device does not need to store a flag if the tile is a null tile.
[0445] Next, the data structure will be described. Fig. 71 is a diagram showing an example of syntax of GPS, which is a parameter set of position information in units of frames. GPS includes at least one of gps_idx indicating a frame number and sps_idx indicating a sequence number.
[0446] The GPS also includes a merged point flag (MergeDuplicatedPointFlag) and tile information (tile_information).
[0447] MergeDuplicatedPointFlag=1 indicates that when a tile is not divided, duplicated points in the point cloud data are merged and there are no duplicated points. When a tile is divided, MergeDuplicatedPointFlag=1 indicates that duplicated points in the tile are merged and there are no duplicated points in all tiles except for null tiles that make up the point cloud data.
[0448] MergeDuplicatedPointFlag=0 indicates that when tile division is not performed, duplicate points in the point cloud data are not merged and there is a possibility of duplicate points. When tile division is performed, MergeDuplicatedPointFlag=0 indicates that duplicate points in the tile are not merged in all tiles except for null tiles that make up the point cloud data and there is a possibility of duplicate points in any tile.
[0449] The tile information (tile_information) indicates information related to tile division. Specifically, the tile information indicates the type of tile division, the number of divisions, the coordinates (position) of each tile, the size of each tile, etc. The tile information may also indicate quantization parameters, null tile information, etc. In addition, in the three-dimensional data decoding device, when the coordinates or size of each tile are known or can be derived, the information indicating the coordinates or size of each tile may be omitted. This makes it possible to reduce the amount of code.
[0450] 72 is a diagram illustrating an example of the syntax of tile information (tile_information). The tile information includes a quantization independent flag (independent_quantization_flag). The quantization independent flag (independent_quantization_flag) is a flag indicating whether a quantization parameter is unified across multiple tiles or is set individually for each tile.
[0451] For example, independent_quantization_flag=1 indicates that the quantization parameters are unified across multiple tiles. In this case, MergeDuplicatedPointFlag and QP_value are indicated in the GPS, and this information is used. Here, QP_value is the quantization parameter used for multiple tiles.
[0452] For example, independent_quantization_flag=2 indicates that a quantization parameter is set for each tile. In this case, if the tile is not a null tile, TileMergeDuplicatedPointFlag and qp_value are indicated in the loop process for each tile. Alternatively, TileMergeDuplicatedPointFlag and qp_value are indicated in the header of the position information. Here, TileMergeDuplicatedPointFlag is a flag indicating whether or not to merge duplicated points in the target tile, which is the tile to be processed. qp_value is the quantization parameter used for the target tile.
[0453] A flag indicating whether MergeDuplicatedPointFlag is unified across multiple tiles or set individually for each tile, and a flag indicating whether the value of QP (quantization parameter) is unified across multiple tiles or set individually for each tile may be provided separately. Alternatively, a flag indicating whether the value of MergeDuplicatedPointFlag or QP is indicated in the GPS or in the header of the location information may be provided. In this way, the quantization parameter can be set independently for each tile.
[0454] 73 is a diagram showing an example of the syntax of node information (node(depth, index)) included in the data of the position information. When MergeDuplicatedPointFlag=0, the node information includes information (num_point_per_leaf) indicating the number of duplicated points of the leaf node. In addition, the corresponding number of attribute information is indicated in the leaf node information included in the data of the attribute information.
[0455] The three-dimensional data encoding process according to this embodiment will be described below. Fig. 74 is a flowchart of the three-dimensional data encoding process according to this embodiment.
[0456] First, the three-dimensional data encoding device determines whether to divide a tile and whether to merge overlapping points (S6201, S6202, S6214). For example, the three-dimensional data encoding device determines whether to divide a tile and whether to merge overlapping points based on an external instruction.
[0457] If tile division is not performed and duplicated points are merged (No in S6201 and Yes in S6202), the three-dimensional data encoding device sets MergeDuplicatedPointFlag=1, indicating that the output point cloud does not contain duplicated points, and stores MergeDuplicatedPointFlag in metadata (additional information) (S6203).
[0458] Next, the three-dimensional data encoding device quantizes the position information of the point group (S6204), and merges overlapping points based on the quantized position information (S6205). Next, the three-dimensional data encoding device encodes an occupancy code (S6206). Next, the three-dimensional data encoding device encodes attribute information for points that do not have overlapping points (S6208).
[0459] On the other hand, if tile division is not performed and duplicate points are not merged (No in S6201 and No in S6202), the three-dimensional data encoding device sets MergeDuplicatedPointFlag = 0, indicating that the output point cloud may contain duplicate points, and stores MergeDuplicatedPointFlag in metadata (additional information) (S6209).
[0460] Next, the three-dimensional data encoding device quantizes the position information of the point group (S6210) and encodes the occupancy code (S6211). The three-dimensional data encoding device also encodes information indicating the number of three-dimensional points contained in each leaf node for all leaf nodes (S6212). Next, the three-dimensional data encoding device encodes attribute information for points that have overlapping points (S6213).
[0461] On the other hand, when dividing into tiles and merging duplicated points (Yes in S6201 and Yes in S6214), the three-dimensional data encoding device sets MergeDuplicatedPointFlag=1, which indicates that there are no duplicated points in all tiles to be output, and stores MergeDuplicatedPointFlag in metadata (additional information) (S6215). Next, the three-dimensional data encoding device divides the point cloud into tiles (S6216).
[0462] Next, the three-dimensional data encoding device quantizes position information of the point group in the target tile to be processed (S6217), and merges overlapping points in the target tile based on the quantized position information (S6218). Next, the three-dimensional data encoding device encodes an occupancy code (S6219), and encodes attribute information for points that do not have overlapping points (S6220).
[0463] If the processing of all tiles has not been completed (No in S6221), the processing of steps S6217 and after is performed on the next tile. If the processing of all tiles has been completed (Yes in S6221), the three-dimensional data encoding device ends the processing.
[0464] On the other hand, if tile division is performed and duplicated points are not merged (Yes in S6201 and No in S6214), the three-dimensional data encoding device sets MergeDuplicatedPointFlag=0, which indicates that the output tile may contain duplicated points, and stores MergeDuplicatedPointFlag in metadata (additional information) (S6222). Next, the three-dimensional data encoding device divides the point cloud into tiles (S6223).
[0465] Next, the three-dimensional data encoding device quantizes position information of the point group in the target tile (S6224), and encodes the occupancy code (S6225). Next, the three-dimensional data encoding device encodes information indicating the number of three-dimensional points included in each leaf node for all leaf nodes (S6226). Next, the three-dimensional data encoding device encodes attribute information for points that have overlapping points (S6227).
[0466] If the processing of all tiles has not been completed (No in S6228), the processing of step S6224 and subsequent steps is performed on the next tile. If the processing of all tiles has been completed (Yes in S6228), the three-dimensional data encoding device ends the processing.
[0467] In addition, in the configuration of the encoding unit, a quantization unit may be disposed before the tile division unit. In other words, tile division may be performed after quantizing all point cloud data. The point cloud data is shifted in position by quantization, and then overlapping points are merged. There is also a case where overlapping points are not merged.
[0468] In this configuration, a flag (independent_quantization_flag) indicating whether the quantization parameters are to be uniform across the tiles or set individually is specified to a value of 1 (uniform across the tiles).
[0469] Fig. 75 is a flowchart of the three-dimensional data encoding process in this case. Note that the process shown in Fig. 75 differs from the process shown in Fig. 74 in the order of the tile division process and the quantization process when tile division is performed (Yes in S6201). The following mainly describes the difference.
[0470] When dividing into tiles and merging duplicated points (Yes in S6201 and Yes in S6214), the three-dimensional data encoding device sets MergeDuplicatedPointFlag=1 (S6215) and quantizes the position information of the point cloud (S6217A). Next, the three-dimensional data encoding device merges duplicated points based on the quantized position information (S6218A). Next, the three-dimensional data encoding device divides the merged point cloud into tiles (S6216A).
[0471] On the other hand, if tile division is performed and duplicated points are not merged (Yes in S6201 and No in S6214), the three-dimensional data encoding device sets MergeDuplicatedPointFlag=0 (S6222) and quantizes the position information of the point cloud (S6224A). Next, the three-dimensional data encoding device divides the quantized point cloud into tiles (S6223A).
[0472] Next, the three-dimensional data decoding process according to this embodiment will be described with reference to FIG 76, which is a flowchart of the three-dimensional data decoding process according to this embodiment.
[0473] First, the three-dimensional data decoding device decodes MergeDuplicatedPointFlag from metadata included in the bitstream (S6231). Next, the three-dimensional data decoding device determines whether or not tile division has been performed (S6232). For example, the three-dimensional data decoding device determines whether or not tile division has been performed based on information included in the bitstream.
[0474] If tile division has not been performed (No in S6232), the three-dimensional data decoding device decodes the occupancy code from the bit stream (S6233). Specifically, the three-dimensional data decoding device generates an octet tree of a certain space (node) using header information and the like included in the bit stream. For example, the three-dimensional data decoding device generates a large space (root node) using the size of the certain space in the x-axis, y-axis, and z-axis directions added to the header information, and divides the space into two in the x-axis, y-axis, and z-axis directions to generate eight small spaces A (nodes A0 to A7) and generate an octet tree. Similarly, the three-dimensional data decoding device further divides each of the nodes A0 to A7 into eight small spaces, and sequentially decodes the occupancy code of each node and the leaf information through the processing of this flow.
[0475] If MergeDuplicatedPointFlag is 0 (Yes in S6234), the three-dimensional data decoding device decodes information indicating the number of three-dimensional points contained in each leaf node for all leaf nodes (S6235). For example, if the large space is 8x8x8, applying octree division three times results in a 1x1x1 node. If this size is the smallest divisible unit (leaf), the three-dimensional data decoding device determines whether or not each leaf node contains a point from the decoded occupancy code of the parent node of the leaf node, and calculates the three-dimensional position of each leaf node.
[0476] After step S6235, or if MergeDuplicatedPointFlag is 1 (No in S6234), the three-dimensional data decoding device calculates the position information (three-dimensional position) of the leaf node using information such as the decoded occupancy code and the division number of the octree (S6236). Next, the three-dimensional data decoding device inversely quantizes the position information (S6237).
[0477] Specifically, the three-dimensional data decoding device calculates position information (three-dimensional position) of the point group by performing inverse quantization using the quantization parameters decoded from the header. For example, the three-dimensional data decoding device calculates the inverse quantization position (x×qx, y×qy, z×qz) by multiplying the three-dimensional position (x, y, z) before inverse quantization by the quantization parameters (qx, qy, qz). Note that the inverse quantization process may be skipped during lossless encoding.
[0478] Next, the three-dimensional data decoding device decodes attribute information related to the three-dimensional point whose position information has been decoded (S6238). When MergeDuplicatedPointFlag=1, one piece of attribute information is linked to each decoded point having a different three-dimensional position after decoding. When MergeDuplicatedPointFlag=0, multiple pieces of different attribute information are decoded and linked to multiple points having the same decoded three-dimensional position.
[0479] On the other hand, if tile division has been performed (Yes in S6232), the three-dimensional data decoding device decodes the occupancy code for each tile (S6239). If MergeDuplicatedPointFlag is 0 (Yes in S6240), the three-dimensional data decoding device decodes information indicating the number of three-dimensional points contained in each leaf node for all leaf nodes in the tile (S6241).
[0480] After step S6241, or if MergeDuplicatedPointFlag is 1 (No in S6240), the three-dimensional data decoding device calculates the position information (three-dimensional position) of the leaf node using information such as the decoded occupancy code and the number of divisions of the octree (S6242).
[0481] Next, the three-dimensional data decoding device dequantizes the position information (three-dimensional positions) of the points in the tile (S6243), and decodes attribute information related to the three-dimensional points whose position information has been decoded (S6244).
[0482] If the processing of all tiles has not been completed (No in S6245), the processing of steps S6239 and after is performed on the next tile. If the processing of all tiles has been completed (Yes in S6245), the three-dimensional data encoding device ends the processing.
[0483] Below, an example of overlapping 3D points between tiles and a coding method for the same will be described. The following are cases where overlapping points occur between tiles. Figures 77 and 78 are diagrams showing examples of tile division.
[0484] As shown in Figure 77, when the dividing unit divides a point cloud into tiles, if the division is performed so that the tile areas shown by solid lines overlap, there will be overlapping points between the tiles in the areas shown by dashed lines after division. Also, as shown in Figure 78, if the division is performed so that the tile areas do not overlap, there will be no overlapping points between the tiles after division.
[0485] In addition, the position of each point is shifted during the quantization of the position information for each tile. This may result in overlapping points within and between tiles. Subsequent merging processes within each tile merge overlapping points within each tile into a single point. However, overlapping points between tiles remain.
[0486] In the example of Figure 77, in addition to the areas where tile regions overlap, new overlapping points may occur near tile boundaries. In the example of Figure 78, overlapping points may occur near tile boundaries.
[0487] If there are overlapping points between tiles, the overlapping points will occur in the point cloud data when the tiles are reconstructed in the three-dimensional data decoding device, which will require extra processing in the case where the overlapping points are not necessary.
[0488] Therefore, the three-dimensional data encoding device stores in the bit stream MergeDuplicatedPointFlag or TileMergeDuplicatedPointFlag, which indicates whether or not there are duplicated points in a tile, and stores in the bit stream UniqueBetweenTilesFlag, which is a flag indicating whether or not there are duplicated points between tiles. As a result, when UniqueBetweenTilesFlag=0, the three-dimensional data decoding device deletes or merges duplicated points, thereby reducing the number of points to be handled and the processing load.
[0489] Furthermore, if overlapping points between tiles occur during tile division or quantization, the three-dimensional data encoding device may subsequently delete or merge the overlapping points between tiles. In this case, the three-dimensional data encoding device stores UniqueBetweenTilesFlag=1, which indicates that no overlapping points occur between tiles, in the bitstream. The three-dimensional data decoding device can determine that merging of overlapping points is not necessary based on UniqueBetweenTilesFlag.
[0490] The three-dimensional data encoding device sets UniqueBetweenTilesFlag to 0 when there is a possibility that overlapping points will occur between tiles, such as when dividing tiles so that the tile areas overlap, or when quantizing each tile. Even when tile areas overlap, if there are no points in the overlapping area to begin with, no overlapping points will occur. In this case as well, the three-dimensional data encoding device may set UniqueBetweenTilesFlag to 0. Furthermore, when quantization is performed, the three-dimensional data encoding device may set UniqueBetweenTilesFlag to 0 when there is a possibility that overlapping points will occur, even in a situation where overlapping points do not necessarily occur.
[0491] 79 is a flowchart of three-dimensional data encoding processing. First, the three-dimensional data encoding device determines whether the tile regions overlap (S6251). If the tile regions do not overlap (No in S6252), the three-dimensional data encoding device determines whether to quantize the tiles individually and merge them (S6252). If the tile regions overlap (Yes in S6251) or if the tiles are to be quantized individually and merged (Yes in S6252), the three-dimensional data encoding device determines whether to merge overlapping points between divided data (S6253).
[0492] If the tiles are quantized individually and no merging is performed (No in S6252), or if overlapping points between split data are to be merged (Yes in S6253), the three-dimensional data encoding device sets UniqueBetweenTilesFlag=1, indicating that there are no overlapping points between tiles (S6254).
[0493] If overlapping points between divided data are not to be merged (No in S6253), the three-dimensional data encoding device sets UniqueBetweenTilesFlag=0, which indicates that there are overlapping points between tiles (S6255).
[0494] Note that a process other than the process shown in Fig. 79 may be used. For example, the three-dimensional data encoding device may actually reconstruct tiles after quantization, search for overlapping points between tiles, and determine whether or not there are overlapping points, and set UniqueBetweenTilesFlag according to the determination result.
[0495] The 3D data encoding device may also store metadata in the bitstream regarding overlapping regions or areas. For example, this metadata may indicate that overlapping points may exist at tile boundaries between tiles. This allows the 3D data decoding device to search tile boundaries and remove overlapping points, thereby reducing processing load.
[0496] Fig. 80 is a block diagram showing the configuration of a three-dimensional data encoding device. As shown in Fig. 80, the three-dimensional data encoding device includes a division unit 6231, a plurality of quantization units 6232A and 6232B, an inter-divided data overlap point merging unit 6233, and a plurality of encoding units 6234A and 6234B.
[0497] The dividing unit 6231 divides the point cloud data into a plurality of tiles to generate a plurality of pieces of divided data. The quantizing units 6232A and 6232B quantize the plurality of pieces of divided data to generate a plurality of pieces of quantized data.
[0498] Each of the multiple quantization units 6232A and 6232B includes a minimum position shift unit 6241, a position information quantization unit 6242, and an intra-divided data overlap point merging unit 6243.
[0499] The minimum position shifter 6241 shifts the point group so that the minimum point with the smallest coordinate value in the point group is shifted to the origin. The position information quantizer 6242 quantizes the position information. The intra-segment data overlapping point merging unit 6243 merges overlapping points within a tile.
[0500] The divided data overlap point merging unit 6233 merges overlap points between tiles. The multiple encoding units 6234A and 6234B generate multiple encoded data by encoding the multiple quantized data after the overlap points between tiles have been merged.
[0501] Also, a configuration in which a quantization unit is placed before the division unit may be used. In other words, the three-dimensional data encoding device may perform tile division after quantizing all point cloud data. In this case, no overlap between tiles occurs during quantization.
[0502] Fig. 81 is a diagram showing an example of GPS syntax. As shown in Fig. 81, the GPS includes an inter-tile overlap point flag (UniqueBetweenTilesFlag). The inter-tile overlap point flag is a flag indicating whether or not there is a possibility that an overlap point exists between tiles.
[0503] 82 is a flowchart of the three-dimensional data decoding process. First, the three-dimensional data decoding device decodes UniqueBetweenTilesFlag and MergeDuplicatedPointFlag from metadata included in the bitstream (S6261). Next, the three-dimensional data decoding device decodes position information and attribute information for each tile, and reconstructs a point group (S6262).
[0504] Next, the three-dimensional data decoding device determines whether or not merging of overlapping points is necessary (S6263). For example, the three-dimensional data decoding device determines whether or not merging is necessary depending on whether or not an application can handle overlapping points, or whether or not it is better to merge overlapping points. Alternatively, the three-dimensional data decoding device may determine to merge overlapping points for the purpose of smoothing or filtering a plurality of pieces of attribute information corresponding to overlapping points, thereby removing noise or improving estimation accuracy.
[0505] If merging of overlapping points is necessary (Yes in S6263), the three-dimensional data decoding device determines whether there is overlap between tiles (existence of overlapping points) (S6264). For example, the three-dimensional data decoding device may determine whether there is overlap between tiles based on the decoding results of UniqueBetweenTilesFlag and MergeDuplicatedPointFlag. This eliminates the need to search for overlapping points in the three-dimensional data decoding device, thereby reducing the processing load of the three-dimensional data decoding device. Note that the three-dimensional data decoding device may determine whether overlapping points exist by searching for overlapping points after reconstructing tiles.
[0506] If there is overlap between tiles (Yes in S6264), the three-dimensional data decoding device merges the overlapping points between the tiles (S6265). Next, the three-dimensional data decoding device merges multiple pieces of overlapping attribute information (S6266).
[0507] After step S6266, or if there is no overlap between tiles (No in S6264), the three-dimensional data decoding device executes the application using a point group without overlapping points (S6267).
[0508] On the other hand, if merging of overlapping points is not necessary (No in S6263), the three-dimensional data decoding device does not merge overlapping points, and executes the application using the point group in which overlapping points exist (S6268).
[0509] Examples of applications will be described below. First, an example of an application that uses a point group with no overlapping points will be described.
[0510] Fig. 83 is a diagram showing an example of an application. The example shown in Fig. 83 shows a use case in which a mobile object traveling from an area of tile A to an area of tile B downloads a map point cloud from a server in real time. The server stores encoded data of map point clouds of multiple overlapping areas. The mobile object has already acquired map information of tile A, and requests the server to acquire map information of tile B, which is located in the direction of movement.
[0511] At that time, the mobile entity determines that the data in the overlapping portion between tile A and tile B is unnecessary, and transmits an instruction to the server to delete the overlapping portion between tile B and tile A contained in tile B. The server deletes the overlapping portion from tile B, and distributes tile B after the deletion to the mobile entity. This makes it possible to reduce the amount of transmitted data and the load of the decoding process.
[0512] The mobile entity may check that there are no overlapping points based on the flag. If the mobile entity has not yet acquired tile A, it requests data from the server that does not delete overlapping points. If the server does not have a function for deleting overlapping points or if it is not clear whether there are overlapping points, the mobile entity may check the distributed data to determine whether there are overlapping points, and may merge the data if there are overlapping points.
[0513] Next, an example of an application using a point cloud with overlapping points will be described. A mobile object uploads map point cloud data acquired by LiDAR to a server in real time. For example, the mobile object uploads data acquired for each tile to the server. In this case, there is an area where tile A and tile B overlap, but the encoding mobile object does not merge the overlapping points between tiles, and sends the data to the server together with a flag indicating that there is an overlap between tiles. The server accumulates the received data as it is, without merging the overlapping data contained in the received data.
[0514] Furthermore, when transmitting or storing point cloud data using a system such as ISOBMFF, MPEG-DASH / MMT, or MPEG-TS, the device may replace a flag included in GPS, indicating whether or not there are overlapping points in a tile or whether or not there are overlapping points between tiles, with a descriptor or metadata in the system layer, and store it in an SI, MPD, moov, or moof box, etc. This allows applications to utilize the functions of the system.
[0515] Furthermore, the three-dimensional data encoding device may divide, for example, tile B into a plurality of slices based on an overlapping area with other tiles, as shown in Fig. 84. In the example shown in Fig. 84, slice 1 is an area that does not overlap with any tile, slice 2 is an area that overlaps with tile A, and slice 3 is an area that overlaps with tile C. This makes it easier to separate desired data from the encoded data.
[0516] The map information may be point cloud data or mesh data. The point cloud data may be tiled for each area and stored in a server.
[0517] 85 is a flowchart showing the flow of processing in the above system. First, a terminal (e.g., a mobile object) detects the movement of the terminal from area A to area B (S6271). Next, the terminal starts acquiring map information of area B (S6272).
[0518] If the terminal has already downloaded the information on area A (Yes in S6273), the terminal instructs the server to obtain data on area B that does not include overlapping points with area A (S6274). The server deletes area A from area B, and transmits the data on area B after deletion to the terminal (S6275). Note that the server may, in response to an instruction from the terminal, encode the data on area B in real time to prevent overlapping points, and transmit it.
[0519] Next, the terminal merges (combines) the map information of area B with the map information of area A, and displays the merged map information (S6276).
[0520] On the other hand, if the terminal has not yet downloaded information about area A (No in S6273), the terminal instructs the server to acquire data about area B, which includes overlapping points with area A (S6277). The server transmits the data about area B to the terminal (S6278). Next, the terminal displays map information about area B, which includes overlapping points with area A (S6279).
[0521] 86 is a flowchart showing another example of operation of the system. The transmitting device (three-dimensional data encoding device) transmits tile data in order (S6281). The transmitting device also adds a flag to the data of the tile to be transmitted, indicating whether the tile to be transmitted overlaps with a tile of the data transmitted immediately before, and transmits the data (S6282).
[0522] The receiving device (three-dimensional data decoding device) determines whether or not the tile of the received data overlaps with a tile of previously received data based on the flag added to the data (S6283). If the tile of the received data overlaps with a tile of previously received data (Yes in S6283), the receiving device deletes or merges the overlapping points (S6284). On the other hand, if the tile of the received data does not overlap with a tile of previously received data (No in S6283), the receiving device does not perform the process of deleting or merging the overlapping points and ends the process. This makes it possible to reduce the processing load of the receiving device and improve the estimation accuracy of attribute information. Note that the receiving device does not need to merge the overlapping points if it is not necessary.
[0523] (Embodiment 7) In this embodiment, a viewpoint-based display method, a random access method for encoded data, a method for encoding point cloud data, and a method for decoding point cloud data will be described.
[0524] With the improvement of sensor performance, it has become possible to obtain high-quality three-dimensional point clouds. However, in order to view high-quality three-dimensional points, a viewing device (viewer) capable of reproducing the high-quality three-dimensional point cloud is required. Specifically, it is desired to be able to display high-quality three-dimensional point clouds (point clouds) with a large amount of data without delay. In this embodiment, a three-dimensional point cloud viewing device (first application) capable of efficiently displaying high-density point cloud data by a scalable method using point cloud compression is described.
[0525] Point cloud compression is performed using multiple data partitioning methods. For example, using Levels of Details (LoD), the resolution required to represent the point cloud data is calculated according to the distance between the virtual camera and the point cloud data. This allows separation or hierarchical organization to be achieved.
[0526] A 3D point cloud viewing device (also referred to as a 3D data decoding device) selects the visible point clouds for rendering, and preferably verifies that all visible point clouds are actual scanned data and not approximations.
[0527] 87 is a block diagram showing a configuration example of a three-dimensional data encoding device. The three-dimensional data encoding device includes a point cloud encoding unit 8701 and a file format generation unit 8702. The point cloud encoding unit 8701 generates encoded data (bit stream) by encoding point cloud data. For example, the point cloud encoding unit 8701 encodes the point cloud data using a position information-based encoding method using an octree, a video-based encoding method, or the like.
[0528] The file format generating unit 8702 changes the encoded data (bit stream) to data in a predetermined file format. For example, the file format is ISOBMFF or MP4. The three-dimensional data encoding device may output encoded data in a file format (for example, transmit it to a three-dimensional data decoding device), or may output encoded data in a bit stream format in the encoding method.
[0529] 88 is a block diagram showing an example of the configuration of a three-dimensional data decoding device 8705. The three-dimensional data decoding device 8705 generates point cloud data by decoding encoded data. Here, the encoded data is, for example, encoded data in a bit stream format or an MP4 format. Note that unencoded point cloud data may also be used.
[0530] All or part of the data in a point cloud is called a brick. This brick may also be called a divided data, tile, or slice. The divided data may be further divided.
[0531] The three-dimensional data decoding device 8705 externally acquires camera viewpoint information indicating the viewpoint (angle) of the camera. The three-dimensional data decoding device 8705 acquires a part or all of the encoded data based on the camera viewpoint information, and generates point cloud data by decoding the acquired encoded data. For example, the camera viewpoint information indicates the position and direction (orientation) of the camera. The three-dimensional data decoding device 8705 then displays the decoded point cloud data.
[0532] The three-dimensional data decoding device 8705 includes a point cloud decoding unit 8706 and a brick decoding control unit 8707. Camera viewpoint information (camera viewing angle) is input to the brick decoding control unit 8707. The brick decoding control unit 8707 selects a brick to be decoded based on the visibility of the brick determined based on the camera viewpoint information. The point cloud decoding unit 8706 decodes the selected brick and outputs the decoded brick.
[0533] The configuration of the three-dimensional data encoding device according to this embodiment will be described below. Fig. 89 is a block diagram showing the configuration of a three-dimensional data encoding device 8710 according to this embodiment. The three-dimensional data encoding device 8710 generates encoded data (encoded stream) by encoding point group data (point cloud). This three-dimensional data encoding device 8710 includes a division unit 8711, a plurality of position information encoding units 8712, a plurality of attribute information encoding units 8713, an additional information encoding unit 8714, a multiplexing unit 8715, and a normal vector generating unit 8716.
[0534] The dividing unit 8711 generates a plurality of pieces of divided data by dividing the point cloud data. Specifically, the dividing unit 8711 generates a plurality of pieces of divided data by dividing the space of the point cloud data into a plurality of subspaces. Here, the subspace is any of bricks, tiles, and slices, or a combination of two or more of bricks, tiles, and slices. More specifically, the point cloud data includes position information, attribute information (color, reflectance, etc.), and additional information. The dividing unit 8711 generates a plurality of pieces of divided position information by dividing the position information, and generates a plurality of pieces of divided attribute information by dividing the attribute information. In addition, the dividing unit 8711 generates additional information related to the division.
[0535] The position information encoding units 8712 generate multiple pieces of encoded position information by encoding the multiple pieces of divided position information. For example, the position information encoding unit 8712 encodes the divided position information using an N-ary tree structure such as an octet tree. Specifically, in the octet tree, the target space is divided into eight nodes (subspaces), and 8-bit information (occupancy code) indicating whether or not a point group is included in each node is generated. In addition, the node including the point group is further divided into eight nodes, and 8-bit information indicating whether or not a point group is included in each of the eight nodes is generated. This process is repeated until the number of point groups included in a predetermined hierarchy or node is equal to or less than a threshold value. For example, the position information encoding units 8712 process the multiple pieces of divided position information in parallel.
[0536] The attribute information encoding unit 8713 generates encoded attribute information, which is encoded data, by encoding attribute information using the configuration information generated by the position information encoding unit 8712. For example, the attribute information encoding unit 8713 determines a reference point (reference node) to be referenced in encoding a target point (target node) to be processed, based on the octree structure generated by the position information encoding unit 8712. For example, the attribute information encoding unit 8713 refers to a node, among peripheral nodes or adjacent nodes, whose parent node in the octree is the same as that of the target node. Note that the method of determining the reference relationship is not limited to this.
[0537] Furthermore, the encoding process of the position information or attribute information may include at least one of a quantization process, a prediction process, and an arithmetic coding process. In this case, the reference means using a reference node to calculate a predicted value of the attribute information, or using a state of the reference node (e.g., occupancy information indicating whether or not a point group is included in the reference node) to determine an encoding parameter. For example, the encoding parameter is a quantization parameter in a quantization process, or a context in an arithmetic coding process.
[0538] The normal vector generation unit 8716 calculates a normal vector for each divided data. Note that the input data does not necessarily have to be divided. In this case, the normal vector generation unit 8716 may calculate a normal vector for each point, instead of a normal vector for each divided data. Alternatively, the normal vector generation unit 8716 may calculate both a normal vector for each divided data and a normal vector for each point.
[0539] The additional information encoding unit 8714 generates encoded additional information by encoding the additional information contained in the point cloud data, the additional information regarding data division generated by the division unit 8711 at the time of division, and the normal vector generated by the normal vector generation unit 8716.
[0540] The multiplexing unit 8715 multiplexes a plurality of pieces of encoding position information, a plurality of pieces of encoding attribute information, and encoding additional information to generate encoded data (encoded stream), and transmits the generated encoded data. In addition, the encoded additional information is used during decoding.
[0541] The configuration of the three-dimensional data decoding device according to this embodiment will be described below. Fig. 90 is a block diagram showing the configuration of a three-dimensional data decoding device 8720. The three-dimensional data decoding device 8720 restores point cloud data by decoding coded data (coded stream) generated by coding point cloud data. This three-dimensional data decoding device 8720 includes a demultiplexing unit 8721, a plurality of position information decoding units 8722, a plurality of attribute information decoding units 8723, an additional information decoding unit 8724, a combining unit 8725, a normal vector extraction unit 8726, a random access control unit 8727, and a selection unit 8728.
[0542] The demultiplexer 8721 demultiplexes the encoded data (encoded stream) to generate a plurality of pieces of encoded position information, a plurality of pieces of encoded attribute information, and encoded additional information. The additional information decoder 8724 decodes the encoded additional information to generate additional information.
[0543] The normal vector extraction unit 8726 extracts a normal vector from the additional information. The random access control unit 8727 determines the divided data to extract based on, for example, the normal vector for each divided data. The selection unit 8728 extracts the multiple divided data (multiple pieces of encoding position information and multiple pieces of encoding attribute information) determined by the random access control unit 8727 from the multiple pieces of divided data (multiple pieces of encoding position information and multiple pieces of encoding attribute information). Note that the selection unit 8728 may extract one divided data.
[0544] The multiple position information decoding units 8722 generate multiple pieces of divided position information by decoding the multiple pieces of encoded position information extracted by the selection unit 8728. For example, the multiple position information decoding units 8722 process the multiple pieces of encoded position information in parallel.
[0545] The multiple attribute information decoding units 8723 generate multiple pieces of divided attribute information by decoding the multiple pieces of encoded attribute information extracted by the selection unit 8728. For example, the multiple attribute information decoding units 8723 process the multiple pieces of encoded attribute information in parallel.
[0546] The combining unit 8725 generates position information by combining a plurality of pieces of divided position information using the additional information The combining unit 8725 generates attribute information by combining a plurality of pieces of divided attribute information using the additional information.
[0547] Next, a first example of generating and encoding a normal vector for each point will be described. FIG. 91 is a diagram showing an example of point cloud data. FIG. 92 is a diagram showing an example of a normal vector for each point. The normal vector can be encoded independently for each three-dimensional group. FIG. 91 and FIG. 92 show a three-dimensional point cloud of a book and a normal vector of the three-dimensional point cloud. As shown in FIG. 92, there are multiple normal vectors extending in the upward, rightward, and forward directions. Here, the surface of the book is flat, and multiple normal vectors of a certain surface extend in the same direction. On the other hand, when the surface is round, the normal vector extends in multiple directions according to the normal of the surface.
[0548] Figure 93 is a diagram showing an example of the syntax of a normal vector in a bit stream. In the normal vector NormalVector[i][face] shown in Figure 93, "i" represents a counter of each 3D point group, and [face] represents the x, y, and z axes representing the 3D point group. In other words, NormalVector represents the magnitude of the normal vector of each axis.
[0549] 94 is a flowchart of a three-dimensional data encoding process. First, the three-dimensional data encoding device encodes position information (geometry) and attribute information for each point (S8701). For example, the three-dimensional data encoding device encodes position information for each point. Furthermore, if attribute information corresponding to a point exists, the three-dimensional data encoding device may encode the attribute information for each point.
[0550] Next, the three-dimensional data encoding device encodes the normal vector (x, y, z) for each point (S8702). The three-dimensional data encoding device may encode the normal vector for each point. The three-dimensional data encoding device may also encode, for example, difference information indicating a difference between the normal vector of the point to be processed and the normal vector of another point. This can reduce the amount of data. The three-dimensional data encoding device may also encode the normal vector by including it in the position information, or may also encode it by including it in the attribute information. The three-dimensional data encoding device may also encode the normal vector independently of the position information and the attribute information. Note that, when multiple normal vectors exist for one point, the three-dimensional data encoding device may encode multiple normal vectors for each point.
[0551] 95 is a flowchart of a three-dimensional data decoding process. First, the three-dimensional data decoding device decodes position information and attribute information for each point from the bit stream (S8706). Next, the three-dimensional data decoding device decodes normal vectors for each point from the bit stream (S8707).
[0552] Note that the processing order shown in Figures 94 and 95 is just an example, and the encoding order and decoding order may be reversed.
[0553] The three-dimensional data encoding device may also reduce the amount of data by encoding the normal vector using the position information or the correlation of the position information. In this case, the three-dimensional data decoding device decodes the normal vector using the position information. By using the above method, the normal vector for each point in the point cloud can be encoded and decoded.
[0554] Next, a second example of generating and encoding a normal vector for each point will be described. As another method of encoding the normal vector for each point, the normal vector is encoded as one of the attribute information. Below, an example of encoding using an attribute information encoding unit or an attribute information decoding unit as one of the attribute information will be described.
[0555] For example, the three-dimensional data encoding device encodes color information as the first attribute information and a normal vector as the second attribute information. FIG. 96 is a diagram showing an example of a bit stream configuration. For example, Attr(0) shown in FIG. 96 is encoded data of the first attribute information, and Attr(1) is encoded data of the second attribute information. Furthermore, metadata related to encoding is stored in a parameter set (APS). The three-dimensional data decoding device decodes the encoded data by referring to the APS corresponding to the encoded data.
[0556] In addition, the SPS stores identification information (attribute_type=Normal Vector) indicating that the second attribute information is a normal vector. If the attribute information is a normal vector, information indicating that the normal vector is data having three elements for each point may be stored in the SPS, etc. In addition, the SPS stores identification information (attribute_type=Color) indicating that the first attribute information is color information.
[0557] 97 is a diagram showing an example of point cloud information having position information, color information, and normal vectors. The three-dimensional data encoding device encodes the uncompressed point cloud data shown in FIG.
[0558] The range of values of the normal vector is a floating-point value from -1 to 1. To facilitate the representation, the three-dimensional data encoding device may convert the floating point to an integer according to the required precision. For example, the three-dimensional data encoding device may convert the floating point to a value from -127 to 128 using an 8-bit representation. That is, the three-dimensional data encoding device may convert the floating point to an integer or a positive integer value. Since the normal vector is treated as one attribute information, different quantization processes can be applied. For example, a different quantization parameter can be used for each attribute information. This allows different precision levels to be realized. For example, the quantization parameter is stored in the APS.
[0559] 98 is a flowchart of the three-dimensional data encoding process. First, the three-dimensional data encoding device encodes position information and attribute information (color information, etc.) for each point (S8711). In addition, the three-dimensional data encoding device encodes the normal vector for each point as attribute information of attribute_type="normal vector" using a predetermined method (S8712).
[0560] 99 is a flowchart of the three-dimensional data decoding process. The three-dimensional data decoding device decodes position information and attribute information for each point from the bit stream (S8716). In addition, the three-dimensional data decoding device decodes normal vectors for each point from the bit stream as attribute information of attribute_type="normal vector" using a predetermined method (S8717).
[0561] Note that the processing order shown in Figures 98 and 99 is just an example, and the encoding order and decoding order may be interchanged.
[0562] Next, an example of generating a normal vector for each data unit including a plurality of points will be described. The three-dimensional data encoding device divides the point cloud data into a plurality of objects or a plurality of regions based on the position information and characteristics of the point cloud. The divided data is, for example, tiles or slices, or hierarchical data. The three-dimensional data encoding device generates a normal vector for this divided data unit, that is, for each data unit including one or more points.
[0563] Here, visibility can be determined by the normal vector representation of the object within the brick. Figures 100 and 101 are diagrams for explaining this process. For example, as shown in Figure 100, the three-dimensional data encoding device divides the normal vector direction into angles spaced at 30° intervals with respect to the horizontal and vertical axes. As a simpler method, as shown in Figure 101, the three-dimensional data encoding device may divide the normal vector into six directions: (0, 0), (0, 90), (0, -90), (90, 0), (-90, 0), and (180, 180).
[0564] The three-dimensional data encoding device may also calculate the effective normal vector using a median, an average, or another more effective algorithm. The three-dimensional data encoding device may also use a representative value as the effective normal vector value, or may use another method.
[0565] In addition, the normal vector for each divided data may show the original x, y, and z values as they are, or may be quantized every 30 degrees as described above, or may be quantized to information every 90 degrees. Quantization can reduce the amount of information.
[0566] Fig. 102 is an example of point cloud data, showing an example of a face object. Fig. 103 is a diagram showing an example of normal vectors in this case. As shown in Fig. 103, the normal vectors of the face object shown in Fig. 102 are oriented in the (0, 0) and (90, 0) directions. The three-dimensional data encoding device can use one bit for each direction to indicate whether or not the normal vector of the object is in that direction.
[0567] In this way, there may be two or more normal vectors for one divided data unit. In this case, multiple normal vectors may be indicated for one divided data unit.
[0568] For example, the data example including a face object shown in Figures 102 and 103 is an example in which the normal vectors of the data are expressed by six different normal vectors for each face in units of 90 degrees. In this example, the two normal vectors in the directions of (0, 0) and (90, 0) are the normal vectors of this divided data.
[0569] As a method of indicating normal vectors, each of the six normal vectors may be represented by 1-bit information. FIG. 104 is a diagram showing an example of this normal vector information. If the divided data has the corresponding normal vector, the 1-bit information is set to a value of 1, and if not, it is set to 0. This allows the amount of information to be reduced by quantizing the data, compared to a method of indicating the x, y, and z values as they are.
[0570] A simpler representation method of normal vectors will be described below. A six-sided cube is used to represent normal vectors and their feasibility (visibility) from a particular camera viewpoint. Figures 105 to 108 are diagrams for explaining this process. Figure 105 shows an example of a six-sided cube. Figures 106, 107, and 108 are diagrams showing front and back faces a and b, left and right faces c and d, and top and bottom faces e and f, respectively. Depending on the object's orientation according to the viewing angle, the normal vector faces at least one or three faces. Six flags of 1 bit each can be used to represent one of the six faces (abcdef) of a cube representing each system. For example, when viewed from the front, (100000), when viewed from the side, (001000), and when viewed from below, (000001) are generated. In this representation, the size is not important, and only the direction is represented. It is also possible for an object to occur for which three faces are specified. Face a is the opposite face of face b, face c is the opposite face of face d, and face e is the opposite face of face f. Therefore, it is impossible to see faces a and b at the same time. In other words, the normal vector can be expressed using three flags (ace).
[0571] In this way, when the camera viewpoint (camera angle) is known in advance, the normal vector information can be expressed in 3 bits. FIG. 109 is a diagram showing the visibility when an object in slice A or slice B is viewed from the direction of face c. Slice A is visible from the direction of face c, so it is expressed as ace=(010). On the other hand, slice B is hidden by slice A when viewed from the direction of face c, so it is expressed as ace=(000).
[0572] Next, a first method of encoding and decoding a normal vector for each brick will be described. Fig. 110 is a diagram showing an example of the configuration of a bit stream in this case. In the example shown in Fig. 110, information on the normal vector is stored in the slice header of the position information in each slice. Note that the information on the normal vector may be stored in the header of the attribute information, or may be stored in metadata independent of the position information and the attribute information.
[0573] Figure 111 is a diagram showing a syntax example of a geometry slice header information of position information. The geometry slice header information of position information includes normal_vector_number, normal_vector_x, normal_vector_y, and normal_vector_z.
[0574] normal_vector_number indicates the number of normal vectors corresponding to the slice data. normal_vector_x, normal_vector_y, and normal_vector_z respectively indicate the elements (x, y, z) of the normal vector corresponding to the slice data.
[0575] In this example, the number of normal_vectors can be changed. The number of normal_vectors indicated is the same as the value of normal_vector_number.
[0576] Note that when the normal vector information is common for all slices, normal_vector_number may be stored in GPS or SPS that can store common information for multiple slices.
[0577] Also, the values of the normal vectors of x, y, and z may be quantized. For example, the three-dimensional data encoding device may quantize the values of the original normal vectors by shifting them by a common bit amount s (bit), and send out the information indicating the bit amount s and the information indicating the quantized normal vectors (normal_vector_x<<s, normal_vector_y<<s, normal_vector_z<<s). This can reduce the bit amount.
[0578] Figure 112 is a diagram showing another syntax example of the geometry slice header of position information. This example shows the simplified (quantized) normal vectors for the six-sided data for each divided data. For each face, it is indicated whether there is a normal vector.
[0579] The slice header of this position information includes is_normal_vector. If there is a normal vector corresponding to the slice data, is_normal_vector is set to 1, and if there is no normal vector, is set to 0. For example, the order of multiple faces is predetermined.
[0580] It should be noted that the quantization accuracy and the number or order of normal vectors are not limited to those described above, and may be fixed or variable.
[0581] 113 is a flowchart of a three-dimensional data encoding process. First, the three-dimensional data encoding device generates a plurality of divided data by dividing point cloud data (S8721). Next, the three-dimensional data encoding device encodes position information and attribute information for each divided data (S8722). Next, the three-dimensional data encoding device stores a normal vector for each divided data in a slice header (S8723).
[0582] 114 is a flowchart of a three-dimensional data decoding process. First, the three-dimensional data decoding device decodes position information and attribute information for each divided data from the bit stream (S8726). Next, the three-dimensional data decoding device decodes a normal vector for each divided data from the slice header for each divided data (S8727). Next, the three-dimensional data decoding device combines multiple divided data (S8728).
[0583] 115 is a flowchart of a three-dimensional data decoding process in the case of partially decoding data. First, the three-dimensional data decoding device decodes a normal vector for each divided data from the slice header for each divided data (S8731). Next, the three-dimensional data decoding device determines the divided data to be decoded based on the normal vector, and decodes the determined divided data (S8732). Next, the decoded divided data are combined (S8733).
[0584] Next, a second method of encoding and decoding normal vectors for each brick will be described. Another method of encoding information of normal vectors is to use metadata (e.g., SEI: Supplemental Enhancement Information). FIG. 116 is a diagram showing an example of a bitstream configuration. As shown in FIG. 116, the SEI may be included in the bitstream, or may be generated as a separate file apart from the main encoded bitstream, depending on how the SEI is implemented in both the encoding device and the decoding device.
[0585] Fig. 117 is a diagram illustrating an example of the syntax of slice information (slice_information) included in SEI. The slice information includes number_of_slice, bounding_box_origin_x, bounding_box_origin_y, bounding_box_origin_z, bounding_box_width, bounding_box_height, bounding_box_depth, normalVector_QP, number_of_normal_vector, normalVector_x, normalVector_y, and normalVector_z.
[0586] number_of_slice indicates the number of divided data. bounding_box_origin_x, bounding_box_origin_y, and bounding_box_origin_z indicate the origin coordinates of the bounding box of the slice data. bounding_box_width, bounding_box_height, and bounding_box_depth indicate the width, height, and depth of the bounding box of the slice data, respectively.
[0587] normalVector_QP indicates quantization scale information or bit shift information when normal_vector is quantized. number_of_normal_vector indicates the number of normal vectors included in the slice data. normalVector_x, normalVector_y, and normalVector_z indicate the components of the normal vector elements (x, y, z), respectively.
[0588] Fig. 118 is a diagram showing another example of slice information included in SEI. The example shown in Fig. 118 is an example showing normal vectors simplified (quantized) into data for six faces for each divided data. Whether or not there is a normal vector is shown for each face.
[0589] This slice information includes is_normal_vector. When there is a normal vector corresponding to the slice data, is_normal_vector is set to 1, and when there is no normal vector, is set to 0. For example, the order of multiple faces is predetermined.
[0590] The slice information may include a flag indicating whether or not the slice information includes information on the bounding box (origin and width, height, and depth) for each slice. In this case, when the flag is on (e.g., 1), the slice information includes information on the bounding box for each slice, and when the flag is off (e.g., 0), the slice information does not include information on the bounding box for each slice. The slice information may also include a flag indicating whether or not the slice information includes information on the normal vector for each slice. In this case, when the flag is on (e.g., 1), the slice information includes information on the normal vector for each slice, and when the flag is off (e.g., 0), the slice information does not include information on the normal vector for each slice.
[0591] Next, random access and partial decoding will be described. A three-dimensional data decoding device independently decodes data for each slice using information for each slice, for example, either or both of bounding box information and normal vectors of the slice.
[0592] 119 is a flowchart of a three-dimensional data decoding process. First, the three-dimensional data decoding device determines slices to be decoded and the decoding order of the slices by a predetermined method (S8741). Next, the three-dimensional data decoding device decodes specific slices in the determined order (S8742).
[0593] Figure 120 is a diagram showing an example of this partial decoding process. For example, the three-dimensional data decoding device receives encoded data divided into slices as shown in (a) of Figure 120. As shown in (b) of Figure 120, the three-dimensional data decoding device decodes the encoded data of some slices and does not decode the encoded data of other slices. Alternatively, as shown in (c) of Figure 120, the three-dimensional data decoding device performs decoding by changing the order of the encoded data.
[0594] Fig. 121 is a diagram showing a configuration example of a three-dimensional data decoding device. As shown in Fig. 121, the three-dimensional data decoding device includes an attribute information decoding unit 8731 and a random access control unit 8732. The attribute information decoding unit 8731 extracts bounding box information and normal vectors for each slice from the encoded data. The random access control unit 8732 determines the number and order of the slices to be decoded based on the bounding box information and normal vectors for each slice and sensor information acquired from outside, such as the camera angle (camera direction) and camera position.
[0595] 122 and 123 are diagrams showing an example of processing by the random access control unit 8732. As shown in FIG. 122, for example, the random access control unit 8732 may calculate distance information indicating the distance from the camera for each slice from the bounding box for each slice and the camera position. Alternatively, as shown in FIG. 123, the random access control unit 8732 may derive visibility information indicating whether an object is visible from the camera for each slice from the normal vector for each slice and the camera angle. Note that the random access control unit 8732 may calculate either the distance information or the visibility information, or may calculate both.
[0596] The visibility information and distance information will be described below. FIG. 124 is a diagram showing an example of the relationship between distance and resolution. For example, what is visible from the camera is decoded (frustum culling). The resolution at which the information is decoded further depends on the distance between the virtual camera and the point cloud data.
[0597] That is, the three-dimensional data decoding device determines whether or not a slice is visible from the camera based on the normal vector and the camera viewpoint (camera angle) of each slice, and decodes the slices that are visible from the camera. Furthermore, the three-dimensional data decoding device may calculate the distance from the camera of the slice to be decoded, and if the distance from the camera is close, decode high-resolution data, and if the distance from the camera is far, decode low-resolution data.
[0598] In this case, the coded data is coded in a hierarchical manner, and the three-dimensional data decoding device can independently decode the low-resolution data. When decoding high-resolution data, the three-dimensional data decoding device further decodes difference information between the low-resolution data and the high-resolution data, and generates high-resolution data by adding the difference information to the low-resolution data. When the coded data is not coded in a hierarchical manner, the three-dimensional data decoding device may not perform this process, or may determine whether or not to perform this process depending on whether the data is coded in a hierarchical manner.
[0599] Next, the determination of visibility using normal vectors will be described. Figure 125 is a diagram showing an example of bricks and normal vectors. In the example shown in Figure 125, two bricks (e.g., slices) on the front side facing the camera (view frustum), that is, bricks whose normal vectors face the camera, are decoded.
[0600] First, the three-dimensional data decoding device determines whether or not one or more normal vectors included in the metadata for each slice data have a normal vector opposite to the camera direction. If the slice data of the target slice has a normal vector opposite to the camera direction, the three-dimensional data decoding device determines that the target slice is visible, and determines that the target slice is to be decoded.
[0601] In addition, the three-dimensional data decoding device may determine that the target slice is invisible (cannot be seen) when another slice exists between the camera and the target slice. Furthermore, the three-dimensional data decoding device may determine whether the target slice is visible or not by determining whether the relationship between the normal vector and the camera direction is within a predetermined angle range, rather than determining whether the normal vector and the camera direction are completely opposite to each other.
[0602] Next, processing using Level of Detail (LoD) will be described. An example of decoding processing according to layers with different resolutions will be described below.
[0603] FIG. 126 is a diagram showing an example of levels (LoD). FIG. 127 is a diagram showing an example of an octree structure. Each brick is divided into layers to control the level of resolution to be decoded. For example, a level is the depth of division when dividing into an octree. As shown in FIG. 126, the number of voxels included in each level is 2. (3×レベル) It should be noted that a different definition may be used for the division method or the number of voxels according to the level.
[0604] By using LoD, the 3D data decoder can realize high-speed visibility judgment and distance calculation. The decoding time affects real-time rendering. By using LoD, intermediate bricks can be displayed, real-time rendering and smooth response can be realized.
[0605] FIG. 128 is a flowchart of a three-dimensional data decoding process using LoD. First, the three-dimensional data decoding device determines the level to be decoded according to the purpose (S8751). Next, the three-dimensional data decoding device decodes the first level (level 0) (S8752). Next, the three-dimensional data decoding device determines whether or not the decoding of all levels to be decoded is completed (S8753). If the decoding of all levels is not completed (No in S8753), the three-dimensional data decoding device decodes the next level (S8754). At this time, the three-dimensional data decoding device may decode the next level using data of the previous level. If the decoding of all levels to be decoded is completed (Yes in S8753), the three-dimensional data decoding device displays the decoded data (S8755).
[0606] In this way, the three-dimensional data decoding device decodes data up to the determined level, and does not decode data after the determined level. This reduces the amount of processing involved in decoding, and improves processing speed. Furthermore, the three-dimensional data decoding device displays data up to the determined level, and does not display data after the determined level. This reduces the amount of processing involved in display, and improves processing speed. The three-dimensional data decoding device may determine the level of a brick to be decoded, for example, based on the distance of the brick from the camera, or whether or not the brick is visible from the camera.
[0607] Next, an implementation example of processing using LoD will be described. FIG. 129 is a flowchart of a three-dimensional data decoding process. First, the three-dimensional data decoding device acquires encoded data (S8761). For example, the encoded data is point cloud data that has been encoded and compressed using an arbitrary encoding method. The encoded data may be in a bit stream format or a file format.
[0608] Next, the three-dimensional data decoding device obtains the normal vector and position information of the brick to be processed from the encoded data (S8762). For example, the three-dimensional data decoding device obtains the normal vector for each brick and the position information of the brick from metadata (SEI or data header) included in the encoded data. The three-dimensional data decoding device may determine the distance between the brick and the camera from the position information of the brick and the information of the camera position. The three-dimensional data decoding device may also determine the visibility of the brick (whether the brick faces the camera direction) from the normal vector and the camera direction.
[0609] Next, the three-dimensional data decoding device determines which brick to decode, and decodes the first level (level 0) of the determined brick (S8763). Fig. 130 is a diagram showing an example of a brick to be decoded. As shown in Fig. 130, the three-dimensional data decoding device decodes all visible bricks at level 0 resolution.
[0610] Next, the three-dimensional data decoding device determines whether or not to decode the next level of each brick according to the position information, and decodes the next level of the brick that is determined to be decoded (S8764). This process is repeated until the decoding process of all levels is completed (S8765). Specifically, the resolution of the brick closer to the position of the virtual camera is set to be high. For example, depending on resources such as memory, levels of decoding are gradually added, with priority given to bricks closer to the camera.
[0611] Figure 131 is a diagram showing an example of the level of decoding target of each brick. As shown in Figure 131, the three-dimensional data decoding device decodes bricks closer to the camera at a higher resolution and bricks farther from the camera at a lower resolution according to the distance from the camera. In addition, the three-dimensional data decoding device does not decode bricks that are not visible.
[0612] When decoding at all levels is completed (Yes in S8765), the three-dimensional data decoding device outputs the obtained three-dimensional point group (S8766).
[0613] So far, a method has been described in which the normal vector and bounding information for each slice data is calculated and encoded in a three-dimensional data encoding device, and the visibility and distance information is calculated in a three-dimensional data decoding device based on the normal vector and bounding information and sensor input information, and the slice to be decoded is determined. Below, an example will be described in which the visibility and distance information according to the camera direction is calculated and encoded in advance in the three-dimensional data encoding device for data for each slice.
[0614] 132 is a diagram illustrating an example of the syntax of a slice header of position information (Geometry slice header information). The slice header of position information includes number_of_angle, view_angle, and visibility.
[0615] The number_of_angle indicates the number of camera angles (camera directions). The view_angle indicates a camera angle, for example, a vector of the camera angle. The visibility indicates whether the slice is visible from the corresponding camera angle. Note that the number of view_angles may be variable or may be a predetermined fixed value. Also, when the number and value of view_angle are predetermined, the view_angle may be omitted.
[0616] Further, although an example of showing visibility according to camera angle has been shown here, as another example, the three-dimensional data encoding device may pre-calculate visibility according to camera position or camera parameters and store the calculated visibility in the encoded data.
[0617] 133 is a flowchart of a three-dimensional data encoding process. First, the three-dimensional data encoding device divides point cloud data into divided data (e.g., slices) (S8771). Next, the three-dimensional data encoding device encodes position information and attribute information for each divided data unit (S8772). In addition, the three-dimensional data encoding device stores visibility information corresponding to the camera angle in metadata for each divided data (S8773).
[0618] 134 is a flowchart of the three-dimensional data decoding process. First, the three-dimensional data decoding device acquires visibility information corresponding to the camera angle from the metadata for each divided data (S8776). Next, the three-dimensional data decoding device determines which divided data are visible from the desired camera angle based on the visibility information, and decodes the visible divided data (S8777).
[0619] Figures 135 and 136 are diagrams showing examples of point cloud data. In these figures, a, c, d, and e represent planes. Therefore, the three-dimensional data encoding device can perform slice division by utilizing the fact that the three-dimensional points of each slice have normal vectors in the same direction. A similar method can also be applied to tile division.
[0620] 137 to 140 are diagrams showing examples of the configuration of a system including a three-dimensional data encoding device, a three-dimensional data decoding device, and a display device.
[0621] In the example shown in Fig. 137, the three-dimensional data encoding device generates encoded data by encoding slice data, normal vectors for each slice, and bounding box information. The three-dimensional data decoding device identifies data to be decoded from the encoded data and sensor information, and generates decoded slice data by decoding the identified data. The display device displays the decoded slice data. In this configuration, the three-dimensional data decoding device can flexibly determine visibility information and whether to decode.
[0622] In the example shown in Fig. 138, the three-dimensional data encoding device generates encoded data by encoding slice data, normal vectors for each slice, and bounding box information. The three-dimensional data decoding device determines the data to be decoded and the order from the encoded data and sensor information, and decodes the determined data in the determined order. In this configuration, the three-dimensional data decoding device can first decode the data to be displayed first (e.g., 3, 4, 5), thereby improving the comfort of display.
[0623] In the example shown in Fig. 139, the three-dimensional data encoding device generates encoded data by encoding slice data and visibility information for each camera angle. The three-dimensional data decoding device identifies data to be decoded from the encoded data information and sensor information, and decodes the identified data. The three-dimensional data decoding device may further determine the decoding order. In this configuration, the three-dimensional data decoding device does not need to calculate visibility information, so the amount of processing performed by the three-dimensional data decoding device can be reduced.
[0624] In the example shown in Fig. 140, the three-dimensional data decoding device notifies the three-dimensional data encoding device of the camera angle or camera position of the three-dimensional data decoding device via communication or the like. The three-dimensional data encoding device calculates visibility information for each slice, determines the data to be encoded and the order, and generates encoded data by encoding the determined data in the determined order. The three-dimensional data decoding device decodes the transmitted slice data as is. In this configuration, the amount of processing and communication bandwidth can be reduced by using an interactive configuration to encode and decode the necessary parts.
[0625] In addition, when the camera position or camera angle changes, the three-dimensional data decoding device may re-determine the slice to be decoded if the amount of change exceeds a predetermined value. In this case, high-speed decoding and display are possible by decoding the difference data other than the data that has already been decoded.
[0626] A method of storing encoded data in a file format such as ISOBMFF will be described below. Fig. 141 is a diagram showing an example of the configuration of a bitstream. Fig. 142 is a diagram showing an example of the configuration of a three-dimensional data encoding device. The three-dimensional data encoding device includes an encoding unit 8741 and a file conversion unit 8742. The encoding unit 8741 generates a bitstream including encoded data and control information by encoding point cloud data. The file conversion unit 8742 converts the bitstream into a file format.
[0627] 143 is a diagram showing an example of the configuration of a three-dimensional data decoding device. The three-dimensional data decoding device includes a file inverse conversion unit 8751 and a decoding unit 8752. The file inverse conversion unit 8751 converts a file format into a bit stream including encoded data and control information. The decoding unit 8752 generates point cloud data by decoding the bit stream.
[0628] Figure 144 is a diagram showing the basic structure of ISOBMFF. Figure 145 is a protocol stack diagram when NAL units common to PCC codecs are stored in ISOBMFF. Here, it is the NAL units of the PCC codec that are stored in ISOBMFF.
[0629] NAL units include NAL units for data and NAL units for metadata. NAL units for data include geometry slice data and attribute slice data. NAL units for metadata include SPS, GPS, APS, and SEI.
[0630] ISOBMFF (ISO based media file format) is a file format standard defined in ISO / IEC14496-12. It specifies a format that can store multiplexed data of various media such as video, audio, and text, and is a media-independent standard.
[0631] The basic unit in ISOBMFF is a box. A box consists of type, length, and data, and a collection of boxes of various types is a file. A file mainly consists of boxes such as ftyp, which indicates the brand of the file using 4CC, moov, which stores metadata such as control information, and mdat, which stores data.
[0632] For example, the method of storing AVC video and HEVC video is specified in ISO / IEC 14496-15. In addition, it is possible to extend the functionality of ISOBMFF to store and transmit PCC encoded data.
[0633] When storing a metadata NAL unit in ISOBMFF, the SEI may be stored in an "mdat box" together with PCC data, or in a "track box" that describes control information related to the stream. When data is packetized and transmitted, the SEI may be stored in a packet header. By indicating the SEI to the system layers, access to attribute information, tiles, and slice data becomes easier and the access speed is improved.
[0634] Next, a method for generating a PCC random access table will be described. The three-dimensional data encoding device generates a random access table using metadata including bounding box information and normal vector information for each slice. Figure 146 is a diagram showing an example of converting a bit stream into a file format.
[0635] The three-dimensional data encoding device stores the slice data in the mdat of the file format. The three-dimensional data encoding device calculates the memory position of the slice data as offset information at the beginning of the file (offsets 1 to 4 in FIG. 146), and includes the calculated offset information in the random access table (PCC random access table).
[0636] Fig. 147 is a diagram showing an example of the syntax of slice information (slice_information). Figs. 148 to 150 are diagrams showing an example of the syntax of a PCC random access table.
[0637] The PCC random access table includes bounding box information (bounding_box_info), normal vector information (normal_vector_info), and offset information (offset), which are stored in slice information (slice_information).
[0638] The three-dimensional data decoder analyzes the PCC random access table to identify the slice to be decoded. The three-dimensional data decoder can access the desired data by obtaining offset information from the PCC random access table.
[0639] As described above, the three-dimensional data encoding device according to this embodiment performs the process shown in Fig. 151. The three-dimensional data encoding device generates a bit stream by encoding position information and one or more pieces of attribute information of each of a plurality of three-dimensional points included in the point cloud data (S8781), and in the encoding (S8781), encodes a normal vector of each of the plurality of three-dimensional points as one piece of attribute information included in the one or more pieces of attribute information.
[0640] According to this, the three-dimensional data encoding device can process the normal vector in the same way as other attribute information by encoding the normal vector as attribute information. Therefore, the three-dimensional data encoding device can reduce the amount of processing. In other words, the three-dimensional data encoding device can encode the normal vector as attribute information without changing the definition of the attribute information, etc.
[0641] For example, in the encoding (S8781), the three-dimensional data encoding device converts a normal vector expressed by a floating point into an integer and then encodes it. This allows the three-dimensional data encoding device to process the normal vector in the same way as other attribute information, for example, when the other attribute information is expressed by an integer.
[0642] For example, the bit stream includes position information and control information (e.g., SPS) common to one or more attribute information, and the control information (e.g., SPS) includes at least one of information indicating that one attribute information included in the one or more attribute information indicates a normal vector (e.g., attribute_type=Normal Vector), or information indicating that the normal vector is data having three elements for each point.
[0643] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor uses the memory to perform the above-mentioned processing.
[0644] Moreover, the three-dimensional data decoding device according to this embodiment performs the process shown in Fig. 152. The three-dimensional data decoding device acquires a bit stream generated by encoding position information and one or more pieces of attribute information of each of a plurality of three-dimensional points included in point cloud data, in which normal vectors of each of the plurality of three-dimensional points are encoded as one piece of attribute information included in the one or more pieces of attribute information (S8786), and acquires the normal vector by decoding the one piece of attribute information from the bit stream (S8787).
[0645] According to this, the three-dimensional data decoding device decodes the normal vector as attribute information, and can process the normal vector in the same way as other attribute information, thereby reducing the amount of processing.
[0646] For example, the three-dimensional data decoding device acquires a normal vector represented by an integer in the acquisition of the normal vector (S8787). This allows the three-dimensional data decoding device to process the normal vector in the same way as other attribute information, for example, when the other attribute information is expressed by an integer.
[0647] For example, the bit stream includes position information and control information (e.g., SPS) common to one or more attribute information, and the control information (e.g., SPS) includes at least one of information indicating that one attribute information included in the one or more attribute information indicates a normal vector (e.g., attribute_type=Normal Vector), or information indicating that the normal vector is data having three elements for each point.
[0648] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor uses the memory to perform the above-mentioned processing.
[0649] Moreover, the three-dimensional data encoding device according to this embodiment performs the process shown in Fig. 153. The three-dimensional data encoding device divides point cloud data into a plurality of divided data (for example, bricks, slices, or tiles) (S8791), and generates a bit stream by encoding the plurality of divided data (S8792). The bit stream includes information indicating the normal vector of each of the plurality of divided data.
[0650] According to this, the three-dimensional data encoding device can reduce the amount of processing and the amount of code by encoding the normal vector for each divided data, compared to the case where the normal vector is encoded for each point. For example, each of the multiple divided data is a random access unit.
[0651] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor uses the memory to perform the above-mentioned processing.
[0652] Moreover, the three-dimensional data decoding device according to this embodiment performs the process shown in Fig. 154. The three-dimensional data decoding device acquires a bit stream generated by encoding a plurality of split data (e.g., bricks, slices, or tiles) generated by splitting point cloud data (S8796), and acquires information indicating the normal vector of each of the plurality of split data from the bit stream (S8797).
[0653] According to this, the three-dimensional data decoding device can reduce the amount of processing by decoding the normal vector for each divided data, compared to decoding the normal vector for each point. For example, each of the multiple divided data is a random access unit.
[0654] For example, the three-dimensional data decoding device further determines split data to be decoded from the multiple split data based on the normal vector, and decodes the split data to be decoded.
[0655] For example, the three-dimensional data decoding device further determines a decoding order for the multiple data segments based on the normal vector, and decodes the multiple data segments in the determined decoding order.
[0656] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor uses the memory to perform the above-mentioned processing.
[0657] (Embodiment 8) Next, a configuration of a three-dimensional data creation device 810 according to this embodiment will be described. Fig. 155 is a block diagram showing an example of the configuration of a three-dimensional data creation device 810 according to this embodiment. This three-dimensional data creation device 810 is mounted on a vehicle, for example. The three-dimensional data creation device 810 transmits and receives three-dimensional data to and from an external traffic monitoring cloud, a leading vehicle, or a following vehicle, and creates and stores three-dimensional data.
[0658] The three-dimensional data creation device 810 includes a data receiving unit 811, a communication unit 812, a receiving control unit 813, a format conversion unit 814, multiple sensors 815, a three-dimensional data creation unit 816, a three-dimensional data synthesis unit 817, a three-dimensional data storage unit 818, a communication unit 819, a transmission control unit 820, a format conversion unit 821, and a data transmission unit 822.
[0659] The data receiving unit 811 receives three-dimensional data 831 from a traffic monitoring cloud or a preceding vehicle. The three-dimensional data 831 includes information such as a point cloud, a visible light image, depth information, sensor position information, or speed information, including an area that cannot be detected by the sensor 815 of the vehicle itself.
[0660] The communication unit 812 communicates with the traffic monitoring cloud or the vehicle ahead, and transmits data transmission requests, etc. to the traffic monitoring cloud or the vehicle ahead.
[0661] The reception control unit 813 exchanges information such as compatible formats with the communication destination via the communication unit 812, and establishes communication with the communication destination.
[0662] The format conversion unit 814 generates three-dimensional data 832 by performing format conversion or the like on the three-dimensional data 831 received by the data receiving unit 811. Furthermore, when the three-dimensional data 831 is compressed or encoded, the format conversion unit 814 performs decompression or decoding processing.
[0663] The multiple sensors 815 are a group of sensors such as LiDAR, a visible light camera, or an infrared camera that acquires information about the outside of the vehicle, and generate sensor information 833. For example, when the sensor 815 is a laser sensor such as LiDAR, the sensor information 833 is three-dimensional data such as a point cloud (point cloud data). Note that the number of sensors 815 does not need to be multiple.
[0664] The three-dimensional data creation unit 816 generates three-dimensional data 834 from the sensor information 833. The three-dimensional data 834 includes information such as a point cloud, a visible light image, depth information, sensor position information, or velocity information.
[0665] The three-dimensional data synthesis unit 817 synthesizes three-dimensional data 834 created based on the host vehicle's sensor information 833 with three-dimensional data 832 created by a traffic monitoring cloud or a preceding vehicle, etc., to construct three-dimensional data 835 that includes the space in front of the preceding vehicle that cannot be detected by the host vehicle's sensor 815.
[0666] The three-dimensional data storage unit 818 stores the generated three-dimensional data 835 and the like.
[0667] The communication unit 819 communicates with the traffic monitoring cloud or the following vehicle, and transmits data transmission requests, etc. to the traffic monitoring cloud or the following vehicle.
[0668] The transmission control unit 820 exchanges information such as compatible formats with the communication destination and establishes communication with the communication destination via the communication unit 819. In addition, the transmission control unit 820 determines a transmission region, which is the space of the three-dimensional data to be transmitted, based on three-dimensional data construction information of the three-dimensional data 832 generated by the three-dimensional data synthesis unit 817 and a data transmission request from the communication destination.
[0669] Specifically, the transmission control unit 820 determines a transmission region including a space in front of the vehicle that cannot be detected by the sensor of the following vehicle in response to a data transmission request from the traffic monitoring cloud or the following vehicle. The transmission control unit 820 also determines the transmission region by judging whether or not a transmittable space or a transmitted space has been updated based on the three-dimensional data construction information. For example, the transmission control unit 820 determines the region specified in the data transmission request and in which the corresponding three-dimensional data 835 exists as the transmission region. Then, the transmission control unit 820 notifies the format conversion unit 821 of the format supported by the communication destination and the transmission region.
[0670] The format conversion unit 821 converts three-dimensional data 836 of the transmission region, among the three-dimensional data 835 stored in the three-dimensional data storage unit 818, into a format supported by the receiving side, thereby generating three-dimensional data 837. Note that the format conversion unit 821 may reduce the amount of data by compressing or encoding the three-dimensional data 837.
[0671] The data transmission unit 822 transmits the three-dimensional data 837 to the traffic monitoring cloud or the following vehicle. The three-dimensional data 837 includes information such as a point cloud, visible light image, depth information, or sensor position information in front of the vehicle, including blind spots of the following vehicle.
[0672] In the above description, the format conversion units 814 and 821 perform format conversion, but the format conversion does not necessarily have to be performed.
[0673] With this configuration, the three-dimensional data creation device 810 externally acquires three-dimensional data 831 of an area that cannot be detected by the sensor 815 of the host vehicle, and generates three-dimensional data 835 by synthesizing the three-dimensional data 831 with three-dimensional data 834 based on sensor information 833 detected by the sensor 815 of the host vehicle. In this way, the three-dimensional data creation device 810 can generate three-dimensional data of an area that cannot be detected by the sensor 815 of the host vehicle.
[0674] In addition, in response to a data transmission request from a traffic monitoring cloud or a following vehicle, the three-dimensional data creation device 810 can transmit three-dimensional data including the space in front of the vehicle that cannot be detected by the sensors of the following vehicle to the traffic monitoring cloud or the following vehicle, etc.
[0675] Next, there will be described a procedure for transmitting three-dimensional data to a following vehicle in the three-dimensional data creation device 810. Fig. 156 is a flowchart showing an example of a procedure for transmitting three-dimensional data to a traffic monitoring cloud or a following vehicle by the three-dimensional data creation device 810.
[0676] First, the three-dimensional data creation device 810 generates and updates three-dimensional data 835 of a space including the space on the road ahead of the vehicle (S801). Specifically, the three-dimensional data creation device 810 constructs three-dimensional data 835 including the space ahead of the vehicle ahead that cannot be detected by the sensor 815 of the vehicle ahead, for example, by combining three-dimensional data 834 created based on the sensor information 833 of the vehicle ahead with three-dimensional data 831 created by the traffic monitoring cloud or the vehicle ahead.
[0677] Next, the three-dimensional data creation device 810 determines whether the three-dimensional data 835 contained in the transmitted space has changed (S802).
[0678] If a change occurs in the three-dimensional data 835 contained in the space that has already been transmitted, such as when a vehicle or person enters the space from outside (Yes in S802), the three-dimensional data creation device 810 transmits three-dimensional data including the three-dimensional data 835 of the space where the change has occurred to the traffic monitoring cloud or a following vehicle (S803).
[0679] The three-dimensional data creation device 810 may transmit the three-dimensional data of the space where a change has occurred in accordance with the transmission timing of the three-dimensional data transmitted at a predetermined interval, or may transmit the three-dimensional data immediately after detecting the change. In other words, the three-dimensional data creation device 810 may transmit the three-dimensional data of the space where a change has occurred in priority over the three-dimensional data transmitted at a predetermined interval.
[0680] In addition, the three-dimensional data creation device 810 may transmit all of the three-dimensional data of the space in which a change has occurred as the three-dimensional data of the space in which a change has occurred, or may transmit only the differences in the three-dimensional data (for example, information on three-dimensional points that have appeared or disappeared, or displacement information of three-dimensional points, etc.).
[0681] Furthermore, the three-dimensional data creation device 810 may transmit metadata regarding the danger avoidance operation of the vehicle, such as a sudden braking warning, to the following vehicle prior to the three-dimensional data of the space where a change has occurred. This allows the following vehicle to quickly recognize the sudden braking of the vehicle ahead, and to start danger avoidance operations, such as deceleration, earlier.
[0682] If there is no change in the three-dimensional data 835 contained in the transmitted space (No in S802), or after step S803, the three-dimensional data creation device 810 transmits the three-dimensional data contained in a space of a specified shape at a distance L ahead of the vehicle to the traffic monitoring cloud or a following vehicle (S804).
[0683] Furthermore, for example, the processes in steps S801 to S804 are repeatedly performed at predetermined time intervals.
[0684] Furthermore, when there is no difference between the three-dimensional data 835 of the space currently being transmitted and the three-dimensional map, the three-dimensional data creation device 810 does not need to transmit the three-dimensional data 837 of the space.
[0685] In this embodiment, the client device transmits sensor information obtained by the sensor to a server or another client device.
[0686] First, the configuration of the system according to this embodiment will be described. Fig. 157 is a diagram showing the configuration of a transmission / reception system for a 3D map and sensor information according to this embodiment. This system includes a server 901 and client devices 902A and 902B. When there is no particular need to distinguish between the client devices 902A and 902B, they are also referred to as client devices 902.
[0687] The client device 902 is, for example, an in-vehicle device mounted on a moving object such as a vehicle. The server 901 is, for example, a traffic monitoring cloud or the like, and is capable of communicating with a plurality of client devices 902.
[0688] The server 901 transmits a three-dimensional map composed of a point cloud to the client device 902. Note that the composition of the three-dimensional map is not limited to a point cloud, and may represent other three-dimensional data such as a mesh structure.
[0689] The client device 902 transmits sensor information acquired by the client device 902 to the server 901. The sensor information includes, for example, at least one of LiDAR acquisition information, a visible light image, an infrared image, a depth image, sensor position information, and speed information.
[0690] Data transmitted between the server 901 and the client device 902 may be compressed to reduce data, or may be left uncompressed to maintain data accuracy. When compressing data, a three-dimensional compression method based on an octree structure, for example, can be used for the point cloud. Also, a two-dimensional image compression method can be used for the visible light image, the infrared image, and the depth image. The two-dimensional image compression method is, for example, MPEG-4 AVC or HEVC standardized by MPEG.
[0691] Furthermore, the server 901 transmits the three-dimensional map managed by the server 901 to the client device 902 in response to a transmission request for the three-dimensional map from the client device 902. Note that the server 901 may transmit the three-dimensional map without waiting for a transmission request for the three-dimensional map from the client device 902. For example, the server 901 may broadcast the three-dimensional map to one or more client devices 902 in a predetermined space. Furthermore, the server 901 may transmit a three-dimensional map suitable for the position of the client device 902 at regular intervals to the client device 902 that has received a transmission request once. Furthermore, the server 901 may transmit the three-dimensional map to the client device 902 every time the three-dimensional map managed by the server 901 is updated.
[0692] The client device 902 issues a request for transmitting a three-dimensional map to the server 901. For example, when the client device 902 wants to estimate its own position while driving, the client device 902 transmits a request for transmitting a three-dimensional map to the server 901.
[0693] In the following cases, the client device 902 may issue a transmission request for a three-dimensional map to the server 901. If the three-dimensional map held by the client device 902 is old, the client device 902 may issue a transmission request for a three-dimensional map to the server 901. For example, if a certain period of time has passed since the client device 902 acquired the three-dimensional map, the client device 902 may issue a transmission request for a three-dimensional map to the server...
Claims
1. encoding, by a processor, the three-dimensional data including the position information and the one or more attribute information; generating, by the processor, control information common to the position information and the one or more pieces of attribute information; generating, by the processor, a bitstream including the encoded three-dimensional data and the control information; the control information includes one or more pieces of attribute type information corresponding to the one or more pieces of attribute information; When the first attribute type information included in the one or more attribute type information includes first information, the first attribute information included in the one or more attribute information relates to a normal vector, The first attribute information corresponds to the first attribute type information. Three-dimensional data encoding method.
2. In the encoding, the processor The normal vector expressed in floating point is converted to an integer and then encoded.
2. The method of claim 1, wherein the three-dimensional data is encoded using a three-dimensional encoding method.
3. When the second attribute type information included in the one or more attribute type information includes second information, the second attribute information included in the one or more attribute information relates to a color, The second attribute information corresponds to the second attribute type information.
2. The method of claim 1, wherein the three-dimensional data is encoded using a three-dimensional encoding method.
4. The three-dimensional data includes point cloud data.
2. The method of claim 1, wherein the three-dimensional data is encoded using a three-dimensional encoding method.
5. The control information is a sequence parameter set.
2. The method of claim 1, wherein the three-dimensional data is encoded using a three-dimensional encoding method.
6. obtaining, by a processor, a bitstream comprising encoded three-dimensional data, the bitstream comprising encoded position information and one or more encoded attribute information; obtaining, by the processor, control information common to the encoding position information and the one or more pieces of encoding attribute information from the bitstream; Decoding, by the processor, the one or more pieces of encoded attribute information; the control information includes one or more pieces of attribute type information corresponding to the one or more pieces of encoded attribute information; When the first attribute type information included in the one or more attribute type information includes first information, the first encoded attribute information included in the one or more encoded attribute information relates to a normal vector, The first encoded attribute information corresponds to the first attribute type information. A method for decoding three-dimensional data.
7. In the decoding, the processor: The first encoded attribute information is decoded to obtain the normal vector.
7. The three-dimensional data decoding method according to claim 6.
8. The obtained normal vector is expressed as an integer.
8. The three-dimensional data decoding method according to claim 7.
9. When the second attribute type information included in the one or more attribute type information includes second information, the second encoded attribute information included in the one or more encoded attribute information relates to color, The second encoded attribute information corresponds to the second attribute type information.
8. The three-dimensional data decoding method according to claim 7.
10. The encoded three-dimensional data includes encoded point cloud data 7. The three-dimensional data decoding method according to claim 6.
11. The control information is a sequence parameter set.
7. The three-dimensional data decoding method according to claim 6.
12. A processor; A memory. The processor uses the memory to: Encoding three-dimensional data including position information and one or more attribute information; generating control information common to the position information and the one or more pieces of attribute information; generating a bitstream including the encoded three-dimensional data and the control information; the control information includes one or more pieces of attribute type information corresponding to the one or more pieces of attribute information; When the first attribute type information included in the one or more attribute type information includes first information, the first attribute information included in the one or more attribute information relates to a normal vector, The first attribute information corresponds to the first attribute type information. Three-dimensional data encoding device.
13. A processor; A memory. The processor uses the memory to: obtaining a bitstream including encoded three-dimensional data including encoded position information and one or more encoded attribute information; obtaining control information common to the encoding position information and the one or more pieces of encoding attribute information from the bit stream; Decoding the one or more pieces of encoded attribute information; the control information includes one or more pieces of attribute type information corresponding to the one or more pieces of encoded attribute information; When the first attribute type information included in the one or more attribute type information includes first information, the first encoded attribute information included in the one or more encoded attribute information relates to a normal vector, The first encoded attribute information corresponds to the first attribute type information. Three-dimensional data decoding device.
Citation Information
Patent Citations
Improved attribute layers and signaling in point cloud coding
JP2022500931A
Method and system for processing, compressing, streaming, and interactive rendering of 3D color image data
US20030038798A1
Point cloud compression using non-cubic projections and masks
US20190087978A1
Transposition structures and methods to accommodate parallel processing in a graphics processing unit (“GPU”)
US7755631B1
Information processing device and method
WO2019065297A1