Three-dimensional data decoding method and three-dimensional data decoding device
The proposed three-dimensional data decoding method addresses the inefficiencies in existing decoding techniques by using visibility information and angle calculations to selectively decode encoded three-dimensional points, thereby enhancing processing efficiency and reducing bandwidth requirements.
Patent Information
- Application Number
- JP2025025400
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-10-18
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2040-10-16
AI Technical Summary
Existing methods for decoding three-dimensional data encoded in point clouds do not efficiently allow for the selection and decoding of specific encoded three-dimensional points, leading to inefficient data processing and transmission.
A three-dimensional data decoding method that acquires encoded data based on visibility information from a sensor and decodes the data by calculating angles between reference points to determine which encoded three-dimensional points to decode, thereby optimizing data processing and transmission.
This method enables efficient decoding of desired encoded three-dimensional data, reducing processing time and bandwidth requirements by selectively decoding only the necessary data based on user viewing angles and focal positions.
Smart Images

Figure 2025071216000001_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to a three-dimensional data decoding method and a three-dimensional data decoding device. [Background technology]
[0002] In the future, devices and services that utilize three-dimensional data are expected to become widespread in a wide range of fields, such as computer vision for autonomous operation of automobiles or robots, map information, surveillance, infrastructure inspection, video distribution, etc. Three-dimensional data is acquired in various ways, such as by distance sensors such as range finders, stereo cameras, or a combination of multiple monocular cameras.
[0003] One method of expressing three-dimensional data is a method called a point cloud, which represents the shape of a three-dimensional structure using a group of points in a three-dimensional space. In a point cloud, the positions and colors of the points are stored. Point clouds are expected to become the mainstream method of expressing three-dimensional data, but point clouds have a very large amount of data. Therefore, when storing or transmitting three-dimensional data, it is essential to compress the amount of data by encoding, just like two-dimensional video images (examples include MPEG-4 AVC or HEVC standardized by MPEG).
[0004] In addition, compression of point clouds is partially supported by public libraries (Point Cloud Library) that perform point cloud-related processing.
[0005] Furthermore, a technique is known that uses three-dimensional map data to search for and display facilities located around a vehicle (see, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0006] [Patent Document 1] International Publication No. 2014 / 020663 Summary of the Invention [Problem to be solved by the invention]
[0007] In encoded three-dimensional data (multiple three-dimensional points), it is desirable to be able to appropriately select and decode desired multiple encoded three-dimensional points from the multiple encoded three-dimensional points.
[0008] The present disclosure aims to provide a three-dimensional data decoding method or a three-dimensional data decoding device that can appropriately select and decode desired encoded three-dimensional points from among a plurality of encoded three-dimensional points. [Means for solving the problem]
[0009] A three-dimensional data decoding method according to one aspect of the present disclosure is a three-dimensional data decoding method executed by a three-dimensional data decoding device, which obtains encoded three-dimensional data based on visibility information regarding visibility from a sensor, and decodes the encoded three-dimensional data. Effect of the Invention
[0010] The present disclosure can provide a three-dimensional data decoding method or a three-dimensional data decoding device that can acquire and decode desired encoded three-dimensional data. [Brief description of the drawings]
[0011] [Figure 1] FIG. 1 is a diagram showing a configuration of a three-dimensional data encoding / decoding system according to the first embodiment. [Diagram 2] FIG. 2 is a diagram illustrating an example of a configuration of point cloud data according to the first embodiment. [Diagram 3] FIG. 3 is a diagram showing an example of a configuration of a data file in which point cloud data information according to the first embodiment is described. [Figure 4] FIG. 4 is a diagram showing types of point cloud data according to the first embodiment. [Diagram 5] FIG. 5 is a diagram showing a configuration of a first encoding unit according to the first embodiment. [Figure 6] FIG. 6 is a block diagram of a first encoding unit according to the first embodiment. [Figure 7] FIG. 7 is a diagram showing a configuration of a first decoding unit according to the first embodiment. [Figure 8] FIG. 8 is a block diagram of a first decoding unit according to the first embodiment. [Figure 9] FIG. 9 is a block diagram of a three-dimensional data encoding device according to the first embodiment. [Figure 10] FIG. 10 is a diagram showing an example of location information according to the first embodiment. [Figure 11] FIG. 11 is a diagram showing an example of an octree representation of position information according to the first embodiment. [Figure 12] FIG. 12 is a block diagram of the three-dimensional data decoding device according to the first embodiment. [Figure 13] FIG. 13 is a block diagram of the attribute information encoding unit according to the first embodiment. [Figure 14] FIG. 14 is a block diagram of the attribute information decoding unit according to the first embodiment. [Figure 15] FIG. 15 is a block diagram showing a configuration of an attribute information encoding unit according to the first embodiment. As shown in FIG. [Figure 16] FIG. 16 is a block diagram of the attribute information encoding unit according to the first embodiment. [Figure 17] FIG. 17 is a block diagram showing a configuration of an attribute information decoding unit according to the first embodiment. As shown in FIG. [Figure 18] FIG. 18 is a block diagram of the attribute information decoding unit according to the first embodiment. As shown in FIG. [Figure 19] FIG. 19 is a diagram showing a configuration of a second encoding unit according to the first embodiment. As shown in FIG. [Figure 20] FIG. 20 is a block diagram of a second encoding unit according to the first embodiment. [Figure 21] FIG. 21 is a diagram illustrating a configuration of a second decoding unit according to the first embodiment. As shown in FIG. [Figure 22] FIG. 22 is a block diagram of a second decoding unit according to the first embodiment. [Figure 23]FIG. 23 is a diagram showing a protocol stack related to PCC encoded data according to the first embodiment. [Figure 24] FIG. 24 is a diagram illustrating configurations of an encoding unit and a multiplexing unit according to the second embodiment. In FIG. [Diagram 25] FIG. 25 is a diagram illustrating an example of a structure of encoded data according to the second embodiment. In FIG. [Figure 26] FIG. 26 is a diagram showing an example of the structure of coded data and an NAL unit according to the second embodiment. [Figure 27] FIG. 27 is a diagram illustrating an example of the semantics of pcc_nal_unit_type according to the second embodiment. [Figure 28] FIG. 28 is a block diagram of a first encoding unit according to the third embodiment. [Figure 29] FIG. 29 is a block diagram of a first decoding unit according to the third embodiment. [Diagram 30] FIG. 30 is a block diagram of a division unit according to the third embodiment. [Diagram 31] FIG. 31 is a diagram showing an example of division into slices and tiles according to the third embodiment. [Diagram 32] FIG. 32 is a diagram showing an example of a division pattern of slices and tiles according to the third embodiment. [Diagram 33] FIG. 33 is a diagram illustrating an example of a dependency relationship according to the third embodiment. [Diagram 34] FIG. 34 is a diagram showing an example of a data decoding order according to the third embodiment. [Diagram 35] FIG. 35 is a flowchart of the encoding process according to the third embodiment. [Diagram 36] FIG. 36 is a block diagram of a coupling unit according to the third embodiment. [Figure 37] FIG. 37 is a diagram showing an example of the structure of coded data and an NAL unit according to the third embodiment. [Figure 38] FIG. 38 is a flowchart of the encoding process according to the third embodiment. [Figure 39] FIG. 39 is a flowchart of the decoding process according to the third embodiment. [Diagram 40] FIG. 40 is a diagram illustrating an example of the syntax of the tile additional information according to the fourth embodiment. [Diagram 41] FIG. 41 is a block diagram of a coding / decoding system according to the fourth embodiment. [Diagram 42] FIG. 42 is a diagram showing an example of the syntax of slice additional information according to the fourth embodiment. [Diagram 43] FIG. 43 is a flowchart of the encoding process according to the fourth embodiment. [Diagram 44] FIG. 44 is a flowchart of the decoding process according to the fourth embodiment. [Diagram 45] FIG. 45 is a diagram showing an example of a division method according to the fifth embodiment. [Figure 46] FIG. 46 is a diagram illustrating an example of division of point cloud data according to the fifth embodiment. [Figure 47] FIG. 47 is a diagram illustrating an example of the syntax of the tile additional information according to the fifth embodiment. [Figure 48] FIG. 48 is a diagram showing an example of index information according to the fifth embodiment. [Figure 49] FIG. 49 is a diagram illustrating an example of a dependency relationship according to the fifth embodiment. [Figure 50] FIG. 50 is a diagram showing an example of transmission data according to the fifth embodiment. In FIG. [Figure 51] FIG. 51 is a diagram showing an example of the structure of a NAL unit according to the fifth embodiment. [Figure 52] FIG. 52 is a diagram illustrating an example of a dependency relationship according to the fifth embodiment. In FIG. [Diagram 53] FIG. 53 is a diagram showing an example of a data decoding order according to the fifth embodiment. [Figure 54] FIG. 54 is a diagram showing an example of a dependency relationship according to the fifth embodiment. In FIG. [Figure 55] FIG. 55 is a diagram showing an example of a data decoding order according to the fifth embodiment. [Figure 56] FIG. 56 is a flowchart of the encoding process according to the fifth embodiment. [Figure 57]FIG. 57 is a flowchart of the decoding process according to the fifth embodiment. [Figure 58] FIG. 58 is a flowchart of the encoding process according to the fifth embodiment. [Figure 59] FIG. 59 is a flowchart of the encoding process according to the fifth embodiment. [Figure 60] FIG. 60 is a diagram showing an example of transmission data and reception data according to the fifth embodiment. In FIG. [Figure 61] FIG. 61 is a flowchart of the decoding process according to the fifth embodiment. [Figure 62] FIG. 62 is a diagram showing an example of transmission data and reception data according to the fifth embodiment. In FIG. [Figure 63] FIG. 63 is a flowchart of the decoding process according to the fifth embodiment. [Figure 64] FIG. 64 is a flowchart of the encoding process according to the fifth embodiment. [Figure 65] FIG. 65 is a diagram showing an example of index information according to the fifth embodiment. [Figure 66] FIG. 66 is a diagram showing an example of a dependency relationship according to the fifth embodiment. In FIG. [Figure 67] FIG. 67 is a diagram showing an example of transmission data according to the fifth embodiment. In FIG. [Figure 68] FIG. 68 is a diagram showing an example of transmission data and reception data according to the fifth embodiment. In FIG. [Figure 69] FIG. 69 is a flowchart of the decoding process according to the fifth embodiment. [Figure 70] FIG. 70 is a diagram illustrating an example of syntax of the GPS according to the sixth embodiment. [Figure 71] FIG. 71 is a flowchart of three-dimensional data decoding processing according to the sixth embodiment. [Figure 72] FIG. 72 is a diagram showing an example of an application according to the sixth embodiment. In FIG. [Figure 73] FIG. 73 is a diagram showing an example of tile division and slice division according to the sixth embodiment. In FIG. [Figure 74]FIG. 74 is a flowchart of processing in the system according to the sixth embodiment. [Figure 75] FIG. 75 is a flowchart of processing in the system according to the sixth embodiment. [Figure 76] FIG. 76 is a block diagram of a three-dimensional data encoding device according to the seventh embodiment. [Figure 77] FIG. 77 is a block diagram of a three-dimensional data decoding device according to the seventh embodiment. [Figure 78] FIG. 78 is a block diagram of a three-dimensional data encoding device according to the seventh embodiment. [Figure 79] FIG. 79 is a block diagram showing a configuration of a three-dimensional data decoding device according to the seventh embodiment. [Figure 80] FIG. 80 is a diagram showing an example of point cloud data according to the seventh embodiment. As shown in FIG. [Figure 81] FIG. 81 is a diagram showing an example of normal vectors for each point according to the seventh embodiment. [Figure 82] FIG. 82 is a diagram illustrating an example of the syntax of a normal vector according to the seventh embodiment. [Figure 83] FIG. 83 is a flowchart of three-dimensional data encoding processing according to the seventh embodiment. [Figure 84] FIG. 84 is a flowchart of three-dimensional data decoding processing according to the seventh embodiment. [Figure 85] FIG. 85 is a diagram showing an example of a configuration of a bit stream according to the seventh embodiment. [Figure 86] FIG. 86 is a diagram showing an example of point cloud information according to the seventh embodiment. In FIG. [Figure 87] FIG. 87 is a flowchart of three-dimensional data encoding processing according to the seventh embodiment. [Figure 88] FIG. 88 is a flowchart of three-dimensional data decoding processing according to the seventh embodiment. [Figure 89] FIG. 89 is a diagram illustrating an example of division of a normal vector according to the seventh embodiment. [Figure 90]FIG. 90 is a diagram showing an example of division of a normal vector according to the seventh embodiment. [Figure 91] FIG. 91 is a diagram illustrating an example of point cloud data according to the seventh embodiment. As shown in FIG. [Figure 92] FIG. 92 is a diagram illustrating an example of a normal vector according to the seventh embodiment. In FIG. [Figure 93] FIG. 93 is a diagram showing an example of information on normal vectors according to the seventh embodiment. In FIG. [Figure 94] FIG. 94 is a diagram showing an example of a cube according to the seventh embodiment. In FIG. [Figure 95] FIG. 95 is a diagram showing an example of a face of a cube according to the seventh embodiment. In FIG. [Figure 96] FIG. 96 is a diagram showing an example of a face of a cube according to the seventh embodiment. In FIG. [Figure 97] FIG. 97 is a diagram showing an example of a face of a cube according to the seventh embodiment. In FIG. [Figure 98] FIG. 98 is a diagram showing an example of visibility of slices according to the seventh embodiment. [Figure 99] FIG. 99 is a diagram showing an example of a bitstream configuration according to the seventh embodiment. [Figure 100] FIG. 100 is a diagram illustrating an example of the syntax of a slice header of position information according to the seventh embodiment. [Figure 101] FIG. 101 is a diagram illustrating an example of the syntax of a slice header of position information according to the seventh embodiment. [Figure 102] FIG. 102 is a flowchart of three-dimensional data encoding processing according to the seventh embodiment. [Figure 103] FIG. 103 is a flowchart of three-dimensional data decoding processing according to the seventh embodiment. [Figure 104] FIG. 104 is a flowchart of three-dimensional data decoding processing according to the seventh embodiment. [Figure 105] FIG. 105 is a diagram showing an example of a configuration of a bitstream according to the seventh embodiment. [Fig. 106] FIG. 106 is a diagram illustrating an example of the syntax of slice information according to the seventh embodiment. [Figure 107] FIG. 107 is a diagram illustrating an example of the syntax of slice information according to the seventh embodiment. [Figure 108] FIG. 108 is a flowchart of three-dimensional data decoding processing according to the seventh embodiment. [Fig. 109] FIG. 109 is a diagram illustrating an example of partial decoding processing according to the seventh embodiment. [Figure 110] FIG. 110 is a diagram showing an example of the configuration of a three-dimensional data decoding device according to the seventh embodiment. [Figure 111] FIG. 111 is a diagram illustrating an example of processing by the random access control unit according to the seventh embodiment. In FIG. [Figure 112] FIG. 112 is a diagram illustrating an example of processing by the random access control unit according to the seventh embodiment. In FIG. [Figure 113] FIG. 113 is a diagram showing an example of the relationship between distance and resolution according to the seventh embodiment. In FIG. [Fig. 114] FIG. 114 is a diagram showing an example of bricks and normal vectors according to the seventh embodiment. [Fig. 115] FIG. 115 is a diagram showing an example of levels according to the seventh embodiment. In FIG. [Fig. 116] FIG. 116 is a diagram showing an example of an octree structure according to the seventh embodiment. [Figure 117] FIG. 117 is a flowchart of three-dimensional data decoding processing according to the seventh embodiment. [Figure 118] FIG. 118 is a flowchart of three-dimensional data decoding processing according to the seventh embodiment. [Figure 119] FIG. 119 is a diagram showing an example of a brick to be decoded according to the seventh embodiment. [Figure 120] FIG. 120 is a diagram showing an example of levels of decoding targets according to the seventh embodiment. In FIG. [Figure 121] FIG. 121 is a diagram illustrating an example of the syntax of a slice header of position information according to the seventh embodiment. [Figure 122] FIG. 122 is a flowchart of three-dimensional data encoding processing according to the seventh embodiment. [Figure 123] FIG. 123 is a flowchart of three-dimensional data decoding processing according to the seventh embodiment. [Figure 124] FIG. 124 is a diagram showing an example of point cloud data according to the seventh embodiment. As shown in FIG. [Fig. 125] FIG. 125 is a diagram showing an example of point cloud data according to the seventh embodiment. As shown in FIG. [Fig. 126] FIG. 126 is a diagram illustrating a configuration example of a system according to the seventh embodiment. In FIG. [Figure 127] FIG. 127 is a diagram illustrating a configuration example of a system according to the seventh embodiment. As shown in FIG. [Figure 128] FIG. 128 is a diagram illustrating a configuration example of a system according to the seventh embodiment. In FIG. [Figure 129] FIG. 129 is a diagram illustrating a configuration example of a system according to the seventh embodiment. As shown in FIG. [Fig. 130] FIG. 130 is a diagram showing an example of a configuration of a bit stream according to the seventh embodiment. [Fig. 131] FIG. 131 is a diagram illustrating an example of the configuration of a three-dimensional data encoding device according to the seventh embodiment. [Fig. 132] FIG. 132 is a diagram showing an example of the configuration of a three-dimensional data decoding device according to the seventh embodiment. [Fig. 133] FIG. 133 is a diagram showing the basic structure of an ISOBMFF according to the seventh embodiment. [Fig. 134] FIG. 134 is a protocol stack diagram for storing a NAL unit common to the PCC codec in the seventh embodiment in an ISOBMFF. [Fig. 135] FIG. 135 is a diagram showing an example of converting a bitstream according to the seventh embodiment into a file format. [Fig. 136] FIG. 136 is a diagram showing an example of the syntax of slice information according to the seventh embodiment. [Fig. 137] FIG. 137 is a diagram showing an example of syntax of a PCC random access table relating to embodiment 7. [Figure 138]FIG. 138 is a diagram showing an example of syntax of a PCC random access table according to embodiment 7. [Figure 139] FIG. 139 is a diagram showing an example of syntax of a PCC random access table relating to embodiment 7. [Fig. 140] FIG. 140 is a flowchart of three-dimensional data encoding processing according to the seventh embodiment. [Fig. 141] FIG. 141 is a flowchart of three-dimensional data decoding processing according to the seventh embodiment. [Fig. 142] FIG. 142 is a flowchart of three-dimensional data encoding processing according to the seventh embodiment. [Fig. 143] FIG. 143 is a flowchart of three-dimensional data decoding processing according to the seventh embodiment. [Fig. 144] FIG. 144 is a block diagram of a three-dimensional data creation device according to the fifth embodiment. [Fig. 145] FIG. 145 is a flowchart of a three-dimensional data creating method according to the fifth embodiment. [Fig. 146] FIG. 146 is a diagram showing a configuration of a system according to the eighth embodiment. In FIG. [Fig. 147] FIG. 147 is a block diagram of a client device according to the eighth embodiment. [Fig. 148] FIG. 148 is a block diagram of a server according to the eighth embodiment. [Figure 149] FIG. 149 is a flowchart of three-dimensional data creation processing by a client device according to the eighth embodiment. [Fig. 150] FIG. 150 is a flowchart of a sensor information transmission process by a client device according to the eighth embodiment. [Fig. 151] FIG. 151 is a flowchart of three-dimensional data creation processing by the server according to the eighth embodiment. [Fig. 152] FIG. 152 is a flowchart of a 3D map transmission process by the server according to the eighth embodiment. [Fig. 153]FIG. 153 is a diagram showing a configuration of a modified example of the system according to the eighth embodiment. In FIG. [Fig. 154] FIG. 154 is a diagram showing the configurations of a server and a client device according to the eighth embodiment. In FIG. [Fig. 155] FIG. 155 is a diagram showing the configurations of a server and a client device according to the eighth embodiment. In FIG. [Fig. 156] FIG. 156 is a flowchart of processing by a client device according to the eighth embodiment. [Fig. 157] FIG. 157 is a diagram illustrating a configuration of a sensor information collecting system according to the eighth embodiment. As shown in FIG. [Fig. 158] FIG. 158 is a diagram illustrating an example of a system according to the eighth embodiment. In FIG. [Fig. 159] FIG. 159 is a diagram showing a modification of the system according to the eighth embodiment. [Fig. 160] FIG. 160 is a flowchart showing an example of application processing according to the eighth embodiment. [Fig. 161] FIG. 161 is a diagram showing the sensor ranges of various sensors according to the eighth embodiment. [Fig. 162] FIG. 162 is a diagram illustrating a configuration example of an autonomous driving system according to the eighth embodiment. [Fig. 163] FIG. 163 is a diagram showing an example of a bitstream configuration according to the eighth embodiment. In FIG. [Fig. 164] FIG. 164 is a flowchart of the point group selection process according to the eighth embodiment. [Fig. 165] FIG. 165 is a diagram showing an example of a screen for the point group selection process according to the eighth embodiment. [Fig. 166] FIG. 166 is a diagram showing an example of a screen for the point group selection process according to the eighth embodiment. [Fig. 167] FIG. 167 is a diagram showing an example of a screen for the point group selection process according to the eighth embodiment. [Fig. 168] FIG. 168 is a diagram for explaining the center positions of tiles according to the ninth embodiment. [Fig. 169]FIG. 169 is a block diagram for explaining the process of decoding an encoded 3D point group according to the ninth embodiment. [Fig. 170] FIG. 170 is a diagram for explaining the process of calculating angle information used by the three-dimensional data decoding device according to the ninth embodiment for determining visibility. [Fig. 171] FIG. 171 is a block diagram showing a configuration for identifying a decoding target according to the ninth embodiment. In FIG. [Fig. 172] FIG. 172 is a flowchart showing a process in which the three-dimensional data decoding device according to the ninth embodiment determines a decoding target. [Fig. 173] FIG. 173 is a diagram for explaining a first example of the decoding process of the three-dimensional data decoding device according to the ninth embodiment. [Fig. 174] FIG. 174 is a diagram for explaining a second example of the decoding process of the three-dimensional data decoding device according to the ninth embodiment. [Fig. 175] FIG. 175 is a diagram for explaining a third example of the decoding process of the three-dimensional data decoding device according to the ninth embodiment. [Fig. 176] FIG. 176 is a diagram for explaining a fourth example of the decoding process of the three-dimensional data decoding device according to the ninth embodiment. [Fig. 177] FIG. 177 is a diagram for explaining the decoding process of the three-dimensional data decoding device according to the first modification of the ninth embodiment. [Fig. 178] FIG. 178 is a diagram for explaining the process of calculating angle information used by the three-dimensional data decoding device according to the first modification of the ninth embodiment for determining visibility. In FIG. [Fig. 179] FIG. 179 is a flowchart showing a process in which the three-dimensional data decoding device according to the second modification of the ninth embodiment determines the resolution based on the angle information. [Fig. 180] FIG. 180 is a flowchart showing the decoding process of the three-dimensional data decoding device according to the ninth embodiment and the modification. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] A three-dimensional data decoding method according to one embodiment of the present disclosure obtains a tile including a plurality of encoded three-dimensional points, calculates the angle between a line segment connecting a predetermined position in the tile and a first reference point and a line segment connecting the first reference point and a second reference point different from the first reference point, determines whether the calculated angle satisfies a predetermined condition, and if it is determined that the angle satisfies the predetermined condition, decodes the plurality of encoded three-dimensional points included in the tile, and if it is determined that the angle does not satisfy the predetermined condition, does not decode the plurality of encoded three-dimensional points included in the tile.
[0013] For example, the three-dimensional data decoding device sequentially decodes a plurality of three-dimensional points that are encoded in the data order contained in a bit stream (encoded data) that includes a plurality of three-dimensional points encoded by a three-dimensional data encoding device. Here, for example, when the three-dimensional data indicates a three-dimensional map, when a user looks at the map, the user often wants to see a map in a specific direction from the user's current position. In such a case, the three-dimensional data decoding device sequentially decodes a plurality of three-dimensional points that are encoded in the data order contained in the bit stream, and causes a display device or the like to sequentially display images indicating the plurality of three-dimensional points in the order in which they were decoded. In this case, it may take time for the location that the user wants to confirm to be displayed on the display device. In addition, depending on the user's current position, there may be three-dimensional points that do not need to be decoded. Therefore, in the three-dimensional data decoding method according to the present disclosure, the line of sight direction of the camera, user, etc. is assumed using the positional relationship between any two points, such as the positions of the camera, user, etc. and the focal positions of the camera, user, etc., and the tiles, and the encoded tiles located in the line of sight direction are decoded. According to this, it is possible to decode encoded tiles located in an area that is likely to be desired by the user. In other words, according to the three-dimensional data decoding method according to the present disclosure, a user can appropriately select and decode desired encoded tiles (i.e., a plurality of three-dimensional points) from among the encoded tiles.
[0014] Also, for example, in the process of determining whether the angle satisfies the specified condition, it is determined whether the calculated angle is less than a specified angle, and in the process of decoding the multiple encoded three-dimensional points contained in the tile, if the calculated angle is less than the specified angle, the multiple encoded three-dimensional points are decoded, and if the calculated angle is equal to or greater than the specified angle, the multiple encoded three-dimensional points are not decoded.
[0015] According to this, by appropriately setting the predetermined angle, tiles located within a range desired by the user in accordance with the viewing angle of the camera, etc. can be appropriately decoded.
[0016] Also, for example, in the process of decoding the encoded three-dimensional points contained in the tile, a resolution of the encoded three-dimensional points is determined based on the angle, and the encoded three-dimensional points are decoded according to the determined resolution.
[0017] For example, the center of the user's field of view is more likely to be important to the user than the outer edge of the user's field of view. Therefore, for example, tiles are decoded so that the resolution of tiles located in the center of the field of view of the camera, user, etc. is increased and the resolution of tiles located in the outer edge of the field of view of the camera, user, etc. is decreased according to the calculated angle. This can reduce the amount of processing by increasing the resolution of tiles that may be relatively important to the user while decreasing the resolution of tiles that are likely to be relatively unimportant to the user.
[0018] Also, for example, the first reference point indicates a position of a camera, and the second reference point indicates a focal position of the camera.
[0019] Alternatively, for example, the second reference point indicates a position of a camera, and the first reference point indicates a focal position of the camera.
[0020] According to these, for example, tiles at positions corresponding to the position and imaging direction of a camera when the camera actually captures an image can be decoded. Therefore, for example, tiles for generating a 3D point image corresponding to an image obtained from the camera can be appropriately selected and decoded.
[0021] Also, for example, in the process of calculating the angle, the angle is calculated by calculating the dot product of a vector extending from the first reference point to the specified position and a vector extending from the first reference point to the second reference point.
[0022] This allows the angle to be calculated through simple processing.
[0023] Furthermore, for example, the predetermined position is the position in the coordinate space in which the tile is located that has the smallest coordinate value, the center position of the tile, or the center of gravity of the tile.
[0024] In other words, it is preferable to adopt a position of the tile that is included in the tile and that can be easily calculated, so that the position of the tile can be specified by a simple process.
[0025] Moreover, a three-dimensional data decoding device according to one embodiment of the present disclosure includes a processor and a memory, and the processor uses the memory to acquire a tile including a plurality of encoded three-dimensional points, calculates an angle between a line segment connecting a predetermined position in the tile and a first reference point and a line segment connecting the first reference point and a second reference point different from the first reference point, determines whether the calculated angle satisfies a predetermined condition, and if it is determined that the angle satisfies the predetermined condition, decodes the plurality of encoded three-dimensional points included in the tile, and if it is determined that the angle does not satisfy the predetermined condition, does not decode the plurality of encoded three-dimensional points included in the tile.
[0026] For example, the three-dimensional data decoding device sequentially decodes a plurality of three-dimensional points that are encoded in the data order contained in a bit stream (encoded data) that includes a plurality of three-dimensional points encoded by a three-dimensional data encoding device. Here, for example, when the three-dimensional data indicates a three-dimensional map, when a user looks at the map, the user often wants to see a map in a specific direction from the user's current position. In such a case, the three-dimensional data decoding device sequentially decodes a plurality of three-dimensional points that are encoded in the data order contained in the bit stream, and causes a display device or the like to sequentially display images indicating the plurality of three-dimensional points in the order in which they were decoded. In this case, it may take time for the location that the user wants to check to be displayed on the display device. In addition, depending on the user's current position, there may be three-dimensional points that do not need to be decoded. Therefore, the three-dimensional data decoding device according to the present disclosure uses the positional relationship between any two points, such as the positions of the camera, the user, etc., and the focal positions of the camera, the user, etc., and the tiles, and assumes the line of sight direction of the camera, the user, etc., and decodes the encoded tiles located in the line of sight direction. According to this, the three-dimensional data decoding device can decode encoded tiles located in an area that is likely to be desired by the user. In other words, according to the three-dimensional data decoding device according to the present disclosure, a user can appropriately select and decode desired encoded tiles (i.e., a plurality of three-dimensional points) from among the encoded tiles.
[0027] Furthermore, these comprehensive or specific aspects may be realized by a system, a method, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM, or may be realized by any combination of the system, the method, the integrated circuit, the computer program, and the recording medium.
[0028] Hereinafter, the embodiments will be described in detail with reference to the drawings. Note that each of the embodiments described below shows a specific example of the present disclosure. The numerical values, shapes, materials, components, arrangement and connection forms of the components, steps, order of steps, and the like shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components that are not described in an independent claim showing a top concept will be described as optional components.
[0029] (Embodiment 1) When using the encoded data of a point cloud in an actual device or service, it is desirable to transmit and receive the necessary information according to the purpose in order to reduce the network bandwidth. However, until now, such a function has not existed in the encoding structure of three-dimensional data, and there has been no encoding method for this purpose.
[0030] In this embodiment, we describe a three-dimensional data encoding method and a three-dimensional data encoding device for providing the function of transmitting and receiving necessary information depending on the application in encoded data of a three-dimensional point cloud, as well as a three-dimensional data decoding method and a three-dimensional data decoding device for decoding the encoded data, a three-dimensional data multiplexing method for multiplexing the encoded data, and a three-dimensional data transmission method for transmitting the encoded data.
[0031] In particular, the first and second encoding methods are currently being considered as methods (encoding schemes) for encoding point cloud data; however, the structure of the encoded data and the method for storing the encoded data in a system format have not been defined, and as things stand, there is a problem that MUX processing (multiplexing) in the encoding unit, or transmission or storage, is not possible.
[0032] Also, there has been no method to date that supports a format in which two codecs, the first encoding method and the second encoding method, are mixed, such as PCC (Point Cloud Compression).
[0033] In this embodiment, a structure of PCC encoded data in which two codecs, a first encoding method and a second encoding method, are mixed, and a method of storing the encoded data in a system format will be described.
[0034] First, the configuration of a three-dimensional data (point cloud data) encoding / decoding system according to this embodiment will be described. Fig. 1 is a diagram showing an example of the configuration of a three-dimensional data encoding / decoding system according to this embodiment. As shown in Fig. 1, the three-dimensional data encoding / decoding system includes a three-dimensional data encoding system 4601, a three-dimensional data decoding system 4602, a sensor terminal 4603, and an external connection unit 4604.
[0035] The three-dimensional data encoding system 4601 generates encoded data or multiplexed data by encoding point cloud data, which is three-dimensional data. The three-dimensional data encoding system 4601 may be a three-dimensional data encoding device realized by a single device, or may be a system realized by multiple devices. The three-dimensional data encoding device may include some of the multiple processing units included in the three-dimensional data encoding system 4601.
[0036] The three-dimensional data encoding system 4601 includes a point cloud data generation system 4611, a presentation unit 4612, an encoding unit 4613, a multiplexing unit 4614, an input / output unit 4615, and a control unit 4616. The point cloud data generation system 4611 includes a sensor information acquisition unit 4617 and a point cloud data generation unit 4618.
[0037] The sensor information acquisition unit 4617 acquires sensor information from the sensor terminal 4603, and outputs the sensor information to the point cloud data generation unit 4618. The point cloud data generation unit 4618 generates point cloud data from the sensor information, and outputs the point cloud data to the encoding unit 4613.
[0038] The presentation unit 4612 presents the sensor information or the point cloud data to the user. For example, the presentation unit 4612 displays information or an image based on the sensor information or the point cloud data.
[0039] The encoding unit 4613 encodes (compresses) the point cloud data, and outputs the obtained encoded data, control information obtained in the encoding process, and other additional information to the multiplexing unit 4614. The additional information includes, for example, sensor information.
[0040] The multiplexing unit 4614 generates multiplexed data by multiplexing the coded data, control information, and additional information input from the coding unit 4613. The format of the multiplexed data is, for example, a file format for storage, or a packet format for transmission.
[0041] The input / output unit 4615 (e.g., a communication unit or an interface) outputs the multiplexed data to the outside. Alternatively, the multiplexed data is stored in a storage unit such as an internal memory. The control unit 4616 (or an application execution unit) controls each processing unit. In other words, the control unit 4616 controls encoding, multiplexing, etc.
[0042] The sensor information may be input to the encoding unit 4613 or the multiplexing unit 4614. Furthermore, the input / output unit 4615 may output the point cloud data or the encoded data directly to the outside.
[0043] The transmission signal (multiplexed data) output from the three-dimensional data encoding system 4601 is input to the three-dimensional data decoding system 4602 via the external connection unit 4604.
[0044] The three-dimensional data decoding system 4602 generates point cloud data, which is three-dimensional data, by decoding the encoded data or multiplexed data. The three-dimensional data decoding system 4602 may be a three-dimensional data decoding device realized by a single device, or may be a system realized by multiple devices. The three-dimensional data decoding device may also include some of the multiple processing units included in the three-dimensional data decoding system 4602.
[0045] The three-dimensional data decoding system 4602 includes a sensor information acquisition unit 4621 , an input / output unit 4622 , a demultiplexing unit 4623 , a decoding unit 4624 , a presentation unit 4625 , a user interface 4626 , and a control unit 4627 .
[0046] The sensor information acquisition unit 4621 acquires sensor information from the sensor terminal 4603 .
[0047] The input / output unit 4622 acquires a transmission signal, decodes multiplexed data (file format or packets) from the transmission signal, and outputs the multiplexed data to the demultiplexer 4623.
[0048] The demultiplexer 4623 obtains the coded data, control information, and additional information from the multiplexed data, and outputs the coded data, control information, and additional information to the decoder 4624.
[0049] The decoding unit 4624 reconstructs the point cloud data by decoding the encoded data.
[0050] The presentation unit 4625 presents the point cloud data to the user. For example, the presentation unit 4625 displays information or an image based on the point cloud data. The user interface 4626 acquires an instruction based on a user's operation. The control unit 4627 (or the application execution unit) controls each processing unit. That is, the control unit 4627 controls demultiplexing, decoding, presentation, etc.
[0051] The input / output unit 4622 may obtain the point cloud data or the encoded data directly from the outside. The presentation unit 4625 may obtain additional information such as sensor information and present information based on the additional information. The presentation unit 4625 may present information based on a user instruction obtained by the user interface 4626.
[0052] The sensor terminal 4603 generates sensor information obtained by a sensor. The sensor terminal 4603 is a terminal equipped with a sensor or a camera, and examples of the sensor terminal include a moving body such as an automobile, a flying object such as an airplane, a mobile terminal, and a camera.
[0053] The sensor information that can be acquired by the sensor terminal 4603 is, for example, (1) the distance between the sensor terminal 4603 and an object, or the reflectance of the object, obtained from a LIDAR, millimeter wave radar, or an infrared sensor, and (2) the distance between a camera and an object, or the reflectance of the object, obtained from a plurality of monocular camera images or stereo camera images. The sensor information may also include the attitude, direction, gyro (angular velocity), position (GPS information or altitude), speed, acceleration, etc. of the sensor. The sensor information may also include temperature, air pressure, humidity, magnetism, etc.
[0054] The external connection unit 4604 is realized by an integrated circuit (LSI or IC), an external storage unit, communication with a cloud server via the Internet, broadcasting, or the like.
[0055] Next, the point cloud data will be described. Fig. 2 is a diagram showing the configuration of the point cloud data. Fig. 3 is a diagram showing an example of the configuration of a data file in which information on the point cloud data is written.
[0056] Point cloud data includes data on multiple points. The data on each point includes position information (three-dimensional coordinates) and attribute information for that position information. A collection of multiple points is called a point cloud. For example, a point cloud may represent the three-dimensional shape of an object.
[0057] Position information such as three-dimensional coordinates is sometimes called geometry. Data for each point may include attribute information of multiple attribute types. The attribute types may be, for example, color or reflectance.
[0058] One piece of attribute information may be associated with one piece of location information, or multiple pieces of attribute information having different attribute types may be associated with one piece of location information, or multiple pieces of attribute information of the same attribute type may be associated with one piece of location information.
[0059] The configuration example of the data file shown in FIG. 3 is an example in which position information and attribute information correspond one-to-one, and shows the position information and attribute information of N points that make up the point cloud data.
[0060] The position information is, for example, information on three axes, x, y, and z. The attribute information is, for example, RGB color information. A typical data file is a ply file.
[0061] Next, the types of point cloud data will be described. Fig. 4 is a diagram showing the types of point cloud data. As shown in Fig. 4, point cloud data includes static objects and dynamic objects.
[0062] A static object is 3D point cloud data at any time (a certain time). A dynamic object is 3D point cloud data that changes over time. Hereinafter, 3D point cloud data at a certain time will be referred to as a PCC frame, or a frame.
[0063] The object may be a point cloud whose area is restricted to a certain extent, such as ordinary video data, or a large-scale point cloud whose area is not restricted, such as map information.
[0064] Also, there may be point cloud data of various densities, and there may be sparse point cloud data and dense point cloud data.
[0065] The details of each processing unit will be described below. The sensor information is acquired by various methods such as a distance sensor such as a LIDAR or a range finder, a stereo camera, or a combination of multiple monocular cameras. The point cloud data generation unit 4618 generates point cloud data based on the sensor information acquired by the sensor information acquisition unit 4617. The point cloud data generation unit 4618 generates position information as point cloud data, and adds attribute information for the position information to the position information.
[0066] The point cloud data generating unit 4618 may process the point cloud data when generating the position information or adding the attribute information. For example, the point cloud data generating unit 4618 may reduce the amount of data by deleting point clouds with overlapping positions. In addition, the point cloud data generating unit 4618 may convert the position information (position shift, rotation, normalization, etc.) or render the attribute information.
[0067] In FIG. 1, the point cloud data generation system 4611 is included in the three-dimensional data encoding system 4601, but it may be provided independently outside the three-dimensional data encoding system 4601.
[0068] The encoding unit 4613 generates encoded data by encoding the point cloud data based on a predefined encoding method. There are two main types of encoding methods: the first is an encoding method using position information, and this encoding method will be referred to as the first encoding method hereinafter; the second is an encoding method using a video codec, and this encoding method will be referred to as the second encoding method hereinafter.
[0069] The decoding unit 4624 decodes the encoded data based on a predefined encoding method to decode the point group data.
[0070] The multiplexing unit 4614 generates multiplexed data by multiplexing the encoded data using an existing multiplexing method. The generated multiplexed data is transmitted or stored. In addition to the PCC encoded data, the multiplexing unit 4614 multiplexes other media such as video, audio, subtitles, applications, and files, or reference time information. The multiplexing unit 4614 may further multiplex attribute information related to sensor information or point cloud data.
[0071] Multiplexing methods or file formats include ISOBMFF, MPEG-DASH, which is an ISOBMFF-based transmission method, MMT, MPEG-2 TS Systems, and RMP.
[0072] The demultiplexer 4623 extracts the PCC encoded data, other media, time information, and the like from the multiplexed data.
[0073] The input / output unit 4615 transmits the multiplexed data using a method suited to a transmission medium or a storage medium, such as broadcasting or communication. The input / output unit 4615 may communicate with other devices via the Internet, or may communicate with a storage unit such as a cloud server.
[0074] The communication protocol used may be http, ftp, TCP, UDP, etc. A PULL type communication method or a PUSH type communication method may be used.
[0075] Either wired transmission or wireless transmission may be used. For wired transmission, Ethernet (registered trademark), USB, RS-232C, HDMI (registered trademark), coaxial cable, etc. are used. For wireless transmission, wireless LAN, Wi-Fi (registered trademark), Bluetooth (registered trademark), millimeter wave, etc. are used.
[0076] As a broadcasting system, for example, DVB-T2, DVB-S2, DVB-C2, ATSC3.0, or ISDB-S3 is used.
[0077] Fig. 5 is a diagram showing a configuration of a first encoding unit 4630 which is an example of the encoding unit 4613 that performs encoding using the first encoding method. Fig. 6 is a block diagram of the first encoding unit 4630. The first encoding unit 4630 generates encoded data (encoded stream) by encoding point cloud data using the first encoding method. The first encoding unit 4630 includes a position information encoding unit 4631, an attribute information encoding unit 4632, an additional information encoding unit 4633, and a multiplexing unit 4634.
[0078] The first encoding unit 4630 has a feature of performing encoding in consideration of a three-dimensional structure. Also, the first encoding unit 4630 has a feature that the attribute information encoding unit 4632 performs encoding using information obtained from the position information encoding unit 4631. The first encoding method is also called GPCC (Geometry based PCC).
[0079] The point cloud data is PCC point cloud data such as a PLY file, or PCC point cloud data generated from sensor information, and includes position information (Position), attribute information (Attribute), and other additional information (MetaData). The position information is input to a position information encoding unit 4631, the attribute information is input to an attribute information encoding unit 4632, and the additional information is input to an additional information encoding unit 4633.
[0080] The position information encoding unit 4631 generates encoded position information (Compressed Geometry) that is encoded data by encoding the position information. For example, the position information encoding unit 4631 encodes the position information using an N-ary tree structure such as an octet tree. Specifically, in the octet tree, the target space is divided into eight nodes (subspaces), and 8-bit information (occupancy code) indicating whether or not a point cloud is included in each node is generated. In addition, the node that includes the point cloud is further divided into eight nodes, and 8-bit information indicating whether or not a point cloud is included in each of the eight nodes is generated. This process is repeated until the number of point clouds included in a predetermined hierarchy or node is equal to or less than a threshold value.
[0081] The attribute information encoding unit 4632 generates encoded attribute information (Compressed Attribute) that is encoded data by encoding using the configuration information generated by the position information encoding unit 4631. For example, the attribute information encoding unit 4632 determines a reference point (reference node) to be referenced in encoding a target point (target node) to be processed based on the octree structure generated by the position information encoding unit 4631. For example, the attribute information encoding unit 4632 refers to a node, among peripheral nodes or adjacent nodes, whose parent node in the octree is the same as that of the target node. Note that the method of determining the reference relationship is not limited to this.
[0082] Furthermore, the encoding process of the attribute information may include at least one of a quantization process, a prediction process, and an arithmetic coding process. In this case, the reference means using a reference node to calculate a predicted value of the attribute information, or using a state of the reference node (e.g., occupancy information indicating whether or not a point group is included in the reference node) to determine a parameter of the encoding. For example, the parameter of the encoding is a quantization parameter in the quantization process, or a context in the arithmetic coding.
[0083] The additional information encoding unit 4633 generates encoded additional information (Compressed MetaData) that is encoded data by encoding compressible data from the additional information.
[0084] The multiplexing unit 4634 multiplexes the encoding position information, the encoding attribute information, the encoding additional information, and other additional information to generate an encoded stream (Compressed Stream) that is encoded data. The generated encoded stream is output to a processing unit of a system layer (not shown).
[0085] Next, a first decoding unit 4640 which is an example of the decoding unit 4624 that performs decoding of the first encoding method will be described. FIG. 7 is a diagram showing a configuration of the first decoding unit 4640. FIG. 8 is a block diagram of the first decoding unit 4640. The first decoding unit 4640 generates point group data by decoding, by the first encoding method, encoded data (encoded stream) encoded by the first encoding method. The first decoding unit 4640 includes a demultiplexing unit 4641, a position information decoding unit 4642, an attribute information decoding unit 4643, and an additional information decoding unit 4644.
[0086] A coded stream (compressed stream), which is coded data, is input to the first decoding unit 4640 from a processing unit in a system layer (not shown).
[0087] The demultiplexer 4641 separates the encoded position information (Compressed Geometry), the encoded attribute information (Compressed Attribute), the encoded additional information (Compressed MetaData), and other additional information from the encoded data.
[0088] The position information decoding unit 4642 generates position information by decoding the encoded position information. For example, the position information decoding unit 4642 restores the position information of a point group represented by three-dimensional coordinates from the encoded position information represented by an N-ary tree structure such as an octree.
[0089] The attribute information decoding unit 4643 decodes the encoded attribute information based on the configuration information generated by the position information decoding unit 4642. For example, the attribute information decoding unit 4643 determines a reference point (reference node) to be referenced in decoding a target point (target node) to be processed based on the octree structure obtained by the position information decoding unit 4642. For example, the attribute information decoding unit 4643 refers to a node, among peripheral nodes or adjacent nodes, whose parent node in the octree is the same as that of the target node. Note that the method of determining the reference relationship is not limited to this.
[0090] In addition, the decoding process of the attribute information may include at least one of an inverse quantization process, a prediction process, and an arithmetic decoding process. In this case, the reference means using a reference node to calculate a predicted value of the attribute information, or using a state of the reference node (e.g., occupancy information indicating whether or not a point group is included in the reference node) to determine a parameter of the decoding. For example, the parameter of the decoding is a quantization parameter in the inverse quantization process, or a context in the arithmetic decoding.
[0091] The additional information decoding unit 4644 generates additional information by decoding the encoded additional information. The first decoding unit 4640 uses the additional information required for decoding the position information and attribute information during decoding, and outputs the additional information required for an application to the outside.
[0092] Next, a configuration example of the position information encoding unit will be described. Fig. 9 is a block diagram of the position information encoding unit 2700 according to this embodiment. The position information encoding unit 2700 includes an octree generating unit 2701, a geometric information calculating unit 2702, a coding table selecting unit 2703, and an entropy encoding unit 2704.
[0093] The octree generating unit 2701 generates, for example, an octree from the input position information, and generates an occupancy code of each node of the octree. The geometric information calculating unit 2702 obtains information indicating whether an adjacent node of the target node is an occupied node. For example, the geometric information calculating unit 2702 calculates occupancy information of the adjacent node (information indicating whether the adjacent node is an occupied node) from the occupancy code of the parent node to which the target node belongs. The geometric information calculating unit 2702 may also store encoded nodes in a list and search for adjacent nodes from the list. The geometric information calculating unit 2702 may also switch adjacent nodes depending on the position of the target node in the parent node.
[0094] The coding table selection unit 2703 selects a coding table to be used for entropy coding of the target node, using the occupancy information of the adjacent nodes calculated by the geometric information calculation unit 2702. For example, the coding table selection unit 2703 may generate a bit string using the occupancy information of the adjacent nodes, and select a coding table for an index number generated from the bit string.
[0095] The entropy coding unit 2704 generates the coding position information and metadata by performing entropy coding on the occupancy code of the target node using the coding table of the selected index number. The entropy coding unit 2704 may add information indicating the selected coding table to the coding position information.
[0096] The octree representation and the scanning order of position information will be described below. Position information (position data) is converted (octreeized) into an octree structure and then encoded. The octree structure is composed of nodes and leaves. Each node has eight nodes or leaves, and each leaf has voxel (VXL) information. Fig. 10 is a diagram showing an example of the structure of position information including multiple voxels. Fig. 11 is a diagram showing an example of the position information shown in Fig. 10 converted into an octree structure. Here, among the leaves shown in Fig. 11, leaves 1, 2, and 3 respectively represent voxels VXL1, VXL2, and VXL3 shown in Fig. 10, and represent a VXL including a point cloud (hereinafter, effective VXL).
[0097] Specifically, node 1 corresponds to the whole space including the position information in FIG. 10. The whole space corresponding to node 1 is divided into eight nodes, and among the eight nodes, the node including the valid VXL is further divided into eight nodes or leaves, and this process is repeated for the number of levels of the tree structure. Here, each node corresponds to a subspace, and has information (occupancy code) indicating at what position the next node or leaf will be located after division as node information. Also, the block at the bottom level is set as a leaf, and the number of point clouds included in the leaf, etc., is held as leaf information.
[0098] Next, a configuration example of the position information decoding unit will be described. Fig. 12 is a block diagram of the position information decoding unit 2710 according to this embodiment. The position information decoding unit 2710 includes an octree generating unit 2711, a geometric information calculating unit 2712, a coding table selecting unit 2713, and an entropy decoding unit 2714.
[0099] The octree generating unit 2711 generates an octree of a certain space (node) using header information or metadata of a bit stream. For example, the octree generating unit 2711 generates a large space (root node) using the size of the x-axis, y-axis, and z-axis directions of a certain space added to the header information, and generates an octree by dividing the space into two in the x-axis, y-axis, and z-axis directions to generate eight small spaces A (nodes A0 to A7). In addition, nodes A0 to A7 are set in order as target nodes.
[0100] The geometric information calculation unit 2712 obtains occupancy information indicating whether an adjacent node of the target node is an occupancy node. For example, the geometric information calculation unit 2712 calculates the occupancy information of the adjacent node from the occupancy code of the parent node to which the target node belongs. The geometric information calculation unit 2712 may also store decoded nodes in a list and search for adjacent nodes from the list. The geometric information calculation unit 2712 may also switch adjacent nodes depending on the position of the target node in the parent node.
[0101] The coding table selection unit 2713 selects a coding table (decoding table) to be used for entropy decoding of the target node using the occupancy information of the adjacent node calculated by the geometric information calculation unit 2712. For example, the coding table selection unit 2713 may generate a bit string using the occupancy information of the adjacent node, and select a coding table for an index number generated from the bit string.
[0102] The entropy decoding unit 2714 generates position information by entropy decoding the occupancy code of the target node using the selected coding table. Note that the entropy decoding unit 2714 may obtain information on the selected coding table by decoding from the bit stream, and entropy decode the occupancy code of the target node using the coding table indicated by the information.
[0103] The configurations of the attribute information encoding unit and the attribute information decoding unit will be described below. Fig. 13 is a block diagram showing an example of the configuration of the attribute information encoding unit A100. The attribute information encoding unit may include a plurality of encoding units that execute different encoding methods. For example, the attribute information encoding unit may switch between the following two methods depending on the use case.
[0104] The attribute information encoding unit A100 includes an LoD attribute information encoding unit A101 and a conversion attribute information encoding unit A102. The LoD attribute information encoding unit A101 classifies each 3D point into a plurality of layers using the position information of the 3D point, predicts the attribute information of the 3D point belonging to each layer, and encodes the prediction residual. Here, each classified layer is called LoD (Level of Detail).
[0105] The transformed attribute information encoding unit A102 encodes the attribute information using RAHT (Region Adaptive Hierarchical Transform). Specifically, the transformed attribute information encoding unit A102 applies RAHT or Haar transform to each piece of attribute information based on the position information of the three-dimensional point to generate high-frequency components and low-frequency components of each layer, and encodes the values using quantization, entropy coding, or the like.
[0106] 14 is a block diagram showing a configuration example of the attribute information decoding unit A110. The attribute information decoding unit may include multiple decoding units that execute different decoding methods. For example, the attribute information decoding unit may switch between the following two methods for decoding based on information included in the header or metadata.
[0107] The attribute information decoding unit A110 includes a LoD attribute information decoding unit A111 and a converted attribute information decoding unit A112. The LoD attribute information decoding unit A111 classifies each 3D point into a plurality of hierarchical levels using the position information of the 3D point, and decodes the attribute value while predicting the attribute information of the 3D point belonging to each hierarchical level.
[0108] The transformed attribute information decoding unit A112 decodes the attribute information using RAHT (Region Adaptive Hierarchical Transform). Specifically, the transformed attribute information decoding unit A112 decodes the attribute value by applying inverse RAHT or inverse Haar transform to high-frequency components and low-frequency components of each attribute value based on the position information of the three-dimensional point.
[0109] FIG. 15 is a block diagram showing a configuration of an attribute information encoding unit 3140 which is an example of the LoD attribute information encoding unit A101.
[0110] The attribute information encoding unit 3140 includes a LoD generation unit 3141, a surrounding search unit 3142, a prediction unit 3143, a prediction residual calculation unit 3144, a quantization unit 3145, an arithmetic encoding unit 3146, an inverse quantization unit 3147, a decoded value generation unit 3148, and a memory 3149.
[0111] The LoD generation unit 3141 generates LoD using the position information of the three-dimensional points.
[0112] The surrounding search unit 3142 searches for nearby 3D points adjacent to each 3D point, using the LoD generation result by the LoD generation unit 3141 and distance information indicating the distance between each 3D point.
[0113] The prediction unit 3143 generates a predicted value of the attribute information of the target 3D point to be encoded.
[0114] The prediction residual calculation unit 3144 calculates (generates) a prediction residual of the predicted value of the attribute information generated by the prediction unit 3143.
[0115] The quantization unit 3145 quantizes the prediction residual of the attribute information calculated by the prediction residual calculation unit 3144 .
[0116] The arithmetic coding unit 3146 arithmetically codes the prediction residuals after being quantized by the quantization unit 3145. The arithmetic coding unit 3146 outputs a bit stream including the arithmetically coded prediction residuals to, for example, a three-dimensional data decoding device.
[0117] Note that the prediction residual may be binarized by, for example, the quantization unit 3145 before being arithmetically coded by the arithmetic coding unit 3146.
[0118] Also, for example, the arithmetic coding unit 3146 may initialize a coding table used for arithmetic coding before arithmetic coding. The arithmetic coding unit 3146 may initialize a coding table used for arithmetic coding for each layer. Also, the arithmetic coding unit 3146 may output information indicating the position of the layer for which the coding table has been initialized, by including it in the bitstream.
[0119] The inverse quantization unit 3147 inverse quantizes the prediction residual after being quantized by the quantization unit 3145 .
[0120] The decoded value generation unit 3148 generates a decoded value by adding the predicted value of the attribute information generated by the prediction unit 3143 and the prediction residual after inverse quantization by the inverse quantization unit 3147.
[0121] The memory 3149 is a memory that stores the decoded values of the attribute information of each 3D point decoded by the decoded value generation unit 3148. For example, when generating a predicted value of a 3D point that has not yet been encoded, the prediction unit 3143 generates the predicted value by using the decoded value of the attribute information of each 3D point stored in the memory 3149.
[0122] 16 is a block diagram of an attribute information encoding unit 6600 which is an example of the transformed attribute information encoding unit A102. The attribute information encoding unit 6600 includes a sorting unit 6601, a Haar transform unit 6602, a quantization unit 6603, an inverse quantization unit 6604, an inverse Haar transform unit 6605, a memory 6606, and an arithmetic encoding unit 6607.
[0123] The sorting unit 6601 generates a Morton code using the position information of the 3D points, and sorts the multiple 3D points in Morton code order. The Haar transform unit 6602 generates coding coefficients by applying a Haar transform to the attribute information. The quantization unit 6603 quantizes the coding coefficients of the attribute information.
[0124] The inverse quantization unit 6604 inversely quantizes the quantized coding coefficients. The inverse Haar transform unit 6605 applies an inverse Haar transform to the coding coefficients. The memory 6606 stores values of attribute information of multiple decoded 3D points. For example, the attribute information of the decoded 3D points stored in the memory 6606 may be used for predicting uncoded 3D points.
[0125] The arithmetic coding unit 6607 calculates ZeroCnt from the coding coefficients after quantization, and arithmetically codes the ZeroCnt. Furthermore, the arithmetic coding unit 6607 arithmetically codes the non-zero coding coefficients after quantization. The arithmetic coding unit 6607 may binarize the coding coefficients before arithmetic coding. Furthermore, the arithmetic coding unit 6607 may generate and code various header information.
[0126] FIG. 17 is a block diagram showing a configuration of an attribute information decoding unit 3150 which is an example of the LoD attribute information decoding unit A111.
[0127] The attribute information decoding unit 3150 includes an LoD generation unit 3151 , a surrounding search unit 3152 , a prediction unit 3153 , an arithmetic decoding unit 3154 , an inverse quantization unit 3155 , a decoded value generation unit 3156 , and a memory 3157 .
[0128] The LoD generation unit 3151 generates LoD using position information of the 3D points decoded by a position information decoding unit (not shown in FIG. 17).
[0129] The surrounding search unit 3152 searches for nearby 3D points adjacent to each 3D point, using the LoD generation result by the LoD generation unit 3151 and distance information indicating the distance between each 3D point.
[0130] The prediction unit 3153 generates a predicted value of the attribute information of the target 3D point to be decoded.
[0131] The arithmetic decoding unit 3154 arithmetically decodes the prediction residual in the bitstream acquired from the attribute information encoding unit 3140 shown in FIG. 15. The arithmetic decoding unit 3154 may initialize a decoding table used for arithmetic decoding. The arithmetic decoding unit 3154 initializes a decoding table used for arithmetic decoding for a layer on which the arithmetic encoding unit 3146 shown in FIG. 15 has performed an encoding process. The arithmetic decoding unit 3154 may initialize a decoding table used for arithmetic decoding for each layer. The arithmetic decoding unit 3154 may initialize the decoding table based on information included in the bitstream and indicating the position of the layer for which the encoding table has been initialized.
[0132] The inverse quantization unit 3155 inverse quantizes the prediction residual arithmetically decoded by the arithmetic decoding unit 3154 .
[0133] The decoded value generation unit 3156 generates a decoded value by adding the predicted value generated by the prediction unit 3153 and the prediction residual after inverse quantization by the inverse quantization unit 3155. The decoded value generation unit 3156 outputs the decoded attribute information data to another device.
[0134] The memory 3157 is a memory that stores the decoded values of the attribute information of each 3D point decoded by the decoded value generation unit 3156. For example, when generating a predicted value of a 3D point that has not yet been decoded, the prediction unit 3153 generates the predicted value by using the decoded value of the attribute information of each 3D point stored in the memory 3157.
[0135] 18 is a block diagram of an attribute information decoding unit 6610 which is an example of the transformed attribute information decoding unit A112. The attribute information decoding unit 6610 includes an arithmetic decoding unit 6611, an inverse quantization unit 6612, an inverse Haar transform unit 6613, and a memory 6614.
[0136] The arithmetic decoding unit 6611 arithmetically decodes the ZeroCnt and the coding coefficients included in the bit stream. Note that the arithmetic decoding unit 6611 may also decode various types of header information.
[0137] The inverse quantization unit 6612 inverse quantizes the arithmetically decoded coding coefficients. The inverse Haar transform unit 6613 applies an inverse Haar transform to the coding coefficients after the inverse quantization. The memory 6614 stores values of attribute information of multiple decoded 3D points. For example, the attribute information of the decoded 3D points stored in the memory 6614 may be used to predict undecoded 3D points.
[0138] Next, a second encoding unit 4650, which is an example of the encoding unit 4613 that performs encoding by the second encoding method, will be described. Fig. 19 is a diagram showing a configuration of the second encoding unit 4650. Fig. 20 is a block diagram of the second encoding unit 4650.
[0139] The second encoding unit 4650 generates encoded data (encoded stream) by encoding the point cloud data by a second encoding method. The second encoding unit 4650 includes an additional information generating unit 4651, a position image generating unit 4652, an attribute image generating unit 4653, a video encoding unit 4654, an additional information encoding unit 4655, and a multiplexing unit 4656.
[0140] The second encoding unit 4650 has a feature of generating a position image and an attribute image by projecting a three-dimensional structure onto a two-dimensional image, and encoding the generated position image and attribute image using an existing video encoding method. The second encoding method is also called VPCC (Video based PCC).
[0141] The point cloud data is PCC point cloud data such as a PLY file, or PCC point cloud data generated from sensor information, and includes position information (Position), attribute information (Attribute), and other additional information (MetaData).
[0142] The additional information generating unit 4651 generates map information of a plurality of two-dimensional images by projecting a three-dimensional structure onto the two-dimensional images.
[0143] The position image generating unit 4652 generates a position image (Geometry Image) based on the position information and the map information generated by the additional information generating unit 4651. This position image is, for example, a distance image in which distance (Depth) is indicated as a pixel value. Note that this distance image may be an image in which a plurality of point clouds are viewed from one viewpoint (an image in which a plurality of point clouds are projected onto one two-dimensional plane), a plurality of images in which a plurality of point clouds are viewed from multiple viewpoints, or a single image in which these multiple images are integrated.
[0144] The attribute image generating unit 4653 generates an attribute image based on the attribute information and the map information generated by the additional information generating unit 4651. This attribute image is, for example, an image in which the attribute information (for example, color (RGB)) is indicated as pixel values. Note that this image may be an image in which multiple point clouds are viewed from one viewpoint (an image in which multiple point clouds are projected onto one two-dimensional plane), or multiple images in which multiple point clouds are viewed from multiple viewpoints, or a single image in which these multiple images are integrated.
[0145] The video encoding unit 4654 generates an encoded position image (Compressed Geometry Image) and an encoded attribute image (Compressed Attribute Image), which are encoded data, by encoding the position image and the attribute image using a video encoding method. Note that any known encoding method may be used as the video encoding method. For example, the video encoding method is AVC, HEVC, or the like.
[0146] The additional information encoding unit 4655 generates encoded additional information (Compressed MetaData) by encoding the additional information and map information included in the point cloud data.
[0147] The multiplexing unit 4656 multiplexes the encoding position image, the encoding attribute image, the encoding additional information, and other additional information to generate an encoded stream (Compressed Stream) that is encoded data. The generated encoded stream is output to a processing unit of a system layer (not shown).
[0148] Next, a second decoding unit 4660 which is an example of the decoding unit 4624 that performs decoding of the second encoding method will be described. FIG. 21 is a diagram showing a configuration of the second decoding unit 4660. FIG. 22 is a block diagram of the second decoding unit 4660. The second decoding unit 4660 generates point group data by decoding, by the second encoding method, encoded data (encoded stream) encoded by the second encoding method. The second decoding unit 4660 includes a demultiplexing unit 4661, a video decoding unit 4662, an additional information decoding unit 4663, a position information generating unit 4664, and an attribute information generating unit 4665.
[0149] A compressed stream, which is encoded data, is input to the second decoding unit 4660 from a processing unit in a system layer (not shown).
[0150] The demultiplexer 4661 separates the encoded position image (Compressed Geometry Image), the encoded attribute image (Compressed Attribute Image), the encoded additional information (Compressed MetaData), and other additional information from the encoded data.
[0151] The video decoding unit 4662 generates a position image and an attribute image by decoding the encoded position image and the encoded attribute image using a video encoding method. Note that any known encoding method may be used as the video encoding method. For example, the video encoding method is AVC or HEVC.
[0152] The additional information decoding unit 4663 decodes the encoded additional information to generate additional information including map information and the like.
[0153] The position information generating unit 4664 generates position information using the position image and map information. The attribute information generating unit 4665 generates attribute information using the attribute image and map information.
[0154] The second decoding unit 4660 uses the additional information necessary for decoding during decoding, and outputs the additional information necessary for an application to the outside.
[0155] Problems with the PCC encoding method will be described below. Fig. 23 is a diagram showing a protocol stack related to PCC encoded data. Fig. 23 shows an example in which other media data such as video (e.g., HEVC) or audio is multiplexed with PCC encoded data and transmitted or stored.
[0156] The multiplexing method and file format have the function of multiplexing various encoded data and transmitting or storing them. In order to transmit or store the encoded data, the encoded data must be converted into the format of the multiplexing method. For example, HEVC specifies a technology for storing encoded data in a data structure called a NAL unit and storing the NAL unit in ISOBMFF.
[0157] Meanwhile, a first encoding method (Codec1) and a second encoding method (Codec2) are currently being considered as methods for encoding point cloud data. However, the structure of the encoded data and the method for storing the encoded data in a system format have not been defined, which poses the problem that MUX processing (multiplexing) in the encoding unit, transmission, and storage cannot be performed as is.
[0158] In the following description, unless a specific encoding method is specified, it refers to either the first encoding method or the second encoding method.
[0159] (Embodiment 2) In this embodiment, the types of encoded data (position information (Geometry), attribute information (Attribute), additional information (Metadata)) generated by the above-mentioned first encoding unit 4630 or second encoding unit 4650, a generation method of the additional information (metadata), and multiplexing processing in the multiplexing unit will be described. Note that the additional information (metadata) may also be referred to as a parameter set or control information.
[0160] In this embodiment, the dynamic object (three-dimensional point cloud data that changes over time) described in Figure 4 will be used as an example, but a similar method may also be used in the case of a static object (three-dimensional point cloud data at any time).
[0161] 24 is a diagram showing configurations of an encoding unit 4801 and a multiplexing unit 4802 included in the three-dimensional data encoding device according to this embodiment. The encoding unit 4801 corresponds to, for example, the first encoding unit 4630 or the second encoding unit 4650 described above. The multiplexing unit 4802 corresponds to the multiplexing unit 4634 or 4656 described above.
[0162] The encoding unit 4801 encodes point cloud data of multiple PCC (Point Cloud Compression) frames, and generates encoded data (Multiple Compressed Data) of multiple pieces of position information, attribute information, and additional information.
[0163] The multiplexing unit 4802 converts data of multiple data types (position information, attribute information, and additional information) into NAL units, thereby converting the data into a data structure that takes into consideration data access in the decoding device.
[0164] 25 is a diagram showing an example of the structure of coded data generated by the coding unit 4801. The arrows in the diagram indicate dependencies related to the decoding of coded data, with the source of the arrow depending on the data at the tip of the arrow. In other words, the decoding device decodes the data at the tip of the arrow, and uses the decoded data to decode the data at the tip of the arrow. In other words, dependency means that the data at the dependency destination is referenced (used) in the processing (encoding, decoding, etc.) of the data at the dependency destination.
[0165] First, the process of generating encoded data of position information will be described. The encoding unit 4801 generates encoded position data (compressed geometry data) for each frame by encoding the position information of each frame. The encoded position data is represented as G(i), where i indicates the frame number, the time of the frame, etc.
[0166] The encoding unit 4801 also generates a position parameter set (GPS(i)) corresponding to each frame. The position parameter set includes parameters that can be used to decode the encoded position data. The encoded position data for each frame depends on the corresponding position parameter set.
[0167] Moreover, encoded position data consisting of multiple frames is defined as a position sequence (Geometry Sequence). The encoding unit 4801 generates a position sequence parameter set (Geometry Sequence PS: also written as position SPS) that stores parameters commonly used in decoding processes for multiple frames in the position sequence. The position sequence depends on the position SPS.
[0168] Next, a process for generating encoded data of attribute information will be described. Encoding unit 4801 generates encoded attribute data (Compressed Attribute Data) for each frame by encoding the attribute information of each frame. The encoded attribute data is represented as A(i). FIG. 25 shows an example in which attribute X and attribute Y exist, and the encoded attribute data of attribute X is represented as AX(i) and the encoded attribute data of attribute Y is represented as AY(i).
[0169] The encoding unit 4801 also generates an attribute parameter set (APS(i)) corresponding to each frame. The attribute parameter set for attribute X is represented as AXPS(i), and the attribute parameter set for attribute Y is represented as AYPS(i). The attribute parameter set includes parameters that can be used to decode the encoded attribute information. The encoded attribute data depends on the corresponding attribute parameter set.
[0170] Moreover, encoded attribute data consisting of multiple frames is defined as an attribute sequence. The encoding unit 4801 generates an attribute sequence parameter set (Attribute Sequence PS: also referred to as attribute SPS) that stores parameters commonly used in decoding processes for multiple frames in the attribute sequence. The attribute sequence depends on the attribute SPS.
[0171] Furthermore, in the first encoding method, the encoded attribute data depends on the encoded position data.
[0172] 25 shows an example in which two types of attribute information (attribute X and attribute Y) exist. When there are two types of attribute information, for example, two encoding units generate respective data and metadata. Also, for example, an attribute sequence is defined for each type of attribute information, and an attribute SPS is generated for each type of attribute information.
[0173] In addition, while FIG. 25 shows an example in which there is one type of position information and two types of attribute information, the present invention is not limited to this, and there may be one type of attribute information, or three or more types. In this case, the encoded data can be generated in a similar manner. In addition, in the case of point cloud data that does not have attribute information, the attribute information may not be necessary. In this case, the encoding unit 4801 does not need to generate a parameter set related to the attribute information.
[0174] Next, a process of generating additional information (metadata) will be described. The encoding unit 4801 generates a PCC stream PS (also written as stream PS), which is a parameter set for the entire PCC stream. The encoding unit 4801 stores parameters that can be commonly used in decoding processes for one or more position sequences and one or more attribute sequences in the stream PS. For example, the stream PS includes identification information indicating the codec of the point cloud data, information indicating the algorithm used for encoding, and the like. The position sequence and the attribute sequence depend on the stream PS.
[0175] Next, an access unit and a GOF will be described. In this embodiment, the concept of an access unit (AU) and a group of frames (GOF) are newly introduced.
[0176] An access unit is a basic unit for accessing data during decoding, and is composed of one or more pieces of data and one or more pieces of metadata. For example, an access unit is composed of position information at the same time and one or more pieces of attribute information. A GOF is a random access unit, and is composed of one or more access units.
[0177] The encoding unit 4801 generates an access unit header (AU Header) as identification information indicating the beginning of an access unit. The encoding unit 4801 stores parameters related to the access unit in the access unit header. For example, the access unit header includes the configuration or information of the encoded data included in the access unit. The access unit header also includes parameters commonly used for data included in the access unit, such as parameters related to the decoding of the encoded data.
[0178] Alternatively, instead of an access unit header, the encoding unit 4801 may generate an access unit delimiter that does not include parameters related to the access unit. This access unit delimiter is used as identification information indicating the beginning of the access unit. The decoding device identifies the beginning of the access unit by detecting the access unit header or the access unit delimiter.
[0179] Next, generation of identification information for the start of a GOF will be described. The encoding unit 4801 generates a GOF header as identification information indicating the start of a GOF. The encoding unit 4801 stores parameters related to the GOF in the GOF header. For example, the GOF header includes the configuration or information of the encoded data included in the GOF. The GOF header also includes parameters commonly used for data included in the GOF, such as parameters related to the decoding of the encoded data.
[0180] Instead of a GOF header, the encoding unit 4801 may generate a GOF delimiter that does not include parameters related to the GOF. This GOF delimiter is used as identification information indicating the beginning of the GOF. The decoding device identifies the beginning of the GOF by detecting the GOF header or the GOF delimiter.
[0181] In PCC encoded data, for example, an access unit is defined as a PCC frame unit, and a decoding device accesses a PCC frame based on identification information at the beginning of the access unit.
[0182] Also, for example, GOF is defined as one random access unit. The decoding device accesses the random access unit based on the identification information at the beginning of the GOF. For example, if PCC frames are not dependent on each other and can be decoded independently, the PCC frames may be defined as the random access unit.
[0183] Note that two or more PCC frames may be assigned to one access unit, and multiple random access units may be assigned to one GOF.
[0184] Furthermore, the encoding unit 4801 may define and generate parameter sets or metadata other than those described above. For example, the encoding unit 4801 may generate SEI (Supplemental Enhancement Information) that stores parameters (optional parameters) that may not necessarily be used during decoding.
[0185] Next, the structure of the coded data and the method of storing the coded data in the NAL unit will be described.
[0186] For example, a data format is defined for each type of encoded data. Figure 26 shows an example of encoded data and NAL units.
[0187] For example, as shown in Fig. 26, the encoded data includes a header and a payload. The encoded data may include length information indicating the length (amount of data) of the encoded data, the header, or the payload. The encoded data may not include a header.
[0188] The header includes, for example, identification information for identifying the data, which indicates, for example, the data type or the frame number.
[0189] The header includes, for example, identification information indicating a reference relationship. This identification information is stored in the header when, for example, there is a dependency between data, and is information for referring to the reference destination from the reference source. For example, the header of the reference destination includes identification information for identifying the data. The header of the reference source includes identification information indicating the reference destination.
[0190] In addition, when the reference destination or the reference source can be identified or derived from other information, the identification information for specifying the data or the identification information indicating the reference relationship may be omitted.
[0191] The multiplexing unit 4802 stores the encoded data in the payload of the NAL unit. The NAL unit header includes pcc_nal_unit_type, which is identification information of the encoded data. Figure 27 shows an example of the semantics of pcc_nal_unit_type.
[0192] 27, when pcc_codec_type is codec 1 (Codec1: first encoding method), values 0 to 10 of pcc_nal_unit_type are assigned to the encoded position data (Geometry), encoded attribute X data (AttributeX), encoded attribute Y data (AttributeY), position PS (Geom.PS), attribute XPS (AttrX.PS), attribute YPS (AttrX.PS), position SPS (Geometry Sequence PS), attribute XSPS (AttributeX Sequence PS), attribute YSPS (AttributeY Sequence PS), AU header (AU Header), and GOF header (GOF Header) in codec 1. Values 11 and above are assigned as spares for codec 1.
[0193] When pcc_codec_type is Codec2 (Codec2: second encoding method), values 0 to 2 of pcc_nal_unit_type are assigned to codec data A (DataA), metadata A (MetaDataA), and metadata B (MetaDataB). Values 3 and above are assigned as spares for Codec2.
[0194] (Embodiment 3) HEVC encoding has data division tools such as slicing or tiling to enable parallel processing in the decoding device, but PCC (Point Cloud Compression) encoding does not yet have such tools.
[0195] In PCC, various data division methods can be considered depending on parallel processing, compression efficiency, and compression algorithms. This section describes the definitions of slices and tiles, the data structure, and the transmission and reception methods.
[0196] 28 is a block diagram showing a configuration of a first encoding unit 4910 included in the three-dimensional data encoding device according to this embodiment. The first encoding unit 4910 generates encoded data (encoded stream) by encoding point cloud data using a first encoding method (GPCC (Geometry based PCC)). The first encoding unit 4910 includes a division unit 4911, a plurality of position information encoding units 4912, a plurality of attribute information encoding units 4913, an additional information encoding unit 4914, and a multiplexing unit 4915.
[0197] The division unit 4911 generates a plurality of divided data by dividing the point cloud data. Specifically, the division unit 4911 generates a plurality of divided data by dividing the space of the point cloud data into a plurality of subspaces. Here, the subspace is one of a tile and a slice, or a combination of a tile and a slice. More specifically, the point cloud data includes position information, attribute information, and additional information. The division unit 4911 divides the position information into a plurality of divided position information, and divides the attribute information into a plurality of divided attribute information. In addition, the division unit 4911 generates additional information related to the division.
[0198] The position information encoding units 4912 generate a plurality of pieces of encoded position information by encoding the plurality of pieces of divided position information. For example, the position information encoding units 4912 process the plurality of pieces of divided position information in parallel.
[0199] The attribute information encoding units 4913 generate a plurality of pieces of encoded attribute information by encoding the plurality of pieces of divided attribute information. For example, the attribute information encoding units 4913 process the plurality of pieces of divided attribute information in parallel.
[0200] The additional information encoding unit 4914 generates encoded additional information by encoding the additional information included in the point cloud data and the additional information related to the data division generated by the division unit 4911 at the time of division.
[0201] The multiplexing unit 4915 multiplexes a plurality of pieces of encoding position information, a plurality of pieces of encoding attribute information, and encoding additional information to generate encoded data (encoded stream), and transmits the generated encoded data. In addition, the encoded additional information is used during decoding.
[0202] 28 shows an example in which there are two position information encoders 4912 and two attribute information encoders 4913, but the number of position information encoders 4912 and two attribute information encoders 4913 may be one, or three or more. Furthermore, multiple pieces of divided data may be processed in parallel within the same chip, such as multiple cores in a CPU, or may be processed in parallel by cores on multiple chips, or may be processed in parallel by multiple cores on multiple chips.
[0203] 29 is a block diagram showing a configuration of the first decoding unit 4920. The first decoding unit 4920 restores the point cloud data by decoding the encoded data (encoded stream) generated by encoding the point cloud data by the first encoding method (GPCC). The first decoding unit 4920 includes a demultiplexing unit 4921, a plurality of position information decoding units 4922, a plurality of attribute information decoding units 4923, an additional information decoding unit 4924, and a combining unit 4925.
[0204] The demultiplexer 4921 demultiplexes the coded data (coded stream) to generate a plurality of pieces of coding position information, a plurality of pieces of coding attribute information, and coded additional information.
[0205] The position information decoders 4922 generate a plurality of pieces of divided position information by decoding the plurality of pieces of encoded position information. For example, the position information decoders 4922 process the plurality of pieces of encoded position information in parallel.
[0206] The multiple attribute information decoding unit 4923 generates multiple pieces of divided attribute information by decoding the multiple pieces of encoded attribute information. For example, the multiple attribute information decoding unit 4923 processes the multiple pieces of encoded attribute information in parallel.
[0207] The multiple additional information decoders 4924 generate additional information by decoding the encoded additional information.
[0208] The combining unit 4925 generates position information by combining a plurality of pieces of divided position information using the additional information The combining unit 4925 generates attribute information by combining a plurality of pieces of divided attribute information using the additional information.
[0209] 29 shows an example in which the number of each of the position information decoding units 4922 and the attribute information decoding units 4923 is two, but the number of each of the position information decoding units 4922 and the attribute information decoding units 4923 may be one, or three or more. Furthermore, multiple pieces of divided data may be processed in parallel within the same chip, such as multiple cores within a CPU, or may be processed in parallel by cores on multiple chips, or may be processed in parallel by multiple cores on multiple chips.
[0210] Next, it is described how the dividing unit 4911 has a configuration. Fig. 30 is a block diagram of the dividing unit 4911. The dividing unit 4911 includes a slice dividing unit 4931, a position information tile dividing unit 4932, and an attribute information tile dividing unit 4933.
[0211] The slice division unit 4931 generates a plurality of slice position information by dividing position information (Position (Geometry)) into slices. The slice division unit 4931 also generates a plurality of slice attribute information by dividing attribute information (Attribute) into slices. The slice division unit 4931 also outputs slice additional information (SliceMetaData) including information related to slice division and information generated in the slice division.
[0212] The position information tile division unit 4932 divides a plurality of slice position information pieces into tiles to generate a plurality of pieces of divided position information pieces (a plurality of pieces of tile position information pieces). In addition, the position information tile division unit 4932 outputs position tile additional information (Geometry Tile MetaData) including information related to the tile division of the position information pieces and information generated in the tile division of the position information pieces.
[0213] The attribute information tile division unit 4933 divides a plurality of slice attribute information pieces into tiles to generate a plurality of pieces of divided attribute information (a plurality of pieces of tile attribute information). In addition, the attribute information tile division unit 4933 outputs attribute tile additional information (Attribute Tile MetaData) including information related to the tile division of the attribute information and information generated in the tile division of the attribute information.
[0214] The number of slices or tiles to be divided is equal to or greater than 1. In other words, it is not necessary to divide the slices or tiles.
[0215] Although an example in which tile division is performed after slice division has been shown here, slice division may be performed after tile division. Furthermore, a new division type may be defined in addition to slices and tiles, and division may be performed using three or more division types.
[0216] A method for dividing point cloud data will be described below. Fig. 31 is a diagram showing an example of division into slices and tiles.
[0217] First, a method of dividing into slices will be described. The dividing unit 4911 divides the three-dimensional point cloud data into arbitrary point clouds in slice units. In the slice division, the dividing unit 4911 does not divide the position information and attribute information constituting a point, but divides the position information and attribute information together. That is, the dividing unit 4911 divides into slices so that the position information and attribute information at an arbitrary point belong to the same slice. In addition, according to these, the number of divisions and the division method may be any method. In addition, the minimum unit of division is a point. For example, the number of divisions of the position information and the attribute information are the same. For example, the three-dimensional point corresponding to the position information after the slice division and the three-dimensional point corresponding to the attribute information are included in the same slice.
[0218] Furthermore, the division unit 4911 generates slice additional information, which is additional information related to the number of divisions and the division method when dividing a slice. The slice additional information is the same as the position information and the attribute information. For example, the slice additional information includes information indicating the reference coordinate position, size, or side length of the bounding box after division. Furthermore, the slice additional information includes information indicating the number of divisions, the division type, etc.
[0219] Next, a tile division method will be described. The division unit 4911 divides the data divided into slices into slice position information (G slices) and slice attribute information (A slices), and divides each of the slice position information and the slice attribute information into tiles.
[0220] Although FIG. 31 shows an example of division using an octree structure, the number of divisions and the division method may be any method.
[0221] Furthermore, the division unit 4911 may divide the position information and the attribute information using different division methods or the same division method. Furthermore, the division unit 4911 may divide a plurality of slices into tiles using different division methods or the same division method.
[0222] Furthermore, the division unit 4911 generates tile additional information related to the number of divisions and the division method when dividing tiles. The tile additional information (position tile additional information and attribute tile additional information) is independent of position information and attribute information. For example, the tile additional information includes information indicating the reference coordinate position, size, or side length of the bounding box after division. The tile additional information also includes information indicating the number of divisions, the division type, etc.
[0223] Next, an example of a method for dividing the point cloud data into slices or tiles will be described. The dividing unit 4911 may use a predetermined method as a method for dividing the point cloud data into slices or tiles, or may adaptively switch the method to be used depending on the point cloud data.
[0224] When dividing into slices, the dividing unit 4911 divides the three-dimensional space collectively based on the position information and the attribute information. For example, the dividing unit 4911 determines the shape of an object and divides the three-dimensional space into slices according to the shape of the object. For example, the dividing unit 4911 extracts objects such as trees or buildings, and divides the space on an object-by-object basis. For example, the dividing unit 4911 divides the space into slices such that one or more objects are entirely included in one slice. Alternatively, the dividing unit 4911 divides one object into multiple slices.
[0225] In this case, the encoding device may change the encoding method for each slice, for example. For example, the encoding device may use a high-quality compression method for a specific object or a specific part of an object. In this case, the encoding device may store information indicating the encoding method for each slice in additional information (metadata).
[0226] Furthermore, the dividing unit 4911 may divide the image into slices based on map information or location information so that each slice corresponds to a predetermined coordinate space.
[0227] When dividing into tiles, the dividing unit 4911 divides the position information and the attribute information independently. For example, the dividing unit 4911 divides a slice into tiles according to the amount of data or the amount of processing. For example, the dividing unit 4911 determines whether the amount of data of a slice (for example, the number of three-dimensional points included in a slice) is greater than a predetermined threshold. If the amount of data of a slice is greater than the threshold, the dividing unit 4911 divides the slice into tiles. If the amount of data of a slice is less than the threshold, the dividing unit 4911 does not divide the slice into tiles.
[0228] For example, the division unit 4911 divides the slice into tiles so that the processing amount or processing time in the decoding device is within a certain range (a predetermined value or less). This makes the processing amount per tile in the decoding device constant, facilitating distributed processing in the decoding device.
[0229] Furthermore, when the processing amount differs between the position information and the attribute information, for example when the processing amount of the position information is greater than the processing amount of the attribute information, dividing section 4911 divides the position information into a number greater than the number of divisions of the attribute information.
[0230] Also, for example, in the case where the decoding device may decode and display the position information quickly and the attribute information may be decoded and displayed slowly later depending on the content, the division unit 4911 may divide the position information into a larger number of parts than the attribute information. This allows the decoding device to process the position information in parallel in a larger number, so that the position information can be processed faster than the attribute information.
[0231] Note that the decoding device does not necessarily need to process sliced or tiled data in parallel, and may determine whether to process the data in parallel depending on the number or capabilities of the decoding processing units.
[0232] By dividing the image in the above manner, it is possible to realize adaptive encoding according to the content or object. In addition, it is possible to realize parallel processing in the decoding process. This improves the flexibility of the point cloud encoding system or the point cloud decoding system.
[0233] Fig. 32 is a diagram showing an example of a division pattern of slices and tiles. DU in the diagram is a data unit (DataUnit) and indicates data of a tile or slice. Each DU includes a slice index (SliceIndex) and a tile index (TileIndex). The numerical value in the upper right of the DU in the diagram indicates the slice index, and the numerical value in the lower left of the DU indicates the tile index.
[0234] In pattern 1, the number of divisions and the division method are the same for G slices and A slices in slice division. In tile division, the number of divisions and the division method for G slices are different from the number of divisions and the division method for A slices. Furthermore, the same number of divisions and the division method are used among multiple G slices. The same number of divisions and the division method are used among multiple A slices.
[0235] In pattern 2, the number of divisions and the division method are the same for G slices and A slices in slice division. In tile division, the number of divisions and the division method for G slices are different from the number of divisions and the division method for A slices. In addition, the number of divisions and the division method are different between multiple G slices. The number of divisions and the division method are different between multiple A slices.
[0236] Next, a method for encoding divided data will be described. A three-dimensional data encoding device (first encoding unit 4910) encodes each of the divided data. When encoding attribute information, the three-dimensional data encoding device generates dependency information indicating which configuration information (position information, additional information, or other attribute information) was used for encoding as additional information. That is, the dependency information indicates, for example, the configuration information of the reference destination (dependency destination). In this case, the three-dimensional data encoding device generates dependency information based on configuration information corresponding to the division shape of the attribute information. Note that the three-dimensional data encoding device may generate dependency information based on configuration information corresponding to a plurality of division shapes.
[0237] The dependency information may be generated by the three-dimensional data encoding device, and the generated dependency information may be sent to the three-dimensional data decoding device. Alternatively, the three-dimensional data decoding device may generate the dependency information, and the three-dimensional data encoding device may not send the dependency information. Also, the dependency relationships used by the three-dimensional data encoding device may be determined in advance, and the three-dimensional data encoding device may not send the dependency information.
[0238] Fig. 33 is a diagram showing an example of the dependency relationship of each data. The tip of the arrow in the diagram indicates the dependency destination, and the start of the arrow indicates the dependency source. The three-dimensional data decoding device decodes data in the order from the dependency destination to the dependency source. Also, data shown by solid lines in the diagram is data that is actually sent, and data shown by dotted lines is data that is not sent.
[0239] Also, in the figure, G indicates position information, and A indicates attribute information. Gs1 indicates position information of slice number 1, and Gs2 indicates position information of slice number 2. Gs1t1 indicates position information of slice number 1 and tile number 1, Gs1t2 indicates position information of slice number 1 and tile number 2, Gs2t1 indicates position information of slice number 2 and tile number 1, and Gs2t2 indicates position information of slice number 2 and tile number 2. Similarly, As1 indicates attribute information of slice number 1, and As2 indicates attribute information of slice number 2. As1t1 indicates attribute information of slice number 1 and tile number 1, As1t2 indicates attribute information of slice number 1 and tile number 2, As2t1 indicates attribute information of slice number 2 and tile number 1, and As2t2 indicates attribute information of slice number 2 and tile number 2.
[0240] Mslice indicates slice additional information, MGtile indicates position tile additional information, and MAtile indicates attribute tile additional information. Ds1t1 indicates dependency relationship information of attribute information As1t1, and Ds2t1 indicates dependency relationship information of attribute information As2t1.
[0241] Furthermore, the three-dimensional data encoding device may rearrange the data in the decoding order so that there is no need to rearrange the data in the three-dimensional data decoding device. Note that the data may be rearranged in the three-dimensional data decoding device, or the data may be rearranged in both the three-dimensional data encoding device and the three-dimensional data decoding device.
[0242] Fig. 34 is a diagram showing an example of the data decoding order. In the example of Fig. 34, decoding is performed in order from the left data. When data has a dependency relationship, the three-dimensional data decoding device decodes the dependent data first. For example, the three-dimensional data encoding device rearranges the data in advance so as to achieve this order, and transmits the data. Any order may be used as long as the dependent data comes first. The three-dimensional data encoding device may also transmit additional information and dependency information before the data.
[0243] Fig. 35 is a flowchart showing the flow of processing by the three-dimensional data encoding device. First, the three-dimensional data encoding device encodes data of a plurality of slices or tiles as described above (S4901). Next, the three-dimensional data encoding device rearranges the data so that the dependent data comes first, as shown in Fig. 34 (S4902). Next, the three-dimensional data encoding device multiplexes (NAL unitizes) the rearranged data (S4903).
[0244] Next, a description will be given of the configuration of the combining unit 4925 included in the first decoding unit 4920. Fig. 36 is a block diagram showing the configuration of the combining unit 4925. The combining unit 4925 includes a position information tile combining unit 4941 (geometry tile combiner), an attribute information tile combining unit 4942 (attribute tile combiner), and a slice combining unit (slice combiner).
[0245] The position information tile combining unit 4941 generates a plurality of slice position information pieces by combining a plurality of pieces of divided position information pieces using the position tile additional information. The attribute information tile combining unit 4942 generates a plurality of slice attribute information pieces by combining a plurality of pieces of divided attribute information pieces using the attribute tile additional information.
[0246] The slice combining unit 4943 generates position information by combining a plurality of slice position information pieces using the slice additional information. Also, the slice combining unit 4943 generates attribute information by combining a plurality of slice attribute information pieces using the slice additional information.
[0247] The number of slices or tiles to be divided is equal to or greater than 1. In other words, the division into slices or tiles does not necessarily have to be performed.
[0248] Although an example in which tile division is performed after slice division has been shown here, slice division may be performed after tile division. Furthermore, a new division type may be defined in addition to slices and tiles, and division may be performed using three or more division types.
[0249] Next, a structure of coded data divided into slices or tiles and a method of storing the coded data in NAL units (multiplexing method) will be described. Fig. 37 is a diagram showing the structure of coded data and a method of storing the coded data in NAL units.
[0250] The encoded data (partition position information and partition attribute information) is stored in the payload of the NAL unit.
[0251] The encoded data includes a header and a payload. The header includes identification information for identifying data included in the payload. This identification information includes, for example, a type of slice division or tile division (slice_type, tile_type), index information for identifying a slice or tile (slice_idx, tile_idx), position information of data (slice or tile), or an address of data (address). The index information for identifying a slice is also referred to as a slice index (SliceIndex). The index information for identifying a tile is also referred to as a tile index (TileIndex). The type of division is, for example, a method based on an object shape as described above, a method based on map information or position information, or a method based on a data amount or a processing amount.
[0252] In addition, all or part of the above information may be stored in one of the headers of the divided position information and the divided attribute information, and may not be stored in the other. For example, when the same division method is used for the position information and the attribute information, the division type (slice_type, tile_type) and index information (slice_idx, tile_idx) are the same for the position information and the attribute information. Therefore, these information may be included in the header of one of the position information and the attribute information. For example, when the attribute information depends on the position information, the position information is processed first. Therefore, these information may be included in the header of the position information, and may not be included in the header of the attribute information. In this case, the three-dimensional data decoding device determines that the dependent attribute information belongs to the same slice or tile as the slice or tile of the dependent position information, for example.
[0253] In addition, additional information related to slice division or tile division (slice additional information, position tile additional information, or attribute tile additional information), and dependency information indicating dependency, etc. may be stored in an existing parameter set (GPS, APS, position SPS, attribute SPS, etc.) and transmitted. When the division method changes for each frame, information indicating the division method may be stored in a parameter set for each frame (GPS or APS, etc.). When the division method does not change within a sequence, information indicating the division method may be stored in a parameter set for each sequence (position SPS or attribute SPS). Furthermore, when the same division method is used for position information and attribute information, information indicating the division method may be stored in a parameter set of a PCC stream (stream PS).
[0254] The above information may be stored in any one of the above parameter sets, or in multiple parameter sets. A parameter set for tile division or slice division may be defined, and the above information may be stored in the parameter set. The information may be stored in a header of the encoded data.
[0255] Furthermore, the header of the encoded data includes identification information indicating a dependency relationship. That is, when there is a dependency relationship between data, the header includes identification information for referring to the dependency from the dependency source. For example, the header of the dependency destination data includes identification information for identifying the data. The header of the dependency source data includes identification information indicating the dependency. Note that, when the identification information for identifying the data, the additional information related to the slice division or tile division, and the identification information indicating the dependency relationship can be identified or derived from other information, these pieces of information may be omitted.
[0256] Next, a flow of the encoding process and the decoding process of the point cloud data according to this embodiment will be described. Fig. 38 is a flowchart of the encoding process of the point cloud data according to this embodiment.
[0257] First, the three-dimensional data encoding device determines the division method to be used (S4911). The division method includes whether or not to perform slice division and whether or not to perform tile division. The division method may also include the number of divisions when performing slice division or tile division, and the type of division. The type of division may be a method based on the object shape as described above, a method based on map information or position information, or a method based on the amount of data or the amount of processing. The division method may be determined in advance.
[0258] When slice division is performed (Yes in S4912), the three-dimensional data encoding device generates a plurality of slice position information and a plurality of slice attribute information by dividing the position information and the attribute information together (S4913). In addition, the three-dimensional data encoding device generates slice additional information related to the slice division. Note that the three-dimensional data encoding device may divide the position information and the attribute information independently.
[0259] If tile division is performed (Yes in S4914), the three-dimensional data encoding device generates a plurality of division position information and a plurality of division attribute information by independently dividing a plurality of slice position information and a plurality of slice attribute information (or position information and attribute information) (S4915). In addition, the three-dimensional data encoding device generates position tile additional information and attribute tile additional information related to the tile division. Note that the three-dimensional data encoding device may divide the slice position information and slice attribute information together.
[0260] Next, the three-dimensional data encoding device generates a plurality of pieces of encoding position information and a plurality of pieces of encoding attribute information by encoding each of the plurality of pieces of division position information and the plurality of pieces of division attribute information (S4916). Also, the three-dimensional data encoding device generates dependency relationship information.
[0261] Next, the three-dimensional data encoding device generates encoded data (encoded stream) by forming (multiplexing) the plurality of pieces of encoding position information, the plurality of pieces of encoding attribute information, and the additional information into NAL units (S4917). In addition, the three-dimensional data encoding device transmits the generated encoded data.
[0262] 39 is a flowchart of a decoding process of point cloud data according to this embodiment. First, the three-dimensional data decoding device analyzes additional information (slice additional information, position tile additional information, and attribute tile additional information) related to the division method included in the encoded data (encoded stream) to determine the division method (S4921). This division method includes whether or not to perform slice division and whether or not to perform tile division. In addition, the division method may include the number of divisions when performing slice division or tile division, the type of division, and the like.
[0263] Next, the three-dimensional data decoding device generates split position information and split attribute information by decoding the multiple pieces of encoded position information and multiple pieces of encoded attribute information contained in the encoded data using dependency information contained in the encoded data (S4922).
[0264] When the additional information indicates that tile division has been performed (Yes in S4923), the three-dimensional data decoding device generates a plurality of slice position information and a plurality of slice attribute information by combining a plurality of pieces of division position information and a plurality of pieces of division attribute information by respective methods based on the position tile additional information and the attribute tile additional information (S4924). Note that the three-dimensional data decoding device may combine a plurality of pieces of division position information and a plurality of pieces of division attribute information by the same method.
[0265] When the additional information indicates that slice division has been performed (Yes in S4925), the three-dimensional data decoding device generates position information and attribute information by combining multiple slice position information and multiple slice attribute information (multiple division position information and multiple division attribute information) in the same manner based on the slice additional information (S4926). Note that the three-dimensional data decoding device may combine multiple slice position information and multiple slice attribute information in different manners.
[0266] In addition, attribute information of tiles or slices (identifiers, area information, address information, position information, etc.) may be stored in other control information, not limited to SEI. For example, the attribute information may be stored in control information indicating the configuration of the entire PCC data, or may be stored in control information for each tile or slice.
[0267] Furthermore, when transmitting PCC data to another device, the three-dimensional data encoding device (three-dimensional data transmission device) may convert control information such as SEI into control information specific to the protocol of that system and display it.
[0268] For example, when the three-dimensional data encoding device converts PCC data including attribute information into ISOBMFF (ISO Base Media File Format), the SEI may be stored in an "mdat box" together with the PCC data, or in a "track box" that describes control information related to the stream. In other words, the three-dimensional data encoding device may store the control information in a table for random access. Also, when the three-dimensional data encoding device packets the PCC data for transmission, the SEI may be stored in a packet header. In this way, by making it possible to acquire attribute information at a system layer, access to the attribute information and the tile data or slice data becomes easier, and the access speed can be improved.
[0269] In addition, in the configuration of the three-dimensional data decoding device, the memory management unit may determine in advance whether or not the information necessary for the decoding process is in the memory, and if the information necessary for the decoding process is not present, the information may be obtained from storage or a network.
[0270] When the three-dimensional data decoding device acquires PCC data from a storage or a network using Pull in a protocol such as MPEG-DASH, the memory management unit may identify attribute information of data necessary for decoding processing based on information from a localization unit or the like, request tiles or slices including the identified attribute information, and acquire the necessary data (PCC stream). Identification of tiles or slices including attribute information may be performed on the storage or network side, or may be performed by the memory management unit. For example, the memory management unit may acquire SEI of all PCC data in advance, and identify tiles or slices based on that information.
[0271] When all PCC data is transmitted from the storage or network using Push in a UDP protocol or the like, the memory management unit may identify attribute information of the data required for the decoding process and tiles or slices based on information from the localization unit or the like, and obtain the desired data by filtering the desired tiles or slices from the transmitted PCC data.
[0272] Furthermore, when acquiring data, the three-dimensional data encoding device may determine whether or not desired data is available, whether or not real-time processing is possible based on the data size, etc., or the communication state, etc. If the three-dimensional data encoding device determines that data acquisition is difficult based on the result of this determination, it may select and acquire another slice or tile with a different priority or amount of data.
[0273] In addition, the three-dimensional data decoding device may transmit information from a localization unit or the like to a cloud server, and the cloud server may determine the necessary information based on that information.
[0274] (Embodiment 4) Next, the tile additional information will be described. The three-dimensional data encoding device generates tile additional information, which is metadata related to a method of dividing tiles, and transmits the generated tile additional information to the three-dimensional data decoding device.
[0275] Fig. 40 is a diagram showing an example of the syntax of the tile additional information (TileMetaData). As shown in Fig. 40, for example, the tile additional information includes division method information (type_of_divide), shape information (topview_shape), overlap flag (tile_overlap_flag), overlap information (type_of_overlap), height information (tile_height), number of tiles (tile_number), and tile position information (global_position, relative_position).
[0276] The division method information (type_of_divide) indicates a division method of tiles. For example, the division method information indicates whether the division method of tiles is based on map information, that is, division based on a top view (top_view), or other (other).
[0277] The shape information (topview_shape) is included in the tile additional information, for example, when the tile division method is division based on a top view. The shape information indicates the shape of the tile when viewed from above. For example, this shape includes a square and a circle. Note that this shape may include an ellipse, a rectangle, or a polygon other than a quadrangle, or may include other shapes. Note that the shape information is not limited to the shape of the tile when viewed from above, and may also indicate the three-dimensional shape of the tile (for example, a cube, a cylinder, etc.).
[0278] The overlap flag (tile_overlap_flag) indicates whether tiles overlap. For example, the overlap flag is included in the tile additional information when the tile division method is division based on a top view. In this case, the overlap flag indicates whether tiles overlap in a top view. Note that the overlap flag may also indicate whether tiles overlap in a three-dimensional space.
[0279] The overlap information (type_of_overlap) is included in the tile additional information when tiles overlap, for example. The overlap information indicates how the tiles overlap, etc. For example, the overlap information indicates the size of the overlapping area, etc.
[0280] Height information (tile_height) indicates the height of a tile. The height information may include information indicating the shape of the tile. For example, if the shape of the tile in top view is rectangular, the information may indicate the lengths of the sides of the rectangle (the vertical length and the horizontal length). Furthermore, if the shape of the tile in top view is circular, the information may indicate the diameter or radius of the circle.
[0281] The height information may indicate the height of each tile, or may indicate a common height for multiple tiles. Multiple height types, such as roads and intersections, may be set in advance, and the height information may indicate the height of each height type and the height type of each tile. Alternatively, the height of each height type may be defined in advance, and the height information may indicate the height type of each tile. In other words, the height of each height type does not have to be indicated by the height information.
[0282] The tile number (tile_number) indicates the number of tiles. The tile additional information may include information indicating the spacing between tiles.
[0283] The tile position information (global_position, relative_position) is information for identifying the position of each tile. For example, the tile position information indicates the absolute coordinates or relative coordinates of each tile.
[0284] Note that some or all of the above information may be provided for each tile, or for each set of tiles (for example, for each frame or for each set of frames).
[0285] The three-dimensional data encoding device may include the tile additional information in SEI (Supplemental Enhancement Information) and transmit it. Alternatively, the three-dimensional data encoding device may store the tile additional information in an existing parameter set (such as PPS, GPS, or APS) and transmit it.
[0286] For example, if the tile additional information changes for each frame, the tile additional information may be stored in a parameter set for each frame (GPS or APS, etc.). If the tile additional information does not change within a sequence, the tile additional information may be stored in a parameter set for each sequence (position SPS or attribute SPS). Furthermore, if the same tile division information is used for position information and attribute information, the tile additional information may be stored in a parameter set of a PCC stream (stream PS).
[0287] Furthermore, the tile additional information may be stored in any one of the above parameter sets, or in multiple parameter sets. Furthermore, the tile additional information may be stored in a header of the encoded data. Furthermore, the tile additional information may be stored in a header of a NAL unit.
[0288] In addition, all or part of the tile additional information may be stored in one of the headers of the divided position information and the divided attribute information, and not in the other. For example, when the same tile additional information is used in the position information and the attribute information, the tile additional information may be included in the header of either the position information or the attribute information. For example, when the attribute information depends on the position information, the position information is processed first. Therefore, the header of the position information may include these tile additional information, and the header of the attribute information may not include the tile additional information. In this case, the three-dimensional data decoding device determines, for example, that the dependent attribute information belongs to the same tile as the tile of the dependent position information.
[0289] The three-dimensional data decoding device reconstructs the point cloud data divided into tiles based on the tile additional information. If there is overlapping point cloud data, the three-dimensional data decoding device identifies the overlapping multiple point cloud data and selects one of them or merges the multiple point cloud data.
[0290] The three-dimensional data decoding device may also perform decoding using tile additional information. For example, when multiple tiles overlap, the three-dimensional data decoding device may perform decoding for each tile, and perform processing (e.g., smoothing or filtering) using the multiple decoded data to generate point cloud data. This may enable highly accurate decoding.
[0291] 41 is a diagram showing an example of the configuration of a system including a three-dimensional data encoding device and a three-dimensional data decoding device. A tile dividing unit 5051 divides point cloud data including position information and attribute information into a first tile and a second tile. The tile dividing unit 5051 also sends tile additional information related to the tile division to a decoding unit 5053 and a tile combining unit 5054.
[0292] The encoding unit 5052 generates encoded data by encoding the first tile and the second tile.
[0293] The decoding unit 5053 restores the first tile and the second tile by decoding the coded data generated by the coding unit 5052. The tile combining unit 5054 restores the point cloud data (position information and attribute information) by combining the first tile and the second tile using the tile additional information.
[0294] Next, slice additional information will be described. The three-dimensional data encoding device generates slice additional information, which is metadata related to a method of dividing a slice, and transmits the generated slice additional information to the three-dimensional data decoding device.
[0295] Fig. 42 is a diagram showing an example of the syntax of slice additional information (SliceMetaData). As shown in Fig. 42, for example, the slice additional information includes division method information (type_of_divide), an overlap flag (slice_overlap_flag), overlap information (type_of_overlap), the number of slices (slice_number), slice position information (global_position, relative_position), and slice size information (slice_bounding_box_size).
[0296] The division method information (type_of_divide) indicates a division method of a slice. For example, the division method information indicates whether the division method of a slice is division based on object information as shown in FIG. 60 (object). The slice additional information may include information indicating a method of object division. For example, this information indicates whether one object is divided into multiple slices or assigned to one slice. This information may also indicate the number of divisions when one object is divided into multiple slices.
[0297] The overlap flag (slice_overlap_flag) indicates whether the slices overlap. The overlap information (type_of_overlap) is included in the slice additional information when the slices overlap, for example. The overlap information indicates how the slices overlap, etc. For example, the overlap information indicates the size of the overlapping area, etc.
[0298] The slice number (slice_number) indicates the number of slices.
[0299] Slice position information (global_position, relative_position) and slice size information (slice_bounding_box_size) are information related to the area of a slice. Slice position information is information for identifying the position of each slice. For example, slice position information indicates absolute coordinates or relative coordinates of each slice. Slice size information (slice_bounding_box_size) indicates the size of each slice. For example, slice size information indicates the size of the bounding box of each slice.
[0300] The three-dimensional data encoding device may include the slice additional information in the SEI and transmit it. Alternatively, the three-dimensional data encoding device may store the slice additional information in an existing parameter set (such as PPS, GPS, or APS) and transmit it.
[0301] For example, when slice additional information changes for each frame, the slice additional information may be stored in a parameter set for each frame (GPS or APS, etc.). When slice additional information does not change within a sequence, the slice additional information may be stored in a parameter set for each sequence (position SPS or attribute SPS). Furthermore, when the same slice division information is used for position information and attribute information, the slice additional information may be stored in a parameter set of a PCC stream (stream PS).
[0302] Furthermore, the slice additional information may be stored in any one of the above parameter sets, or in a plurality of parameter sets. Furthermore, the slice additional information may be stored in a header of the encoded data. Furthermore, the slice additional information may be stored in a header of the NAL unit.
[0303] In addition, all or a part of the slice additional information may be stored in one of the headers of the division position information and the header of the division attribute information, and may not be stored in the other. For example, when the same slice additional information is used in the position information and the attribute information, the slice additional information may be included in the header of one of the position information and the attribute information. For example, when the attribute information depends on the position information, the position information is processed first. Therefore, the header of the position information may include these slice additional information, and the header of the attribute information may not include the slice additional information. In this case, the three-dimensional data decoding device determines that the dependent attribute information belongs to the same slice as the slice of the dependent position information, for example.
[0304] The three-dimensional data decoding device reconstructs the point cloud data divided into slices based on the slice additional information. If there is overlapping point cloud data, the three-dimensional data decoding device identifies the overlapping multiple point cloud data and selects one of them or merges the multiple point cloud data.
[0305] The three-dimensional data decoding device may also perform decoding using slice additional information. For example, when multiple slices overlap, the three-dimensional data decoding device may perform decoding for each slice, and perform processing (e.g., smoothing or filtering) using the multiple decoded data to generate point cloud data. This may enable highly accurate decoding.
[0306] FIG. 43 is a flowchart of three-dimensional data encoding processing, including processing for generating tile additional information, performed by the three-dimensional data encoding device according to this embodiment.
[0307] First, the three-dimensional data encoding device determines a tile division method (S5031). Specifically, the three-dimensional data encoding device determines whether to use a division method based on a top view (top_view) or another method (other) as a tile division method. The three-dimensional data encoding device also determines the shape of the tile when using a division method based on a top view. The three-dimensional data encoding device also determines whether the tile overlaps with other tiles.
[0308] If the tile division method determined in step S5031 is a division method based on a top view (Yes in S5032), the three-dimensional data encoding device describes in the tile additional information that the tile division method is a division method based on a top view (top_view) (S5033).
[0309] On the other hand, if the tile division method determined in step S5031 is other than a division method based on a top view (No in S5032), the three-dimensional data encoding device describes in the tile additional information that the tile division method is other than a division method based on a top view (top_view) (S5034).
[0310] Furthermore, if the shape of the tile viewed from above determined in step S5031 is a square (square in S5035), the three-dimensional data encoding device records in the tile additional information that the shape of the tile viewed from above is a square (S5036).On the other hand, if the shape of the tile viewed from above determined in step S5031 is a circle (circle in S5035), the three-dimensional data encoding device records in the tile additional information that the shape of the tile viewed from above is a circle (S5037).
[0311] Next, the three-dimensional data encoding device determines whether the tile overlaps with other tiles (S5038). If the tile overlaps with other tiles (Yes in S5038), the three-dimensional data encoding device records in the tile additional information that the tile overlaps (S5039). On the other hand, if the tile does not overlap with other tiles (No in S5038), the three-dimensional data encoding device records in the tile additional information that the tile does not overlap (S5040).
[0312] Next, the three-dimensional data encoding device divides the tiles based on the tile division method determined in step S5031, encodes each tile, and transmits the generated encoded data and tile additional information (S5041).
[0313] FIG. 44 is a flowchart of three-dimensional data decoding processing using tile additional information by the three-dimensional data decoding device according to this embodiment.
[0314] First, the 3D data decoding device analyzes the tile additional information included in the bitstream (S5051).
[0315] If the tile additional information indicates that the tile does not overlap with other tiles (No in S5052), the 3D data decoding device generates point cloud data for each tile by decoding each tile (S5053). Next, the 3D data decoding device reconstructs point cloud data from the point cloud data for each tile based on the tile division method and tile shape indicated in the tile additional information (S5054).
[0316] On the other hand, if the tile additional information indicates that the tile overlaps with other tiles (Yes in S5052), the three-dimensional data decoding device generates point cloud data for each tile by decoding each tile. The three-dimensional data decoding device also identifies overlapping parts of the tiles based on the tile additional information (S5055). Note that the three-dimensional data decoding device may perform decoding processing for the overlapping parts using multiple pieces of overlapping information. Next, the three-dimensional data decoding device reconstructs point cloud data from the point cloud data of each tile based on the tile division method, tile shape, and overlap information indicated in the tile additional information (S5056).
[0317] Modifications and the like regarding slices will be described below. The three-dimensional data encoding device may transmit information indicating the type (road, building, tree, etc.) or attribute (dynamic information, static information, etc.) of an object as additional information. Alternatively, encoding parameters may be predefined according to the object, and the three-dimensional data encoding device may notify the three-dimensional data decoding device of the encoding parameters by sending the type or attribute of the object.
[0318] The following methods may be used for the coding order and transmission order of slice data. For example, the three-dimensional data encoding device may encode slice data in the order of data that is easiest to recognize or cluster. Alternatively, the three-dimensional data encoding device may encode slice data in the order of slice data that has been clustered first. Furthermore, the three-dimensional data encoding device may transmit the encoded slice data in the order of decreasing priority for decoding in an application. For example, when the priority of decoding dynamic information is high, the three-dimensional data encoding device may transmit the slice data in the order of decreasing priority for decoding dynamic information.
[0319] Furthermore, when the order of the encoded data differs from the order of the decoding priority, the three-dimensional data encoding device may rearrange the encoded data before transmitting it. Furthermore, when storing the encoded data, the three-dimensional data encoding device may store the encoded data after rearranging it.
[0320] An application (a three-dimensional data decoding device) requests a server (a three-dimensional data encoding device) to transmit slices containing desired data. The server transmits slice data required by the application, but does not need to transmit unnecessary slice data.
[0321] An application requests the server to send tiles containing desired data. The server sends the tile data required by the application, but does not need to send unnecessary tile data.
[0322] (Embodiment 5) In this embodiment, processing of division units (for example, tiles or slices) that do not include points will be described. First, a method of dividing point cloud data will be described.
[0323] In video coding standards such as HEVC, data exists for every pixel in a two-dimensional image, so even if a two-dimensional space is divided into multiple data regions, data exists in every data region. On the other hand, in coding of three-dimensional point cloud data, the points that are elements of the point cloud data are themselves data, and there is a possibility that data does not exist in some regions.
[0324] There are various methods for spatially dividing point cloud data, but the division methods can be classified according to whether a division unit (for example, a tile or slice), which is a divided data unit, always contains one or more point data.
[0325] A division method in which all of the division units contain one or more point data is called a first division method. For example, the first division method is a method in which the point cloud data is divided while considering the encoding processing time or the size of the encoded data. In this case, the number of points in each division unit is roughly equal.
[0326] Fig. 45 is a diagram showing an example of a division method. For example, as a first division method, a method of dividing points belonging to the same space into two identical spaces as shown in (a) of Fig. 45 may be used. Also, as shown in (b) of Fig. 45, a space may be divided into a plurality of subspaces (division units) such that each division unit includes a point.
[0327] These methods are point-aware divisions, so every division unit always contains at least one point.
[0328] A division method in which multiple division units may include one or more division units that do not contain point data is called a second division method. For example, as the second division method, a method of equally dividing a space can be used, as shown in (c) of FIG. 45. In this case, a point does not necessarily exist in a division unit. In other words, there are cases in which a point does not exist in a division unit.
[0329] When a three-dimensional data encoding device divides point cloud data, it may indicate in division additional information (metadata) related to the division (e.g., tile additional information or slice additional information) whether (1) a division method was used in which all of the multiple division units include one or more point data, (2) a division method was used in which the multiple division units have one or more division units that do not include point data, or (3) a division method was used in which the multiple division units may have one or more division units that do not include point data, and send out the division additional information.
[0330] The three-dimensional data encoding device may indicate the above information as the type of division method. Also, the three-dimensional data encoding device may perform division using a predetermined division method and not transmit the division additional information. In that case, the three-dimensional data encoding device indicates in advance whether the division method is the first division method or the second division method.
[0331] The second division method and an example of generating and transmitting encoded data will be described below. Note that, although tile division will be described as an example of a method of dividing a three-dimensional space, the following technique can also be applied to a division method using a division unit other than tiles. For example, tile division may be read as slice division.
[0332] Fig. 46 is a diagram showing an example of dividing point cloud data into six tiles. Fig. 46 shows an example in which the minimum unit is a point, and shows an example in which position information (Geometry) and attribute information (Attribute) are divided together. Note that the same applies when the position information and attribute information are divided by different division methods or division numbers, when there is no attribute information, and when there are multiple pieces of attribute information.
[0333] In the example shown in Figure 46, after tile division, there are tiles (#1, #2, #4, #6) that contain points and tiles (#3, #5) that do not contain points. Tiles that do not contain points are called null tiles.
[0334] Note that the division is not limited to six tiles, and any division method may be used. For example, the division unit may be a cube, or may be a non-cubic shape such as a rectangular parallelepiped or a cylinder. The multiple division units may have the same shape, or may include different shapes. In addition, a predetermined method may be used as the division method, or a different method may be used for each predetermined unit (e.g., PCC frame).
[0335] In this division method, when point cloud data is divided into tiles, if there is no data in a tile, a bitstream is generated that includes information indicating that the tile is a null tile.
[0336] Hereinafter, a method for transmitting null tiles and a method for signaling null tiles will be described. The three-dimensional data encoding device may generate, for example, the following information as additional information (metadata) related to data division, and transmit the generated information. Fig. 47 is a diagram showing an example of the syntax of tile additional information (TileMetaData). The tile additional information includes division method information (type_of_divide), division method null information (type_of_divide_null), number of tile divisions (number_of_tiles), and tile null flag (tile_null_flag).
[0337] The division method information (type_of_divide) is information related to the division method or division type. For example, the division method information indicates one or more division methods or division types. For example, the division method includes top view division and equal division. Note that when there is one division method definition, the division method information does not need to be included in the tile additional information.
[0338] The division method null information (type_of_divide_null) is information indicating whether the division method used is the first division method or the second division method described below. Here, the first division method is a division method in which all of the multiple division units always contain one or more point data. The second division method is a division method in which the multiple division units include one or more division units that do not contain point data, or in which the multiple division units may include one or more division units that do not contain point data.
[0339] The tile additional information may include, as division information for the entire tile, at least one of: (1) information indicating the number of divisions of the tile (number_of_tiles) or information for specifying the number of divisions of the tile, (2) information indicating the number of null tiles or information for specifying the number of null tiles, and (3) information indicating the number of tiles other than null tiles or information for specifying the number of tiles other than null tiles. The tile additional information may include, as division information for the entire tile, information indicating the shape of the tile or whether the tiles overlap.
[0340] Furthermore, the tile additional information indicates the division information for each tile in order. For example, the order of tiles is predetermined for each division method and is known in the three-dimensional data encoding device and the three-dimensional data decoding device. Note that, if the order of tiles is not predetermined, the three-dimensional data encoding device may send information indicating the order to the three-dimensional data decoding device.
[0341] The division information for each tile includes a tile null flag (tile_null_flag) that is a flag indicating whether or not data (points) exist in the tile. Note that if there is no data in the tile, the tile null flag may be included as the tile division information.
[0342] Furthermore, if the tile is not a null tile, the tile additional information includes division information for each tile (position information (e.g., coordinates of the origin (origin_x, origin_y, origin_z)) and height information of the tile, etc.). If the tile is a null tile, the tile additional information does not include division information for each tile.
[0343] For example, when slice division information for each tile is stored in the division information for each tile, the three-dimensional data encoding device does not need to store slice division information for null tiles in the additional information.
[0344] In this example, the number of tile divisions (number_of_tiles) indicates the number of tiles including null tiles. Fig. 48 is a diagram showing an example of tile index information (idx). In the example shown in Fig. 48, index information is also assigned to null tiles.
[0345] Next, a data structure and a transmission method of coded data including null tiles will be described. Figures 49 to 51 are diagrams showing data structures in the case where position information and attribute information are divided into six tiles and data does not exist in the third and fifth tiles.
[0346] Fig. 49 is a diagram showing an example of the dependency of each data. The tip of the arrow in the diagram indicates the dependency destination, and the base of the arrow indicates the dependency source. Also, in the diagram, Gtn (n is 1 to 6) indicates the position information of the tile number n, Atn indicates the attribute information of the tile number n, and Mtile indicates the tile additional information.
[0347] Fig. 50 is a diagram showing an example of the structure of transmission data, which is encoded data transmitted from a three-dimensional data encoding device. Fig. 51 is a diagram showing the structure of encoded data and a method of storing the encoded data in an NAL unit.
[0348] As shown in FIG. 51, the headers of the data of the position information (division position information) and the attribute information (division attribute information) each include tile index information (tile_idx).
[0349] Also, as shown in Structure 1 of Fig. 50, the three-dimensional data encoding device may not transmit position information or attribute information constituting a null tile. Alternatively, as shown in Structure 2 of Fig. 50, the three-dimensional data encoding device may transmit information indicating that the tile is a null tile as data of the null tile. For example, the three-dimensional data encoding device may indicate that the type of the data is a null tile in the tile_type stored in the header of the NAL unit or in the header in the payload (nal_unit_payload) of the NAL unit, and transmit the header. Note that the following description will be given assuming Structure 1.
[0350] In structure 1, when a null tile exists, the value of the tile index information (tile_idx) included in the header of the position information data or attribute information data in the transmission data is not consecutive and has gaps.
[0351] Furthermore, when there is a dependency between data, the three-dimensional data encoding device transmits the data so that the referenced data can be decoded before the referenced data. Note that the attribute information tiles have a dependency on the position information tiles. The attribute information and position information that have a dependency are assigned the same tile index number.
[0352] In addition, the tile additional information related to the tile division may be stored in both the parameter set of the position information (GPS) and the parameter set of the attribute information (APS), or in either one of them. When the tile additional information is stored in one of the GPS and the APS, reference information indicating the referenced GPS or APS may be stored in the other of the GPS and the APS. Furthermore, when the tile division method differs between the position information and the attribute information, different tile additional information is stored in the GPS and the APS, respectively. Furthermore, when the tile division method is the same for a sequence (multiple PCC frames), the tile additional information may be stored in the GPS, the APS, or the SPS (sequence parameter set).
[0353] For example, when tile additional information is stored in both the GPS and the APS, the tile additional information of position information is stored in the GPS, and the tile additional information of attribute information is stored in the APS. When tile additional information is stored in common information such as an SPS, tile additional information commonly used in the position information and the attribute information may be stored, or the tile additional information of position information and the tile additional information of attribute information may be stored separately.
[0354] A combination of tile division and slice division will be described below. First, a data structure and data transmission in the case where tile division is performed after slice division will be described.
[0355] Fig. 52 is a diagram showing an example of the dependency relationship of each piece of data when dividing into tiles after dividing into slices. The tip of the arrow in the diagram indicates the dependency destination, and the base of the arrow indicates the dependency source. In addition, data shown with a solid line in the diagram is data that is actually transmitted, and data shown with a dotted line is data that is not transmitted.
[0356] Also in the figure, G indicates position information, and A indicates attribute information. Gs1 indicates position information of slice number 1, and Gs2 indicates position information of slice number 2. Gs1t1 indicates position information of slice number 1 and tile number 1, and Gs2t2 indicates position information of slice number 2 and tile number 2. Similarly, As1 indicates attribute information of slice number 1, and As2 indicates attribute information of slice number 2. As1t1 indicates attribute information of slice number 1 and tile number 1, and As2t1 indicates attribute information of slice number 2 and tile number 1.
[0357] Mslice indicates slice additional information, MGtile indicates position tile additional information, and MAtile indicates attribute tile additional information. Ds1t1 indicates dependency relationship information of attribute information As1t1, and Ds2t1 indicates dependency relationship information of attribute information As2t1.
[0358] The three-dimensional data encoding device does not need to generate and transmit position information and attribute information related to null tiles.
[0359] In addition, even if the number of tile divisions is the same for all slices, the number of tiles generated and transmitted between slices may differ. For example, when the number of tile divisions differs between position information and attribute information, null tiles may exist in either position information or attribute information, and not exist in the other. In the example shown in FIG. 52, position information (Gs1) of slice 1 is divided into two tiles, Gs1t1 and Gs1t2, of which Gs1t2 is a null tile. On the other hand, attribute information (As1) of slice 1 is not divided, and one As1t1 exists, and no null tiles exist.
[0360] Furthermore, the three-dimensional data encoding device generates and transmits dependency information of the attribute information when data exists at least in the tile of the attribute information, regardless of whether or not a null tile is included in the slice of the position information. For example, when the three-dimensional data encoding device stores information on slice division for each tile in division information for each slice included in slice additional information related to slice division, the three-dimensional data encoding device stores information on whether the tile is a null tile in this information.
[0361] Fig. 53 is a diagram showing an example of the data decoding order. In the example of Fig. 53, decoding is performed in order from the left data. When data has a dependency relationship, the three-dimensional data decoding device decodes the dependent data first. For example, the three-dimensional data encoding device rearranges the data in advance so as to achieve this order, and transmits the data. Any order may be used as long as the dependent data comes first. The three-dimensional data encoding device may also transmit additional information and dependency information before the data.
[0362] Next, a data structure and data transmission in the case where slice division is performed after tile division will be described.
[0363] Fig. 54 is a diagram showing an example of the dependency relationship of each piece of data when dividing into slices after dividing into tiles. The tip of the arrow in the diagram indicates the dependency destination, and the base of the arrow indicates the dependency source. In addition, data shown with a solid line in the diagram is data that is actually sent, and data shown with a dotted line is data that is not sent.
[0364] In the figure, G indicates position information, and A indicates attribute information. Gt1 indicates position information of tile number 1. Gt1s1 indicates position information of tile number 1 and slice number 1, and Gt1s2 indicates position information of tile number 1 and slice number 2. Similarly, At1 indicates attribute information of tile number 1, and At1s1 indicates attribute information of tile number 1 and slice number 1.
[0365] Mtile indicates tile additional information, MGslice indicates position slice additional information, and MAslice indicates attribute slice additional information. Dt1s1 indicates dependency relationship information of attribute information At1s1, and Dt2s1 indicates dependency relationship information of attribute information At2s1.
[0366] The three-dimensional data encoding device does not divide null tiles into slices, and does not need to generate and transmit position information, attribute information, and dependency information of attribute information related to null tiles.
[0367] FIG. 55 is a diagram showing an example of the data decoding order. In the example of FIG. 55, decoding is performed in order from the left data. When data has a dependency relationship, the three-dimensional data decoding device decodes the dependent data first. For example, the three-dimensional data encoding device rearranges the data in advance so as to achieve this order, and transmits the data. Any order may be used as long as the dependent data comes first. Furthermore, the three-dimensional data encoding device may transmit the additional information and dependency information before the data.
[0368] Next, a flow of a process of dividing and combining point cloud data will be described. Note that although an example of dividing into tiles and slices will be described here, a similar method can be applied to dividing other spaces.
[0369] 56 is a flowchart of a three-dimensional data encoding process including a data division process by a three-dimensional data encoding device. First, the three-dimensional data encoding device determines the division method to be used (S5101). Specifically, the three-dimensional data encoding device determines whether to use the first division method or the second division method. For example, the three-dimensional data encoding device may determine the division method based on a designation from a user or an external device (e.g., a three-dimensional data decoding device), or may determine the division method according to input point cloud data. In addition, the division method to be used may be predetermined.
[0370] Here, the first division method is a division method in which all of the multiple division units (tiles or slices) always contain one or more point data, and the second division method is a division method in which the multiple division units include one or more division units that do not include point data, or there is a possibility that the multiple division units include one or more division units that do not include point data.
[0371] If the determined division method is the first division method (first division method in S5102), the three-dimensional data encoding device describes that the division method used is the first division method in the division additional information (e.g., tile additional information or slice additional information), which is metadata related to data division (S5103).Then, the three-dimensional data encoding device encodes all division units (S5104).
[0372] On the other hand, if the determined division method is the second division method (second division method in S5102), the three-dimensional data encoding device describes in the division additional information that the division method used is the second division method (S5105).Then, the three-dimensional data encoding device encodes the division units, excluding the division units that do not include point data (e.g., null tiles), among the multiple division units (S5106).
[0373] 57 is a flowchart of a three-dimensional data decoding process including a data splicing process by a three-dimensional data decoding device. First, the three-dimensional data decoding device refers to the partitioning additional information included in the bitstream, and determines whether the partitioning method used is the first partitioning method or the second partitioning method (S5111).
[0374] If the division method used is the first division method (first division method in S5112), the three-dimensional data decoding device receives the encoded data of all division units and generates decoded data of all division units by decoding the received encoded data (S5113). Next, the three-dimensional data decoding device reconstructs a three-dimensional point cloud using the decoded data of all division units (S5114). For example, the three-dimensional data decoding device reconstructs a three-dimensional point cloud by combining multiple division units.
[0375] On the other hand, if the division method used is the second division method (second division method in S5112), the three-dimensional data decoding device receives the coded data of the division units including point data and the coded data of the division units not including point data, and generates decoded data by decoding the coded data of the received division units (S5115). Note that, if a division unit not including point data is not sent, the three-dimensional data decoding device does not need to receive and decode a division unit not including point cloud data. Next, the three-dimensional data decoding device reconstructs a three-dimensional point cloud using the decoded data of the division units including point data (S5116). For example, the three-dimensional data decoding device reconstructs a three-dimensional point cloud by combining a plurality of division units.
[0376] Other methods of dividing point cloud data will be described below. When dividing a space evenly as shown in (c) of Fig. 45, there are cases where no points exist in the divided spaces. In this case, the three-dimensional data encoding device combines the space where no points exist with another space where points exist. In this way, the three-dimensional data encoding device can form multiple division units so that all division units include one or more points.
[0377] 58 is a flowchart of data division in this case. First, the three-dimensional data encoding device divides the data in a specific method (S5121). For example, the specific method is the second division method described above.
[0378] Next, the three-dimensional data encoding device determines whether or not a point is included in the target division unit, which is the division unit to be processed (S5122). If a point is included in the target division unit (Yes in S5122), the three-dimensional data encoding device encodes the target division unit (S5123). On the other hand, if a point is not included in the target division unit (No in S5122), the three-dimensional data encoding device combines the target division unit with other division units that include points, and encodes the combined division unit (S5124). In other words, the three-dimensional data encoding device encodes the target division unit together with other division units that include points.
[0379] Although the example of performing the determination and combination for each division unit has been described here, the processing method is not limited to this. For example, the three-dimensional data encoding device may determine whether or not each of the multiple division units includes a point, combine the multiple division units so that there is no division unit that does not include a point, and encode each of the multiple division units after the combination.
[0380] Next, a method for transmitting data including null tiles will be described. If a target tile to be processed is a null tile, the three-dimensional data encoding device does not transmit data of the target tile. Figure 59 is a flowchart of the data transmission process.
[0381] First, the three-dimensional data encoding device determines a tile division method, and divides the point cloud data into tiles using the determined division method (S5131).
[0382] Next, the three-dimensional data encoding device determines whether or not the target tile is a null tile (S5132), that is, whether or not there is no data in the target tile.
[0383] If the target tile is a null tile (Yes in S5132), the three-dimensional data encoding device indicates that the target tile is a null tile in the tile additional information, and does not indicate information about the target tile (such as the tile's position and size) (S5133).In addition, the three-dimensional data encoding device does not send out the target tile (S5134).
[0384] On the other hand, if the target tile is not a null tile (No in S5132), the three-dimensional data encoding device indicates in the tile additional information that the target tile is not a null tile and indicates information for each tile (S5135).The three-dimensional data encoding device also outputs the target tile (S5136).
[0385] In this way, by not including information about null tiles in the tile additional information, the amount of information in the tile additional information can be reduced.
[0386] A method for decoding coded data including null tiles will be described below. First, a process for when there is no packet loss will be described.
[0387] 60 is a diagram showing an example of transmission data, which is encoded data transmitted from a three-dimensional data encoding device, and reception data input to a three-dimensional data decoding device. Note that, here, a system environment without packet loss is assumed, and the reception data is the same as the transmission data.
[0388] In the case of a system environment without packet loss, the three-dimensional data decoding device receives all of the transmitted data. Fig. 61 is a flowchart of the process performed by the three-dimensional data decoding device.
[0389] First, the three-dimensional data decoding device refers to the tile additional information (S5141), and determines whether or not each tile is a null tile (S5142).
[0390] If the tile additional information indicates that the target tile is not a null tile (No in S5142), the 3D data decoding device determines that the target tile is not a null tile and decodes the target tile (S5143). Next, the 3D data decoding device obtains tile information (tile position information (origin coordinates, etc.) and size, etc.) from the tile additional information, and reconstructs 3D data by combining multiple tiles using the obtained information (S5144).
[0391] On the other hand, if the tile additional information indicates that the target tile is not a null tile (Yes in S5142), the 3D data decoding device determines that the target tile is a null tile and does not decode the target tile (S5145).
[0392] The three-dimensional data decoding device may determine that missing data is a null tile by sequentially analyzing index information indicated in the header of the encoded data. The three-dimensional data decoding device may also combine a determination method using tile additional information and a determination method using index information.
[0393] Next, the process when packet loss occurs will be described. Fig. 62 is a diagram showing an example of data sent from a three-dimensional data encoding device and data received by a three-dimensional data decoding device. Here, a system environment with packet loss is assumed.
[0394] In a system environment with packet loss, the three-dimensional data decoding device may not be able to receive all of the transmitted data. In this example, the packets Gt2 and At2 are lost.
[0395] 63 is a flowchart of the process of the three-dimensional data decoding device in this case. First, the three-dimensional data decoding device analyzes the continuity of the index information indicated in the header of the encoded data (S5151), and determines whether the index number of the target tile exists (S5152).
[0396] If the index number of the target tile exists (Yes in S5152), the 3D data decoding device determines that the target tile is not a null tile, and performs a decoding process for the target tile (S5153). Next, the 3D data decoding device obtains tile information (tile position information (origin coordinates, etc.) and size, etc.) from the tile additional information, and reconstructs 3D data by combining multiple tiles using the obtained information (S5154).
[0397] On the other hand, if index information for the target tile does not exist (No in S5152), the 3D data decoding device determines whether or not the target tile is a null tile by referring to the tile additional information (S5155).
[0398] If the target tile is not a null tile (No in S5156), the three-dimensional data decoding device determines that the target tile has been lost (packet loss) and performs error decoding processing (S5157). The error decoding processing is, for example, processing that attempts to decode the original data assuming that the data exists. In this case, the three-dimensional data decoding device may reproduce the three-dimensional data and perform reconstruction of the three-dimensional data (S5154).
[0399] On the other hand, if the target tile is a null tile (Yes in S5156), the 3D data decoding device assumes that the target tile is a null tile and does not perform the decoding process or reconstruct the 3D data (S5158).
[0400] Next, a coding method in which null tiles are not explicitly indicated will be described. The three-dimensional data coding device may generate coded data and additional information in the following manner.
[0401] The three-dimensional data encoder does not indicate information about null tiles in the tile additional information. The three-dimensional data encoder adds index numbers of tiles other than null tiles to the data header. The three-dimensional data encoder does not send null tiles.
[0402] In this case, the number of tile divisions (number_of_tiles) indicates the number of divisions excluding null tiles. The three-dimensional data encoding device may store information indicating the number of null tiles separately in the bit stream. The three-dimensional data encoding device may also indicate information about null tiles in the additional information, or may indicate part of the information about null tiles.
[0403] 64 is a flowchart of the three-dimensional data encoding process by the three-dimensional data encoding device in this case. First, the three-dimensional data encoding device determines a tile division method, and divides the point cloud data into tiles using the determined division method (S5161).
[0404] Next, the three-dimensional data encoding device determines whether the target tile is a null tile (S5162), that is, whether there is no data in the target tile.
[0405] If the target tile is not a null tile (No in S5162), the three-dimensional data encoding device adds index information of tiles other than null tiles to the data header (S5163).Then, the three-dimensional data encoding device outputs the target tile (S5164).
[0406] On the other hand, if the target tile is a null tile (Yes in S5162), the three-dimensional data encoding device does not add index information of the target tile to the data header, and does not transmit the target tile.
[0407] Fig. 65 is a diagram showing an example of index information (idx) added to a data header. As shown in Fig. 65, index information for null tiles is not added, and consecutive numbers are added to tiles other than null tiles.
[0408] Fig. 66 is a diagram showing an example of the dependency of each data. The tip of the arrow in the diagram indicates the dependency destination, and the base of the arrow indicates the dependency source. Also, in the diagram, Gtn (n is 1 to 4) indicates the position information of the tile number n, Atn indicates the attribute information of the tile number n, and Mtile indicates the tile additional information.
[0409] FIG. 67 is a diagram showing an example of the structure of transmission data, which is encoded data transmitted from the three-dimensional data encoding device.
[0410] The following describes a decoding method when null tiles are not explicitly indicated. Figure 68 is a diagram showing an example of data sent from a three-dimensional data encoding device and data received by a three-dimensional data decoding device. Here, a system environment with packet loss is assumed.
[0411] 69 is a flowchart of the process of the three-dimensional data decoding device in this case. First, the three-dimensional data decoding device analyzes the tile index information indicated in the header of the encoded data, and determines whether the index number of the target tile exists. In addition, the three-dimensional data decoding device obtains the number of divisions of the tile from the tile additional information (S5171).
[0412] If the index number of the target tile exists (Yes in S5172), the 3D data decoding device performs a decoding process for the target tile (S5173). Next, the 3D data decoding device obtains tile information (tile position information (origin coordinates, etc.) and size, etc.) from the tile additional information, and reconstructs 3D data by combining multiple tiles using the obtained information (S5175).
[0413] On the other hand, if the index number of the target tile does not exist (No in S5172), the three-dimensional data decoding device determines that the target tile is a packet loss, and performs error decoding processing (S5174). In addition, the three-dimensional data decoding device determines that a space that does not exist in the data is a null tile, and reconstructs the three-dimensional data.
[0414] In addition, by explicitly indicating null tiles, the three-dimensional data encoding device can properly determine that no points exist within a tile, and that this is not due to measurement error or data loss due to data processing, etc., or packet loss.
[0415] The three-dimensional data encoding device may use both a method of explicitly indicating null packets and a method of not explicitly indicating null packets. In this case, the three-dimensional data encoding device may indicate information indicating whether or not null packets are explicitly indicated in the tile additional information. Also, whether or not null packets are explicitly indicated may be determined in advance according to the type of division method, and the three-dimensional data encoding device may indicate whether or not null packets are explicitly indicated by indicating the type of division method.
[0416] In addition, in Figure 47 etc., an example is shown in which the tile additional information indicates information related to all tiles, but the tile additional information may indicate information on some of the multiple tiles, or may indicate null tile information on some of the multiple tiles.
[0417] Also, an example has been described in which information about split data, such as information about whether or not split data (tiles) exists, is stored in the tile additional information, but some or all of this information may be stored in a parameter set or stored as data. When this information is stored as data, for example, nal_unit_type, which means information indicating whether or not split data exists, may be defined, and this information may be stored in the NAL unit. Also, this information may be stored in both the additional information and the data.
[0418] (Embodiment 6) The process of performing quantization for each tile will be described below.
[0419] Fig. 70 is a diagram showing an example of GPS syntax. As shown in Fig. 70, the GPS includes an inter-tile overlap point flag (UniqueBetweenTilesFlag). The inter-tile overlap point flag is a flag indicating whether or not there is a possibility that an overlap point exists between tiles.
[0420] 71 is a flowchart of the three-dimensional data decoding process. First, the three-dimensional data decoding device decodes UniqueBetweenTilesFlag and MergeDuplicatedPointFlag from metadata included in the bitstream (S6261). Next, the three-dimensional data decoding device decodes position information and attribute information for each tile, and reconstructs a point group (S6262).
[0421] Next, the three-dimensional data decoding device determines whether or not merging of overlapping points is necessary (S6263). For example, the three-dimensional data decoding device determines whether or not merging is necessary depending on whether or not an application can handle overlapping points, or whether or not it is better to merge overlapping points. Alternatively, the three-dimensional data decoding device may determine to merge overlapping points for the purpose of smoothing or filtering a plurality of pieces of attribute information corresponding to overlapping points, thereby removing noise or improving estimation accuracy.
[0422] If merging of overlapping points is necessary (Yes in S6263), the three-dimensional data decoding device determines whether there is overlap between tiles (existence of overlapping points) (S6264). For example, the three-dimensional data decoding device may determine whether there is overlap between tiles based on the decoding results of UniqueBetweenTilesFlag and MergeDuplicatedPointFlag. This eliminates the need to search for overlapping points in the three-dimensional data decoding device, thereby reducing the processing load of the three-dimensional data decoding device. Note that the three-dimensional data decoding device may determine whether overlapping points exist by searching for overlapping points after reconstructing tiles.
[0423] If there is overlap between tiles (Yes in S6264), the three-dimensional data decoding device merges the overlapping points between the tiles (S6265). Next, the three-dimensional data decoding device merges multiple pieces of overlapping attribute information (S6266).
[0424] After step S6266, or if there is no overlap between tiles (No in S6264), the three-dimensional data decoding device executes the application using a point group without overlapping points (S6267).
[0425] On the other hand, if merging of overlapping points is not necessary (No in S6263), the three-dimensional data decoding device does not merge overlapping points, and executes the application using the point group in which overlapping points exist (S6268).
[0426] Examples of applications will be described below. First, an example of an application that uses a point group with no overlapping points will be described.
[0427] Fig. 72 is a diagram showing an example of an application. The example shown in Fig. 72 shows a use case in which a mobile object traveling from an area of tile A to an area of tile B downloads a map point cloud from a server in real time. The server stores encoded data of map point clouds of multiple overlapping areas. The mobile object has already acquired map information of tile A, and requests the server to acquire map information of tile B, which is located in the direction of movement.
[0428] At that time, the mobile entity determines that the data in the overlapping portion between tile A and tile B is unnecessary, and transmits an instruction to the server to delete the overlapping portion between tile B and tile A contained in tile B. The server deletes the overlapping portion from tile B, and distributes tile B after the deletion to the mobile entity. This makes it possible to reduce the amount of transmitted data and the load of the decoding process.
[0429] The mobile entity may check that there are no overlapping points based on the flag. If the mobile entity has not yet acquired tile A, it requests data from the server that does not delete overlapping points. If the server does not have a function for deleting overlapping points or if it is not clear whether there are overlapping points, the mobile entity may check the distributed data to determine whether there are overlapping points, and may merge the data if there are overlapping points.
[0430] Next, an example of an application using a point cloud with overlapping points will be described. A mobile object uploads map point cloud data acquired by LiDAR to a server in real time. For example, the mobile object uploads data acquired for each tile to the server. In this case, there is an area where tile A and tile B overlap, but the encoding mobile object does not merge the overlapping points between tiles, and sends the data to the server together with a flag indicating that there is an overlap between tiles. The server accumulates the received data as it is, without merging the overlapping data contained in the received data.
[0431] Furthermore, when transmitting or storing point cloud data using a system such as ISOBMFF, MPEG-DASH / MMT, or MPEG-TS, the device may replace a flag included in GPS, indicating whether or not there are overlapping points in a tile or whether or not there are overlapping points between tiles, with a descriptor or metadata in the system layer, and store it in an SI, MPD, moov, or moof box, etc. This allows applications to utilize the functions of the system.
[0432] Furthermore, the three-dimensional data encoding device may divide, for example, tile B into a plurality of slices based on an overlapping area with other tiles, as shown in Fig. 73. In the example shown in Fig. 73, slice 1 is an area that does not overlap with any tile, slice 2 is an area that overlaps with tile A, and slice 3 is an area that overlaps with tile C. This makes it easier to separate desired data from the encoded data.
[0433] The map information may be point cloud data or mesh data. The point cloud data may be tiled for each area and stored in a server.
[0434] 74 is a flowchart showing the flow of processing in the above system. First, a terminal (e.g., a mobile object) detects the movement of the terminal from area A to area B (S6271). Next, the terminal starts acquiring map information of area B (S6272).
[0435] If the terminal has already downloaded the information on area A (Yes in S6273), the terminal instructs the server to obtain data on area B that does not include overlapping points with area A (S6274). The server deletes area A from area B, and transmits the data on area B after deletion to the terminal (S6275). Note that the server may, in response to an instruction from the terminal, encode the data on area B in real time to prevent overlapping points, and transmit it.
[0436] Next, the terminal merges (combines) the map information of area B with the map information of area A, and displays the merged map information (S6276).
[0437] On the other hand, if the terminal has not yet downloaded information about area A (No in S6273), the terminal instructs the server to acquire data about area B, which includes overlapping points with area A (S6277). The server transmits the data about area B to the terminal (S6278). Next, the terminal displays map information about area B, which includes overlapping points with area A (S6279).
[0438] 75 is a flowchart showing another example of operation of the system. The transmitting device (three-dimensional data encoding device) transmits tile data in order (S6281). The transmitting device also adds a flag to the data of the tile to be transmitted, indicating whether the tile to be transmitted overlaps with a tile of the data transmitted immediately before, and transmits the data (S6282).
[0439] The receiving device (three-dimensional data decoding device) determines whether or not the tile of the received data overlaps with a tile of previously received data based on the flag added to the data (S6283). If the tile of the received data overlaps with a tile of previously received data (Yes in S6283), the receiving device deletes or merges the overlapping points (S6284). On the other hand, if the tile of the received data does not overlap with a tile of previously received data (No in S6283), the receiving device does not perform the process of deleting or merging the overlapping points and ends the process. This makes it possible to reduce the processing load of the receiving device and improve the estimation accuracy of attribute information. Note that the receiving device does not need to merge the overlapping points if it is not necessary.
[0440] (Embodiment 7) In this embodiment, a viewpoint-based display method, a random access method for encoded data, a method for encoding point cloud data, and a method for decoding point cloud data will be described.
[0441] With the improvement of sensor performance, it has become possible to obtain high-quality three-dimensional point clouds. However, in order to view high-quality three-dimensional points, a viewing device (viewer) capable of reproducing the high-quality three-dimensional point cloud is required. Specifically, it is desired to be able to display high-quality three-dimensional point clouds (point clouds) with a large amount of data without delay. In this embodiment, a three-dimensional point cloud viewing device (first application) capable of efficiently displaying high-density point cloud data by a scalable method using point cloud compression is described.
[0442] Point cloud compression is performed using multiple data partitioning methods. For example, using Levels of Details (LoD), the resolution required to represent the point cloud data is calculated according to the distance between the virtual camera and the point cloud data. This allows separation or hierarchical organization to be achieved.
[0443] A 3D point cloud viewing device (also referred to as a 3D data decoding device) selects the visible point clouds for rendering, and preferably verifies that all visible point clouds are actual scanned data and not approximations.
[0444] 76 is a block diagram showing an example of the configuration of a three-dimensional data encoding device. The three-dimensional data encoding device includes a point cloud encoding unit 8701 and a file format generating unit 8702. The point cloud encoding unit 8701 generates encoded data (bit stream) by encoding point cloud data. For example, the point cloud encoding unit 8701 encodes the point cloud data using a position information-based encoding method using an octree, a video-based encoding method, or the like.
[0445] The file format generating unit 8702 changes the encoded data (bit stream) to data in a predetermined file format. For example, the file format is ISOBMFF or MP4. The three-dimensional data encoding device may output encoded data in a file format (for example, transmit it to a three-dimensional data decoding device), or may output encoded data in a bit stream format in the encoding method.
[0446] 77 is a block diagram showing an example of the configuration of a three-dimensional data decoding device 8705. The three-dimensional data decoding device 8705 generates point cloud data by decoding encoded data. Here, the encoded data is, for example, encoded data in a bit stream format or an MP4 format. Note that unencoded point cloud data may be used.
[0447] All or part of the data in a point cloud is called a brick. This brick may also be called a divided data, tile, or slice. The divided data may be further divided.
[0448] The three-dimensional data decoding device 8705 externally acquires camera viewpoint information indicating the viewpoint (angle) of the camera. The three-dimensional data decoding device 8705 acquires a part or all of the encoded data based on the camera viewpoint information, and generates point cloud data by decoding the acquired encoded data. For example, the camera viewpoint information indicates the position and direction (orientation) of the camera. The three-dimensional data decoding device 8705 then displays the decoded point cloud data.
[0449] The three-dimensional data decoding device 8705 includes a point cloud decoding unit 8706 and a brick decoding control unit 8707. Camera viewpoint information (camera viewing angle) is input to the brick decoding control unit 8707. The brick decoding control unit 8707 selects a brick to be decoded based on the visibility of the brick determined based on the camera viewpoint information. The point cloud decoding unit 8706 decodes the selected brick and outputs the decoded brick.
[0450] The configuration of the three-dimensional data encoding device according to this embodiment will be described below. Fig. 78 is a block diagram showing the configuration of a three-dimensional data encoding device 8710 according to this embodiment. The three-dimensional data encoding device 8710 generates encoded data (encoded stream) by encoding point group data (point cloud). This three-dimensional data encoding device 8710 includes a division unit 8711, a plurality of position information encoding units 8712, a plurality of attribute information encoding units 8713, an additional information encoding unit 8714, a multiplexing unit 8715, and a normal vector generating unit 8716.
[0451] The dividing unit 8711 generates a plurality of pieces of divided data by dividing the point cloud data. Specifically, the dividing unit 8711 generates a plurality of pieces of divided data by dividing the space of the point cloud data into a plurality of subspaces. Here, the subspace is any of bricks, tiles, and slices, or a combination of two or more of bricks, tiles, and slices. More specifically, the point cloud data includes position information, attribute information (color, reflectance, etc.), and additional information. The dividing unit 8711 generates a plurality of pieces of divided position information by dividing the position information, and generates a plurality of pieces of divided attribute information by dividing the attribute information. In addition, the dividing unit 8711 generates additional information related to the division.
[0452] The position information encoding units 8712 generate multiple pieces of encoded position information by encoding the multiple pieces of divided position information. For example, the position information encoding unit 8712 encodes the divided position information using an N-ary tree structure such as an octet tree. Specifically, in the octet tree, the target space is divided into eight nodes (subspaces), and 8-bit information (occupancy code) indicating whether or not a point group is included in each node is generated. In addition, the node including the point group is further divided into eight nodes, and 8-bit information indicating whether or not a point group is included in each of the eight nodes is generated. This process is repeated until the number of point groups included in a predetermined hierarchy or node is equal to or less than a threshold value. For example, the position information encoding units 8712 process the multiple pieces of divided position information in parallel.
[0453] The attribute information encoding unit 8713 generates encoded attribute information, which is encoded data, by encoding attribute information using the configuration information generated by the position information encoding unit 8712. For example, the attribute information encoding unit 8713 determines a reference point (reference node) to be referenced in encoding a target point (target node) to be processed, based on the octree structure generated by the position information encoding unit 8712. For example, the attribute information encoding unit 8713 refers to a node, among peripheral nodes or adjacent nodes, whose parent node in the octree is the same as that of the target node. Note that the method of determining the reference relationship is not limited to this.
[0454] Furthermore, the encoding process of the position information or attribute information may include at least one of a quantization process, a prediction process, and an arithmetic coding process. In this case, the reference means using a reference node to calculate a predicted value of the attribute information, or using a state of the reference node (e.g., occupancy information indicating whether or not a point group is included in the reference node) to determine an encoding parameter. For example, the encoding parameter is a quantization parameter in a quantization process, or a context in an arithmetic coding process.
[0455] The normal vector generation unit 8716 calculates a normal vector for each divided data. Note that the input data does not necessarily have to be divided. In this case, the normal vector generation unit 8716 may calculate a normal vector for each point, instead of a normal vector for each divided data. Alternatively, the normal vector generation unit 8716 may calculate both a normal vector for each divided data and a normal vector for each point.
[0456] The additional information encoding unit 8714 generates encoded additional information by encoding the additional information contained in the point cloud data, the additional information regarding data division generated by the division unit 8711 at the time of division, and the normal vector generated by the normal vector generation unit 8716.
[0457] The multiplexing unit 8715 multiplexes a plurality of pieces of encoding position information, a plurality of pieces of encoding attribute information, and encoding additional information to generate encoded data (encoded stream), and transmits the generated encoded data. In addition, the encoded additional information is used during decoding.
[0458] The configuration of the three-dimensional data decoding device according to this embodiment will be described below. Fig. 79 is a block diagram showing the configuration of a three-dimensional data decoding device 8720. The three-dimensional data decoding device 8720 restores point cloud data by decoding coded data (coded stream) generated by coding point cloud data. This three-dimensional data decoding device 8720 includes a demultiplexing unit 8721, a plurality of position information decoding units 8722, a plurality of attribute information decoding units 8723, an additional information decoding unit 8724, a combining unit 8725, a normal vector extraction unit 8726, a random access control unit 8727, and a selection unit 8728.
[0459] The demultiplexer 8721 demultiplexes the encoded data (encoded stream) to generate a plurality of pieces of encoded position information, a plurality of pieces of encoded attribute information, and encoded additional information. The additional information decoder 8724 decodes the encoded additional information to generate additional information.
[0460] The normal vector extraction unit 8726 extracts a normal vector from the additional information. The random access control unit 8727 determines the divided data to extract based on, for example, the normal vector for each divided data. The selection unit 8728 extracts the multiple divided data (multiple pieces of encoding position information and multiple pieces of encoding attribute information) determined by the random access control unit 8727 from the multiple pieces of divided data (multiple pieces of encoding position information and multiple pieces of encoding attribute information). Note that the selection unit 8728 may extract one divided data.
[0461] The multiple position information decoding units 8722 generate multiple pieces of divided position information by decoding the multiple pieces of encoded position information extracted by the selection unit 8728. For example, the multiple position information decoding units 8722 process the multiple pieces of encoded position information in parallel.
[0462] The multiple attribute information decoding units 8723 generate multiple pieces of divided attribute information by decoding the multiple pieces of encoded attribute information extracted by the selection unit 8728. For example, the multiple attribute information decoding units 8723 process the multiple pieces of encoded attribute information in parallel.
[0463] The combining unit 8725 generates position information by combining a plurality of pieces of divided position information using the additional information The combining unit 8725 generates attribute information by combining a plurality of pieces of divided attribute information using the additional information.
[0464] Next, a first example of generating and encoding a normal vector for each point will be described. FIG. 80 is a diagram showing an example of point cloud data. FIG. 81 is a diagram showing an example of a normal vector for each point. The normal vector can be encoded independently for each three-dimensional group. FIG. 80 and FIG. 81 show a three-dimensional point cloud of a book and a normal vector of the three-dimensional point cloud. As shown in FIG. 81, there are multiple normal vectors extending in the upward, right, and forward directions. Here, the surface of the book is a plane, and multiple normal vectors of a certain surface extend in the same direction. On the other hand, when the surface is round, the normal vector extends in multiple directions according to the normal of the surface.
[0465] Figure 82 is a diagram showing an example of the syntax of a normal vector in a bit stream. In the normal vector NormalVector[i][face] shown in Figure 82, "i" represents a counter of each 3D point group, and [face] represents the x, y, and z axes representing the 3D point group. In other words, NormalVector represents the magnitude of the normal vector of each axis.
[0466] 83 is a flowchart of a three-dimensional data encoding process. First, the three-dimensional data encoding device encodes position information (geometry) and attribute information for each point (S8701). For example, the three-dimensional data encoding device encodes position information for each point. Furthermore, if attribute information corresponding to a point exists, the three-dimensional data encoding device may encode the attribute information for each point.
[0467] Next, the three-dimensional data encoding device encodes the normal vector (x, y, z) for each point (S8702). The three-dimensional data encoding device may encode the normal vector for each point. The three-dimensional data encoding device may also encode, for example, difference information indicating a difference between the normal vector of the point to be processed and the normal vector of another point. This can reduce the amount of data. The three-dimensional data encoding device may also encode the normal vector by including it in the position information, or may also encode it by including it in the attribute information. The three-dimensional data encoding device may also encode the normal vector independently of the position information and the attribute information. Note that, when multiple normal vectors exist for one point, the three-dimensional data encoding device may encode multiple normal vectors for each point.
[0468] 84 is a flowchart of a three-dimensional data decoding process. First, the three-dimensional data decoding device decodes position information and attribute information for each point from the bit stream (S8706). Next, the three-dimensional data decoding device decodes normal vectors for each point from the bit stream (S8707).
[0469] Note that the processing order shown in Figures 83 and 84 is just an example, and the encoding order and decoding order may be interchanged.
[0470] The three-dimensional data encoding device may also reduce the amount of data by encoding the normal vector using the position information or the correlation of the position information. In this case, the three-dimensional data decoding device decodes the normal vector using the position information. By using the above method, the normal vector for each point in the point cloud can be encoded and decoded.
[0471] Next, a second example of generating and encoding a normal vector for each point will be described. As another method of encoding the normal vector for each point, the normal vector is encoded as one of the attribute information. Below, an example of encoding using an attribute information encoding unit or an attribute information decoding unit as one of the attribute information will be described.
[0472] For example, the three-dimensional data encoding device encodes color information as the first attribute information and a normal vector as the second attribute information. FIG. 85 is a diagram showing an example of a bit stream configuration. For example, Attr(0) shown in FIG. 85 is encoded data of the first attribute information, and Attr(1) is encoded data of the second attribute information. Furthermore, metadata related to encoding is stored in a parameter set (APS). The three-dimensional data decoding device decodes the encoded data by referring to the APS corresponding to the encoded data.
[0473] In addition, the SPS stores identification information (attribute_type=Normal Vector) indicating that the second attribute information is a normal vector. If the attribute information is a normal vector, information indicating that the normal vector is data having three elements for each point may be stored in the SPS, etc. In addition, the SPS stores identification information (attribute_type=Color) indicating that the first attribute information is color information.
[0474] 86 is a diagram showing an example of point group information having position information, color information, and normal vectors. The three-dimensional data encoding device encodes the uncompressed point group data shown in FIG.
[0475] The range of values of the normal vector is a floating-point value from -1 to 1. To facilitate the representation, the three-dimensional data encoding device may convert the floating point to an integer according to the required precision. For example, the three-dimensional data encoding device may convert the floating point to a value from -127 to 128 using an 8-bit representation. That is, the three-dimensional data encoding device may convert the floating point to an integer or a positive integer value. Since the normal vector is treated as one attribute information, different quantization processes can be applied. For example, a different quantization parameter can be used for each attribute information. This allows different precision levels to be realized. For example, the quantization parameter is stored in the APS.
[0476] 87 is a flowchart of a three-dimensional data encoding process. First, the three-dimensional data encoding device encodes position information and attribute information (color information, etc.) for each point (S8711). In addition, the three-dimensional data encoding device encodes the normal vector for each point as attribute information of attribute_type="normal vector" using a predetermined method (S8712).
[0477] 88 is a flowchart of the three-dimensional data decoding process. The three-dimensional data decoding device decodes position information and attribute information for each point from the bit stream (S8716). In addition, the three-dimensional data decoding device decodes normal vectors for each point from the bit stream as attribute information of attribute_type="normal vector" using a predetermined method (S8717).
[0478] Note that the processing order shown in Figures 87 and 88 is just an example, and the encoding order and decoding order may be interchanged.
[0479] Next, an example of generating a normal vector for each data unit including a plurality of points will be described. The three-dimensional data encoding device divides the point cloud data into a plurality of objects or a plurality of regions based on the position information and characteristics of the point cloud. The divided data is, for example, tiles or slices, or hierarchical data. The three-dimensional data encoding device generates a normal vector for this divided data unit, that is, for each data unit including one or more points.
[0480] Here, visibility can be determined by the normal vector representation of the object within the brick. Figures 89 and 90 are diagrams for explaining this process. For example, as shown in Figure 89, the three-dimensional data encoding device divides the normal vector direction into angles spaced at 30° intervals with respect to the horizontal and vertical axes. As a simpler method, as shown in Figure 90, the three-dimensional data encoding device may divide the normal vector into six directions: (0, 0), (0, 90), (0, -90), (90, 0), (-90, 0), and (180, 180).
[0481] The three-dimensional data encoding device may also calculate the effective normal vector using a median, an average, or another more effective algorithm. The three-dimensional data encoding device may also use a representative value as the effective normal vector value, or may use another method.
[0482] In addition, the normal vector for each divided data may show the original x, y, and z values as they are, or may be quantized every 30 degrees as described above, or may be quantized to information every 90 degrees. Quantization can reduce the amount of information.
[0483] Fig. 91 is an example of point cloud data, showing an example of a face object. Fig. 92 is a diagram showing an example of normal vectors in this case. As shown in Fig. 92, the normal vectors of the face object shown in Fig. 91 are oriented in the (0, 0) and (90, 0) directions. The three-dimensional data encoding device can use one bit for each direction to indicate whether or not the normal vector of the object is in that direction.
[0484] In this way, there may be two or more normal vectors for one divided data unit. In this case, multiple normal vectors may be indicated for one divided data unit.
[0485] For example, the data example including a face object shown in Figures 91 and 92 is an example in which the normal vectors of the data are expressed by six different normal vectors for each face in units of 90 degrees. In this example, the two normal vectors in the directions of (0, 0) and (90, 0) are the normal vectors of this divided data.
[0486] As a method of indicating normal vectors, each of the six normal vectors may be represented by 1-bit information. FIG. 93 is a diagram showing an example of this normal vector information. If the divided data has the corresponding normal vector, the 1-bit information is set to a value of 1, and if not, it is set to 0. This allows the amount of information to be reduced by quantizing the data, compared to a method of indicating the values of x, y, and z as they are.
[0487] A simpler representation of normal vectors is described below. A six-sided cube is used to represent normal vectors and their feasibility (visibility) from a particular camera viewpoint. Figures 94 to 97 are diagrams for explaining this process. Figure 94 shows an example of a six-sided cube. Figures 95, 96, and 97 are diagrams showing front and back faces a and b, left and right faces c and d, and top and bottom faces e and f, respectively. Depending on the object's orientation according to the viewing angle, the normal vector faces at least one or three faces. Six flags, each of 1 bit, can be used to represent one of the six faces (abcdef) of a cube representing each system. For example, when viewed from the front, (100000), when viewed from the side, (001000), and when viewed from below, (000001) are generated. In this representation, the size is not important, and only the direction is represented. It is also possible for an object to occur for which three faces are specified. Face a is the opposite face of face b, face c is the opposite face of face d, and face e is the opposite face of face f. Therefore, it is impossible to see faces a and b at the same time. In other words, the normal vector can be expressed using three flags (ace).
[0488] In this way, when the camera viewpoint (camera angle) is known in advance, the normal vector information can be expressed in 3 bits. FIG. 98 is a diagram showing the visibility when an object in slice A or slice B is viewed from the direction of face c. Since slice A is visible from the direction of face c, it is expressed as ace=(010). On the other hand, since slice B is hidden by slice A when viewed from the direction of face c, it is expressed as ace=(000).
[0489] Next, a first method of encoding and decoding a normal vector for each brick will be described. FIG. 99 is a diagram showing an example of the configuration of a bit stream in this case. In the example shown in FIG. 99, information on the normal vector is stored in the slice header of the position information in each slice. Note that the information on the normal vector may be stored in the header of the attribute information, or may be stored in metadata independent of the position information and the attribute information.
[0490] FIG. 100 is a diagram showing an example of the syntax of a geometry slice header information of position information. The geometry slice header information of position information includes normal_vector_number, normal_vector_x, normal_vector_y, and normal_vector_z.
[0491] normal_vector_number indicates the number of normal vectors corresponding to the slice data. normal_vector_x, normal_vector_y, and normal_vector_z respectively indicate the elements (x, y, z) of the normal vector corresponding to the slice data.
[0492] In this example, the number of normal_vectors can be changed. The number of normal_vectors shown is equal to normal_vector_number.
[0493] Note that when the normal vector information is common for all slices, normal_vector_number may be stored in GPS or SPS that can store common information for multiple slices.
[0494] Also, the values of the normal vectors of x, y, and z may be quantized. For example, the three-dimensional data encoding device may quantize the values of the original normal vectors by shifting them by a common bit amount s (bit), and send out the information indicating the bit amount s and the information indicating the quantized normal vectors (normal_vector_x << s, normal_vector_y << s, normal_vector_z << s). This can reduce the bit amount.
[0495] FIG. 101 is a diagram showing another example of the syntax of the geometry slice header information of position information. This example shows the simplified (quantized) normal vectors for the six-sided data for each divided data. For each face, it is indicated whether there is a normal vector or not.
[0496] The slice header of this position information includes is_normal_vector. If there is a normal vector corresponding to the slice data, is_normal_vector is set to 1, and if there is no normal vector, is set to 0. For example, the order of multiple faces is predetermined.
[0497] It should be noted that the quantization accuracy and the number or order of normal vectors are not limited to those described above, and may be fixed or variable.
[0498] 102 is a flowchart of a three-dimensional data encoding process. First, the three-dimensional data encoding device generates a plurality of divided data by dividing point cloud data (S8721). Next, the three-dimensional data encoding device encodes position information and attribute information for each divided data (S8722). Next, the three-dimensional data encoding device stores a normal vector for each divided data in a slice header (S8723).
[0499] 103 is a flowchart of a three-dimensional data decoding process. First, the three-dimensional data decoding device decodes position information and attribute information for each divided data from the bit stream (S8726). Next, the three-dimensional data decoding device decodes a normal vector for each divided data from the slice header for each divided data (S8727). Next, the three-dimensional data decoding device combines multiple divided data (S8728).
[0500] 104 is a flowchart of a three-dimensional data decoding process in the case of partially decoding data. First, the three-dimensional data decoding device decodes a normal vector for each divided data from the slice header for each divided data (S8731). Next, the three-dimensional data decoding device determines the divided data to be decoded based on the normal vector, and decodes the determined divided data (S8732). Next, the decoded divided data are combined (S8733).
[0501] Next, a second method of encoding and decoding normal vectors for each brick will be described. Another method of encoding information of normal vectors is to use metadata (e.g., SEI: Supplemental Enhancement Information). Figure 105 is a diagram showing an example of a bitstream configuration. As shown in Figure 105, the SEI may be included in the bitstream, or may be generated as a separate file apart from the main encoded bitstream, depending on how the SEI is implemented in both the encoding device and the decoding device.
[0502] Fig. 106 is a diagram illustrating an example of the syntax of slice information (slice_information) included in SEI. The slice information includes number_of_slice, bounding_box_origin_x, bounding_box_origin_y, bounding_box_origin_z, bounding_box_width, bounding_box_height, bounding_box_depth, normalVector_QP, number_of_normal_vector, normalVector_x, normalVector_y, and normalVector_z.
[0503] number_of_slice indicates the number of divided data. bounding_box_origin_x, bounding_box_origin_y, and bounding_box_origin_z indicate the origin coordinates of the bounding box of the slice data. bounding_box_width, bounding_box_height, and bounding_box_depth indicate the width, height, and depth of the bounding box of the slice data, respectively.
[0504] normalVector_QP indicates quantization scale information or bit shift information when normal_vector is quantized. number_of_normal_vector indicates the number of normal vectors included in the slice data. normalVector_x, normalVector_y, and normalVector_z indicate the components of the normal vector elements (x, y, z), respectively.
[0505] Fig. 107 is a diagram showing another example of slice information included in SEI. The example shown in Fig. 107 is an example showing normal vectors simplified (quantized) into data for six faces for each divided data. Whether or not there is a normal vector is shown for each face.
[0506] This slice information includes is_normal_vector. When there is a normal vector corresponding to the slice data, is_normal_vector is set to 1, and when there is no normal vector, is set to 0. For example, the order of multiple faces is predetermined.
[0507] The slice information may include a flag indicating whether or not the slice information includes information on the bounding box (origin and width, height, and depth) for each slice. In this case, when the flag is on (e.g., 1), the slice information includes information on the bounding box for each slice, and when the flag is off (e.g., 0), the slice information does not include information on the bounding box for each slice. The slice information may also include a flag indicating whether or not the slice information includes information on the normal vector for each slice. In this case, when the flag is on (e.g., 1), the slice information includes information on the normal vector for each slice, and when the flag is off (e.g., 0), the slice information does not include information on the normal vector for each slice.
[0508] Next, random access and partial decoding will be described. A three-dimensional data decoding device independently decodes data for each slice using information for each slice, for example, either or both of bounding box information and normal vectors of the slice.
[0509] 108 is a flowchart of a three-dimensional data decoding process. First, the three-dimensional data decoding device determines slices to be decoded and the decoding order of the slices by a predetermined method (S8741). Next, the three-dimensional data decoding device decodes specific slices in the determined order (S8742).
[0510] Figure 109 is a diagram showing an example of this partial decoding process. For example, the three-dimensional data decoding device receives encoded data divided into slices as shown in (a) of Figure 109. As shown in (b) of Figure 109, the three-dimensional data decoding device decodes the encoded data of some slices and does not decode the encoded data of the other slices. Alternatively, as shown in (c) of Figure 109, the three-dimensional data decoding device performs decoding by changing the order of the encoded data.
[0511] Fig. 110 is a diagram showing a configuration example of a three-dimensional data decoding device. As shown in Fig. 110, the three-dimensional data decoding device includes an attribute information decoding unit 8731 and a random access control unit 8732. The attribute information decoding unit 8731 extracts bounding box information and normal vectors for each slice from the encoded data. The random access control unit 8732 determines the number and order of the slices to be decoded based on the bounding box information and normal vectors for each slice and sensor information acquired from outside, such as the camera angle (camera direction) and camera position.
[0512] 111 and 112 are diagrams showing an example of processing by the random access control unit 8732. As shown in FIG. 111, for example, the random access control unit 8732 may calculate distance information indicating the distance from the camera for each slice from the bounding box for each slice and the camera position. Alternatively, as shown in FIG. 112, the random access control unit 8732 may derive visibility information indicating whether an object is visible from the camera for each slice from the normal vector for each slice and the camera angle. Note that the random access control unit 8732 may calculate either the distance information or the visibility information, or may calculate both.
[0513] The visibility information and the distance information will be described below. FIG. 113 is a diagram showing an example of the relationship between distance and resolution. For example, what is visible from the camera is decoded (frustum culling). The resolution at which the information is decoded further depends on the distance between the virtual camera and the point cloud data.
[0514] That is, the three-dimensional data decoding device determines whether or not a slice is visible from the camera based on the normal vector and the camera viewpoint (camera angle) of each slice, and decodes the slices that are visible from the camera. Furthermore, the three-dimensional data decoding device may calculate the distance from the camera of the slice to be decoded, and if the distance from the camera is close, decode high-resolution data, and if the distance from the camera is far, decode low-resolution data.
[0515] In this case, the coded data is coded in a hierarchical manner, and the three-dimensional data decoding device can independently decode the low-resolution data. When decoding high-resolution data, the three-dimensional data decoding device further decodes difference information between the low-resolution data and the high-resolution data, and generates high-resolution data by adding the difference information to the low-resolution data. When the coded data is not coded in a hierarchical manner, the three-dimensional data decoding device may not perform this process, or may determine whether or not to perform this process depending on whether the data is coded in a hierarchical manner.
[0516] Next, the determination of visibility using normal vectors will be described. Figure 114 is a diagram showing an example of bricks and normal vectors. In the example shown in Figure 114, two bricks (e.g., slices) on the front side facing the camera (view frustum), that is, bricks whose normal vectors point toward the camera, are decoded.
[0517] First, the three-dimensional data decoding device determines whether or not one or more normal vectors included in the metadata for each slice data have a normal vector opposite to the camera direction. If the slice data of the target slice has a normal vector opposite to the camera direction, the three-dimensional data decoding device determines that the target slice is visible, and determines that the target slice is to be decoded.
[0518] In addition, the three-dimensional data decoding device may determine that the target slice is invisible (cannot be seen) when another slice exists between the camera and the target slice. Furthermore, the three-dimensional data decoding device may determine whether the target slice is visible or not by determining whether the relationship between the normal vector and the camera direction is within a predetermined angle range, rather than determining whether the normal vector and the camera direction are completely opposite to each other.
[0519] Next, processing using Level of Detail (LoD) will be described. An example of decoding processing according to layers with different resolutions will be described below.
[0520] FIG. 115 is a diagram showing an example of levels (LoD). FIG. 116 is a diagram showing an example of an octree structure. Each brick is divided into layers to control the level of resolution to be decoded. For example, a level is the depth of division when dividing into an octree. As shown in FIG. 115, the number of voxels included in each level may be specified as 2 (3×level). Note that a different definition may be used for the division method or the number of voxels according to the level.
[0521] By using LoD, the 3D data decoder can realize high-speed visibility judgment and distance calculation. The decoding time affects real-time rendering. By using LoD, intermediate bricks can be displayed, real-time rendering and smooth response can be realized.
[0522] FIG. 117 is a flowchart of a three-dimensional data decoding process using LoD. First, the three-dimensional data decoding device determines the level to be decoded according to the purpose (S8751). Next, the three-dimensional data decoding device decodes the first level (level 0) (S8752). Next, the three-dimensional data decoding device determines whether or not the decoding of all levels to be decoded is completed (S8753). If the decoding of all levels is not completed (No in S8753), the three-dimensional data decoding device decodes the next level (S8754). At this time, the three-dimensional data decoding device may decode the next level using data of the previous level. If the decoding of all levels to be decoded is completed (Yes in S8753), the three-dimensional data decoding device displays the decoded data (S8755).
[0523] In this way, the three-dimensional data decoding device decodes data up to the determined level, and does not decode data after the determined level. This reduces the amount of processing involved in decoding, and improves processing speed. Furthermore, the three-dimensional data decoding device displays data up to the determined level, and does not display data after the determined level. This reduces the amount of processing involved in display, and improves processing speed. The three-dimensional data decoding device may determine the level of a brick to be decoded, for example, based on the distance of the brick from the camera, or whether or not the brick is visible from the camera.
[0524] Next, an implementation example of processing using LoD will be described. FIG. 118 is a flowchart of a three-dimensional data decoding process. First, the three-dimensional data decoding device acquires encoded data (S8761). For example, the encoded data is point cloud data that has been encoded and compressed using an arbitrary encoding method. The encoded data may be in a bit stream format or a file format.
[0525] Next, the three-dimensional data decoding device obtains the normal vector and position information of the brick to be processed from the encoded data (S8762). For example, the three-dimensional data decoding device obtains the normal vector for each brick and the position information of the brick from metadata (SEI or data header) included in the encoded data. The three-dimensional data decoding device may determine the distance between the brick and the camera from the position information of the brick and the information of the camera position. The three-dimensional data decoding device may also determine the visibility of the brick (whether the brick faces the camera direction) from the normal vector and the camera direction.
[0526] Next, the three-dimensional data decoding device determines which brick to decode, and decodes the first level (level 0) of the determined brick (S8763). Fig. 119 is a diagram showing an example of a brick to be decoded. As shown in Fig. 119, the three-dimensional data decoding device decodes all visible bricks at level 0 resolution.
[0527] Next, the three-dimensional data decoding device determines whether or not to decode the next level of each brick according to the position information, and decodes the next level of the brick that is determined to be decoded (S8764). This process is repeated until the decoding process of all levels is completed (S8765). Specifically, the resolution of the brick closer to the position of the virtual camera is set to be high. For example, depending on resources such as memory, levels of decoding are gradually added, with priority given to bricks closer to the camera.
[0528] Fig. 120 is a diagram showing an example of the level of the decoding target of each brick. As shown in Fig. 120, the three-dimensional data decoding device decodes bricks closer to the camera at a higher resolution and bricks farther from the camera at a lower resolution according to the distance from the camera. In addition, the three-dimensional data decoding device does not decode bricks that are not visible.
[0529] When decoding at all levels is completed (Yes in S8765), the three-dimensional data decoding device outputs the obtained three-dimensional point group (S8766).
[0530] So far, a method has been described in which the normal vector and bounding information for each slice data is calculated and encoded in a three-dimensional data encoding device, and the visibility and distance information is calculated in a three-dimensional data decoding device based on the normal vector and bounding information and sensor input information, and the slice to be decoded is determined. Below, an example will be described in which the visibility and distance information according to the camera direction is calculated and encoded in advance in the three-dimensional data encoding device for data for each slice.
[0531] 121 is a diagram illustrating an example of the syntax of a slice header of position information (Geometry slice header information). The slice header of position information includes number_of_angle, view_angle, and visibility.
[0532] The number_of_angle indicates the number of camera angles (camera directions). The view_angle indicates a camera angle, for example, a vector of the camera angle. The visibility indicates whether the slice is visible from the corresponding camera angle. Note that the number of view_angles may be variable or may be a predetermined fixed value. Also, when the number and value of view_angle are predetermined, the view_angle may be omitted.
[0533] Further, although an example of showing visibility according to camera angle has been shown here, as another example, the three-dimensional data encoding device may pre-calculate visibility according to camera position or camera parameters and store the calculated visibility in the encoded data.
[0534] 122 is a flowchart of a three-dimensional data encoding process. First, the three-dimensional data encoding device divides point cloud data into divided data (e.g., slices) (S8771). Next, the three-dimensional data encoding device encodes position information and attribute information for each divided data unit (S8772). In addition, the three-dimensional data encoding device stores visibility information corresponding to the camera angle in metadata for each divided data (S8773).
[0535] 123 is a flowchart of the three-dimensional data decoding process. First, the three-dimensional data decoding device acquires visibility information corresponding to the camera angle from the metadata for each divided data (S8776). Next, the three-dimensional data decoding device determines which divided data are visible from the desired camera angle based on the visibility information, and decodes the visible divided data (S8777).
[0536] 124 and 125 are diagrams showing examples of point cloud data. In the figures, a, c, d, and e represent planes. Therefore, the three-dimensional data encoding device can perform slice division by utilizing the fact that the three-dimensional points of each slice have normal vectors in the same direction. A similar method can also be applied to tile division.
[0537] 126 to 129 are diagrams showing examples of the configuration of a system including a three-dimensional data encoding device, a three-dimensional data decoding device, and a display device.
[0538] In the example shown in Fig. 126, the three-dimensional data encoding device generates encoded data by encoding slice data, normal vectors for each slice, and bounding box information. The three-dimensional data decoding device identifies data to be decoded from the encoded data and sensor information, and generates decoded slice data by decoding the identified data. The display device displays the decoded slice data. In this configuration, the three-dimensional data decoding device can flexibly determine visibility information and whether to decode.
[0539] In the example shown in Fig. 127, the three-dimensional data encoding device generates encoded data by encoding slice data, normal vectors for each slice, and bounding box information. The three-dimensional data decoding device determines the data to be decoded and the order from the encoded data and sensor information, and decodes the determined data in the determined order. In this configuration, the three-dimensional data decoding device can first decode the data to be displayed first (e.g., 3, 4, 5), thereby improving the comfort of display.
[0540] In the example shown in Fig. 128, the three-dimensional data encoding device generates encoded data by encoding slice data and visibility information for each camera angle. The three-dimensional data decoding device identifies data to be decoded from the encoded data information and sensor information, and decodes the identified data. The three-dimensional data decoding device may further determine the decoding order. In this configuration, the three-dimensional data decoding device does not need to calculate visibility information, so the amount of processing performed by the three-dimensional data decoding device can be reduced.
[0541] In the example shown in Fig. 129, the three-dimensional data decoding device notifies the three-dimensional data encoding device of the camera angle or camera position of the three-dimensional data decoding device via communication or the like. The three-dimensional data encoding device calculates visibility information for each slice, determines the data to be encoded and the order, and generates encoded data by encoding the determined data in the determined order. The three-dimensional data decoding device decodes the transmitted slice data as is. In this configuration, the amount of processing and communication bandwidth can be reduced by using an interactive configuration to encode and decode the necessary parts.
[0542] In addition, when the camera position or camera angle changes, the three-dimensional data decoding device may re-determine the slice to be decoded if the amount of change exceeds a predetermined value. In this case, high-speed decoding and display are possible by decoding the difference data other than the data that has already been decoded.
[0543] A method of storing encoded data in a file format such as ISOBMFF will be described below. Fig. 130 is a diagram showing an example of the configuration of a bitstream. Fig. 131 is a diagram showing an example of the configuration of a three-dimensional data encoding device. The three-dimensional data encoding device includes an encoding unit 8741 and a file conversion unit 8742. The encoding unit 8741 generates a bitstream including encoded data and control information by encoding point cloud data. The file conversion unit 8742 converts the bitstream into a file format.
[0544] 132 is a diagram showing a configuration example of a three-dimensional data decoding device. The three-dimensional data decoding device includes a file inverse conversion unit 8751 and a decoding unit 8752. The file inverse conversion unit 8751 converts a file format into a bit stream including encoded data and control information. The decoding unit 8752 generates point cloud data by decoding the bit stream.
[0545] Fig. 133 is a diagram showing the basic structure of ISOBMFF. Fig. 134 is a protocol stack diagram when NAL units common to PCC codecs are stored in ISOBMFF. Here, it is the NAL units of the PCC codec that are stored in ISOBMFF.
[0546] NAL units include NAL units for data and NAL units for metadata. NAL units for data include geometry slice data and attribute slice data. NAL units for metadata include SPS, GPS, APS, and SEI.
[0547] ISOBMFF (ISO based media file format) is a file format standard defined in ISO / IEC14496-12. It specifies a format that can store multiplexed data of various media such as video, audio, and text, and is a media-independent standard.
[0548] The basic unit in ISOBMFF is a box. A box consists of type, length, and data, and a collection of boxes of various types is a file. A file mainly consists of boxes such as ftyp, which indicates the brand of the file using 4CC, moov, which stores metadata such as control information, and mdat, which stores data.
[0549] For example, the method of storing AVC video and HEVC video is specified in ISO / IEC 14496-15. In addition, it is possible to extend the functionality of ISOBMFF to store and transmit PCC encoded data.
[0550] When storing a metadata NAL unit in ISOBMFF, the SEI may be stored in an "mdat box" together with PCC data, or in a "track box" that describes control information related to the stream. When data is packetized and transmitted, the SEI may be stored in a packet header. By indicating the SEI to the system layers, access to attribute information, tiles, and slice data becomes easier and the access speed is improved.
[0551] Next, a method for generating a PCC random access table will be described. The three-dimensional data encoding device generates a random access table using metadata including bounding box information and normal vector information for each slice. Figure 135 is a diagram showing an example of converting a bit stream into a file format.
[0552] The three-dimensional data encoding device stores the slice data in the mdat of the file format. The three-dimensional data encoding device calculates the memory position of the slice data as offset information at the beginning of the file (offsets 1 to 4 in FIG. 135), and includes the calculated offset information in the random access table (PCC random access table).
[0553] Fig. 136 is a diagram showing an example of the syntax of slice information (slice_information). Figs. 137 to 139 are diagrams showing an example of the syntax of a PCC random access table.
[0554] The PCC random access table includes bounding box information (bounding_box_info), normal vector information (normal_vector_info), and offset information (offset), which are stored in slice information (slice_information).
[0555] The three-dimensional data decoder analyzes the PCC random access table to identify the slice to be decoded. The three-dimensional data decoder can access the desired data by obtaining offset information from the PCC random access table.
[0556] As described above, the three-dimensional data encoding device according to this embodiment performs the process shown in Fig. 140. The three-dimensional data encoding device generates a bit stream by encoding position information and one or more pieces of attribute information of each of a plurality of three-dimensional points included in the point cloud data (S8781), and in the encoding (S8781), encodes a normal vector of each of the plurality of three-dimensional points as one piece of attribute information included in the one or more pieces of attribute information.
[0557] According to this, the three-dimensional data encoding device can process the normal vector in the same way as other attribute information by encoding the normal vector as attribute information. Therefore, the three-dimensional data encoding device can reduce the amount of processing. In other words, the three-dimensional data encoding device can encode the normal vector as attribute information without changing the definition of the attribute information, etc.
[0558] For example, in the encoding (S8781), the three-dimensional data encoding device converts a normal vector expressed by a floating point into an integer and then encodes it. This allows the three-dimensional data encoding device to process the normal vector in the same way as other attribute information, for example, when the other attribute information is expressed by an integer.
[0559] For example, the bit stream includes position information and control information (e.g., SPS) common to one or more attribute information, and the control information (e.g., SPS) includes at least one of information indicating that one attribute information included in the one or more attribute information indicates a normal vector (e.g., attribute_type=Normal Vector), or information indicating that the normal vector is data having three elements for each point.
[0560] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor uses the memory to perform the above-mentioned processing.
[0561] Moreover, the three-dimensional data decoding device according to this embodiment performs the process shown in Fig. 141. The three-dimensional data decoding device acquires a bit stream generated by encoding position information and one or more pieces of attribute information of each of a plurality of three-dimensional points included in point cloud data, in which normal vectors of each of the plurality of three-dimensional points are encoded as one piece of attribute information included in the one or more pieces of attribute information (S8786), and acquires the normal vector by decoding the one piece of attribute information from the bit stream (S8787).
[0562] According to this, the three-dimensional data decoding device decodes the normal vector as attribute information, and can process the normal vector in the same way as other attribute information, thereby reducing the amount of processing.
[0563] For example, the three-dimensional data decoding device acquires a normal vector represented by an integer in the acquisition of the normal vector (S8787). This allows the three-dimensional data decoding device to process the normal vector in the same way as other attribute information, for example, when the other attribute information is expressed by an integer.
[0564] For example, the bit stream includes position information and control information (e.g., SPS) common to one or more attribute information, and the control information (e.g., SPS) includes at least one of information indicating that one attribute information included in the one or more attribute information indicates a normal vector (e.g., attribute_type=Normal Vector), or information indicating that the normal vector is data having three elements for each point.
[0565] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor uses the memory to perform the above-mentioned processing.
[0566] Moreover, the three-dimensional data encoding device according to this embodiment performs the process shown in Fig. 142. The three-dimensional data encoding device divides point cloud data into a plurality of divided data (for example, bricks, slices, or tiles) (S8791), and generates a bit stream by encoding the plurality of divided data (S8792). The bit stream includes information indicating the normal vector of each of the plurality of divided data.
[0567] According to this, the three-dimensional data encoding device can reduce the amount of processing and the amount of code by encoding the normal vector for each divided data, compared to the case where the normal vector is encoded for each point. For example, each of the multiple divided data is a random access unit.
[0568] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor uses the memory to perform the above-mentioned processing.
[0569] Moreover, the three-dimensional data decoding device according to this embodiment performs the process shown in Fig. 143. The three-dimensional data decoding device acquires a bit stream generated by encoding a plurality of split data (e.g., bricks, slices, or tiles) generated by splitting point cloud data (S8796), and acquires information indicating the normal vector of each of the plurality of split data from the bit stream (S8797).
[0570] According to this, the three-dimensional data decoding device can reduce the amount of processing by decoding the normal vector for each divided data, compared to decoding the normal vector for each point. For example, each of the multiple divided data is a random access unit.
[0571] For example, the three-dimensional data decoding device further determines split data to be decoded from the multiple split data based on the normal vector, and decodes the split data to be decoded.
[0572] For example, the three-dimensional data decoding device further determines a decoding order for the multiple data segments based on the normal vector, and decodes the multiple data segments in the determined decoding order.
[0573] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor uses the memory to perform the above-mentioned processing.
[0574] (Embodiment 8) Next, a configuration of a three-dimensional data creation device 810 according to this embodiment will be described. Fig. 144 is a block diagram showing an example of the configuration of a three-dimensional data creation device 810 according to this embodiment. This three-dimensional data creation device 810 is mounted on a vehicle, for example. The three-dimensional data creation device 810 transmits and receives three-dimensional data to and from an external traffic monitoring cloud, a leading vehicle, or a following vehicle, and creates and stores three-dimensional data.
[0575] The three-dimensional data creation device 810 includes a data receiving unit 811, a communication unit 812, a receiving control unit 813, a format conversion unit 814, multiple sensors 815, a three-dimensional data creation unit 816, a three-dimensional data synthesis unit 817, a three-dimensional data storage unit 818, a communication unit 819, a transmission control unit 820, a format conversion unit 821, and a data transmission unit 822.
[0576] The data receiving unit 811 receives three-dimensional data 831 from a traffic monitoring cloud or a preceding vehicle. The three-dimensional data 831 includes information such as a point cloud, a visible light image, depth information, sensor position information, or speed information, including an area that cannot be detected by the sensor 815 of the vehicle itself.
[0577] The communication unit 812 communicates with the traffic monitoring cloud or the vehicle ahead, and transmits data transmission requests, etc. to the traffic monitoring cloud or the vehicle ahead.
[0578] The reception control unit 813 exchanges information such as compatible formats with the communication destination via the communication unit 812, and establishes communication with the communication destination.
[0579] The format conversion unit 814 generates three-dimensional data 832 by performing format conversion or the like on the three-dimensional data 831 received by the data receiving unit 811. Furthermore, when the three-dimensional data 831 is compressed or encoded, the format conversion unit 814 performs decompression or decoding processing.
[0580] The multiple sensors 815 are a group of sensors such as LiDAR, a visible light camera, or an infrared camera that acquires information about the outside of the vehicle, and generate sensor information 833. For example, when the sensor 815 is a laser sensor such as LiDAR, the sensor information 833 is three-dimensional data such as a point cloud (point cloud data). Note that the number of sensors 815 does not need to be multiple.
[0581] The three-dimensional data creation unit 816 generates three-dimensional data 834 from the sensor information 833. The three-dimensional data 834 includes information such as a point cloud, a visible light image, depth information, sensor position information, or velocity information.
[0582] The three-dimensional data synthesis unit 817 synthesizes three-dimensional data 834 created based on the host vehicle's sensor information 833 with three-dimensional data 832 created by a traffic monitoring cloud or a preceding vehicle, etc., to construct three-dimensional data 835 that includes the space in front of the preceding vehicle that cannot be detected by the host vehicle's sensor 815.
[0583] The three-dimensional data storage unit 818 stores the generated three-dimensional data 835 and the like.
[0584] The communication unit 819 communicates with the traffic monitoring cloud or the following vehicle, and transmits data transmission requests, etc. to the traffic monitoring cloud or the following vehicle.
[0585] The transmission control unit 820 exchanges information such as compatible formats with the communication destination and establishes communication with the communication destination via the communication unit 819. In addition, the transmission control unit 820 determines a transmission region, which is the space of the three-dimensional data to be transmitted, based on three-dimensional data construction information of the three-dimensional data 832 generated by the three-dimensional data synthesis unit 817 and a data transmission request from the communication destination.
[0586] Specifically, the transmission control unit 820 determines a transmission region including a space in front of the vehicle that cannot be detected by the sensor of the following vehicle in response to a data transmission request from the traffic monitoring cloud or the following vehicle. The transmission control unit 820 also determines the transmission region by judging whether or not a transmittable space or a transmitted space has been updated based on the three-dimensional data construction information. For example, the transmission control unit 820 determines the region specified in the data transmission request and in which the corresponding three-dimensional data 835 exists as the transmission region. Then, the transmission control unit 820 notifies the format conversion unit 821 of the format supported by the communication destination and the transmission region.
[0587] The format conversion unit 821 converts three-dimensional data 836 of the transmission region, among the three-dimensional data 835 stored in the three-dimensional data storage unit 818, into a format supported by the receiving side, thereby generating three-dimensional data 837. Note that the format conversion unit 821 may reduce the amount of data by compressing or encoding the three-dimensional data 837.
[0588] The data transmission unit 822 transmits the three-dimensional data 837 to the traffic monitoring cloud or the following vehicle. The three-dimensional data 837 includes information such as a point cloud, visible light image, depth information, or sensor position information in front of the vehicle, including blind spots of the following vehicle.
[0589] In the above description, the format conversion units 814 and 821 perform format conversion, but the format conversion does not necessarily have to be performed.
[0590] With this configuration, the three-dimensional data creation device 810 externally acquires three-dimensional data 831 of an area that cannot be detected by the sensor 815 of the host vehicle, and generates three-dimensional data 835 by synthesizing the three-dimensional data 831 with three-dimensional data 834 based on sensor information 833 detected by the sensor 815 of the host vehicle. In this way, the three-dimensional data creation device 810 can generate three-dimensional data of an area that cannot be detected by the sensor 815 of the host vehicle.
[0591] In addition, in response to a data transmission request from a traffic monitoring cloud or a following vehicle, the three-dimensional data creation device 810 can transmit three-dimensional data including the space in front of the vehicle that cannot be detected by the sensors of the following vehicle to the traffic monitoring cloud or the following vehicle, etc.
[0592] Next, there will be described a procedure for transmitting three-dimensional data to a following vehicle in the three-dimensional data creation device 810. Fig. 145 is a flowchart showing an example of a procedure for transmitting three-dimensional data to a traffic monitoring cloud or a following vehicle by the three-dimensional data creation device 810.
[0593] First, the three-dimensional data creation device 810 generates and updates three-dimensional data 835 of a space including the space on the road ahead of the vehicle (S801). Specifically, the three-dimensional data creation device 810 constructs three-dimensional data 835 including the space ahead of the vehicle ahead that cannot be detected by the sensor 815 of the vehicle ahead, for example, by combining three-dimensional data 834 created based on the sensor information 833 of the vehicle ahead with three-dimensional data 831 created by the traffic monitoring cloud or the vehicle ahead.
[0594] Next, the three-dimensional data creation device 810 determines whether the three-dimensional data 835 contained in the transmitted space has changed (S802).
[0595] If a change occurs in the three-dimensional data 835 contained in the space that has already been transmitted, such as when a vehicle or person enters the space from outside (Yes in S802), the three-dimensional data creation device 810 transmits three-dimensional data including the three-dimensional data 835 of the space where the change has occurred to the traffic monitoring cloud or a following vehicle (S803).
[0596] The three-dimensional data creation device 810 may transmit the three-dimensional data of the space where a change has occurred in accordance with the transmission timing of the three-dimensional data transmitted at a predetermined interval, or may transmit the three-dimensional data immediately after detecting the change. In other words, the three-dimensional data creation device 810 may transmit the three-dimensional data of the space where a change has occurred in priority over the three-dimensional data transmitted at a predetermined interval.
[0597] In addition, the three-dimensional data creation device 810 may transmit all of the three-dimensional data of the space in which a change has occurred as the three-dimensional data of the space in which a change has occurred, or may transmit only the differences in the three-dimensional data (for example, information on three-dimensional points that have appeared or disappeared, or displacement information of three-dimensional points, etc.).
[0598] Furthermore, the three-dimensional data creation device 810 may transmit metadata regarding the danger avoidance operation of the vehicle, such as a sudden braking warning, to the following vehicle prior to the three-dimensional data of the space where a change has occurred. This allows the following vehicle to quickly recognize the sudden braking of the vehicle ahead, and to start danger avoidance operations, such as deceleration, earlier.
[0599] If there is no change in the three-dimensional data 835 contained in the transmitted space (No in S802), or after step S803, the three-dimensional data creation device 810 transmits the three-dimensional data contained in a space of a specified shape at a distance L ahead of the vehicle to the traffic monitoring cloud or a following vehicle (S804).
[0600] Furthermore, for example, the processes in steps S801 to S804 are repeatedly performed at predetermined time intervals.
[0601] Furthermore, when there is no difference between the three-dimensional data 835 of the space currently being transmitted and the three-dimensional map, the three-dimensional data creation device 810 does not need to transmit the three-dimensional data 837 of the space.
[0602] In this embodiment, the client device transmits sensor information obtained by the sensor to a server or another client device.
[0603] First, the configuration of the system according to this embodiment will be described. Fig. 146 is a diagram showing the configuration of a transmission / reception system for a 3D map and sensor information according to this embodiment. This system includes a server 901 and client devices 902A and 902B. When there is no particular need to distinguish between the client devices 902A and 902B, they are also referred to as client devices 902.
[0604] The client device 902 is, for example, an in-vehicle device mounted on a moving object such as a vehicle. The server 901 is, for example, a traffic monitoring cloud or the like, and is capable of communicating with a plurality of client devices 902.
[0605] The server 901 transmits a three-dimensional map composed of a point cloud to the client device 902. Note that the composition of the three-dimensional map is not limited to a point cloud, and may represent other three-dimensional data such as a mesh structure.
[0606] The client device 902 transmits sensor information acquired by the client device 902 to the server 901. The sensor information includes, for example, at least one of LiDAR acquisition information, a visible light image, an infrared image, a depth image, sensor position information, and speed information.
[0607] Data transmitted between the server 901 and the client device 902 may be compressed to reduce data, or may be left uncompressed to maintain data accuracy. When compressing data, a three-dimensional compression method based on an octree structure, for example, can be used for the point cloud. Also, a two-dimensional image compression method can be used for the visible light image, the infrared image, and the depth image. The two-dimensional image compression method is, for example, MPEG-4 AVC or HEVC standardized by MPEG.
[0608] Furthermore, the server 901 transmits the three-dimensional map managed by the server 901 to the client device 902 in response to a transmission request for the three-dimensional map from the client device 902. Note that the server 901 may transmit the three-dimensional map without waiting for a transmission request for the three-dimensional map from the client device 902. For example, the server 901 may broadcast the three-dimensional map to one or more client devices 902 in a predetermined space. Furthermore, the server 901 may transmit a three-dimensional map suitable for the position of the client device 902 at regular intervals to the client device 902 that has received a transmission request once. Furthermore, the server 901 may transmit the three-dimensional map to the client device 902 every time the three-dimensional map managed by the server 901 is updated.
[0609] The client device 902 issues a request for transmitting a three-dimensional map to the server 901. For example, when the client device 902 wants to estimate its own position while driving, the client device 902 transmits a request for transmitting a three-dimensional map to the server 901.
[0610] In the following cases, the client device 902 may issue a transmission request for a three-dimensional map to the server 901. If the three-dimensional map held by the client device 902 is old, the client device 902 may issue a transmission request for a three-dimensional map to the server 901. For example, if a certain period of time has passed since the client device 902 acquired the three-dimensional map, the client device 902 may issue a transmission request for a three-dimensional map to the server 901.
[0611] The client device 902 may issue a request to send a three-dimensional map to the server 901 a certain time before the client device 902 leaves the space shown in the three-dimensional map held by the client device 902. For example, when the client device 902 is within a predetermined distance from the boundary of the space shown in the three-dimensional map held by the client device 902, the client device 902 may issue a request to send a three-dimensional map to the server 901. In addition, when the movement route and movement speed of the client device 902 are known, the time when the client device 902 will leave the space shown in the three-dimensional map held by the client device 902 may be predicted based on these.
[0612] If the error in aligning the three-dimensional data created by the client device 902 from sensor information with the three-dimensional map is equal to or greater than a certain level, the client device 902 may issue a request to the server 901 to transmit the three-dimensional map.
[0613] The client device 902 transmits the sensor information to the server 901 in response to a request to transmit the sensor information transmitted from the server 901. The client device 902 may transmit the sensor information to the server 901 without waiting for a request to transmit the sensor information from the server 901. For example, once the client device 902 receives a request to transmit the sensor information from the server 901, the client device 902 may transmit the sensor information to the server 901 periodically for a certain period of time. In addition, when an error occurs in aligning three-dimensional data created by the client device 902 based on the sensor information with the three-dimensional map obtained from the server 901, the client device 902 may determine that a change may have occurred in the three-dimensional map around the client device 902 and transmit a message to that effect along with the sensor information to the server 901.
[0614] The server 901 issues a transmission request for sensor information to the client device 902. For example, the server 901 receives location information of the client device 902, such as a GPS, from the client device 902. When the server 901 determines, based on the location information of the client device 902, that the client device 902 is approaching a space with little information in the three-dimensional map managed by the server 901, it issues a transmission request for sensor information to the client device 902 in order to generate a new three-dimensional map. The server 901 may also issue a transmission request for sensor information when it is desired to update the three-dimensional map, when it is desired to check road conditions such as during snowfall or a disaster, when it is desired to check traffic congestion conditions, or when it is desired to check incident and accident conditions, etc.
[0615] Furthermore, the client device 902 may set the amount of data of the sensor information to be transmitted to the server 901 depending on the communication state or bandwidth at the time of receiving a request to transmit the sensor information from the server 901. Setting the amount of data of the sensor information to be transmitted to the server 901 means, for example, increasing or decreasing the amount of the data itself or appropriately selecting a compression method.
[0616] 147 is a block diagram showing a configuration example of a client device 902. The client device 902 receives a three-dimensional map composed of a point cloud or the like from the server 901, and estimates the self-position of the client device 902 from three-dimensional data created based on sensor information of the client device 902. In addition, the client device 902 transmits the acquired sensor information to the server 901.
[0617] The client device 902 includes a data receiving unit 1011, a communication unit 1012, a reception control unit 1013, a format conversion unit 1014, multiple sensors 1015, a three-dimensional data creation unit 1016, a three-dimensional image processing unit 1017, a three-dimensional data storage unit 1018, a format conversion unit 1019, a communication unit 1020, a transmission control unit 1021, and a data transmission unit 1022.
[0618] The data receiving unit 1011 receives a three-dimensional map 1031 from the server 901. The three-dimensional map 1031 is data including a point cloud such as a WLD or a SWLD. The three-dimensional map 1031 may include either compressed data or uncompressed data.
[0619] The communication unit 1012 communicates with the server 901, and transmits data transmission requests (for example, a request to transmit a three-dimensional map) to the server 901.
[0620] The reception control unit 1013 exchanges information such as compatible formats with the communication destination via the communication unit 1012, and establishes communication with the communication destination.
[0621] The format conversion unit 1014 generates a three-dimensional map 1032 by performing format conversion and the like on the three-dimensional map 1031 received by the data receiving unit 1011. Furthermore, if the three-dimensional map 1031 is compressed or encoded, the format conversion unit 1014 performs decompression or decoding processing. Note that if the three-dimensional map 1031 is uncompressed data, the format conversion unit 1014 does not perform decompression or decoding processing.
[0622] The multiple sensors 1015 are a group of sensors, such as LiDAR, a visible light camera, an infrared camera, or a depth sensor, that acquire information about the outside of the vehicle in which the client device 902 is mounted, and generate sensor information 1033. For example, when the sensor 1015 is a laser sensor such as LiDAR, the sensor information 1033 is three-dimensional data such as a point cloud (point cloud data). Note that the number of sensors 1015 does not need to be multiple.
[0623] The three-dimensional data creation unit 1016 creates three-dimensional data 1034 of the surroundings of the vehicle based on the sensor information 1033. For example, the three-dimensional data creation unit 1016 creates point cloud data with color information of the surroundings of the vehicle using information acquired by the LiDAR and a visible light image acquired by a visible light camera.
[0624] The three-dimensional image processing unit 1017 performs self-position estimation processing of the vehicle using a three-dimensional map 1032 such as a received point cloud and three-dimensional data 1034 of the surroundings of the vehicle generated from the sensor information 1033. The three-dimensional image processing unit 1017 may generate three-dimensional data 1035 of the surroundings of the vehicle by combining the three-dimensional map 1032 and the three-dimensional data 1034, and perform self-position estimation processing using the generated three-dimensional data 1035.
[0625] The three-dimensional data storage unit 1018 stores a three-dimensional map 1032, three-dimensional data 1034, and three-dimensional data 1035, and the like.
[0626] The format conversion unit 1019 generates sensor information 1037 by converting the sensor information 1033 into a format supported by the receiving side. The format conversion unit 1019 may reduce the amount of data by compressing or encoding the sensor information 1037. Furthermore, the format conversion unit 1019 may omit processing when format conversion is not necessary. Furthermore, the format conversion unit 1019 may control the amount of data to be transmitted in accordance with a designated transmission range.
[0627] The communication unit 1020 communicates with the server 901, and receives data transmission requests (sensor information transmission requests) and the like from the server 901.
[0628] The transmission control unit 1021 exchanges information such as compatible formats with the communication destination via the communication unit 1020, and establishes communication.
[0629] The data transmission unit 1022 transmits the sensor information 1037 to the server 901. The sensor information 1037 includes information acquired by a plurality of sensors 1015, such as information acquired by a LiDAR, a luminance image acquired by a visible light camera, an infrared image acquired by an infrared camera, a depth image acquired by a depth sensor, sensor position information, and speed information.
[0630] Next, the configuration of the server 901 will be described. Fig. 148 is a block diagram showing an example of the configuration of the server 901. The server 901 receives sensor information transmitted from the client device 902, and creates three-dimensional data based on the received sensor information. The server 901 updates a three-dimensional map managed by the server 901 using the created three-dimensional data. In addition, the server 901 transmits the updated three-dimensional map to the client device 902 in response to a transmission request for the three-dimensional map from the client device 902.
[0631] The server 901 includes a data receiving unit 1111, a communication unit 1112, a receiving control unit 1113, a format conversion unit 1114, a three-dimensional data creation unit 1116, a three-dimensional data synthesis unit 1117, a three-dimensional data storage unit 1118, a format conversion unit 1119, a communication unit 1120, a transmission control unit 1121, and a data transmission unit 1122.
[0632] The data receiving unit 1111 receives sensor information 1037 from the client device 902. The sensor information 1037 includes, for example, information acquired by a LiDAR, a luminance image acquired by a visible light camera, an infrared image acquired by an infrared camera, a depth image acquired by a depth sensor, sensor position information, and speed information.
[0633] The communication unit 1112 communicates with the client device 902 and transmits a data transmission request (for example, a request to transmit sensor information) to the client device 902 .
[0634] The reception control unit 1113 exchanges information such as compatible formats with the communication destination via the communication unit 1112, and establishes communication.
[0635] When the received sensor information 1037 is compressed or encoded, the format conversion unit 1114 performs decompression or decoding processing to generate the sensor information 1132. Note that, when the sensor information 1037 is non-compressed data, the format conversion unit 1114 does not perform decompression or decoding processing.
[0636] The three-dimensional data creation unit 1116 creates three-dimensional data 1134 of the periphery of the client device 902 based on the sensor information 1132. For example, the three-dimensional data creation unit 1116 creates point cloud data with color information of the periphery of the client device 902 using information acquired by the LiDAR and visible light images acquired by a visible light camera.
[0637] A three-dimensional data synthesis unit 1117 synthesizes three-dimensional data 1134 created based on the sensor information 1132 with a three-dimensional map 1135 managed by the server 901, thereby updating the three-dimensional map 1135.
[0638] The three-dimensional data storage unit 1118 stores a three-dimensional map 1135 and the like.
[0639] The format conversion unit 1119 generates the three-dimensional map 1031 by converting the three-dimensional map 1135 into a format supported by the receiving side. The format conversion unit 1119 may reduce the amount of data by compressing or encoding the three-dimensional map 1135. Furthermore, the format conversion unit 1119 may omit processing when format conversion is not necessary. Furthermore, the format conversion unit 1119 may control the amount of data to be transmitted in accordance with a designated transmission range.
[0640] The communication unit 1120 communicates with the client device 902 and receives a data transmission request (a request to transmit a three-dimensional map) or the like from the client device 902 .
[0641] A transmission control unit 1121 exchanges information such as compatible formats with a communication destination via a communication unit 1120, and establishes communication.
[0642] The data transmission unit 1122 transmits the three-dimensional map 1031 to the client device 902. The three-dimensional map 1031 is data including a point cloud such as a WLD or a SWLD. The three-dimensional map 1031 may include either compressed data or uncompressed data.
[0643] Next, a description will be given of the operation flow of the client device 902. Fig. 149 is a flowchart showing the operation of the client device 902 when acquiring a three-dimensional map.
[0644] First, the client device 902 requests the server 901 to transmit a three-dimensional map (such as a point cloud) (S1001). At this time, the client device 902 may also transmit position information of the client device 902 obtained by GPS or the like, thereby requesting the server 901 to transmit a three-dimensional map related to the position information.
[0645] Next, the client device 902 receives the three-dimensional map from the server 901 (S1002). If the received three-dimensional map is compressed data, the client device 902 decodes the received three-dimensional map to generate an uncompressed three-dimensional map (S1003).
[0646] Next, the client device 902 creates three-dimensional data 1034 of the periphery of the client device 902 from sensor information 1033 obtained by the multiple sensors 1015 (S1004). Next, the client device 902 estimates its own position using the three-dimensional map 1032 received from the server 901 and the three-dimensional data 1034 created from the sensor information 1033 (S1005).
[0647] 150 is a flowchart showing an operation of the client device 902 when transmitting sensor information. First, the client device 902 receives a request to transmit sensor information from the server 901 (S1011). The client device 902 that has received the transmission request transmits sensor information 1037 to the server 901 (S1012). Note that, when the sensor information 1033 includes a plurality of pieces of information obtained by a plurality of sensors 1015, the client device 902 may generate the sensor information 1037 by compressing each piece of information using a compression method suitable for each piece of information.
[0648] Next, the operation flow of the server 901 will be described. Fig. 151 is a flowchart showing the operation when the server 901 acquires sensor information. First, the server 901 requests the client device 902 to transmit sensor information (S1021). Next, the server 901 receives the sensor information 1037 transmitted from the client device 902 in response to the request (S1022). Next, the server 901 creates three-dimensional data 1134 using the received sensor information 1037 (S1023). Next, the server 901 reflects the created three-dimensional data 1134 in the three-dimensional map 1135 (S1024).
[0649] Fig. 152 is a flowchart showing the operation of the server 901 when transmitting a three-dimensional map. First, the server 901 receives a request to transmit a three-dimensional map from the client device 902 (S1031). The server 901 that has received the request to transmit the three-dimensional map transmits the three-dimensional map 1031 to the client device 902 (S1032). At this time, the server 901 may extract a three-dimensional map of the vicinity according to the position information of the client device 902, and transmit the extracted three-dimensional map. The server 901 may also compress the three-dimensional map composed of a point cloud using, for example, a compression method based on an octet tree structure, and transmit the compressed three-dimensional map.
[0650] A modification of this embodiment will now be described.
[0651] The server 901 creates three-dimensional data 1134 of the vicinity of the position of the client device 902 using the sensor information 1037 received from the client device 902. Next, the server 901 calculates the difference between the three-dimensional data 1134 and the three-dimensional map 1135 by matching the created three-dimensional data 1134 with a three-dimensional map 1135 of the same area managed by the server 901. If the difference is equal to or greater than a predetermined threshold, the server 901 determines that some abnormality has occurred in the vicinity of the client device 902. For example, when ground subsidence or the like occurs due to a natural disaster such as an earthquake, it is considered that a large difference occurs between the three-dimensional map 1135 managed by the server 901 and the three-dimensional data 1134 created based on the sensor information 1037.
[0652] The sensor information 1037 may include information indicating at least one of the type of sensor, the performance of the sensor, and the model number of the sensor. In addition, a class ID or the like according to the performance of the sensor may be added to the sensor information 1037. For example, when the sensor information 1037 is information acquired by LiDAR, it is possible to assign an identifier to the performance of the sensor, such as class 1 for a sensor that can acquire information with an accuracy of several millimeters, class 2 for a sensor that can acquire information with an accuracy of several centimeters, and class 3 for a sensor that can acquire information with an accuracy of several meters. In addition, the server 901 may estimate the performance information of the sensor from the model number of the client device 902. For example, when the client device 902 is mounted on a vehicle, the server 901 may determine the spec information of the sensor from the model of the vehicle. In this case, the server 901 may acquire information on the model of the vehicle in advance, or the information may be included in the sensor information. In addition, the server 901 may use the acquired sensor information 1037 to switch the degree of correction for the three-dimensional data 1134 created using the sensor information 1037. For example, when the sensor performance is high accuracy (class 1), the server 901 does not perform correction on the three-dimensional data 1134. When the sensor performance is low accuracy (class 3), the server 901 applies correction according to the accuracy of the sensor to the three-dimensional data 1134. For example, the server 901 increases the degree (strength) of correction as the accuracy of the sensor decreases.
[0653] The server 901 may simultaneously issue requests to send sensor information to multiple client devices 902 in a certain space. When the server 901 receives multiple pieces of sensor information from multiple client devices 902, the server 901 does not need to use all of the sensor information to create the three-dimensional data 1134, and may select the sensor information to be used according to the performance of the sensor, for example. For example, when updating the three-dimensional map 1135, the server 901 may select high-precision sensor information (class 1) from the multiple pieces of sensor information received, and create the three-dimensional data 1134 using the selected sensor information.
[0654] The server 901 is not limited to a server such as a traffic monitoring cloud, but may be another client device (mounted in a vehicle). Figure 153 is a diagram showing the system configuration in this case.
[0655] For example, client device 902C issues a request to send sensor information to nearby client device 902A and acquires the sensor information from client device 902A. Client device 902C then creates three-dimensional data using the acquired sensor information of client device 902A and updates the three-dimensional map of client device 902C. This allows client device 902C to generate a three-dimensional map of the space that can be acquired from client device 902A by taking advantage of the performance of client device 902C. For example, it is considered that such a case occurs when client device 902C has high performance.
[0656] In this case, the client device 902A that provided the sensor information is given the right to obtain the high-precision 3D map generated by the client device 902C. The client device 902A receives the high-precision 3D map from the client device 902C in accordance with the right.
[0657] In addition, the client device 902C may issue a request to transmit sensor information to multiple nearby client devices 902 (client device 902A and client device 902B). If the sensor of the client device 902A or the client device 902B is high performance, the client device 902C can create three-dimensional data using the sensor information obtained by the high performance sensor.
[0658] 154 is a block diagram showing the functional configuration of a server 901 and a client device 902. The server 901 includes, for example, a 3D map compression / decoding processing unit 1201 that compresses and decodes a 3D map, and a sensor information compression / decoding processing unit 1202 that compresses and decodes sensor information.
[0659] The client device 902 includes a three-dimensional map decoding processing unit 1211 and a sensor information compression processing unit 1212. The three-dimensional map decoding processing unit 1211 receives encoded data of the compressed three-dimensional map and decodes the encoded data to obtain the three-dimensional map. The sensor information compression processing unit 1212 compresses the sensor information itself instead of the three-dimensional data created from the acquired sensor information, and transmits the encoded data of the compressed sensor information to the server 901. With this configuration, the client device 902 only needs to internally hold a processing unit (device or LSI) that performs processing to decode the three-dimensional map (point cloud, etc.), and does not need to internally hold a processing unit that performs processing to compress the three-dimensional data of the three-dimensional map (point cloud, etc.). This makes it possible to reduce the cost and power consumption of the client device 902.
[0660] As described above, the client device 902 according to this embodiment is mounted on a moving object, and creates three-dimensional data 1034 of the surroundings of the moving object from sensor information 1033 indicating the surrounding conditions of the moving object, obtained by a sensor 1015 mounted on the moving object. The client device 902 estimates the self-position of the moving object using the created three-dimensional data 1034. The client device 902 transmits the acquired sensor information 1033 to the server 901 or another client device 902.
[0661] According to this, the client device 902 transmits the sensor information 1033 to the server 901 or the like. This may reduce the amount of data to be transmitted compared to the case of transmitting three-dimensional data. Also, since the client device 902 does not need to perform processing such as compression or encoding of the three-dimensional data, the amount of processing by the client device 902 can be reduced. Therefore, the client device 902 can reduce the amount of data to be transmitted or simplify the device configuration.
[0662] Moreover, the client device 902 further transmits a request to the server 901 to transmit a three-dimensional map, and receives a three-dimensional map 1031 from the server 901. The client device 902 estimates its own location using the three-dimensional data 1034 and the three-dimensional map 1032.
[0663] The sensor information 1033 includes at least one of information obtained by a laser sensor, a luminance image, an infrared image, a depth image, sensor position information, and sensor speed information.
[0664] Furthermore, the sensor information 1033 includes information indicating the performance of the sensor.
[0665] Furthermore, the client device 902 encodes or compresses the sensor information 1033, and transmits the encoded or compressed sensor information 1037 to the server 901 or another client device 902. This allows the client device 902 to reduce the amount of data to be transmitted.
[0666] For example, the client device 902 includes a processor and a memory, and the processor uses the memory to perform the above-mentioned processing.
[0667] Furthermore, the server 901 according to this embodiment is capable of communicating with a client device 902 mounted on the mobile object, and receives sensor information 1037 indicating the surrounding conditions of the mobile object, obtained by a sensor 1015 mounted on the mobile object, from the client device 902. The server 901 creates three-dimensional data 1134 of the surroundings of the mobile object from the received sensor information 1037.
[0668] According to this, the server 901 creates three-dimensional data 1134 using the sensor information 1037 transmitted from the client device 902. This may reduce the amount of data transmitted compared to when the client device 902 transmits three-dimensional data. In addition, since the client device 902 does not need to perform processing such as compression or encoding of the three-dimensional data, the amount of processing by the client device 902 can be reduced. Therefore, the server 901 can reduce the amount of data transmitted or simplify the device configuration.
[0669] Moreover, the server 901 further transmits a request to the client device 902 to transmit the sensor information.
[0670] Furthermore, the server 901 further updates a three-dimensional map 1135 using the created three-dimensional data 1134 , and transmits the three-dimensional map 1135 to the client device 902 in response to a transmission request for the three-dimensional map 1135 from the client device 902 .
[0671] The sensor information 1037 includes at least one of information obtained by a laser sensor, a luminance image, an infrared image, a depth image, sensor position information, and sensor speed information.
[0672] Additionally, the sensor information 1037 includes information indicating the performance of the sensor.
[0673] Moreover, the server 901 further corrects the three-dimensional data in accordance with the performance of the sensor, thereby enabling the three-dimensional data creation method to improve the quality of the three-dimensional data.
[0674] Furthermore, in receiving the sensor information, the server 901 receives a plurality of pieces of sensor information 1037 from a plurality of client devices 902, and selects the sensor information 1037 to be used for creating the three-dimensional data 1134 based on a plurality of pieces of information indicating the performance of the sensors included in the plurality of pieces of sensor information 1037. This allows the server 901 to improve the quality of the three-dimensional data 1134.
[0675] Furthermore, the server 901 decodes or expands the received sensor information 1037, and creates three-dimensional data 1134 from the decoded or expanded sensor information 1132. This allows the server 901 to reduce the amount of data to be tr...
Claims
1. A three-dimensional data decoding method executed by a three-dimensional data decoding device, comprising: Obtaining encoded three-dimensional data based on visual information regarding visibility from the sensor; Decoding the encoded three-dimensional data. A method for decoding three-dimensional data.
2. The visibility information is determined based on viewpoint information of the sensor. The three-dimensional data decoding method according to claim 1 .
3. The viewpoint information is determined based on a position of the sensor, an orientation of the sensor, or a parameter of the sensor. The three-dimensional data decoding method according to claim 2.
4. the visibility information is determined by view frustum culling; The three-dimensional data decoding method further includes selecting a divided data item to be decoded; The segmented data is located in front of the sensor's viewing frustum and is part of the encoded three-dimensional data. The three-dimensional data decoding method according to claim 1 .
5. The viewpoint information is stored in metadata of the encoded three-dimensional data. The three-dimensional data decoding method according to claim 2.
6. A processor; A memory, The processor uses the memory to: Obtaining encoded three-dimensional data based on visual information regarding visibility from the sensor; Decoding the encoded three-dimensional data. Three-dimensional data decoding device.
Citation Information
Patent Citations
A system for encoding multiple videos acquired from moving objects in a scene by multiple fixed cameras
JP2007519285A
Out-of-core point rendering with dynamic shapes
US20180137671A1
Point cloud compression using non-cubic projections and masks
US20190087978A1
Encoding device, decoding device, encoding method, and decoding method
WO2017204185A1
Image processing apparatus and image processing method
WO2019003953A1