Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device

The method encodes three-dimensional data by region with connection information, facilitating efficient decoding of desired points, addressing inefficiencies in three-dimensional data processing and display.

JP2025111630AActive Publication Date: 2025-07-30PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025071581
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-10-03
Filing Date
2025-04-23
Publication Date
2025-07-30
Estimated Expiration
2040-10-02

AI Technical Summary

Technical Problem

Existing three-dimensional data encoding methods struggle to efficiently select and decode desired encoded three-dimensional points, leading to inefficiencies in data processing and display.

Method used

A method that encodes three-dimensional points by region, generating connection information with tile and association information based on visibility relationships, allowing for selective decoding of desired points.

Benefits of technology

Enables efficient selection and decoding of desired three-dimensional points, reducing processing time and resource usage by prioritizing relevant data for display.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025111630000001_ABST
    Figure 2025111630000001_ABST
Patent Text Reader

Abstract

To provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, and a three-dimensional data decoding device that can appropriately select and decode a plurality of desired encoded three-dimensional points.SOLUTION: A three-dimensional data encoding method includes: encoding three-dimensional points each located in any one of a plurality of regions, for each region (S9401); generating connectivity information generated based on a relationship between a predetermined region among the plurality of regions and regions other than the predetermined region among the plurality of regions, the connectivity information including (i) tile information indicating values each uniquely assigned to the regions and (ii) relation information indicating that the predetermined region and other regions are related with the tile information (S9402); and generating a bitstream including the generated connectivity information and the encoded three-dimensional points (S9403). The relationship is based on visibility based on a viewpoint in the predetermined region.SELECTED DRAWING: Figure 206
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding apparatus, and a three-dimensional data decoding apparatus.

Background Art

[0002] In the future, the spread of devices or services that utilize three-dimensional data is expected in a wide range of fields such as computer vision, map information, monitoring, infrastructure inspection, or video distribution for autonomous operation of automobiles or robots. Three-dimensional data is acquired by various methods such as a distance sensor such as a range finder, a stereo camera, or a combination of a plurality of monocular cameras.

[0003] As one of the methods for expressing three-dimensional data, there is a method called point cloud that represents the shape of a three-dimensional structure by a point group in a three-dimensional space. In the point cloud, the position and color of the point group are stored. Although the point cloud is expected to become mainstream as a method for expressing three-dimensional data, the point group has a very large data volume. Therefore, in the accumulation or transmission of three-dimensional data, as with two-dimensional moving images (for example, MPEG-4 AVC or HEVC standardized by MPEG), compression of the data volume by encoding is essential.

[0004] Also, regarding the compression of the point cloud, it is partially supported by a publicly available library (Point Cloud Library) that performs point cloud-related processing.

[0005] Also, a technique for searching and displaying facilities located around a vehicle using three-dimensional map data is known (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0006]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0007] In the encoded three-dimensional data (a plurality of three-dimensional points), among the plurality of encoded three-dimensional points, it is desired that a plurality of desired encoded three-dimensional points can be appropriately selected and decoded.

[0008] The present disclosure aims to provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device that can appropriately select and decode a plurality of desired encoded three-dimensional points among the plurality of encoded three-dimensional points.

Means for Solving the Problems

[0009] A three-dimensional data encoding method according to an aspect of the present disclosure is an encoding method executed by an encoding device, which encodes a plurality of three-dimensional points each located in one of a plurality of regions for each region, and connection information generated based on a relationship between a predetermined region among the plurality of regions and a plurality of other regions other than the predetermined region among the plurality of regions, the connection information including: (i) tile information indicating values uniquely assigned to each of the plurality of regions, and (ii) association information indicating that there is an association between the predetermined region and the other regions by the tile information, and generates a bitstream including the generated connection information and the plurality of encoded three-dimensional points, the relationship being a relationship based on visibility with respect to a viewpoint within the predetermined region.

[0010] The three-dimensional data decoding method according to one aspect of the present disclosure is a decoding method executed by a decoding device, and includes obtaining a bit stream including a plurality of three-dimensional points each located in one of a plurality of regions and encoded for each region, and obtaining connection information generated based on a relationship between a predetermined region among the plurality of regions and a plurality of other regions other than the predetermined region among the plurality of regions, the connection information including (i) tile information indicating values uniquely assigned to each of the plurality of regions, and (ii) association information indicating that there is an association between the predetermined region and the other regions based on the tile information, decoding the plurality of encoded three-dimensional points based on the obtained connection information, and the relationship being a relationship based on visibility with respect to a viewpoint within the predetermined region.

Advantages of the Invention

[0011] The present disclosure can provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device that can appropriately select and decode a desired plurality of encoded three-dimensional points among the plurality of encoded three-dimensional points.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Figure 36

Figure 37

Figure 38

Figure 39

Figure 40

Figure 41

Figure 42

Figure 43

Figure 44

Figure 45

Figure 46

Figure 47

Figure 48

Figure 49

Figure 50

Figure 51

Figure 52

Figure 53

Figure 54

Figure 55

Figure 56

Figure 57

Figure 58

Figure 59

Figure 60

Figure 61

Figure 62

Figure 63

Figure 64

Figure 65

Figure 66

Figure 67

Figure 68

Figure 69

Figure 70

Figure 71

Figure 72

Figure 73

Figure 74

Figure 75

Figure 76

Figure 77

Figure 78

Figure 79

Figure 80

Figure 81

Figure 82

Figure 83

Figure 84

Figure 85

Figure 86

Figure 87

Figure 88

Figure 89

Figure 90

Figure 91

Figure 92

Figure 93

Figure 94

Figure 95

Figure 96

Figure 97

Figure 98

Figure 99

Figure 100

Figure 101

Figure 102

Figure 103

Figure 104

Figure 105

Figure 106

Figure 107

Figure 108

Figure 109

Figure 110

Figure 111

Figure 112

Figure 113

Figure 114

Figure 115

Figure 116

Figure 117

Figure 118

Figure 119

Figure 120

Figure 121

Figure 122

Figure 123

Figure 124

Figure 125

Figure 126

Figure 127

Figure 128

Figure 129

Figure 130

Figure 131

Figure 132

Figure 133

Figure 134

Figure 135

Figure 136

Figure 137

Figure 138

Figure 139

Figure 140

Figure 141

Figure 142

Figure 143

Figure 144

Figure 145

Figure 146

Figure 147

Figure 148

Figure 149

Figure 150

Figure 151

Figure 152

Figure 153

【Figure 15FIG. 154 is a diagram showing the configuration of the server and the client device according to Embodiment 8. ​ FIG. 155 is a diagram showing the configuration of the server and the client device according to Embodiment 8. ​ FIG. 156 is a flowchart of the processing by the client device according to Embodiment 8. ​ FIG. 157 is a diagram showing the configuration of the sensor information collection system according to Embodiment 8. ​ FIG. 158 is a diagram showing an example of the system according to Embodiment 8. ​ FIG. 159 is a diagram showing a modified example of the system according to Embodiment 8. ​ FIG. 160 is a flowchart showing an example of the application processing according to Embodiment 8. ​ FIG. 161 is a diagram showing the sensor ranges of various sensors according to Embodiment 8. ​ FIG. 162 is a diagram showing a configuration example of the automatic driving system according to Embodiment 8. ​ FIG. 163 is a diagram showing a configuration example of the bit stream according to Embodiment 8. ​ FIG. 164 is a flowchart of the point cloud selection process according to Embodiment 8. ​ FIG. 165 is a diagram showing an example of the screen of the point cloud selection process according to Embodiment 8. ​ FIG. 166 is a diagram showing an example of the screen of the point cloud selection process according to Embodiment 8. ​ FIG. 167 is a diagram showing an example of the screen of the point cloud selection process according to Embodiment 8. ​ FIG. 168 is a diagram for explaining a first example of the connectivity determination method according to Embodiment 9. ​ FIG. 169 is a diagram for explaining a second example of the connectivity determination method according to Embodiment 9. ​FIG. 170 is a block diagram for explaining an example of a method for a three-dimensional data decoding apparatus to determine connectivity and / or the strength of connectivity. ​ FIG. 171 is a block diagram for explaining another example of a method for a three-dimensional data decoding apparatus to determine connectivity and / or the strength of connectivity. ​ FIG. 172 is a block diagram showing the configuration of the three-dimensional data encoding apparatus according to Embodiment 9. ​ FIG. 173 is a block diagram showing the configuration of the three-dimensional data decoding apparatus. ​ FIG. 174 is a diagram for explaining the signaling of additional information including connection information. ​ FIG. 175 is a diagram for explaining a first example of another determination method for connectivity according to Embodiment 9. ​ FIG. 176 is a diagram for explaining another example of the first example of another determination method for connectivity according to Embodiment 9. ​ FIG. 177 is a diagram showing a syntax example of connection information in the example shown in FIG. 176. ​ FIG. 178 is a diagram for explaining a second example of another determination method for connectivity according to Embodiment 9. ​ FIG. 179 is a diagram showing a syntax example of connection information in the example shown in FIG. 178. ​ FIG. 180 is a diagram for explaining a third example of another determination method for connectivity according to Embodiment 9. ​ FIG. 181 is a diagram showing a first example of the container type. ​ FIG. 182 is a diagram showing a second example of the container type. ​ FIG. 183 is a diagram showing a third example of the container type. ​ FIG. 184 is a diagram showing a fourth example of the container type.

Figure 185

Figure 186

Figure 187

Figure 188

Figure 189

Figure 190

Figure 191

Figure 192

Figure 193

Figure 194

Figure 195

Figure 196

Figure 197

Figure 198

Figure 199

Figure 200

Figure 201

Figure 202

Figure 203

Figure 204

Figure 205

Figure 206

Figure 207

[0013] A three-dimensional data encoding method according to an aspect of the present disclosure encodes a plurality of three-dimensional points each located in one of a plurality of regions for each region, and generates connection information based on a relationship between a predetermined region among the plurality of regions and a plurality of other regions other than the predetermined region among the plurality of regions, the connection information including: (i) tile information indicating a value uniquely assigned to each of the plurality of regions; and (ii) association information indicating that there is an association between the predetermined region and the other regions by the tile information, and generates a bitstream including the generated connection information and the encoded plurality of three-dimensional points.

[0014] For example, a three-dimensional data decoding device sequentially decodes a plurality of three-dimensional points encoded in the order of data included in a bit stream. Here, for example, when the three-dimensional data represents a three-dimensional map, when a user views the map, the user often wants to view the map around the center with the user's current position as the center. In such a case, if the three-dimensional data decoding device sequentially decodes a plurality of three-dimensional points encoded in the order of data included in the bit stream and causes a display device or the like to sequentially display an image showing the plurality of three-dimensional points on the display device in the decoded order, it may take time until the location that the user wants to confirm is displayed on the display device. Also, depending on the user's current position, there may be three-dimensional points that do not need to be decoded. Therefore, the three-dimensional data encoding device generates connection information including tile information indicating values uniquely assigned to each of the plurality of regions based on the relationship between the plurality of regions, and association information indicating that a predetermined region and other regions are related by the tile information. According to this, for example, when the three-dimensional data decoding device receives information indicating the user's current position from a device owned by the user, it determines a predetermined region based on the information, and sequentially decodes three-dimensional points encoded based on the connection information from the determined predetermined region. According to this, the three-dimensional data decoding device can sequentially decode from the encoded three-dimensional points located in the region where the user is likely to desire. That is, according to the three-dimensional data encoding method according to the present disclosure, among the plurality of encoded three-dimensional points, a desired plurality of encoded three-dimensional points can be appropriately selected and decoded.

[0015] Also, for example, the relationship is a positional relationship between a predetermined region among the plurality of regions and the plurality of other regions other than the predetermined region among the plurality of regions.

[0016] Also, for example, in the generation of the connection information, the connection information including the association information indicating to decode the encoded three-dimensional points located in other regions that are in contact with or overlap the predetermined region among the plurality of other regions is generated.

[0017] For example, when the three-dimensional data represents a three-dimensional map, other regions that are in contact with or overlap a predetermined region are likely to be a continuation from the map of the predetermined region. Therefore, according to this, the three-dimensional data decoding device can more appropriately select and decode a plurality of desired encoded three-dimensional points among the plurality of encoded three-dimensional points.

[0018] Also, for example, in the generation of the connection information, the connection information including the related information indicating decoding of the encoded three-dimensional points located in other regions located in the direction as viewed from the predetermined region based on the direction information indicating the direction from the predetermined region is generated.

[0019] For example, when the three-dimensional data represents a three-dimensional map, even in other regions that are not in contact with or overlap a predetermined region corresponding to the user's current position, there may be regions that the user wants to quickly check. For example, when the user is moving, the user is likely to want to check the map of the region located in the traveling direction. Therefore, according to this, since it is possible to determine whether to decode a plurality of three-dimensional points encoded according to the direction, the three-dimensional data decoding device can more appropriately select and decode a plurality of desired encoded three-dimensional points among the plurality of encoded three-dimensional points.

[0020] Also, for example, in the generation of the connection information, the connection information including the related information indicating that the closer the distance to the predetermined region of the other regions, the earlier the order of the regions for decoding the encoded three-dimensional points among the plurality of other regions is generated.

[0021] The closer a region is to a predetermined region, the more likely it is to be related to the predetermined region. Therefore, according to this, the three-dimensional data decoding device can more appropriately select and decode a plurality of desired encoded three-dimensional points among the plurality of encoded three-dimensional points.

[0022] Further, for example, in generating the connection information, based on the relationship, it is determined to which of a plurality of predetermined groups each of the plurality of other regions belongs, and the connection information including group information indicating the determined predetermined group is generated.

[0023] For example, when the three-dimensional data represents a three-dimensional map, depending on the current position of the user, even if the regions are different from each other among the plurality of regions, there may be cases where the priorities desired by the user are about the same. In such a case, for example, if there is information grouping regions with about the same priority desired by the user, the three-dimensional data decoding device can determine whether to decode / not decode the three-dimensional points encoded for each group, or the order of decoding. Therefore, according to this, the processing amount for the three-dimensional data decoding device to determine whether to decode / not decode the encoded three-dimensional points or the order of decoding can be reduced.

[0024] Further, a three-dimensional data decoding method according to an aspect of the present disclosure obtains a bitstream including a plurality of three-dimensional points each located in one of a plurality of regions, the plurality of three-dimensional points encoded for each region, and connection information generated based on the relationship between a predetermined region among the plurality of regions and a plurality of other regions other than the predetermined region among the plurality of regions, the connection information including: (i) tile information indicating values uniquely assigned to each of the plurality of regions, and (ii) association information indicating that there is an association between the predetermined region and the other regions by the tile information, and selectively decodes the plurality of encoded three-dimensional points for each region based on the obtained connection information.

[0025] For example, a three-dimensional data decoder sequentially decodes a plurality of three-dimensional points encoded in the order of data included in a bit stream. Here, for example, when the three-dimensional data represents a three-dimensional map, when a user views the map, the user often wants to view the map around the center with the user's current position as the center. In such a case, if the three-dimensional data decoder sequentially decodes a plurality of three-dimensional points encoded in the order of data included in the bit stream and causes a display device or the like to sequentially display an image showing the plurality of three-dimensional points in the decoded order on the display device, it may take time until the location that the user wants to confirm is displayed on the display device. Also, depending on the user's current position, there may be three-dimensional points that do not need to be decoded. Therefore, the three-dimensional data encoding device generates connection information including tile information indicating values uniquely assigned to each of the plurality of regions based on the relationship between the plurality of regions, and connection information indicating that a predetermined region and other regions are related by the tile information. According to this, for example, when the three-dimensional data decoder receives information indicating the user's current position from a device owned by the user, the three-dimensional data decoder determines a predetermined region based on the information and can sequentially decode three-dimensional points encoded based on the connection information from the determined predetermined region. According to this, the three-dimensional data decoder can sequentially decode from the encoded three-dimensional points located in the region where the user is likely to desire. That is, according to the three-dimensional data decoding method according to the present disclosure, among the plurality of encoded three-dimensional points, a plurality of desired encoded three-dimensional points can be appropriately selected and decoded.

[0026] Also, for example, the relationship is a positional relationship between a predetermined region among the plurality of regions and the plurality of other regions other than the predetermined region among the plurality of regions.

[0027] Also, for example, in the acquisition of the connection information, the connection information including the relevant information indicating to decode the encoded three-dimensional points located in the other regions that are in contact with or overlap the predetermined region among the plurality of other regions is acquired.

[0028] For example, when the three-dimensional data represents a three-dimensional map, other regions that are in contact with or overlap a predetermined region are likely to be a continuation from the map of the predetermined region. Therefore, according to this, among the plurality of encoded three-dimensional points, the desired plurality of encoded three-dimensional points can be selected and decoded more appropriately.

[0029] Also, for example, in the acquisition of the connection information, the connection information is obtained, which includes the related information generated based on the orientation information indicating the orientation from the predetermined region, and the related information indicating decoding of the encoded three-dimensional points located in other regions located in the orientation as viewed from the predetermined region.

[0030] For example, when the three-dimensional data represents a three-dimensional map, even in other regions that are not in contact with or overlap a predetermined region corresponding to the current position of the user, there may be regions that the user wants to check quickly. For example, when the user is moving, the user is likely to want to check the map of the region located in the traveling direction. Therefore, according to this, it is possible to determine whether to decode a plurality of three-dimensional points encoded according to the orientation, so that among the plurality of encoded three-dimensional points, the desired plurality of encoded three-dimensional points can be selected and decoded more appropriately.

[0031] Also, for example, in the acquisition of the connection information, the connection information is obtained, which includes the related information indicating that among the plurality of other regions, the closer the distance to the predetermined region, the earlier the order of the region for decoding the encoded three-dimensional points.

[0032] Regions closer to the predetermined region are more likely to be related to the predetermined region. Therefore, according to this, among the plurality of encoded three-dimensional points, the desired plurality of encoded three-dimensional points can be selected and decoded more appropriately.

[0033] Also, for example, in the acquisition of the connection information, the connection information is obtained, which includes group information indicating which of the plurality of predetermined groups each of the plurality of other regions belongs to based on the relationship.

[0034] For example, when the three-dimensional data represents a three-dimensional map, depending on the user's current position, even if the regions are different from each other among a plurality of regions, the priorities desired by the user may be the same. In such a case, for example, if there is information grouping the regions with the same priority desired by the user, it is possible to determine whether to decode / not decode the three-dimensional points encoded for each group, or the order of decoding. Therefore, according to this, it is possible to reduce the processing amount for determining whether to decode / not decode the encoded three-dimensional points, or for determining the decoding order.

[0035] Also, for example, the connection information is included in the bit stream, and in obtaining the connection information, the connection information included in the bit stream is obtained.

[0036] According to this, it is possible to determine whether to decode / not decode the three-dimensional points encoded using the connection information included in the bit stream, or the decoding order.

[0037] Also, for example, in obtaining the connection information, the connection information is obtained by generating the connection information based on a plurality of encoded three-dimensional points included in the bit stream.

[0038] According to this, even when the connection information is not included in the bit stream, it is possible to determine whether to decode / not decode the encoded three-dimensional points, or the decoding order.

[0039] In addition, a three-dimensional data encoding device according to one aspect of the present disclosure includes a processor and a memory. The processor uses the memory to encode a plurality of three-dimensional points, each located in one of a plurality of regions, for each region, and generates connection information based on the relationship between a predetermined region among the plurality of regions and a plurality of other regions other than the predetermined region among the plurality of regions. The connection information includes (i) tile information indicating values uniquely assigned to each of the plurality of regions, and (ii) association information indicating that there is an association between the predetermined region and the other regions by the tile information. The processor generates a bit stream including the generated connection information and the encoded plurality of three-dimensional points.

[0040] For example, a three-dimensional data decoding device sequentially decodes a plurality of three-dimensional points encoded in the order of data included in a bit stream. Here, for example, when the three-dimensional data represents a three-dimensional map, when a user views the map, the user often wants to view the map around the center with the user's current position as the center. In such a case, assume that the three-dimensional data decoding device sequentially decodes a plurality of three-dimensional points encoded in the order of data included in the bit stream, and causes a display device or the like to sequentially display, in the decoded order, an image showing the plurality of three-dimensional points on the display device. In this way, it may take time until the location that the user wants to check is displayed on the display device. Also, depending on the user's current position, there may be three-dimensional points that do not need to be decoded. Therefore, the three-dimensional data encoding device generates connection information including tile information indicating values uniquely assigned to each of the plurality of regions based on the relationship of the plurality of regions, and connection information indicating that a predetermined region is related to other regions by the tile information. According to this, for example, when the three-dimensional data decoding device receives information indicating the user's current position from a device owned by the user, the three-dimensional data decoding device determines a predetermined region based on the information, and sequentially decodes three-dimensional points encoded based on the connection information from the determined predetermined region. According to this, the three-dimensional data decoding device can sequentially decode from the encoded three-dimensional points located in the region where the user is likely to desire. That is, according to the three-dimensional data encoding device according to the present disclosure, among the plurality of encoded three-dimensional points, a desired plurality of encoded three-dimensional points can be appropriately selected and decoded.

[0041] In addition, a three-dimensional data decoding apparatus according to an aspect of the present disclosure includes a processor and a memory. The processor uses the memory to acquire a bit stream including a plurality of three-dimensional points encoded for each region, and connection information generated based on the relationship between a predetermined region among the plurality of regions and a plurality of other regions other than the predetermined region among the plurality of regions. The connection information includes (i) tile information indicating values uniquely assigned to each of the plurality of regions, and (ii) association information indicating that there is an association between the predetermined region and the other regions by the tile information. Based on the acquired connection information, the plurality of encoded three-dimensional points are selectively decoded for each region.

[0042] For example, a three-dimensional data decoding device sequentially decodes a plurality of three-dimensional points encoded in the order of data included in a bit stream. Here, for example, when the three-dimensional data represents a three-dimensional map, when a user views the map, the user often wants to view the map around the center with the user's current position as the center. In such a case, if the three-dimensional data decoding device sequentially decodes a plurality of three-dimensional points encoded in the order of data included in the bit stream and causes a display device or the like to sequentially display, in the decoded order, images showing the plurality of three-dimensional points on the display device, it may take time until the location that the user wants to check is displayed on the display device. Also, depending on the user's current position, there may be three-dimensional points that do not need to be decoded. Therefore, the three-dimensional data encoding device generates connection information including tile information indicating values uniquely assigned to each of a plurality of regions based on the relationship between the plurality of regions, and connection information indicating that there is a relevance between a predetermined region and other regions based on the tile information. According to this, the three-dimensional data decoding device, for example, when receiving information indicating the user's current position from a device owned by the user, determines a predetermined region based on the information, and can sequentially decode three-dimensional points encoded based on the connection information from the determined predetermined region. According to this, the three-dimensional data decoding device can sequentially decode from the encoded three-dimensional points located in the region that the user is likely to desire. That is, according to the three-dimensional data decoding device according to the present disclosure, among the plurality of encoded three-dimensional points, a desired plurality of encoded three-dimensional points can be appropriately selected and decoded.

[0043] Note that these general or specific aspects may be implemented by a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.

[0044] Hereinafter, embodiments will be specifically described with reference to the drawings. Note that all the embodiments described below are specific examples of the present disclosure. The numerical values, shapes, materials, components, arrangement positions and connection forms of the components, steps, order of steps, etc. shown in the following embodiments are merely examples and are not intended to limit the present disclosure. In addition, among the components in the following embodiments, the components not described in the independent claims indicating the most general concept are described as optional components.

[0045] (Embodiment 1) When using the encoded data of the point cloud in an actual device or service, it is desirable to transmit and receive the information necessary according to the application in order to suppress the network bandwidth. However, until now, such a function has not existed in the encoding structure of three-dimensional data, and there has also been no encoding method therefor.

[0046] In the present embodiment, a three-dimensional data encoding method and a three-dimensional data encoding device for providing a function of transmitting and receiving information necessary according to the application in the encoded data of a three-dimensional point cloud, a three-dimensional data decoding method and a three-dimensional data decoding device for decoding the encoded data, a three-dimensional data multiplexing method for multiplexing the encoded data, and a three-dimensional data transmission method for transmitting the encoded data will be described.

[0047] In particular, currently, as encoding methods (encoding schemes) for point cloud data, a first encoding method and a second encoding method are being studied, but the configuration of the encoded data and the method of storing the encoded data in the system format are not defined, and there is a problem that MUX processing (multiplexing), transmission, or storage cannot be performed in the encoding unit as it is.

[0048] Also, a method for supporting a format in which two codecs of a first encoding method and a second encoding method are mixed, such as PCC (Point Cloud Compression), has not existed until now.

[0049] In this embodiment, the configuration of PCC encoded data in which two codecs, a first encoding method and a second encoding method, are mixed, and a method for storing the encoded data in a system format will be described.

[0050] First, the configuration of a three-dimensional data (point cloud data) encoding / decoding system according to this embodiment will be described. FIG. 1 is a diagram showing a configuration example of a three-dimensional data encoding / decoding system according to this embodiment. As shown in FIG. 1, the three-dimensional data encoding / decoding system includes a three-dimensional data encoding system 4601, a three-dimensional data decoding system 4602, a sensor terminal 4603, and an external connection unit 4604.

[0051] The three-dimensional data encoding system 4601 generates encoded data or multiplexed data by encoding point cloud data, which is three-dimensional data. Note that the three-dimensional data encoding system 4601 may be a three-dimensional data encoding device realized by a single device, or may be a system realized by a plurality of devices. Further, the three-dimensional data encoding device may include a part of a plurality of processing units included in the three-dimensional data encoding system 4601.

[0052] The three-dimensional data encoding system 4601 includes a point cloud data generation system 4611, a presentation unit 4612, an encoding unit 4613, a multiplexing unit 4614, an input / output unit 4615, and a control unit 4616. The point cloud data generation system 4611 includes a sensor information acquisition unit 4617 and a point cloud data generation unit 4618.

[0053] The sensor information acquisition unit 4617 acquires sensor information from the sensor terminal 4603 and outputs the sensor information to the point cloud data generation unit 4618. The point cloud data generation unit 4618 generates point cloud data from the sensor information and outputs the point cloud data to the encoding unit 4613.

[0054] The presentation unit 4612 presents sensor information or point cloud data to the user. For example, the presentation unit 4612 displays information or an image based on the sensor information or point cloud data.

[0055] The symbolization unit 4613 encodes (compresses) the point cloud data, and outputs the obtained encoded data, the control information obtained in the encoding process, and other additional information to the multiplexing unit 4614. The additional information includes, for example, sensor information.

[0056] The multiplexing unit 4614 generates multiplexed data by multiplexing the encoded data, the control information, and the additional information input from the symbolization unit 4613. The format of the multiplexed data is, for example, a file format for storage or a packet format for transmission.

[0057] The input / output unit 4615 (for example, a communication unit or an interface) outputs the multiplexed data to the outside. Alternatively, the multiplexed data is stored in a storage unit such as an internal memory. The control unit 4616 (or the application execution unit) controls each processing unit. That is, the control unit 4616 performs control such as encoding and multiplexing.

[0058] Note that the sensor information may be input to the symbolization unit 4613 or the multiplexing unit 4614. Also, the input / output unit 4615 may output the point cloud data or the encoded data as it is to the outside.

[0059] The transmission signal (multiplexed data) output from the three-dimensional data encoding system 4601 is input to the three-dimensional data decoding system 4602 via the external connection unit 4604.

[0060] The three-dimensional data decoding system 4602 generates point cloud data, which is three-dimensional data, by decoding the encoded data or the multiplexed data. Note that the three-dimensional data decoding system 4602 may be a three-dimensional data decoding device realized by a single device, or may be a system realized by a plurality of devices. Also, the three-dimensional data decoding device may include a part of a plurality of processing units included in the three-dimensional data decoding system 4602.

[0061] The three-dimensional data decoding system 4602 includes a sensor information acquisition unit 4621, an input / output unit 4622, a de-multiplexing unit 4623, a decoding unit 4624, a presentation unit 4625, a user interface 4626, and a control unit 4627.

[0062] The sensor information acquisition unit 4621 acquires sensor information from the sensor terminal 4603.

[0063] The input / output unit 4622 acquires a transmission signal, decodes multiplexed data (file format or packet) from the transmission signal, and outputs the multiplexed data to the de-multiplexing unit 4623.

[0064] The de-multiplexing unit 4623 acquires encoded data, control information, and additional information from the multiplexed data, and outputs the encoded data, control information, and additional information to the decoding unit 4624.

[0065] The decoding unit 4624 reconstructs point cloud data by decoding the encoded data.

[0066] The presentation unit 4625 presents the point cloud data to the user. For example, the presentation unit 4625 displays information or an image based on the point cloud data. The user interface 4626 acquires an instruction based on the user's operation. The control unit 4627 (or application execution unit) controls each processing unit. That is, the control unit 4627 performs control such as de-multiplexing, decoding, and presentation.

[0067] Note that the input / output unit 4622 may directly acquire point cloud data or encoded data from the outside. Also, the presentation unit 4625 may acquire additional information such as sensor information and present information based on the additional information. Further, the presentation unit 4625 may perform presentation based on the user's instruction acquired by the user interface 4626.

[0068] The sensor terminal 4603 generates sensor information, which is information obtained by the sensor. The sensor terminal 4603 is a terminal equipped with a sensor or a camera, and examples include a moving body such as an automobile, a flying object such as an airplane, a mobile terminal, or a camera.

[0069] The sensor information that can be acquired by the sensor terminal 4603 is, for example, (1) the distance between the sensor terminal 4603 and the object, or the reflectivity of the object, obtained from a LIDAR, millimeter-wave radar, or infrared sensor, (2) the distance between the camera and the object or the reflectivity of the object obtained from a plurality of monocular camera images or stereo camera images, etc. Further, the sensor information may include the attitude, orientation, gyro (angular velocity), position (GPS information or altitude), speed, or acceleration, etc. of the sensor. Further, the sensor information may include the temperature, atmospheric pressure, humidity, or magnetism, etc.

[0070] The external connection unit 4604 is realized by an integrated circuit (LSI or IC), an external storage unit, communication with a cloud server via the Internet, or broadcasting, etc.

[0071] Next, the point cloud data will be described. FIG. 2 is a diagram showing the configuration of the point cloud data. FIG. 3 is a diagram showing a configuration example of a data file in which the information of the point cloud data is described.

[0072] The point cloud data includes data of a plurality of points. The data of each point includes position information (three-dimensional coordinates) and attribute information for the position information. A collection of such points is called a point cloud. For example, the point cloud shows the three-dimensional shape of an object.

[0073] Position information (Position) such as three-dimensional coordinates may also be called geometry. Further, the data of each point may include attribute information (attribute) of a plurality of attribute types. The attribute types are, for example, color or reflectivity, etc.

[0074] One piece of attribute information may be associated with one piece of position information, or attribute information having a plurality of different attribute types may be associated with one piece of position information. Further, a plurality of pieces of attribute information of the same attribute type may be associated with one piece of position information.

[0075] The configuration example of the data file shown in FIG. 3 is an example where the position information and the attribute information correspond one-to-one, and shows the position information and the attribute information of N points constituting the point cloud data.

[0076] The position information is, for example, information on three axes of x, y, and z. The attribute information is, for example, color information of RGB. As a typical data file, there is a ply file or the like.

[0077] Next, the types of point cloud data will be described. FIG. 4 is a diagram showing the types of point cloud data. As shown in FIG. 4, the point cloud data includes a static object and a dynamic object.

[0078] The static object is three-dimensional point cloud data at an arbitrary time (a certain moment). The dynamic object is three-dimensional point cloud data that changes over time. Hereinafter, the three-dimensional point cloud data at a certain moment is referred to as a PCC frame or a frame.

[0079] The object may be a point cloud with a limited area like ordinary video data, or a large-scale point cloud with an unlimited area like map information.

[0080] Also, there is point cloud data with various densities, and there may be sparse point cloud data and dense point cloud data.

[0081] Hereinafter, the details of each processing unit will be described. The sensor information is acquired by various methods such as a distance sensor such as LIDAR or a range finder, a stereo camera, or a combination of a plurality of monocular cameras. The point cloud data generation unit 4618 generates point cloud data based on the sensor information obtained by the sensor information acquisition unit 4617. The point cloud data generation unit 4618 generates position information as the point cloud data, and adds attribute information for the position information to the position information.

[0082] When generating position information or adding attribute information, the point cloud data generation unit 4618 may process the point cloud data. For example, the point cloud data generation unit 4618 may reduce the data volume by deleting overlapping point clouds. In addition, the point cloud data generation unit 4618 may convert the position information (such as position shift, rotation, or normalization), or may render the attribute information.

[0083] Note that in FIG. 1, the point cloud data generation system 4611 is included in the three-dimensional data encoding system 4601, but may be provided independently outside the three-dimensional data encoding system 4601.

[0084] The encoding unit 4613 generates encoded data by encoding the point cloud data based on a predefined encoding method. There are roughly two types of encoding methods as follows. The first is an encoding method using position information, which will be described as the first encoding method hereinafter. The second is an encoding method using a video codec, which will be described as the second encoding method hereinafter.

[0085] The decoding unit 4624 decodes the point cloud data by decoding the encoded data based on a predefined encoding method.

[0086] The multiplexing unit 4614 generates multiplexed data by multiplexing the encoded data using an existing multiplexing method. The generated multiplexed data is transmitted or stored. In addition to the PCC encoded data, the multiplexing unit 4614 multiplexes other media such as video, audio, subtitles, applications, files, or reference time information. Further, the multiplexing unit 4614 may further multiplex sensor information or attribute information related to the point cloud data.

[0087] Examples of the multiplexing method or file format include ISOBMFF, MPEG-DASH which is an ISOBMFF-based transmission method, MMT, MPEG-2 TS Systems, RMP, etc.

[0088] The inverse multiplexing unit 4623 extracts PCC-encoded data, other media, time information, etc. from the multiplexed data.

[0089] The input / output unit 4615 transmits the multiplexed data using a method suitable for the medium for transmission such as broadcasting or communication or the medium for storage. The input / output unit 4615 may communicate with other devices via the Internet or may communicate with a storage unit such as a cloud server.

[0090] As the communication protocol, http, ftp, TCP, UDP, etc. are used. A PULL-type communication method may be used or a PUSH-type communication method may be used.

[0091] Either wired transmission or wireless transmission may be used. As the wired transmission, Ethernet (registered trademark), USB, RS-232C, HDMI (registered trademark), or coaxial cable, etc. are used. As the wireless transmission, wireless LAN, Wi-Fi (registered trademark), Bluetooth (registered trademark), or millimeter wave, etc. are used.

[0092] Also, as the broadcasting method, for example, DVB-T2, DVB-S2, DVB-C2, ATSC 3.0, or ISDB-S3, etc. are used.

[0093] FIG. 5 is a diagram showing the configuration of a first encoding unit 4630 which is an example of the encoding unit 4613 that performs encoding according to the first encoding method. FIG. 6 is a block diagram of the first encoding unit 4630. The first encoding unit 4630 generates encoded data (encoded stream) by encoding the point cloud data according to the first encoding method. This first encoding unit 4630 includes a position information encoding unit 4631, an attribute information encoding unit 4632, an additional information encoding unit 4633, and a multiplexing unit 4634.

[0094] The first encoding unit 4630 has the feature of performing encoding while being aware of the three-dimensional structure. Also, the first encoding unit 4630 has the feature that the attribute information encoding unit 4632 performs encoding using the information obtained from the position information encoding unit 4631. The first encoding method is also called GPCC (Geometry based PCC).

[0095] The point cloud data is PCC point cloud data such as a PLY file or PCC point cloud data generated from sensor information, and includes position information (Position), attribute information (Attribute), and other additional information (MetaData). The position information is input to the position information encoding unit 4631, the attribute information is input to the attribute information encoding unit 4632, and the additional information is input to the additional information encoding unit 4633.

[0096] The position information encoding unit 4631 generates encoded position information (Compressed Geometry), which is encoded data, by encoding the position information. For example, the position information encoding unit 4631 encodes the position information using an N-ary tree structure such as an octree. Specifically, in an octree, the target space is divided into eight nodes (sub-spaces), and 8-bit information (occupancy code) indicating whether each node contains a point cloud is generated. Also, the node containing the point cloud is further divided into eight nodes, and 8-bit information indicating whether each of the eight nodes contains a point cloud is generated. This process is repeated until it reaches a threshold of the number of point clouds included in a predetermined hierarchy or node.

[0097] The attribute information encoding unit 4632 generates encoded attribute information (Compressed Attribute), which is encoded data, by encoding using the configuration information generated by the position information encoding unit 4631. For example, the attribute information encoding unit 4632 determines a reference point (reference node) to be referred to in the encoding of a target point (target node) to be processed based on the octree structure generated by the position information encoding unit 4631. For example, the attribute information encoding unit 4632 refers to a node among the peripheral nodes or adjacent nodes whose parent node in the octree is the same as the target node. Note that the method for determining the reference relationship is not limited to this.

[0098] Also, the encoding process of the attribute information may include at least one of quantization processing, prediction processing, and arithmetic encoding processing. In this case, reference means using the reference node to calculate the predicted value of the attribute information, or using the state of the reference node (for example, occupancy information indicating whether the reference node includes a point cloud) to determine the encoding parameter. For example, the encoding parameter is a quantization parameter in quantization processing or a context in arithmetic encoding.

[0099] The additional information encoding unit 4633 generates encoded additional information (Compressed MetaData), which is encoded data, by encoding compressible data among the additional information.

[0100] The multiplexing unit 4634 generates an encoded stream (Compressed Stream), which is encoded data, by multiplexing the encoded position information, encoded attribute information, encoded additional information, and other additional information. The generated encoded stream is output to a processing unit in a system layer (not shown).

[0101] Next, a first decoding unit 4640, which is an example of a decoding unit 4624 that decodes using the first encoding method, will be described. FIG. 7 is a diagram showing the configuration of the first decoding unit 4640. FIG. 8 is a block diagram of the first decoding unit 4640. The first decoding unit 4640 generates point cloud data by decoding encoded data (encoded stream) encoded by the first encoding method using the first encoding method. This first decoding unit 4640 includes a demultiplexing unit 4641, a position information decoding unit 4642, an attribute information decoding unit 4643, and an additional information decoding unit 4644.

[0102] An encoded stream (Compressed Stream), which is encoded data, is input to the first decoding unit 4640 from a processing unit in a system layer (not shown).

[0103] The demultiplexing unit 4641 separates encoded position information (Compressed Geometry), encoded attribute information (Compressed Attribute), encoded additional information (Compressed MetaData), and other additional information from the encoded data.

[0104] The position information decoding unit 4642 generates position information by decoding the encoded position information. For example, the position information decoding unit 4642 restores the position information of the point cloud represented by three-dimensional coordinates from the encoded position information represented by an N-ary tree structure such as an octree.

[0105] The attribute information decoding unit 4643 decodes the encoded attribute information based on the configuration information generated by the position information decoding unit 4642. For example, the attribute information decoding unit 4643 determines a reference point (reference node) to be referred to in decoding a target point (target node) to be processed based on the octree structure obtained by the position information decoding unit 4642. For example, the attribute information decoding unit 4643 refers to a node in which the parent node in the octree is the same as the target node among the surrounding nodes or adjacent nodes. Note that the method for determining the reference relationship is not limited to this.

[0106] Further, the decoding process of the attribute information may include at least one of inverse quantization processing, prediction processing, and arithmetic decoding processing. In this case, reference means using a reference node to calculate a predicted value of the attribute information, or using the state of the reference node (for example, occupancy information indicating whether a point group is included in the reference node) to determine decoding parameters. For example, the decoding parameters are quantization parameters in inverse quantization processing or contexts in arithmetic decoding, etc.

[0107] The additional information decoding unit 4644 generates additional information by decoding the encoded additional information. Also, the first decoding unit 4640 uses the additional information required for the decoding process of the position information and the attribute information during decoding and outputs the additional information required for the application to the outside.

[0108] Next, a configuration example of the position information encoding unit will be described. FIG. 9 is a block diagram of the position information encoding unit 2700 according to the present embodiment. The position information encoding unit 2700 includes an octree generation unit 2701, a geometric information calculation unit 2702, an encoding table selection unit 2703, and an entropy encoding unit 2704.

[0109] The octree generation unit 2701 generates, for example, an octree from the input position information and generates an occupancy code for each node of the octree. The geometric information calculation unit 2702 acquires information indicating whether an adjacent node of the target node is an occupied node. For example, the geometric information calculation unit 2702 calculates occupancy information (information indicating whether an adjacent node is an occupied node) of the adjacent node from the occupancy code of the parent node to which the target node belongs. Also, the geometric information calculation unit 2702 may save the encoded nodes in a list and search for adjacent nodes from within that list. Note that the geometric information calculation unit 2702 may switch adjacent nodes according to the position within the parent node of the target node.

[0110] The symbolic table selection unit 2703 selects a coding table to be used for entropy coding of the target node by using the occupancy information of adjacent nodes calculated by the geometric information calculation unit 2702. For example, the symbolic table selection unit 2703 may generate a bit string by using the occupancy information of adjacent nodes, and select a coding table with an index number generated from the bit string.

[0111] The entropy coding unit 2704 generates coding position information and metadata by performing entropy coding on the occupancy code of the target node by using the coding table with the selected index number. The entropy coding unit 2704 may add information indicating the selected coding table to the coding position information.

[0112] Hereinafter, the octree representation and the scanning order of position information will be described. The position information (position data) is coded after being converted (octreed) into an octree structure. The octree structure is composed of nodes and leaves. Each node has eight nodes or leaves, and each leaf has voxel (VXL) information. FIG. 10 is a diagram showing an example of the structure of position information including a plurality of voxels. FIG. 11 is a diagram showing an example of converting the position information shown in FIG. 10 into an octree structure. Here, among the leaves shown in FIG. 11, leaves 1, 2, and 3 represent voxels VXL1, VXL2, and VXL3 shown in FIG. 10, respectively, and represent VXLs (hereinafter, valid VXLs) including point clouds.

[0113] Specifically, node 1 corresponds to the entire space including the position information in FIG. 10. The entire space corresponding to node 1 is divided into eight nodes, and among the eight nodes, the nodes including valid VXLs are further divided into eight nodes or leaves, and this process is repeated for the hierarchy of the tree structure. Here, each node corresponds to a subspace and has information (occupancy code) indicating at which position after division the next node or leaf is located as node information. Also, the lowermost block is set as a leaf, and the number of point clouds included in the leaf, etc. are held as leaf information.

[0114] Next, a configuration example of the position information decoding unit will be described. FIG. 12 is a block diagram of the position information decoding unit 2710 according to the present embodiment. The position information decoding unit 2710 includes an octree generation unit 2711, a geometric information calculation unit 2712, a coding table selection unit 2713, and an entropy decoding unit 2714.

[0115] The octree generation unit 2711 generates an octree of a certain space (node) using the header information or metadata of the bit stream. For example, the octree generation unit 2711 generates a large space (root node) using the sizes in the x-axis, y-axis, and z-axis directions of a certain space added to the header information, and divides the space into two parts in the x-axis, y-axis, and z-axis directions respectively to generate eight small spaces A (nodes A0 to A7), thereby generating an octree. Also, nodes A0 to A7 are sequentially set as target nodes.

[0116] The geometric information calculation unit 2712 acquires occupancy information indicating whether an adjacent node of the target node is an occupied node. For example, the geometric information calculation unit 2712 calculates the occupancy information of the adjacent node from the occupancy code of the parent node to which the target node belongs. Also, the geometric information calculation unit 2712 may save the decoded nodes in a list and search for adjacent nodes from within the list. Note that the geometric information calculation unit 2712 may switch adjacent nodes according to the position within the parent node of the target node.

[0117] The coding table selection unit 2713 selects a coding table (decoding table) to be used for entropy decoding of the target node using the occupancy information of the adjacent node calculated by the geometric information calculation unit 2712. For example, the coding table selection unit 2713 may generate a bit string using the occupancy information of the adjacent node and select the coding table of the index number generated from the bit string.

[0118] The entropy decoding unit 2714 generates position information by entropy-decoding the occupancy code of the target node using the selected encoding table. Note that the entropy decoding unit 2714 may decode and obtain the information of the selected encoding table from the bit stream, and entropy-decode the occupancy code of the target node using the encoding table indicated by the information.

[0119] Hereinafter, the configurations of the attribute information encoding unit and the attribute information decoding unit will be described. FIG. 13 is a block diagram showing a configuration example of the attribute information encoding unit A100. The attribute information encoding unit may include a plurality of encoding units that execute different encoding methods. For example, the attribute information encoding unit may switch and use the following two methods according to the use case.

[0120] The attribute information encoding unit A100 includes a LoD attribute information encoding unit A101 and a transformed attribute information encoding unit A102. The LoD attribute information encoding unit A101 classifies each three-dimensional point into a plurality of layers using the position information of the three-dimensional points, predicts the attribute information of the three-dimensional points belonging to each layer, and encodes the prediction residual. Here, each classified layer is called LoD (Level of Detail).

[0121] The transformed attribute information encoding unit A102 encodes the attribute information using RAHT (Region Adaptive Hierarchical Transform). Specifically, the transformed attribute information encoding unit A102 generates high-frequency components and low-frequency components of each layer by applying RAHT or Haar transform to each attribute information based on the position information of the three-dimensional points, and encodes their values using quantization, entropy encoding, etc.

[0122] FIG. 14 is a block diagram showing a configuration example of the attribute information decoding unit A110. The attribute information decoding unit may include a plurality of decoding units that execute different decoding methods. For example, the attribute information decoding unit may switch and decode based on the information included in the header and metadata using the following two methods.

[0123] The attribute information decoding unit A110 includes a LoD attribute information decoding unit A111 and a transformed attribute information decoding unit A112. The LoD attribute information decoding unit A111 classifies each three-dimensional point into a plurality of layers using the position information of the three-dimensional points, and decodes the attribute values while predicting the attribute information of the three-dimensional points belonging to each layer.

[0124] The transformed attribute information decoding unit A112 decodes the attribute information using RAHT (Region Adaptive Hierarchical Transform). Specifically, the transformed attribute information decoding unit A112 decodes the attribute values by applying inverse RAHT or inverse Haar transform to the high-frequency component and the low-frequency component of each attribute value based on the position information of the three-dimensional points.

[0125] FIG. 15 is a block diagram showing the configuration of an attribute information encoding unit 3140, which is an example of the LoD attribute information encoding unit A101.

[0126] The attribute information encoding unit 3140 includes a LoD generation unit 3141, a surrounding search unit 3142, a prediction unit 3143, a prediction residual calculation unit 3144, a quantization unit 3145, an arithmetic encoding unit 3146, an inverse quantization unit 3147, a decoded value generation unit 3148, and a memory 3149.

[0127] The LoD generation unit 3141 generates LoD using the position information of the three-dimensional points.

[0128] The surrounding search unit 3142 searches for neighboring three-dimensional points adjacent to each three-dimensional point using the LoD generation result by the LoD generation unit 3141 and distance information indicating the distance between each three-dimensional point.

[0129] The prediction unit 3143 generates a predicted value of the attribute information of the target three-dimensional point to be encoded.

[0130] The prediction residual calculation unit 3144 calculates (generates) the prediction residual of the predicted value of the attribute information generated by the prediction unit 3143.

[0131] The quantization unit 3145 quantizes the prediction residual of the attribute information calculated by the prediction residual calculation unit 3144.

[0132] The arithmetic coding unit 3146 arithmetically codes the prediction residual after quantization by the quantization unit 3145. The arithmetic coding unit 3146 outputs a bit stream including the arithmetically coded prediction residual to, for example, a three-dimensional data decoding device.

[0133] Note that the prediction residual may be binarized by, for example, the quantization unit 3145 before being arithmetically coded by the arithmetic coding unit 3146.

[0134] Also, for example, the arithmetic coding unit 3146 may initialize the coding table used for arithmetic coding before arithmetic coding. The arithmetic coding unit 3146 may initialize the coding table for each layer. Further, the arithmetic coding unit 3146 may output information indicating the position of the layer for which the coding table is initialized, included in the bit stream.

[0135] The inverse quantization unit 3147 inverse quantizes the prediction residual after quantization by the quantization unit 3145.

[0136] The decoded value generation unit 3148 generates a decoded value by adding the predicted value of the attribute information generated by the prediction unit 3143 and the prediction residual after inverse quantization by the inverse quantization unit 3147.

[0137] The memory 3149 is a memory that stores the decoded values of the attribute information of each three-dimensional point decoded by the decoded value generation unit 3148. For example, when the prediction unit 3143 generates a predicted value of a three-dimensional point that has not yet been coded, the prediction unit 3143 generates a predicted value using the decoded values of the attribute information of each three-dimensional point stored in the memory 3149.

[0138] FIG. 16 is a block diagram of an attribute information encoding unit 6600 which is an example of the conversion attribute information encoding unit A102. The attribute information encoding unit 6600 includes a sorting unit 6601, a Haar transform unit 6602, a quantization unit 6603, an inverse quantization unit 6604, an inverse Haar transform unit 6605, a memory 6606, and an arithmetic encoding unit 6607.

[0139] The sorting unit 6601 generates a Morton code using the position information of three-dimensional points and sorts a plurality of three-dimensional points in the order of the Morton code. The Haar transform unit 6602 generates encoded coefficients by applying a Haar transform to the attribute information. The quantization unit 6603 quantizes the encoded coefficients of the attribute information.

[0140] The inverse quantization unit 6604 inverse-quantizes the quantized encoded coefficients. The inverse Haar transform unit 6605 applies an inverse Haar transform to the encoded coefficients. The memory 6606 stores the values of the attribute information of a plurality of decoded three-dimensional points. For example, the decoded attribute information of the three-dimensional points stored in the memory 6606 may be used for prediction of non-encoded three-dimensional points.

[0141] The arithmetic encoding unit 6607 calculates ZeroCnt from the quantized encoded coefficients and arithmetically encodes ZeroCnt. Further, the arithmetic encoding unit 6607 arithmetically encodes the non-zero quantized encoded coefficients. The arithmetic encoding unit 6607 may binarize the encoded coefficients before arithmetic encoding. Further, the arithmetic encoding unit 6607 may generate and encode various header information.

[0142] FIG. 17 is a block diagram showing the configuration of an attribute information decoding unit 3150 which is an example of the LoD attribute information decoding unit A111.

[0143] The attribute information decoding unit 3150 includes a LoD generation unit 3151, a surrounding search unit 3152, a prediction unit 3153, an arithmetic decoding unit 3154, an inverse quantization unit 3155, a decoded value generation unit 3156, and a memory 3157.

[0144] The LoD generation unit 3151 generates LoD using the position information of the three-dimensional points decoded by a position information decoding unit (not shown in FIG. 17).

[0145] The surrounding search unit 3152 searches for neighboring three-dimensional points adjacent to each three-dimensional point using the LoD generation result by the LoD generation unit 3151 and distance information indicating the distance between each three-dimensional point.

[0146] The prediction unit 3153 generates a predicted value of the attribute information of the target three-dimensional point to be decoded.

[0147] The arithmetic decoding unit 3154 arithmetically decodes the prediction residual in the bit stream obtained from the attribute information encoding unit 3140 shown in FIG. 15. Note that the arithmetic decoding unit 3154 may initialize the decoding table used for arithmetic decoding. The arithmetic decoding unit 3154 initializes the decoding table used for arithmetic decoding for the layer on which the arithmetic encoding unit 3146 shown in FIG. 15 performed the encoding process. The arithmetic decoding unit 3154 may initialize the decoding table for each layer. Also, the arithmetic decoding unit 3154 may initialize the decoding table based on the information indicating the position of the layer in which the encoding table was initialized, included in the bit stream.

[0148] The inverse quantization unit 3155 inverse quantizes the prediction residual arithmetically decoded by the arithmetic decoding unit 3154.

[0149] The decoded value generation unit 3156 generates a decoded value by adding the predicted value generated by the prediction unit 3153 and the prediction residual after being inverse quantized by the inverse quantization unit 3155. The decoded value generation unit 3156 outputs the decoded attribute information data to another device.

[0150] The memory 3157 is a memory that stores the decoded values of the attribute information of each three-dimensional point decoded by the decoded value generation unit 3156. For example, when the prediction unit 3153 generates a predicted value of a three-dimensional point that has not yet been decoded, it generates the predicted value using the decoded values of the attribute information of each three-dimensional point stored in the memory 3157.

[0151] FIG. 18 is a block diagram of an attribute information decoding unit 6610, which is an example of the conversion attribute information decoding unit A112. The attribute information decoding unit 6610 includes an arithmetic decoding unit 6611, an inverse quantization unit 6612, an inverse Haar transform unit 6613, and a memory 6614.

[0152] The arithmetic decoding unit 6611 arithmetically decodes ZeroCnt and the encoded coefficients included in the bit stream. Note that the arithmetic decoding unit 6611 may decode various header information.

[0153] The inverse quantization unit 6612 inverse quantizes the arithmetically decoded encoded coefficients. The inverse Haar transform unit 6613 applies an inverse Haar transform to the encoded coefficients after inverse quantization. The memory 6614 stores the values of the attribute information of a plurality of decoded three-dimensional points. For example, the decoded attribute information of the three-dimensional points stored in the memory 6614 may be used for prediction of the non-decoded three-dimensional points.

[0154] Next, a second encoding unit 4650, which is an example of an encoding unit 4613 that performs encoding using the second encoding method, will be described. FIG. 19 is a diagram showing the configuration of the second encoding unit 4650. FIG. 20 is a block diagram of the second encoding unit 4650.

[0155] The second encoding unit 4650 generates encoded data (encoded stream) by encoding the point cloud data using the second encoding method. This second encoding unit 4650 includes an additional information generation unit 4651, a position image generation unit 4652, an attribute image generation unit 4653, a video encoding unit 4654, an additional information encoding unit 4655, and a multiplexing unit 4656.

[0156] The second encoding unit 4650 is characterized by generating a position image and an attribute image by projecting a three-dimensional structure onto a two-dimensional image, and encoding the generated position image and attribute image using an existing video encoding method. The second encoding method is also called VPCC (Video based PCC).

[0157] The point cloud data is PCC point cloud data such as a PLY file or PCC point cloud data generated from sensor information, and includes position information (Position), attribute information (Attribute), and other additional information (MetaData).

[0158] The additional information generation unit 4651 generates map information of a plurality of two-dimensional images by projecting the three-dimensional structure onto a two-dimensional image.

[0159] The position image generation unit 4652 generates a position image (Geometry Image) based on the position information and the map information generated by the additional information generation unit 4651. This position image is, for example, a distance image in which the distance (Depth) is indicated as a pixel value. Note that this distance image may be an image obtained by viewing a plurality of point clouds from one viewpoint (an image obtained by projecting a plurality of point clouds onto one two-dimensional plane), or may be a plurality of images obtained by viewing a plurality of point clouds from a plurality of viewpoints, or may be one integrated image of these plurality of images.

[0160] The attribute image generation unit 4653 generates an attribute image based on the attribute information and the map information generated by the additional information generation unit 4651. This attribute image is, for example, an image in which the attribute information (e.g., color (RGB)) is indicated as a pixel value. Note that this image may be an image obtained by viewing a plurality of point clouds from one viewpoint (an image obtained by projecting a plurality of point clouds onto one two-dimensional plane), or may be a plurality of images obtained by viewing a plurality of point clouds from a plurality of viewpoints, or may be one integrated image of these plurality of images.

[0161] The video encoding unit 4654 generates an encoded position image (Compressed Geometry Image) and an encoded attribute image (Compressed Attribute Image), which are encoded data, by encoding the position image and the attribute image using a video encoding method. Note that any known encoding method may be used as the video encoding method. For example, the video encoding method may be AVC or HEVC, etc.

[0162] The additional information encoding unit 4655 generates encoded additional information (Compressed MetaData) by encoding the additional information included in the point cloud data and map information or the like.

[0163] The multiplexing unit 4656 generates an encoded stream (Compressed Stream), which is encoded data, by multiplexing the encoded position image, the encoded attribute image, the encoded additional information, and other additional information. The generated encoded stream is output to a processing unit in a system layer (not shown).

[0164] Next, a second decoding unit 4660, which is an example of a decoding unit 4624 that decodes the second encoding method, will be described. FIG. 21 is a diagram showing the configuration of the second decoding unit 4660. FIG. 22 is a block diagram of the second decoding unit 4660. The second decoding unit 4660 generates point cloud data by decoding the encoded data (encoded stream) encoded by the second encoding method using the second encoding method. This second decoding unit 4660 includes a demultiplexing unit 4661, a video decoding unit 4662, an additional information decoding unit 4663, a position information generation unit 4664, and an attribute information generation unit 4665.

[0165] An encoded stream (Compressed Stream), which is encoded data, is input to the second decoding unit 4660 from a processing unit in a system layer (not shown).

[0166] The demultiplexing unit 4661 separates the encoded position image (Compressed Geometry Image), the encoded attribute image (Compressed Attribute Image), the encoded additional information (Compressed MetaData), and other additional information from the encoded data.

[0167] The video decoding unit 4662 generates a position image and an attribute image by decoding the encoded position image and the encoded attribute image using a video encoding method. Note that any known encoding method may be used as the video encoding method. For example, the video encoding method is AVC or HEVC or the like.

[0168] The additional information decoding unit 4663 generates additional information including map information and the like by decoding the encoded additional information.

[0169] The position information generation unit 4664 generates position information using the position image and the map information. The attribute information generation unit 4665 generates attribute information using the attribute image and the map information.

[0170] The second decoding unit 4660 uses the additional information necessary for decoding during decoding and outputs the additional information necessary for the application to the outside.

[0171] Hereinafter, the problems in the PCC encoding method will be described. FIG. 23 is a diagram showing a protocol stack related to PCC encoded data. FIG. 23 shows an example in which data of other media such as video (for example, HEVC) or audio is multiplexed with the PCC encoded data and transmitted or stored.

[0172] The multiplexing method and the file format have functions for multiplexing various encoded data and transmitting or storing it. In order to transmit or store the encoded data, the encoded data must be converted into the format of the multiplexing method. For example, in HEVC, a technique is defined in which the encoded data is stored in a data structure called a NAL unit, and the NAL unit is stored in ISOBMFF.

[0173] On the other hand, currently, a first encoding method (Codec1) and a second encoding method (Codec2) are being considered as encoding methods for point cloud data, but the configuration of the encoded data and the method of storing the encoded data in the system format are not defined, and there is a problem that MUX processing (multiplexing), transmission, and storage in the encoding unit cannot be performed as it is.

[0174] In the following, unless otherwise specified for a specific encoding method, either the first encoding method or the second encoding method is assumed.

[0175] (Embodiment 2) In this embodiment, the type of encoded data (Geometry, Attribute, Metadata) generated by the above-described first encoding unit 4630 or second encoding unit 4650, the method for generating the additional information (metadata), and the multiplexing process in the multiplexing unit will be described. Note that the additional information (metadata) may also be referred to as a parameter set or control information.

[0176] In this embodiment, the dynamic object (three-dimensional point cloud data that changes over time) described with reference to FIG. 4 will be used as an example for explanation, but the same method may also be used for a static object (three-dimensional point cloud data at an arbitrary time).

[0177] FIG. 24 is a diagram showing the configuration of an encoding unit 4801 and a multiplexing unit 4802 included in the three-dimensional data encoding apparatus according to this embodiment. The encoding unit 4801 corresponds to, for example, the above-described first encoding unit 4630 or second encoding unit 4650. The multiplexing unit 4802 corresponds to the above-described multiplexing unit 4634 or 4656.

[0178] The encoding unit 4801 encodes the point cloud data of a plurality of PCC (Point Cloud Compression) frames and generates encoded data (Multiple Compressed Data) of a plurality of pieces of position information, attribute information, and additional information.

[0179] The multiplexing unit 4802 converts the data into a data configuration considering data access in the decoding apparatus by NAL unitizing the data of a plurality of data types (position information, attribute information, and additional information).

[0180] FIG. 25 is a diagram showing a configuration example of the encoded data generated by the encoding unit 4801. The arrows in the figure indicate the dependency relationships related to the decoding of the encoded data, and the origin of the arrow depends on the data at the tip of the arrow. That is, the decoding device decodes the data at the tip of the arrow and uses the decoded data to decode the data at the origin of the arrow. In other words, to depend means that the data at the destination is referenced (used) in the processing (encoding or decoding, etc.) of the data at the source of the dependency.

[0181] First, the generation process of the encoded data of the position information will be described. The encoding unit 4801 generates frame-by-frame encoded position data (Compressed Geometry Data) by encoding the position information of each frame. Also, the encoded position data is represented by G(i). Here, i indicates the frame number, or the time of the frame, etc.

[0182] Also, the encoding unit 4801 generates a position parameter set (GPS(i)) corresponding to each frame. The position parameter set includes parameters that can be used for decoding the encoded position data. Also, the encoded position data for each frame depends on the corresponding position parameter set.

[0183] Also, the encoded position data composed of a plurality of frames is defined as a position sequence (Geometry Sequence). The encoding unit 4801 generates a position sequence parameter set (Geometry Sequence PS, also denoted as position SPS) that stores parameters commonly used for the decoding process for a plurality of frames in the position sequence. The position sequence depends on the position SPS.

[0184] Next, the generation process of the encoded data of the attribute information will be described. The encoding unit 4801 generates encoded attribute data (Compressed Attribute Data) for each frame by encoding the attribute information of each frame. Also, the encoded attribute data is represented by A(i). Further, in FIG. 25, an example in which there are attribute X and attribute Y is shown, the encoded attribute data of attribute X is represented by AX(i), and the encoded attribute data of attribute Y is represented by AY(i).

[0185] Also, the encoding unit 4801 generates an attribute parameter set (APS(i)) corresponding to each frame. Also, the attribute parameter set of attribute X is represented by AXPS(i), and the attribute parameter set of attribute Y is represented by AYPS(i). The attribute parameter set includes parameters that can be used for decoding the encoded attribute information. The encoded attribute data depends on the corresponding attribute parameter set.

[0186] Also, the encoded attribute data composed of a plurality of frames is defined as an attribute sequence. The encoding unit 4801 generates an attribute sequence parameter set (also denoted as Attribute Sequence PS: Attribute SPS) that stores parameters commonly used for the decoding process for a plurality of frames in the attribute sequence. The attribute sequence depends on the Attribute SPS.

[0187] Also, in the first encoding method, the encoded attribute data depends on the encoding position data.

[0188] Also, FIG. 25 shows an example in the case where there are two types of attribute information (attribute X and attribute Y). When there are two types of attribute information, for example, two encoding units generate respective data and metadata. Also, for example, an attribute sequence is defined for each type of attribute information, and an Attribute SPS is generated for each type of attribute information.

[0189] Note that in FIG. 25, an example is shown where there is one type of position information and two types of attribute information. However, this is not the only case. The attribute information may be one type or three or more types. Even in this case, the encoded data can be generated in the same way. Also, in the case of point cloud data without attribute information, the attribute information may not be present. In that case, the encoding unit 4801 does not necessarily need to generate a parameter set related to the attribute information.

[0190] Next, the generation process of additional information (metadata) will be described. The encoding unit 4801 generates a PCC stream PS (also referred to as stream PS), which is a parameter set for the entire PCC stream. The encoding unit 4801 stores in the stream PS parameters that can be commonly used for the decoding process of one or more position sequences and one or more attribute sequences. For example, the stream PS includes identification information indicating the codec of the point cloud data, information indicating the algorithm used for encoding, and the like. The position sequence and the attribute sequence depend on the stream PS.

[0191] Next, the access unit and the GOF will be described. In the present embodiment, the concepts of a new access unit (Access Unit: AU) and a GOF (Group of Frame) are introduced.

[0192] The access unit is a basic unit for accessing data during decoding and is composed of one or more data and one or more metadata. For example, the access unit is composed of position information at the same time and one or more pieces of attribute information. The GOF is a random access unit and is composed of one or more access units.

[0193] The symbolization unit 4801 generates an access unit header (AU Header) as identification information indicating the start of an access unit. The symbolization unit 4801 stores parameters related to the access unit in the access unit header. For example, the access unit header includes the configuration or information of the encoded data included in the access unit. Also, the access unit header includes parameters commonly used for the data included in the access unit, such as parameters related to the decoding of the encoded data.

[0194] Note that the symbolization unit 4801 may generate an access unit delimiter that does not include parameters related to the access unit instead of the access unit header. This access unit delimiter is used as identification information indicating the start of the access unit. The decoding device identifies the start of the access unit by detecting the access unit header or the access unit delimiter.

[0195] Next, the generation of the identification information at the start of the GOF will be described. The symbolization unit 4801 generates a GOF header (GOF Header) as identification information indicating the start of the GOF. The symbolization unit 4801 stores parameters related to the GOF in the GOF header. For example, the GOF header includes the configuration or information of the encoded data included in the GOF. Also, the GOF header includes parameters commonly used for the data included in the GOF, such as parameters related to the decoding of the encoded data.

[0196] Note that the symbolization unit 4801 may generate a GOF delimiter that does not include parameters related to the GOF instead of the GOF header. This GOF delimiter is used as identification information indicating the start of the GOF. The decoding device identifies the start of the GOF by detecting the GOF header or the GOF delimiter.

[0197] In the PCC encoded data, for example, an access unit is defined in units of PCC frames. The decoding device accesses the PCC frame based on the identification information at the start of the access unit.

[0198] Also, for example, a GOF is defined as one random access unit. The decoding device accesses the random access unit based on the identification information at the head of the GOF. For example, if PCC frames are independent of each other and can be decoded alone, the PCC frames may be defined as random access units.

[0199] Note that two or more PCC frames may be assigned to one access unit, or a plurality of random access units may be assigned to one GOF.

[0200] Also, the encoding unit 4801 may define and generate parameter sets or metadata other than those described above. For example, the encoding unit 4801 may generate SEI (Supplemental Enhancement Information) that stores parameters (optional parameters) that may not necessarily be used during decoding.

[0201] Next, the configuration of the encoded data and the method of storing the encoded data in the NAL unit will be described.

[0202] For example, a data format is defined for each type of encoded data. FIG. 26 is a diagram showing an example of encoded data and NAL units.

[0203] For example, as shown in FIG. 26, the encoded data includes a header and a payload. Note that the encoded data may include length information indicating the length (data amount) of the encoded data, the header, or the payload. Also, the encoded data may not include a header.

[0204] The header includes, for example, identification information for specifying the data. This identification information indicates, for example, the data type or the frame number.

[0205] The header includes, for example, identification information indicating a reference relationship. This identification information is stored in the header when there is a dependency between data, for example, and is information for referring from a reference source to a reference destination. For example, the header of the reference destination includes identification information for specifying the data. The header of the reference source includes identification information indicating the reference destination.

[0206] Note that when the reference destination or reference source can be identified or derived from other information, the identification information for specifying the data or the identification information indicating the reference relationship may be omitted.

[0207] The multiplexing unit 4802 stores the encoded data in the payload of the NAL unit. The NAL unit header includes pcc_nal_unit_type which is identification information of the encoded data. FIG. 27 is a diagram showing an example of the semantics of pcc_nal_unit_type.

[0208] As shown in FIG. 27, when pcc_codec_type is Codec1 (Codec1: the first encoding method), the values 0 to 10 of pcc_nal_unit_type are assigned to the encoded position data (Geometry), encoded attribute X data (AttributeX), encoded attribute Y data (AttributeY), position PS (Geom.PS), attribute XPS (AttrX.PS), attribute YPS (AttrX.PS), position SPS (Geometry Sequence PS), attribute XSPS (AttributeX Sequence PS), attribute YSPS (AttributeY Sequence PS), AU header (AU Header), and GOF header (GOF Header) in Codec1. Also, the values 11 and later are assigned for future use in Codec1.

[0209] When the pcc_codec_type is Codec2 (the second encoding method), the values 0 to 2 of pcc_nal_unit_type are assigned to the codec's Data A, MetaData A, and MetaData B. Also, values 3 and above are assigned to the reserve of Codec2.

[0210] (Embodiment 3) In HEVC encoding, there are data splitting tools such as slices or tiles to enable parallel processing in the decoder, but there are none in PCC (Point Cloud Compression) encoding yet.

[0211] In PCC, various data splitting methods can be considered depending on parallel processing, compression efficiency, and compression algorithms. Here, the definitions of slices and tiles, data structures, and transmission and reception methods will be described.

[0212] FIG. 28 is a block diagram showing the configuration of a first encoding unit 4910 included in the three-dimensional data encoding apparatus according to the present embodiment. The first encoding unit 4910 generates encoded data (encoded stream) by encoding point cloud data using a first encoding method (GPCC (Geometry based PCC)). The first encoding unit 4910 includes a splitting unit 4911, a plurality of position information encoding units 4912, a plurality of attribute information encoding units 4913, an additional information encoding unit 4914, and a multiplexing unit 4915.

[0213] The splitting unit 4911 generates a plurality of split data by splitting the point cloud data. Specifically, the splitting unit 4911 generates a plurality of split data by splitting the space of the point cloud data into a plurality of sub-spaces. Here, the sub-space is either one of a tile and a slice, or a combination of a tile and a slice. More specifically, the point cloud data includes position information, attribute information, and additional information. The splitting unit 4911 splits the position information into a plurality of split position information and splits the attribute information into a plurality of split attribute information. Also, the splitting unit 4911 generates additional information regarding the split.

[0214] The plurality of position information encoding units 4912 generate a plurality of encoded position information by encoding a plurality of divided position information. For example, the plurality of position information encoding units 4912 process the plurality of divided position information in parallel.

[0215] The plurality of attribute information encoding units 4913 generate a plurality of encoded attribute information by encoding a plurality of divided attribute information. For example, the plurality of attribute information encoding units 4913 process the plurality of divided attribute information in parallel.

[0216] The additional information encoding unit 4914 generates encoded additional information by encoding the additional information included in the point cloud data and the additional information regarding data division generated at the time of division by the division unit 4911.

[0217] The multiplexing unit 4915 generates encoded data (encoded stream) by multiplexing the plurality of encoded position information, the plurality of encoded attribute information, and the encoded additional information, and transmits the generated encoded data. Also, the encoded additional information is used at the time of decoding.

[0218] Note that in FIG. 28, examples in which the number of the position information encoding unit 4912 and the attribute information encoding unit 4913 is two each are shown, but the number of the position information encoding unit 4912 and the attribute information encoding unit 4913 may be one each, or may be three or more. Also, the plurality of divided data may be processed in parallel within the same chip like a plurality of cores in the CPU, may be processed in parallel by the cores of a plurality of chips, or may be processed in parallel by the plurality of cores of a plurality of chips.

[0219] FIG. 29 is a block diagram showing the configuration of the first decoding unit 4920. The first decoding unit 4920 restores the point cloud data by decoding the encoded data (encoded stream) generated by encoding the point cloud data by the first encoding method (GPCC). This first decoding unit 4920 includes a demultiplexing unit 4921, a plurality of position information decoding units 4922, a plurality of attribute information decoding units 4923, an additional information decoding unit 4924, and a combining unit 4925.

[0220] The inverse multiplexing unit 4921 generates a plurality of encoded position information, a plurality of encoded attribute information, and encoded additional information by inverse multiplexing the encoded data (encoded stream).

[0221] The plurality of position information decoding units 4922 generate a plurality of divided position information by decoding the plurality of encoded position information. For example, the plurality of position information decoding units 4922 perform parallel processing on the plurality of encoded position information.

[0222] The plurality of attribute information decoding units 4923 generate a plurality of divided attribute information by decoding the plurality of encoded attribute information. For example, the plurality of attribute information decoding units 4923 perform parallel processing on the plurality of encoded attribute information.

[0223] The plurality of additional information decoding units 4924 generate additional information by decoding the encoded additional information.

[0224] The combining unit 4925 generates position information by combining the plurality of divided position information using the additional information. The combining unit 4925 generates attribute information by combining the plurality of divided attribute information using the additional information.

[0225] Note that in FIG. 29, an example is shown where the number of the position information decoding units 4922 and the attribute information decoding units 4923 is two each, but the number of the position information decoding units 4922 and the attribute information decoding units 4923 may be one each or three or more. Also, the plurality of divided data may be processed in parallel within the same chip like a plurality of cores in the CPU, or may be processed in parallel by the cores of a plurality of chips, or may be processed in parallel by the plurality of cores of a plurality of chips.

[0226] Next, the configuration of the dividing unit 4911 will be described. FIG. 30 is a block diagram of the dividing unit 4911. The dividing unit 4911 includes a slice dividing unit 4931 (Slice Divider), a geometry tile dividing unit 4932 (Geometry Tile Divider), and an attribute tile dividing unit 4933 (Attribute Tile Divider).

[0227] The slice division unit 4931 generates a plurality of slice position information by dividing the position information (Position(Geometry)) into slices. Further, the slice division unit 4931 generates a plurality of slice attribute information by dividing the attribute information (Attribute) into slices. Further, the slice division unit 4931 outputs slice additional information (SliceMetaData) including information related to slice division and information generated in slice division.

[0228] The position information tile division unit 4932 generates a plurality of divided position information (a plurality of tile position information) by dividing the plurality of slice position information into tiles. Further, the position information tile division unit 4932 outputs position tile additional information (Geometry Tile MetaData) including information related to the tile division of the position information and information generated in the tile division of the position information.

[0229] The attribute information tile division unit 4933 generates a plurality of divided attribute information (a plurality of tile attribute information) by dividing the plurality of slice attribute information into tiles. Further, the attribute information tile division unit 4933 outputs attribute tile additional information (Attribute Tile MetaData) including information related to the tile division of the attribute information and information generated in the tile division of the attribute information.

[0230] Note that the number of slices or tiles to be divided is 1 or more. That is, it is not necessary to perform slice or tile division.

[0231] Also, here, an example in which tile division is performed after slice division is shown, but slice division may be performed after tile division. Further, a new division type may be defined in addition to slices and tiles, and division may be performed with three or more division types.

[0232] Hereinafter, a method for dividing point cloud data will be described. FIG. 31 is a diagram showing an example of slice and tile division.

[0233] First, the method of slice division will be described. The division unit 4911 divides the three-dimensional point cloud data into arbitrary point clouds in units of slices. In slice division, the division unit 4911 does not divide the position information and the attribute information that constitute the points, but divides the position information and the attribute information together. That is, the division unit 4911 performs slice division so that the position information and the attribute information at any point belong to the same slice. Note that according to these, the number of divisions and the division method may be any method. Also, the minimum unit of division is a point. For example, the number of divisions of the position information and the attribute information is the same. For example, the three-dimensional points corresponding to the position information after slice division and the three-dimensional points corresponding to the attribute information are included in the same slice.

[0234] Also, the division unit 4911 generates slice additional information, which is additional information related to the number of divisions and the division method during slice division. The slice additional information is the same for the position information and the attribute information. For example, the slice additional information includes information indicating the reference coordinate position, size, or side length of the bounding box after division. Also, the slice additional information includes information indicating the number of divisions, the division type, and the like.

[0235] Next, the method of tile division will be described. The division unit 4911 divides the data divided by slice into slice position information (G slice) and slice attribute information (A slice), and divides the slice position information and the slice attribute information into tiles respectively.

[0236] Note that FIG. 31 shows an example of division in an octree structure, but the number of divisions and the division method may be any method.

[0237] Also, the division unit 4911 may divide the position information and the attribute information by different division methods, or by the same division method. Also, the division unit 4911 may divide a plurality of slices into tiles by different division methods, or by the same division method.

[0238] In addition, the splitting unit 4911 generates tile addition information related to the number of splits and the splitting method during tile splitting. The tile addition information (position tile addition information and attribute tile addition information) is independent of the position information and the attribute information. For example, the tile addition information includes information indicating the reference coordinate position, size, or side length of the bounding box after splitting. In addition, the tile addition information includes information indicating the number of splits, the split type, and the like.

[0239] Next, an example of a method for splitting point cloud data into slices or tiles will be described. The splitting unit 4911 may use a predetermined method as the method for slice or tile splitting, or may adaptively switch the method to be used according to the point cloud data.

[0240] During slice splitting, the splitting unit 4911 divides the three-dimensional space all at once with respect to the position information and the attribute information. For example, the splitting unit 4911 determines the shape of the object and divides the three-dimensional space into slices according to the shape of the object. For example, the splitting unit 4911 extracts an object such as a tree or a building and performs splitting in object units. For example, the splitting unit 4911 performs slice splitting so that the whole of one or a plurality of objects is included in one slice. Or, the splitting unit 4911 divides one object into a plurality of slices.

[0241] In this case, the encoding device may change the encoding method for each slice, for example. For example, the encoding device may use a high-quality compression method for a specific object or a specific part of the object. In this case, the encoding device may store information indicating the encoding method for each slice in the additional information (metadata).

[0242] In addition, the splitting unit 4911 may perform slice splitting so that each slice corresponds to a predetermined coordinate space based on the map information or the position information.

[0243] When performing tile division, the division unit 4911 divides the position information and the attribute information independently. For example, the division unit 4911 divides the slices into tiles according to the data volume or the processing volume. For example, the division unit 4911 determines whether the data volume of the slice (for example, the number of three-dimensional points included in the slice) is more than a predetermined threshold value. When the data volume of the slice is more than the threshold value, the division unit 4911 divides the slice into tiles. When the data volume of the slice is less than the threshold value, the division unit 4911 does not divide the slice into tiles.

[0244] For example, the division unit 4911 divides the slices into tiles so that the processing volume or the processing time in the decoding device is within a certain range (equal to or less than a predetermined value). Thereby, the processing volume per tile in the decoding device becomes constant, and parallel processing in the decoding device becomes easy.

[0245] In addition, when the processing volumes of the position information and the attribute information are different, for example, when the processing volume of the position information is larger than the processing volume of the attribute information, the division unit 4911 increases the number of divisions of the position information compared to the number of divisions of the attribute information.

[0246] Also, for example, when, depending on the content, in the decoding device, the position information may be decoded and displayed quickly, and the attribute information may be decoded and displayed slowly later, the division unit 4911 may increase the number of divisions of the position information compared to the number of divisions of the attribute information. Thereby, the decoding device can increase the parallelism of the position information, so that the processing of the position information can be made faster than the processing of the attribute information.

[0247] Note that the decoding device does not necessarily need to perform parallel processing on the sliced or tiled data, and may determine whether to perform parallel processing according to the number or the capabilities of the decoding processing units.

[0248] By dividing in the above-described manner, adaptive encoding according to the content or the object can be realized. Also, parallel processing in the decoding process can be realized. Thereby, the flexibility of the point cloud encoding system or the point cloud decoding system is improved.

[0249] FIG. 32 is a diagram showing an example of a pattern of slice and tile division. In the figure, DU is a data unit (DataUnit) and represents the data of a tile or a slice. Each DU also includes a slice index (SliceIndex) and a tile index (TileIndex). The numerical value at the upper right of the DU in the figure indicates the slice index, and the numerical value at the lower left of the DU indicates the tile index.

[0250] In Pattern 1, in slice division, the number of divisions and the division method are the same for the G slice and the A slice. In tile division, the number of divisions and the division method for the G slice are different from those for the A slice. Also, the same number of divisions and the same division method are used among multiple G slices. The same number of divisions and the same division method are used among multiple A slices.

[0251] In Pattern 2, in slice division, the number of divisions and the division method are the same for the G slice and the A slice. In tile division, the number of divisions and the division method for the G slice are different from those for the A slice. Also, the number of divisions and the division method are different among multiple G slices. The number of divisions and the division method are different among multiple A slices.

[0252] Next, the encoding method for the divided data will be described. The three-dimensional data encoding device (first encoding unit 4910) encodes the divided data respectively. When encoding the attribute information, the three-dimensional data encoding device generates dependency information indicating based on which configuration information (position information, additional information, or other attribute information) the encoding is performed as additional information. That is, the dependency information indicates, for example, the configuration information of the reference destination (dependency destination). In this case, the three-dimensional data encoding device generates the dependency information based on the configuration information corresponding to the division shape of the attribute information. Note that the three-dimensional data encoding device may generate the dependency information based on the configuration information corresponding to a plurality of division shapes.

[0253] Dependency relationship information is generated by a three-dimensional data encoding device, and the generated dependency relationship information may be sent to a three-dimensional data decoding device. Alternatively, the three-dimensional data decoding device may generate the dependency relationship information, and the three-dimensional data encoding device may not need to send the dependency relationship information. Also, the dependency relationships used by the three-dimensional data encoding device may be determined in advance, and the three-dimensional data encoding device may not need to send the dependency relationship information.

[0254] FIG. 33 is a diagram showing an example of the dependency relationship of each data. The tip of the arrow in the figure indicates the dependency destination, and the origin of the arrow indicates the dependency source. The three-dimensional data decoding device decodes the data in the order from the dependency destination to the dependency source. Also, the data shown by the solid line in the figure is the data actually sent, and the data shown by the dotted line is the data not sent.

[0255] Also, in the same figure, G indicates position information, and A indicates attribute information. Gs1 indicates the position information of slice number 1, and Gs2 indicates the position information of slice number 2. Gs1t1 indicates the position information of slice number 1 and tile number 1, Gs1t2 indicates the position information of slice number 1 and tile number 2, Gs2t1 indicates the position information of slice number 2 and tile number 1, and Gs2t2 indicates the position information of slice number 2 and tile number 2. Similarly, As1 indicates the attribute information of slice number 1, and As2 indicates the attribute information of slice number 2. As1t1 indicates the attribute information of slice number 1 and tile number 1, As1t2 indicates the attribute information of slice number 1 and tile number 2, As2t1 indicates the attribute information of slice number 2 and tile number 1, and As2t2 indicates the attribute information of slice number 2 and tile number 2.

[0256] Mslice indicates slice addition information, MGtile indicates position tile addition information, and MAtile indicates attribute tile addition information. Ds1t1 indicates the dependency relationship information of the attribute information As1t1, and Ds2t1 indicates the dependency relationship information of the attribute information As2t1.

[0257] Further, the three-dimensional data encoding device may rearrange the data in the decoding order so that the three-dimensional data decoding device does not need to rearrange the data. Note that the three-dimensional data decoding device may rearrange the data, or the data may be rearranged by both the three-dimensional data encoding device and the three-dimensional data decoding device.

[0258] FIG. 34 is a diagram showing an example of the decoding order of data. In the example of FIG. 34, decoding is performed in order from the left data. The three-dimensional data decoding device decodes the data of the dependent destination first among the data in a dependency relationship. For example, the three-dimensional data encoding device rearranges and sends out the data in advance so as to be in this order. Note that any order may be used as long as the data of the dependent destination comes first. Further, the three-dimensional data encoding device may send out the additional information and the dependency relationship information before the data.

[0259] FIG. 35 is a flowchart showing the processing flow by the three-dimensional data encoding device. First, the three-dimensional data encoding device encodes the data of a plurality of slices or tiles as described above (S4901). Next, the three-dimensional data encoding device rearranges the data so that the data of the dependent destination comes first as shown in FIG. 34 (S4902). Next, the three-dimensional data encoding device multiplexes (NAL unitizes) the rearranged data (S4903).

[0260] Next, the configuration of the combining unit 4925 included in the first decoding unit 4920 will be described. FIG. 36 is a block diagram showing the configuration of the combining unit 4925. The combining unit 4925 includes a position information tile combining unit 4941 (Geometry Tile Combiner), an attribute information tile combining unit 4942 (Attribute Tile Combiner), and a slice combining unit (Slice Combiner).

[0261] The position information tile combining unit 4941 generates a plurality of slice position information by combining a plurality of divided position information using the position tile additional information. The attribute information tile combining unit 4942 generates a plurality of slice attribute information by combining a plurality of divided attribute information using the attribute tile additional information.

[0262] The slice combining unit 4943 generates position information by combining a plurality of slice position information using the slice additional information. Also, the slice combining unit 4943 generates attribute information by combining a plurality of slice attribute information using the slice additional information.

[0263] Note that the number of slices or tiles to be divided is 1 or more. That is, the slices or tiles may not be divided.

[0264] Also, here, an example in which tile division is performed after slice division is shown, but slice division may be performed after tile division. Also, in addition to slices and tiles, a new division type may be defined, and division may be performed using three or more division types.

[0265] Next, the configuration of the sliced or tiled encoded data and the method of storing the encoded data in the NAL unit (multiplexing method) will be described. FIG. 37 is a diagram showing the configuration of the encoded data and the method of storing the encoded data in the NAL unit.

[0266] The encoded data (divided position information and divided attribute information) is stored in the payload of the NAL unit.

[0267] The symbolized data includes a header and a payload. The header includes identification information for identifying the data included in the payload. This identification information includes, for example, the type of slice division or tile division (slice_type, tile_type), index information for identifying a slice or tile (slice_idx, tile_idx), position information of the data (slice or tile), or the address of the data, etc. The index information for identifying a slice is also referred to as a slice index (SliceIndex). The index information for identifying a tile is also referred to as a tile index (TileIndex). Also, the type of division is, for example, a method based on an object shape as described above, a method based on map information or position information, or a method based on the amount of data or the amount of processing, etc.

[0268] Note that all or part of the above information may be stored in one of the header of the division position information and the header of the division attribute information, and may not be stored in the other. For example, when the same division method is used for the position information and the attribute information, the type of division (slice_type, tile_type) and the index information (slice_idx, tile_idx) are the same for the position information and the attribute information. Therefore, these information may be included in the header of one of the position information and the attribute information. For example, when the attribute information depends on the position information, the position information is processed first. Therefore, these information may be included in the header of the position information, and may not be included in the header of the attribute information. In this case, the three-dimensional data decoding device determines, for example, that the attribute information of the dependency source belongs to the same slice or tile as the slice or tile of the position information of the dependency destination.

[0269] In addition, additional information related to slice division or tile division (slice additional information, position tile additional information, or attribute tile additional information), dependency information indicating dependencies, etc. may be stored in an existing parameter set (such as GPS, APS, position SPS, or attribute SPS) and sent. When the division method changes for each frame, information indicating the division method may be stored in a parameter set for each frame (such as GPS or APS). When the division method does not change within a sequence, information indicating the division method may be stored in a parameter set for each sequence (position SPS or attribute SPS). Furthermore, when the same division method is used for position information and attribute information, information indicating the division method may be stored in a parameter set for the PCC stream (stream PS).

[0270] Also, the above information may be stored in any of the above parameter sets or in multiple parameter sets. Alternatively, a parameter set for tile division or slice division may be defined, and the above information may be stored in the parameter set. Also, these information may be stored in the header of the encoded data.

[0271] Also, the header of the encoded data includes identification information indicating dependencies. That is, when there is a dependency between data, the header includes identification information for referring from the dependent source to the dependent destination. For example, the header of the dependent destination data includes identification information for specifying the data. The header of the dependent source data includes identification information indicating the dependent destination. Note that if the identification information for specifying the data, the additional information related to slice division or tile division, and the identification information indicating dependencies are distinguishable or derivable from other information, these information may be omitted.

[0272] Next, the flow of the encoding process and decoding process of the point cloud data according to the present embodiment will be described. FIG. 38 is a flowchart of the encoding process of the point cloud data according to the present embodiment.

[0273] First, the three-dimensional data encoding device determines the splitting method to be used (S4911). This splitting method includes whether to perform slice splitting or tile splitting. Further, the splitting method may include the number of splits when performing slice splitting or tile splitting, and the type of split, etc. The type of split is a method based on the object shape as described above, a method based on map information or position information, or a method based on data volume or processing volume, etc. Note that the splitting method may be predetermined.

[0274] When slice splitting is performed (Yes in S4912), the three-dimensional data encoding device generates a plurality of slice position information and a plurality of slice attribute information by splitting the position information and the attribute information together (S4913). Further, the three-dimensional data encoding device generates slice addition information related to slice splitting. Note that the three-dimensional data encoding device may split the position information and the attribute information independently.

[0275] When tile splitting is performed (Yes in S4914), the three-dimensional data encoding device generates a plurality of split position information and a plurality of split attribute information by independently splitting a plurality of slice position information and a plurality of slice attribute information (or position information and attribute information) (S4915). Further, the three-dimensional data encoding device generates position tile addition information and attribute tile addition information related to tile splitting. Note that the three-dimensional data encoding device may split the slice position information and the slice attribute information together.

[0276] Next, the three-dimensional data encoding device generates a plurality of encoded position information and a plurality of encoded attribute information by encoding each of the plurality of split position information and the plurality of split attribute information (S4916). Further, the three-dimensional data encoding device generates dependency information.

[0277] Next, the three-dimensional data encoding device generates encoded data (encoded stream) by NAL unitizing (multiplexing) the plurality of encoded position information, the plurality of encoded attribute information, and the addition information (S4917). Further, the three-dimensional data encoding device transmits the generated encoded data.

[0278] Figure 39 is a flowchart of the decoding process of the point cloud data according to the present embodiment. First, the three-dimensional data decoding device determines the splitting method (S4921) by analyzing the additional information related to the splitting method (slice additional information, position tile additional information, and attribute tile additional information) included in the encoded data (encoded stream). This splitting method includes whether to perform slice splitting or tile splitting. Further, the splitting method may include the number of splits and the type of split when performing slice splitting or tile splitting.

[0279] Next, the three-dimensional data decoding device generates split position information and split attribute information by decoding a plurality of encoded position information and a plurality of encoded attribute information included in the encoded data using the dependency information included in the encoded data (S4922).

[0280] When it is shown by the additional information that tile splitting is performed (Yes in S4923), the three-dimensional data decoding device generates a plurality of slice position information and a plurality of slice attribute information by combining the plurality of split position information and the plurality of split attribute information in their respective methods based on the position tile additional information and the attribute tile additional information (S4924). Note that the three-dimensional data decoding device may combine the plurality of split position information and the plurality of split attribute information in the same method.

[0281] When it is shown by the additional information that slice splitting is performed (Yes in S4925), the three-dimensional data decoding device generates position information and attribute information by combining the plurality of slice position information and the plurality of slice attribute information (the plurality of split position information and the plurality of split attribute information) in the same method based on the slice additional information (S4926). Note that the three-dimensional data decoding device may combine the plurality of slice position information and the plurality of slice attribute information in different methods.

[0282] Note that the attribute information (identifiers, area information, address information, position information, etc.) of tiles or slices may be stored not only in SEI but also in other control information. For example, the attribute information may be stored in control information indicating the overall configuration of the PCC data, or may be stored in control information for each tile or slice.

[0283] In addition, when transmitting PCC data to other devices, the three-dimensional data encoding device (three-dimensional data transmission device) may convert control information such as SEI into control information specific to the protocol of the system and indicate it.

[0284] For example, when converting PCC data including attribute information into ISOBMFF (ISO Base Media File Format), the three-dimensional data encoding device may store the SEI together with the PCC data in the "mdat box", or may store it in the "track box" that describes control information related to the stream. That is, the three-dimensional data encoding device may store the control information in a table for random access. Also, when packetizing and transmitting PCC data, the three-dimensional data encoding device may store the SEI in the packet header. By making the attribute information accessible at the system layer in this way, access to the attribute information and tile data or slice data becomes easier, and the access speed can be improved.

[0285] Note that in the configuration of the three-dimensional data decoding device, the memory management unit may determine in advance whether the information required for the decoding process is in the memory, and if the information required for the decoding process is not available, the information may be obtained from the storage or network.

[0286] When the three-dimensional data decoding device acquires PCC data from a storage or a network using Pull in a protocol such as MPEG-DASH, the memory management unit may identify the attribute information of the data necessary for the decoding process based on information from the localization unit or the like, request a tile or a slice including the identified attribute information, and acquire the necessary data (PCC stream). The identification of a tile or a slice including the attribute information may be performed on the storage or network side, or may be performed by the memory management unit. For example, the memory management unit may acquire the SEI of all PCC data in advance, and identify a tile or a slice based on the information.

[0287] When all PCC data is transmitted from a storage or a network using Push in a protocol such as UDP, the memory management unit may identify the attribute information of the data necessary for the decoding process and a tile or a slice based on information from the localization unit or the like, and filter out a desired tile or slice from the transmitted PCC data to acquire the desired data.

[0288] In addition, when acquiring data, the three-dimensional data encoding device may determine whether there is desired data, whether real-time processing is possible based on the data size or the like, or the communication state or the like. When the three-dimensional data encoding device determines based on this determination result that it is difficult to acquire data, it may select and acquire another slice or tile with a different priority or data amount.

[0289] In addition, the three-dimensional data decoding device may transmit information from the localization unit or the like to a cloud server, and the cloud server may determine necessary information based on the information.

[0290] (Embodiment 4) Next, tile additional information will be described. The three-dimensional data encoding device generates tile additional information, which is metadata regarding a tile division method, and transmits the generated tile additional information to the three-dimensional data decoding device.

[0291] Figure 40 is a diagram showing an example of the syntax of tile additional information (TileMetaData). As shown in Figure 40, for example, the tile additional information includes division method information (type_of_divide), shape information (topview_shape), overlap flag (tile_overlap_flag), overlap information (type_of_overlap), height information (tile_height), number of tiles (tile_number), and tile position information (global_position, relative_position).

[0292] The division method information (type_of_divide) indicates the division method of the tile. For example, the division method information indicates whether the division method of the tile is division based on map information, that is, division based on top view (top_view), or other.

[0293] The shape information (topview_shape) is included in the tile additional information when, for example, the division method of the tile is division based on top view. The shape information indicates the shape of the tile in top view. For example, this shape includes a square and a circle. Note that this shape may include a polygon other than an ellipse, rectangle, or quadrilateral, or other shapes. Note that the shape information is not limited to the shape of the tile in top view, and may also indicate the three-dimensional shape of the tile (for example, a cube and a cylinder, etc.).

[0294] The overlap flag (tile_overlap_flag) indicates whether the tiles overlap. For example, the overlap flag is included in the tile additional information when the division method of the tile is division based on top view. In this case, the overlap flag indicates whether the tiles overlap in top view. Note that the overlap flag may also indicate whether the tiles overlap in three-dimensional space.

[0295] Overlap information (type_of_overlap) is included in the tile additional information when, for example, tiles overlap. The overlap information indicates how the tiles overlap. For example, the overlap information indicates the size of the overlapping area, etc.

[0296] Height information (tile_height) indicates the height of the tile. Note that the height information may include information indicating the shape of the tile. For example, when the shape of the tile in a top view is a rectangle, this information may indicate the lengths of the sides of the rectangle (vertical length and horizontal length). Also, when the shape of the tile in a top view is a circle, this information may indicate the diameter or radius of the circle.

[0297] Also, the height information may indicate the height of each tile, or may indicate a common height for a plurality of tiles. Also, a plurality of height types such as roads and intersections may be set in advance, and the height information may indicate the height of each height type and the height type of each tile. Alternatively, the height of each height type may be predefined, and the height information may indicate the height type of each tile. That is, the height of each height type may not be indicated by the height information.

[0298] The number of tiles (tile_number) indicates the number of tiles. Note that the tile additional information may include information indicating the interval between tiles.

[0299] Tile position information (global_position, relative_position) is information for specifying the position of each tile. For example, the tile position information indicates the absolute coordinates or relative coordinates of each tile.

[0300] Note that some or all of the above information may be provided for each tile, or may be provided for every plurality of tiles (for example, for each frame or every plurality of frames).

[0301] The three-dimensional data encoding device may send the tile additional information included in SEI (Supplemental Enhancement Information). Or, the three-dimensional data encoding device may store the tile additional information in an existing parameter set (PPS, GPS, or APS, etc.) and then send it.

[0302] For example, when the tile additional information changes for each frame, the tile additional information may be stored in a parameter set (such as GPS or APS) for each frame. When the tile additional information does not change within a sequence, the tile additional information may be stored in a parameter set (position SPS or attribute SPS) for each sequence. Furthermore, when the same tile division information is used for position information and attribute information, the tile additional information may be stored in a parameter set (stream PS) of the PCC stream.

[0303] Also, the tile additional information may be stored in any one of the above parameter sets, or may be stored in a plurality of parameter sets. Also, the tile additional information may be stored in the header of the encoded data. Also, the tile additional information may be stored in the header of the NAL unit.

[0304] Also, all or part of the tile additional information may be stored in one of the headers of the division position information and the division attribute information, and may not be stored in the other. For example, when the same tile additional information is used for position information and attribute information, the tile additional information may be included in one of the headers of the position information and the attribute information. For example, when the attribute information depends on the position information, the position information is processed first. Therefore, these tile additional information may be included in the header of the position information, and the tile additional information may not be included in the header of the attribute information. In this case, the three-dimensional data decoding device determines, for example, that the attribute information of the dependency source belongs to the same tile as the tile of the position information of the dependency destination.

[0305] The three-dimensional data decoding device reconstructs the tile-divided point cloud data based on the tile addition information. When there is overlapping point cloud data, the three-dimensional data decoding device identifies the multiple overlapping point cloud data and selects any one of them or merges the multiple point cloud data.

[0306] Also, the three-dimensional data decoding device may perform decoding using the tile addition information. For example, when multiple tiles overlap, the three-dimensional data decoding device performs decoding for each tile, performs processing (such as smoothing or filtering) using the multiple decoded data, and may generate point cloud data. This may enable highly accurate decoding.

[0307] FIG. 41 is a diagram showing a configuration example of a system including a three-dimensional data encoding device and a three-dimensional data decoding device. The tile division unit 5051 divides the point cloud data including position information and attribute information into a first tile and a second tile. Also, the tile division unit 5051 sends the tile addition information related to the tile division to the decoding unit 5053 and the tile combination unit 5054.

[0308] The encoding unit 5052 generates encoded data by encoding the first tile and the second tile.

[0309] The decoding unit 5053 restores the first tile and the second tile by decoding the encoded data generated by the encoding unit 5052. The tile combination unit 5054 restores the point cloud data (position information and attribute information) by combining the first tile and the second tile using the tile addition information.

[0310] Next, the slice addition information will be described. The three-dimensional data encoding device generates slice addition information, which is metadata regarding the slice division method, and sends the generated slice addition information to the three-dimensional data decoding device.

[0311] FIG. 42 is a diagram showing an example of the syntax of slice additional information (SliceMetaData). As shown in FIG. 42, for example, the slice additional information includes division method information (type_of_divide), an overlap flag (slice_overlap_flag), overlap information (type_of_overlap), the number of slices (slice_number), slice position information (global_position, relative_position), and slice size information (slice_bounding_box_size).

[0312] The division method information (type_of_divide) indicates the division method of the slice. For example, the division method information indicates whether the division method of the slice is division based on object information as shown in FIG. 60 (object). Note that the slice additional information may include information indicating the method of object division. For example, this information indicates whether to divide one object into multiple slices or assign it to one slice. Further, this information may indicate the number of divisions when dividing one object into multiple slices.

[0313] The overlap flag (slice_overlap_flag) indicates whether the slices overlap. The overlap information (type_of_overlap) is included in the slice additional information, for example, when the slices overlap. The overlap information indicates how the slices overlap. For example, the overlap information indicates the size of the overlapping area.

[0314] The number of slices (slice_number) indicates the number of slices.

[0315] Slice position information (global_position, relative_position) and slice size information (slice_bounding_box_size) are information regarding the area of a slice. The slice position information is information for specifying the position of each slice. For example, the slice position information indicates the absolute coordinates or relative coordinates of each slice. The slice size information (slice_bounding_box_size) indicates the size of each slice. For example, the slice size information indicates the size of the bounding box of each slice.

[0316] The three-dimensional data encoding device may send out the slice additional information included in the SEI. Alternatively, the three-dimensional data encoding device may store the slice additional information in an existing parameter set (PPS, GPS, or APS, etc.) and then send it out.

[0317] For example, when the slice additional information changes for each frame, the slice additional information may be stored in a parameter set (GPS or APS, etc.) for each frame. When the slice additional information does not change within a sequence, the slice additional information may be stored in a parameter set (position SPS or attribute SPS) for each sequence. Furthermore, when the same slice division information is used for position information and attribute information, the slice additional information may be stored in a parameter set (stream PS) of the PCC stream.

[0318] Also, the slice additional information may be stored in any one of the above parameter sets, or may be stored in a plurality of parameter sets. Also, the slice additional information may be stored in the header of the encoded data. Also, the slice additional information may be stored in the header of the NAL unit.

[0319] In addition, all or part of the slice addition information may be stored in one of the headers of the division position information and the division attribute information, and may not be stored in the other. For example, when the same slice addition information is used for the position information and the attribute information, the slice addition information may be included in one of the headers of the position information and the attribute information. For example, when the attribute information depends on the position information, the position information is processed first. Therefore, these slice addition information may be included in the header of the position information, and the slice addition information may not be included in the header of the attribute information. In this case, the three-dimensional data decoding device determines, for example, that the attribute information of the dependency source belongs to the same slice as the slice of the position information of the dependency destination.

[0320] The three-dimensional data decoding device reconstructs the point cloud data sliced based on the slice addition information. When there is duplicate point cloud data, the three-dimensional data decoding device identifies the multiple duplicate point cloud data and selects one of them, or merges the multiple point cloud data.

[0321] In addition, the three-dimensional data decoding device may perform decoding using the slice addition information. For example, when multiple slices overlap, the three-dimensional data decoding device performs decoding for each slice, performs processing (such as smoothing or filtering) using the multiple decoded data, and may generate point cloud data. This may enable more accurate decoding.

[0322] FIG. 43 is a flowchart of three-dimensional data encoding processing including tile addition information generation processing by the three-dimensional data encoding device according to the present embodiment.

[0323] First, the three-dimensional data encoding device determines a tile division method (S5031). Specifically, the three-dimensional data encoding device determines whether to use a division method based on a top view (top_view) or another method (other) as the tile division method. In addition, when using the division method based on the top view, the three-dimensional data encoding device determines the shape of the tile. In addition, the three-dimensional data encoding device determines whether the tile overlaps with other tiles.

[0324] If the tile splitting method determined in step S5031 is a splitting method based on the top view (Yes in S5032), the three-dimensional data encoding device describes in the tile additional information that the tile splitting method is a splitting method based on the top view (top_view) (S5033).

[0325] On the other hand, if the tile splitting method determined in step S5031 is other than the splitting method based on the top view (No in S5032), the three-dimensional data encoding device describes in the tile additional information that the tile splitting method is other than the splitting method based on the top view (top_view) (S5034).

[0326] Also, if the shape of the tile in the top view determined in step S5031 is a square (square in S5035), the three-dimensional data encoding device describes in the tile additional information that the shape of the tile in the top view is a square (S5036). On the other hand, if the shape of the tile in the top view determined in step S5031 is a circle (circle in S5035), the three-dimensional data encoding device describes in the tile additional information that the shape of the tile in the top view is a circle (S5037).

[0327] Next, the three-dimensional data encoding device determines whether the tile overlaps with other tiles (S5038). If the tile overlaps with other tiles (Yes in S5038), the three-dimensional data encoding device describes in the tile additional information that the tile overlaps (S5039). On the other hand, if the tile does not overlap with other tiles (No in S5038), the three-dimensional data encoding device describes in the tile additional information that the tile does not overlap (S5040).

[0328] Next, the three-dimensional data encoding device splits the tile based on the tile splitting method determined in step S5031, encodes each tile, and sends out the generated encoded data and tile additional information (S5041).

[0329] FIG. 44 is a flowchart of three-dimensional data decoding processing using tile addition information by the three-dimensional data decoding apparatus according to the present embodiment.

[0330] First, the three-dimensional data decoding apparatus analyzes tile addition information included in the bit stream (S5051).

[0331] When it is shown by the tile addition information that the tile does not overlap with other tiles (No in S5052), the three-dimensional data decoding apparatus generates point cloud data for each tile by decoding each tile (S5053). Next, the three-dimensional data decoding apparatus reconstructs the point cloud data from the point cloud data for each tile based on the tile division method and the tile shape indicated by the tile addition information (S5054).

[0332] On the other hand, when it is shown by the tile addition information that the tile overlaps with other tiles (Yes in S5052), the three-dimensional data decoding apparatus generates point cloud data for each tile by decoding each tile. Further, the three-dimensional data decoding apparatus specifies the overlapping portion of the tiles based on the tile addition information (S5055). Note that the three-dimensional data decoding apparatus may perform decoding processing using a plurality of overlapping pieces of information for the overlapping portion. Next, the three-dimensional data decoding apparatus reconstructs the point cloud data from the point cloud data for each tile based on the tile division method, the tile shape, and the overlapping information indicated by the tile addition information (S5056).

[0333] Hereinafter, a modification example and the like regarding slices will be described. The three-dimensional data encoding apparatus may transmit information indicating the type (road, building, tree, etc.) or attribute (dynamic information, static information, etc.) of the object as addition information. Alternatively, encoding parameters may be defined in advance according to the object, and the three-dimensional data encoding apparatus may notify the three-dimensional data decoding apparatus of the encoding parameters by sending out the type or attribute of the object.

[0334] The following methods may be used for the encoding order and transmission order of slice data. For example, the three-dimensional data encoding device may encode the slice data in order from data that is easy to recognize or cluster objects. Or, the three-dimensional data encoding device may perform encoding in order from the slice data for which clustering has ended earlier. Also, the three-dimensional data encoding device may transmit the encoded slice data in order. Or, the three-dimensional data encoding device may transmit the slice data in order of the priority of decoding in the application. For example, when the priority of decoding dynamic information is high, the three-dimensional data encoding device may transmit the slice data in order from the slices grouped by dynamic information.

[0335] Also, when the order of the encoded data and the order of the decoding priority are different, the three-dimensional data encoding device may rearrange the encoded data and then transmit it. Also, when accumulating the encoded data, the three-dimensional data encoding device may rearrange the encoded data and then accumulate it.

[0336] The application (three-dimensional data decoding device) requests the server (three-dimensional data encoding device) to transmit the slices including the desired data. The server may transmit the slice data required by the application and may not transmit the unnecessary slice data.

[0337] The application requests the server to transmit the tiles including the desired data. The server may transmit the tile data required by the application and may not transmit the unnecessary tile data.

[0338] (Embodiment 5) In this embodiment, the processing of a division unit (for example, a tile or a slice) that does not include points will be described. First, the method of dividing the point cloud data will be described.

[0339] In video coding standards such as HEVC, since data exists for all pixels of a two-dimensional image, even when the two-dimensional space is divided into a plurality of data regions, data exists in all data regions. On the other hand, in the coding of three-dimensional point cloud data, the points themselves, which are elements of the point cloud data, are data, and there is a possibility that data does not exist in some regions.

[0340] There are various methods for spatially dividing point cloud data, and the dividing methods can be classified according to whether the dividing unit (for example, a tile or a slice), which is the divided data unit, always contains one or more point data.

[0341] A dividing method in which all of a plurality of dividing units contain one or more point data is called a first dividing method. As a first dividing method, for example, there is a method of dividing point cloud data while taking into account the processing time of coding or the size of the coded data. In this case, the number of points in each dividing unit is approximately equal.

[0342] FIG. 45 is a diagram showing an example of a dividing method. For example, as a first dividing method, as shown in FIG. 45(a), a method of dividing points belonging to the same space into two identical spaces may be used. Also, as shown in FIG. 45(b), the space may be divided into a plurality of sub-spaces (dividing units) so that each dividing unit contains points.

[0343] Since these methods are divisions that take points into account, one or more points are always included in all dividing units.

[0344] A dividing method in which there may be one or more dividing units that do not contain point data among a plurality of dividing units is called a second dividing method. For example, as a second dividing method, as shown in FIG. 45(c), a method of equally dividing the space can be used. In this case, points do not necessarily exist in the dividing unit. That is, there may be a case where no points exist in the dividing unit.

[0345] When the three-dimensional data encoding device divides point cloud data, it may indicate whether (1) a division method in which all of a plurality of division units contain one or more point data is used, (2) a division method in which one or more of the plurality of division units do not contain point data is used, or (3) a division method in which one or more of the plurality of division units may not contain point data, in division additional information (metadata), which is additional information related to the division (for example, tile additional information or slice additional information), and send out the division additional information.

[0346] Note that the three-dimensional data encoding device may indicate the above information as the type of division method. Further, the three-dimensional data encoding device may perform division by a predetermined division method and may not send out the division additional information. In that case, the three-dimensional data encoding device explicitly indicates in advance whether the division method is the first division method or the second division method.

[0347] Hereinafter, an example of the second division method and generation and transmission of encoded data will be described. Hereinafter, tile division will be described as an example of a division method in a three-dimensional space, but it is not necessary to be tile division, and the following method can also be applied to a division method with a division unit different from a tile. For example, tile division may be read as slice division.

[0348] FIG. 46 is a diagram showing an example of dividing point cloud data into six tiles. FIG. 46 shows an example in which the minimum unit is a point, and shows an example of dividing position information (Geometry) and attribute information (Attribute) together. Note that the same applies when the position information and the attribute information are divided by different division methods or numbers of divisions, when there is no attribute information, and when there are multiple pieces of attribute information.

[0349] In the example shown in FIG. 46, after tile division, there are tiles (#1, #2, #4, #6) that contain points in the tile and tiles (#3, #5) that do not contain points in the tile. Tiles that do not contain points in the tile are called null tiles.

[0350] Note that not only when dividing into six tiles, but any division method may be used. For example, the division unit may be a cube, or a shape other than a cube such as a rectangular parallelepiped or a cylinder. The plurality of division units may have the same shape or may include different shapes. Also, as the division method, a predetermined method may be used, or different methods may be used for each predetermined unit (for example, a PCC frame).

[0351] In this division method, when the point cloud data is divided into tiles, if there is no data in a tile, a bit stream including information indicating that the tile is a null tile is generated.

[0352] Hereinafter, the method of sending null tiles and the method of signaling null tiles will be described. The three-dimensional data encoding device may generate, for example, the following information as additional information (metadata) related to data division, and send the generated information. FIG. 47 is a diagram showing an example of the syntax of tile additional information (TileMetaData). The tile additional information includes division method information (type_of_divide), division method null information (type_of_divide_null), number of tile divisions (number_of_tiles), and tile null flag (tile_null_flag).

[0353] The division method information (type_of_divide) is information related to the division method or division type. For example, the division method information indicates one or more division methods or division types. For example, as the division method, there are top view (top_view) division and equal division. Note that if there is only one definition of the division method, the division method information may not be included in the tile additional information.

[0354] The division method null information (type_of_divide_null) is information indicating whether the division method used is the following first division method or the second division method. Here, the first division method is a division method in which each of a plurality of division units always contains one or more point data. The second division method is a division method in which there is one or more division units that do not contain point data among a plurality of division units, or a division method in which there may be one or more division units that do not contain point data among a plurality of division units.

[0355] In addition, the tile addition information may include at least one of (1) information indicating the number of divisions of the tile (number_of_tiles), or information for specifying the number of divisions of the tile, (2) information indicating the number of null tiles, or information for specifying the number of null tiles, and (3) information indicating the number of tiles other than the null tiles, or information for specifying the number of tiles other than the null tiles, as the division information of the entire tile. The tile addition information may also include information indicating the shape of the tile or whether the tiles overlap as the division information of the entire tile.

[0356] In addition, the tile addition information sequentially indicates the division information for each tile. For example, the order of the tiles is predetermined for each division method and is known in the three-dimensional data encoding device and the three-dimensional data decoding device. If the order of the tiles is not predetermined, the three-dimensional data encoding device may send information indicating the order to the three-dimensional data decoding device.

[0357] The division information for each tile includes a tile null flag (tile_null_flag) that is a flag indicating whether data (points) exist in the tile. Note that when there is no data in the tile, the tile null flag may be included as the tile division information.

[0358] Also, when the tile is not a null tile, the tile additional information includes per-tile division information (position information (e.g., coordinates of the origin (origin_x, origin_y, origin_z)), tile height information, etc.). Also, when the tile is a null tile, the tile additional information does not include per-tile division information.

[0359] For example, when storing per-tile slice division information in the per-tile division information, the three-dimensional data encoding device may not store the slice division information of the null tile in the additional information.

[0360] Note that in this example, the number of tiles (number_of_tiles) indicates the number of tiles including the null tile. FIG. 48 is a diagram showing an example of the index information (idx) of the tiles. In the example shown in FIG. 48, the index information is also assigned to the null tile.

[0361] Next, the data structure and transmission method of the encoded data including the null tile will be described. FIGS. 49 to 51 are diagrams showing the data structure when the position information and the attribute information are divided into six tiles and there is no data in the third and fifth tiles.

[0362] FIG. 49 is a diagram showing an example of the dependency relationship of each data. The tip of the arrow in the figure indicates the dependency destination, and the origin of the arrow indicates the dependency source. Also, in the same figure, Gtn (n is 1 to 6) indicates the position information of tile number n, and Atn indicates the attribute information of tile number n. Mtile indicates the tile additional information.

[0363] FIG. 50 is a diagram showing a configuration example of the transmission data which is the encoded data transmitted from the three-dimensional data encoding device. Also, FIG. 51 is a diagram showing the configuration of the encoded data and the method of storing the encoded data in the NAL unit.

[0364] As shown in FIG. 51, the tile index information (tile_idx) is included in the headers of the data of the position information (division position information) and the attribute information (division attribute information), respectively.

[0365] Also, as shown in Structure 1 of FIG. 50, the three-dimensional data encoding device may not send the position information or attribute information that constitutes the null tile. Or, as shown in Structure 2 of FIG. 50, the three-dimensional data encoding device may send information indicating that the tile is a null tile as the data of the null tile. For example, the three-dimensional data encoding device may describe that the data type is a null tile in the tile_type stored in the header of the NAL unit or in the header within the nal_unit_payload of the NAL unit, and send the header. Hereinafter, the description will be made on the premise of Structure 1.

[0366] In Structure 1, when there is a null tile, in the transmitted data, the value of the tile index information (tile_idx) included in the header of the position information data or the attribute information data is missing and not continuous.

[0367] Also, when there is a dependency relationship between data, the three-dimensional data encoding device sends the data so that the reference destination data can be decoded earlier than the reference source data. Note that the tile of the attribute information has a dependency relationship with the tile of the position information. The same tile index number is added to the attribute information and the position information with a dependency relationship.

[0368] Note that the tile addition information related to tile division may be stored in both the parameter set (GPS) of the position information and the parameter set (APS) of the attribute information, or may be stored in either one of them. When the tile addition information is stored in one of GPS and APS, reference information indicating the reference destination GPS or APS may be stored in the other of GPS and APS. Also, when the tile division method is different between the position information and the attribute information, different tile addition information is stored in each of GPS and APS. Also, when the tile division method is the same in a sequence (multiple PCC frames), the tile addition information may be stored in GPS, APS, or SPS (sequence parameter set).

[0369] For example, when tile-added information is stored in both GPS and APS, the tile-added information of the position information is stored in GPS, and the tile-added information of the attribute information is stored in APS. Also, when tile-added information is stored in common information such as SPS, the tile-added information commonly used for both the position information and the attribute information may be stored, or the tile-added information of the position information and the tile-added information of the attribute information may be stored respectively.

[0370] Hereinafter, the combination of tile division and slice division will be described. First, the data configuration and data transmission when tile division is performed after slice division will be described.

[0371] FIG. 52 is a diagram showing an example of the dependency relationship of each data when tile division is performed after slice division. The tip of the arrow in the figure indicates the dependency destination, and the origin of the arrow indicates the dependency source. Also, the data indicated by the solid line in the figure is the data actually transmitted, and the data indicated by the dotted line is the data not transmitted.

[0372] Also, in the same figure, G indicates position information, and A indicates attribute information. Gs1 indicates the position information of slice number 1, and Gs2 indicates the position information of slice number 2. Gs1t1 indicates the position information of slice number 1 and tile number 1, and Gs2t2 indicates the position information of slice number 2 and tile number 2. Similarly, As1 indicates the attribute information of slice number 1, and As2 indicates the attribute information of slice number 2. As1t1 indicates the attribute information of slice number 1 and tile number 1, and As2t1 indicates the attribute information of slice number 2 and tile number 1.

[0373] Mslice indicates slice-added information, MGtile indicates position tile-added information, and MAtile indicates attribute tile-added information. Ds1t1 indicates the dependency relationship information of the attribute information As1t1, and Ds2t1 indicates the dependency relationship information of the attribute information As2t1.

[0374] The three-dimensional data encoding device may not generate and transmit the position information and the attribute information related to the null tile.

[0375] Also, even when the number of tile divisions is the same for all slices, the number of tiles generated and transmitted between slices may be different. For example, when the number of tile divisions for position information and attribute information is different, there may be a null tile in either the position information or the attribute information, but not in the other. In the example shown in FIG. 52, the position information (Gs1) of slice 1 is divided into two tiles, Gs1t1 and Gs1t2, where Gs1t2 is a null tile. On the other hand, the attribute information (As1) of slice 1 is not divided and there is one As1t1, and there is no null tile.

[0376] Also, regardless of whether a null tile is included in the slice of the position information, when there is data in at least the tile of the attribute information, the three-dimensional data encoding device generates and transmits dependency information of the attribute information. For example, when the three-dimensional data encoding device stores the slice division information for each tile in the slice addition information related to slice division, this information stores information on whether the tile is a null tile.

[0377] FIG. 53 is a diagram showing an example of the decoding order of data. In the example of FIG. 53, decoding is performed in order from the left data. The three-dimensional data decoding device decodes the data that is the dependency destination first among the data in a dependency relationship. For example, the three-dimensional data encoding device rearranges and transmits the data in advance so as to be in this order. Note that any order may be used as long as the data that is the dependency destination comes first. Also, the three-dimensional data encoding device may transmit the additional information and the dependency information before the data.

[0378] Next, the data configuration and data transmission when slice division is performed after tile division will be described.

[0379] FIG. 54 is a diagram showing an example of the dependency relationship of each data when slice division is performed after tile division. The tip of the arrow in the figure indicates the dependency destination, and the origin of the arrow indicates the dependency source. Also, the data indicated by the solid line in the figure is the data actually transmitted, and the data indicated by the dotted line is the data not transmitted.

[0380] Also, in the same figure, G indicates position information, and A indicates attribute information. Gt1 indicates the position information of tile number 1. Gt1s1 indicates the position information of tile number 1 and slice number 1, and Gt1s2 indicates the position information of tile number 1 and slice number 2. Similarly, At1 indicates the attribute information of tile number 1, and At1s1 indicates the attribute information of tile number 1 and slice number 1.

[0381] Mtile indicates tile addition information, MGslice indicates position slice addition information, and MAslice indicates attribute slice addition information. Dt1s1 indicates the dependency relationship information of the attribute information At1s1, and Dt2s1 indicates the dependency relationship information of the attribute information At2s1.

[0382] The three-dimensional data encoding device does not perform slice division on null tiles. Also, it may not be necessary to generate and transmit the position information, attribute information, and dependency relationship information of the attribute information related to the null tiles.

[0383] FIG. 55 is a diagram showing an example of the decoding order of data. In the example of FIG. 55, decoding is performed in order from the left data. The three-dimensional data decoding device decodes the data with a dependency relationship first from the dependent data. For example, the three-dimensional data encoding device rearranges the data in advance and transmits it in this order. Note that any order may be used as long as the dependent data comes first. Also, the three-dimensional data encoding device may transmit the additional information and the dependency relationship information before the data.

[0384] Next, the flow of the division process and the combination process of the point cloud data will be described. Here, an example of tile division and slice division will be described, but the same method can be applied to other spatial divisions.

[0385] FIG. 56 is a flowchart of a three-dimensional data encoding process including a data splitting process by a three-dimensional data encoding device. First, the three-dimensional data encoding device determines a splitting method to be used (S5101). Specifically, the three-dimensional data encoding device determines whether to use the first splitting method or the second splitting method. For example, the three-dimensional data encoding device may determine the splitting method based on a specification from a user or an external device (e.g., a three-dimensional data decoding device), or may determine the splitting method according to the input point cloud data. Also, the splitting method to be used may be predetermined.

[0386] Here, the first splitting method is a splitting method in which each of a plurality of splitting units (tiles or slices) always contains one or more point data. The second splitting method is a splitting method in which there is one or more splitting units that do not contain point data among the plurality of splitting units, or a splitting method in which there may be one or more splitting units that do not contain point data among the plurality of splitting units.

[0387] When the determined splitting method is the first splitting method (the first splitting method in S5102), the three-dimensional data encoding device describes that the splitting method used for the splitting additional information (e.g., tile additional information or slice additional information), which is metadata related to data splitting, is the first splitting method (S5103). Then, the three-dimensional data encoding device encodes all the splitting units (S5104).

[0388] On the other hand, when the determined splitting method is the second splitting method (the second splitting method in S5102), the three-dimensional data encoding device describes that the splitting method used for the splitting additional information is the second splitting method (S5105). Then, the three-dimensional data encoding device encodes the splitting units excluding the splitting units that do not contain point data (e.g., null tiles) among the plurality of splitting units (S5106).

[0389] FIG. 57 is a flowchart of three-dimensional data decoding processing including data combining processing by a three-dimensional data decoding apparatus. First, the three-dimensional data decoding apparatus refers to the division addition information included in the bit stream and determines whether the division method used is the first division method or the second division method (S5111).

[0390] When the division method used is the first division method (the first division method in S5112), the three-dimensional data decoding apparatus receives the encoded data of all division units, and generates the decoded data of all division units by decoding the received encoded data (S5113). Next, the three-dimensional data decoding apparatus reconstructs the three-dimensional point cloud using the decoded data of all division units (S5114). For example, the three-dimensional data decoding apparatus reconstructs the three-dimensional point cloud by combining a plurality of division units.

[0391] On the other hand, when the division method used is the second division method (the second division method in S5112), the three-dimensional data decoding apparatus receives the encoded data of the division units including point data and the encoded data of the division units not including point data, and generates decoded data by decoding the received encoded data of the division units (S5115). Note that when the division units not including point data are not transmitted, the three-dimensional data decoding apparatus does not have to receive and decode the division units not including point cloud data. Next, the three-dimensional data decoding apparatus reconstructs the three-dimensional point cloud using the decoded data of the division units including point data (S5116). For example, the three-dimensional data decoding apparatus reconstructs the three-dimensional point cloud by combining a plurality of division units.

[0392] Hereinafter, other division methods of point cloud data will be described. When the space is evenly divided as shown in FIG. 45(c), there may be a case where there are no points in the divided space. In this case, the three-dimensional data encoding apparatus combines the space where there are no points with other spaces where there are points. Thereby, the three-dimensional data encoding apparatus can form a plurality of division units such that all division units include one or more points.

[0393] FIG. 58 is a flowchart of data division in this case. First, the three-dimensional data encoding device divides the data in a specific method (S5121). For example, the specific method is the second division method described above.

[0394] Next, the three-dimensional data encoding device determines whether a point is included in a target division unit that is a division unit to be processed (S5122). If a point is included in the target division unit (Yes in S5122), the three-dimensional data encoding device encodes the target division unit (S5123). On the other hand, if no point is included in the target division unit (No in S5122), the three-dimensional data encoding device combines the target division unit and another division unit including a point, and encodes the combined division unit (S5124). That is, the three-dimensional data encoding device encodes the target division unit together with another division unit including a point.

[0395] Here, an example of performing determination and combination for each division unit has been described, but the processing method is not limited to this. For example, the three-dimensional data encoding device may determine whether a point is included in each of a plurality of division units, perform combination so that there is no division unit not including a point, and encode each of the plurality of combined division units.

[0396] Next, a data sending method including a null tile will be described. When a target tile that is a tile to be processed is a null tile, the three-dimensional data encoding device does not send the data of the target tile. FIG. 59 is a flowchart of data sending processing.

[0397] First, the three-dimensional data encoding device determines a tile division method, and divides the point cloud data into tiles using the determined division method (S5131).

[0398] Next, the three-dimensional data encoding device determines whether the target tile is a null tile (S5132). That is, the three-dimensional data encoding device determines whether there is no data in the target tile.

[0399] When the target tile is a null tile (Yes in S5132), the three-dimensional data encoding device indicates in the tile additional information that the target tile is a null tile and does not indicate the information of the target tile (such as the position and size of the tile) (S5133). Further, the three-dimensional data encoding device does not send out the target tile (S5134).

[0400] On the other hand, when the target tile is not a null tile (No in S5132), the three-dimensional data encoding device indicates in the tile additional information that the target tile is not a null tile and indicates the information for each tile (S5135). Further, the three-dimensional data encoding device sends out the target tile (S5136).

[0401] In this way, by not including the information of the null tile in the tile additional information, the amount of information of the tile additional information can be reduced.

[0402] Hereinafter, a method for decoding encoded data including null tiles will be described. First, the processing in the case of no packet loss will be described.

[0403] FIG. 60 is a diagram showing an example of transmission data which is encoded data transmitted from a three-dimensional data encoding device and reception data input to a three-dimensional data decoding device. Here, it is assumed that the system environment has no packet loss, and the reception data is the same as the transmission data.

[0404] In the case of a system environment with no packet loss, the three-dimensional data decoding device receives all of the transmission data. FIG. 61 is a flowchart of the processing by the three-dimensional data decoding device.

[0405] First, the three-dimensional data decoding device refers to the tile additional information (S5141) and determines whether each tile is a null tile (S5142).

[0406] When it is shown that the target tile is not a null tile in the tile addition information (No in S5142), the three-dimensional data decoding device determines that the target tile is not a null tile and decodes the target tile (S5143). Next, the three-dimensional data decoding device acquires tile information (tile position information (such as origin coordinates) and size, etc.) from the tile addition information, and reconstructs three-dimensional data by combining a plurality of tiles using the acquired information (S5144).

[0407] On the other hand, when it is shown that the target tile is not a null tile in the tile addition information (Yes in S5142), the three-dimensional data decoding device determines that the target tile is a null tile and does not decode the target tile (S5145).

[0408] Note that the three-dimensional data decoding device may determine that the missing data is a null tile by sequentially analyzing the index information shown in the header of the encoded data. Also, the three-dimensional data decoding device may combine the determination method using tile addition information and the determination method using index information.

[0409] Next, the processing in the case of packet loss will be described. FIG. 62 is a diagram showing an example of transmitted data sent from the three-dimensional data encoding device and received data input to the three-dimensional data decoding device. Here, a system environment with packet loss is assumed.

[0410] In a system environment with packet loss, the three-dimensional data decoding device may not be able to receive all of the transmitted data. In this example, the packets of Gt2 and At2 are lost.

[0411] FIG. 63 is a flowchart of the processing of the three-dimensional data decoding device in this case. First, the three-dimensional data decoding device analyzes the continuity of the index information shown in the header of the encoded data (S5151), and determines whether the index number of the target tile exists (S5152).

[0412] If the index number of the target tile exists (Yes in S5152), the three-dimensional data decoding device determines that the target tile is not a null tile and performs the decoding process for the target tile (S5153). Next, the three-dimensional data decoding device acquires tile information (tile position information (such as origin coordinates) and size, etc.) from the tile addition information, and reconstructs the three-dimensional data by combining a plurality of tiles using the acquired information (S5154).

[0413] On the other hand, if the index information of the target tile does not exist (No in S5152), the three-dimensional data decoding device determines whether the target tile is a null tile by referring to the tile addition information (S5155).

[0414] If the target tile is not a null tile (No in S5156), the three-dimensional data decoding device determines that the target tile is lost (packet loss) and performs error decoding processing (S5157). The error decoding process is, for example, a process of attempting to decode the original data as if there was data. In this case, the three-dimensional data decoding device may reproduce the three-dimensional data and perform the reconstruction of the three-dimensional data (S5154).

[0415] On the other hand, if the target tile is a null tile (Yes in S5156), the three-dimensional data decoding device does not perform the decoding process and the reconstruction of the three-dimensional data assuming that the target tile is a null tile (S5158).

[0416] Next, a coding method when not explicitly indicating a null tile will be described. The three-dimensional data coding device may generate coded data and additional information by the following method.

[0417] The three-dimensional data coding device does not indicate null tile information in the tile addition information. The three-dimensional data coding device assigns the index numbers of tiles excluding the null tile to the data header. The three-dimensional data coding device does not send out the null tile.

[0418] In this case, the number of tiles indicates the number of divisions excluding null tiles. Note that the three-dimensional data encoding device may separately store information indicating the number of null tiles in the bit stream. Further, the three-dimensional data encoding device may indicate information regarding null tiles in the additional information, or may indicate some of the information regarding null tiles.

[0419] FIG. 64 is a flowchart of three-dimensional data encoding processing by the three-dimensional data encoding device in this case. First, the three-dimensional data encoding device determines a tile division method, and divides the point cloud data into tiles using the determined division method (S5161).

[0420] Next, the three-dimensional data encoding device determines whether the target tile is a null tile (S5162). That is, the three-dimensional data encoding device determines whether there is no data in the target tile.

[0421] When the target tile is not a null tile (No in S5162), the three-dimensional data encoding device assigns index information of the tiles excluding the null tile to the data header (S5163). Then, the three-dimensional data encoding device sends out the target tile (S5164).

[0422] On the other hand, when the target tile is a null tile (Yes in S5162), the three-dimensional data encoding device assigns index information of the target tile to the data header and does not send out the target tile.

[0423] FIG. 65 is a diagram showing an example of the index information (idx) added to the data header. As shown in FIG. 65, index information of the null tile is not added, and consecutive numbers are added to the tiles other than the null tile.

[0424] FIG. 66 is a diagram showing an example of the dependency relationship of each data. The tip of the arrow in the figure indicates the dependency destination, and the origin of the arrow indicates the dependency source. Also, in the same figure, Gtn (n is 1 to 4) indicates the position information of tile number n, and Atn indicates the attribute information of tile number n. Mtile indicates tile addition information.

[0425] FIG. 67 is a diagram showing a configuration example of transmission data which is encoded data sent from a three-dimensional data encoding device.

[0426] Hereinafter, a decoding method in the case where null tiles are not explicitly shown will be described. FIG. 68 is a diagram showing an example of transmission data sent from a three-dimensional data encoding device and reception data input to a three-dimensional data decoding device. Here, a system environment with packet loss is assumed.

[0427] FIG. 69 is a flowchart of the processing of the three-dimensional data decoding device in this case. First, the three-dimensional data decoding device analyzes the index information of the tile indicated in the header of the encoded data, and determines whether the index number of the target tile exists. Also, the three-dimensional data decoding device acquires the number of tile divisions from the tile addition information (S5171).

[0428] When the index number of the target tile exists (Yes in S5172), the three-dimensional data decoding device performs decoding processing of the target tile (S5173). Next, the three-dimensional data decoding device acquires tile information (tile position information (origin coordinates, etc.) and size, etc.) from the tile addition information, and reconstructs three-dimensional data by combining a plurality of tiles using the acquired information (S5175).

[0429] On the other hand, when the index number of the target tile does not exist (No in S5172), the three-dimensional data decoding device determines that the target tile is a packet loss, and performs error decoding processing (S5174). Also, the three-dimensional data decoding device determines that a space that does not exist in the data is a null tile, and reconstructs three-dimensional data.

[0430] In addition, the three-dimensional data encoding device can appropriately determine that there are no points in the tile by explicitly indicating null tiles, rather than measurement errors, data loss due to data processing, etc., or packet loss.

[0431] Note that the three-dimensional data encoding device may use a method of explicitly indicating null packets and a method of not explicitly indicating null packets in combination. In that case, the three-dimensional data encoding device may indicate in the tile additional information information indicating whether or not to explicitly indicate null packets. Also, depending on the type of splitting method, it may be determined in advance whether or not to explicitly indicate null packets, and the three-dimensional data encoding device may indicate whether or not to explicitly indicate null packets by indicating the type of splitting method.

[0432] Also, in FIG. 47 and the like, an example in which information related to all tiles is shown in the tile additional information has been shown, but in the tile additional information, information on some of the plurality of tiles may be shown, or information on null tiles of some of the plurality of tiles may be shown.

[0433] Also, an example in which information related to split data, such as information on whether or not there is split data (tiles), is stored in the tile additional information has been described, but some or all of these information may be stored in the parameter set or may be stored as data. When these information are stored as data, for example, nal_unit_type that means information indicating whether or not there is split data may be defined, and these information may be stored in the NAL unit. Also, these information may be stored in both the additional information and the data.

[0434] (Embodiment 6) Hereinafter, the process of performing quantization for each tile will be described.

[0435] FIG. 70 is a diagram showing a syntax example of GPS. As shown in FIG. 70, GPS includes a UniqueBetweenTilesFlag. The UniqueBetweenTilesFlag is a flag indicating whether there may be duplicate points between tiles.

[0436] FIG. 71 is a flowchart of three-dimensional data decoding processing. First, the three-dimensional data decoding device decodes the UniqueBetweenTilesFlag and the MergeDuplicatedPointFlag from the metadata included in the bit stream (S6261). Next, the three-dimensional data decoding device decodes the position information and the attribute information for each tile and reconstructs the point cloud (S6262).

[0437] Next, the three-dimensional data decoding device determines whether it is necessary to merge duplicate points (S6263). For example, the three-dimensional data decoding device determines whether merging is necessary according to whether the application can handle duplicate points or whether it is better to merge duplicate points. Alternatively, the three-dimensional data decoding device may smooth or filter a plurality of attribute information corresponding to duplicate points and determine to merge the duplicate points for the purpose of noise removal or improvement of estimation accuracy.

[0438] If it is necessary to merge duplicate points (Yes in S6263), the three-dimensional data decoding device determines whether there is duplication between tiles (whether there are duplicate points) (S6264). For example, the three-dimensional data decoding device may determine the presence or absence of duplication between tiles based on the decoding results of the UniqueBetweenTilesFlag and the MergeDuplicatedPointFlag. Thereby, the search for duplicate points in the three-dimensional data decoding device becomes unnecessary, and the processing load of the three-dimensional data decoding device can be reduced. Note that the three-dimensional data decoding device may also determine whether there are duplicate points by searching for duplicate points after reconstructing the tiles.

[0439] If there is overlap between tiles (Yes in S6264), the three-dimensional data decoding device merges the overlapping points between the tiles (S6265). Next, the three-dimensional data decoding device merges a plurality of overlapping attribute information (S6266).

[0440] After step S6266, or if there is no overlap between tiles (No in S6264), the three-dimensional data decoding device executes an application using the point cloud without overlapping points (S6267).

[0441] On the other hand, if merging of overlapping points is not necessary (No in S6263), the three-dimensional data decoding device does not perform merging of overlapping points and executes an application using the point cloud with overlapping points (S6268).

[0442] Hereinafter, examples of applications will be described. First, an example of an application using a point cloud without overlapping points will be described.

[0443] FIG. 72 is a diagram showing an example of an application. The example shown in FIG. 72 shows a use case in which a moving object traveling from the area of tile A to the area of tile B downloads map point clouds from a server in real time. The server stores encoded data of map point clouds of a plurality of overlapping areas. The moving object has already acquired the map information of tile A and requests the server to acquire the map information of tile B located in the moving direction.

[0444] At this time, the moving object determines that the data of the overlapping portion between tile A and tile B is unnecessary and transmits an instruction to the server to delete the overlapping portion between tile B and tile A included in tile B. The server deletes the overlapping portion from tile B and distributes the deleted tile B to the moving object. Thereby, reduction of the transmission data amount and reduction of the load of the decoding process can be realized.

[0445] Note that the mobile object may confirm that there are no duplicate points based on the flag. Also, if the mobile object has not acquired Tile A yet, it requests the server for data without deleting the overlapping part. Further, if the server does not have a function to delete duplicate points, or if it is not known whether there are duplicate points, the mobile object may check the delivered data to determine whether there are duplicate points, and if there are duplicate points, perform merging.

[0446] Next, an example of an application using a point cloud with duplicate points will be described. The mobile object uploads the map point cloud data acquired by LiDAR to the server in real time. For example, the mobile object uploads the data acquired for each tile to the server. In this case, there is an overlapping area between Tile A and Tile B, but the mobile object on the encoding side does not merge the duplicate points between the tiles and sends the data to the server together with a flag indicating that there is an overlap between the tiles. The server accumulates the received data as it is without merging the duplicate data included in the received data.

[0447] Also, when transmitting or storing point cloud data using a system such as ISOBMFF, MPEG-DASH / MMT, or MPEG-TS, the device may replace the flag indicating whether there are duplicate points within a tile or between tiles, which is included in GPS, with a descriptor or metadata in the system layer and store it in an SI, MPD, moov, or moof box, etc. Thereby, the application can utilize the functions of the system.

[0448] Also, as shown in FIG. 73, for example, the three-dimensional data encoding device may divide Tile B into a plurality of slices based on the overlapping area with other tiles. In the example shown in FIG. 73, slice 1 is an area that does not overlap with any tile, slice 2 is an area that overlaps with Tile A, and slice 3 is an area that overlaps with Tile C. This facilitates the separation of the desired data from the encoded data.

[0449] Alternatively, the map information may be point cloud data or mesh data. The point cloud data may be tiled for each area and stored in the server.

[0450] FIG. 74 is a flowchart showing the processing flow in the above system. First, the terminal (e.g., a mobile object) detects the movement from area A to area B of the terminal (S6271). Next, the terminal starts acquiring the map information of area B (S6272).

[0451] If the terminal has already downloaded the information of area A (Yes in S6273), the terminal instructs the server to acquire the data of area B that does not include the overlapping points with area A (S6274). The server deletes area A from area B and transmits the data of area B after deletion to the terminal (S6275). Note that the server may encode and transmit the data of area B in real time according to the instruction from the terminal so that no overlapping points occur.

[0452] Next, the terminal merges (combines) the map information of area B with the map information of area A and displays the merged map information (S6276).

[0453] On the other hand, if the terminal has not downloaded the information of area A (No in S6273), the terminal instructs the server to acquire the data of area B that includes the overlapping points with area A (S6277). The server transmits the data of area B to the terminal (S6278). Next, the terminal displays the map information of area B that includes the overlapping points with area A (S6279).

[0454] FIG. 75 is a flowchart showing another operation example in the system. The transmission device (3D data encoding device) transmits the data of the tiles in order (S6281). Also, the transmission device adds a flag indicating whether the data of the tile to be transmitted overlaps with the tile of the data transmitted one tile before to the data of the tile to be transmitted, and sends out the data (S6282).

[0455] The receiving device (3D data decoding device) determines whether the tile of the received data overlaps with the tile of the previously received data based on the flag added to the data (S6283). If the tile of the received data overlaps with the tile of the previously received data (Yes in S6283), the receiving device deletes or merges the overlapping points (S6284). On the other hand, if the tile of the received data does not overlap with the tile of the previously received data (No in S6283), the receiving device does not perform the process of deleting or merging the overlapping points and ends the process. Thereby, reduction of the processing load of the receiving device and improvement of the estimation accuracy of the attribute information can be realized. Note that if merging of overlapping points is not necessary, the receiving device does not have to perform the merging.

[0456] (Embodiment 7) In this embodiment, a display method based on a viewpoint, a random access method for encoded data, an encoding method for point cloud data, and a decoding method in an application using point clouds will be described.

[0457] With the improvement of sensor performance, it has become possible to obtain high-quality three-dimensional point clouds. However, in order to view high-quality three-dimensional points, a viewing device (viewer) that can reproduce this high-quality three-dimensional point cloud is required. Specifically, it is desired that a high-quality three-dimensional point cloud (point cloud) with a large amount of data can be displayed without delay. In this embodiment, a three-dimensional point cloud viewing device (first application) that can efficiently display high-density point cloud data by a scalable method using point cloud compression will be described.

[0458] Point cloud compression is performed by a plurality of data partitioning methods. For example, using LoD (Levels of Details), the resolution required to represent the point cloud data is calculated according to the distance between the virtual camera and the point cloud data. Thereby, separation or stratification is realized.

[0459] The three-dimensional point cloud viewing device (also referred to as a three-dimensional data decoding device) selects a visible point cloud for rendering. At this time, it is preferable for the three-dimensional data decoding device to confirm that all visible point clouds are data that have been actually scanned rather than approximated.

[0460] FIG. 76 is a block diagram showing a configuration example of a three-dimensional data encoding device. The three-dimensional data encoding device includes a point cloud encoding unit 8701 and a file format generation unit 8702. The point cloud encoding unit 8701 generates encoded data (bitstream) by encoding point cloud data. For example, the point cloud encoding unit 8701 encodes point cloud data using a position information-based encoding method using an octree, or a video-based encoding method, etc.

[0461] The file format generation unit 8702 changes the encoded data (bitstream) into data in a predetermined file format. For example, the file format is ISOBMFF or MP4, etc. Note that the three-dimensional data encoding device may output encoded data in file format (for example, transmit it to a three-dimensional data decoding device), or may output encoded data in bitstream format in the encoding method.

[0462] FIG. 77 is a block diagram showing a configuration example of a three-dimensional data decoding device 8705. The three-dimensional data decoding device 8705 generates point cloud data by decoding encoded data. Here, the encoded data is, for example, encoded data in bitstream format or MP4 format. Note that non-encoded point cloud data may also be used.

[0463] All data groups or some data groups in the point cloud are called bricks. Note that this brick may also be called divided data, a tile, or a slice. The divided data may be further divided.

[0464] The three-dimensional data decoding device 8705 acquires camera viewpoint information indicating the viewpoint (angle) of the camera from the outside. The three-dimensional data decoding device 8705 acquires part or all of the encoded data based on the camera viewpoint information, and generates point cloud data by decoding the acquired encoded data. For example, the camera viewpoint information indicates the position and direction (orientation) of the camera. Then, the three-dimensional data decoding device 8705 displays the decoded point cloud data.

[0465] The three-dimensional data decoding device 8705 includes a point cloud decoding unit 8706 and a brick decoding control unit 8707. The camera viewpoint information (camera field of view angle) is input to the brick decoding control unit 8707. The brick decoding control unit 8707 selects the bricks to be decoded based on the visibility of the bricks determined based on the camera viewpoint information. The point cloud decoding unit 8706 decodes the selected bricks and outputs the decoded bricks.

[0466] Hereinafter, the configuration of the three-dimensional data encoding device according to the present embodiment will be described. FIG. 78 is a block diagram showing the configuration of the three-dimensional data encoding device 8710 according to the present embodiment. The three-dimensional data encoding device 8710 generates encoded data (encoded stream) by encoding point cloud data (point cloud). This three-dimensional data encoding device 8710 includes a division unit 8711, a plurality of position information encoding units 8712, a plurality of attribute information encoding units 8713, an additional information encoding unit 8714, a multiplexing unit 8715, and a normal vector generation unit 8716.

[0467] The splitting unit 8711 generates a plurality of split data by splitting the point cloud data. Specifically, the splitting unit 8711 generates a plurality of split data by splitting the space of the point cloud data into a plurality of sub-spaces. Here, the sub-space is any one of a brick, a tile, and a slice, or a combination of two or more of a brick, a tile, and a slice. More specifically, the point cloud data includes position information, attribute information (such as color or reflectance), and additional information. The splitting unit 8711 generates a plurality of split position information by splitting the position information, and generates a plurality of split attribute information by splitting the attribute information. Further, the splitting unit 8711 generates additional information regarding the splitting.

[0468] The plurality of position information encoding units 8712 generate a plurality of encoded position information by encoding the plurality of split position information. For example, the position information encoding unit 8712 encodes the split position information using an N-ary tree structure such as an octree. Specifically, in an octree, the target space is divided into eight nodes (sub-spaces), and 8-bit information (occupancy code) indicating whether or not a point cloud is included in each node is generated. Further, the node including the point cloud is further divided into eight nodes, and 8-bit information indicating whether or not a point cloud is included in each of the eight nodes is generated. This process is repeated until the number of point clouds included in a predetermined hierarchy or node becomes less than or equal to a threshold value. For example, the plurality of position information encoding units 8712 perform parallel processing on the plurality of split position information.

[0469] The attribute information encoding unit 8713 generates encoded attribute information, which is encoded data, by encoding the attribute information using the configuration information generated by the position information encoding unit 8712. For example, the attribute information encoding unit 8713 determines a reference point (reference node) to be referred to in the encoding of the target point (target node) to be processed based on the octree structure generated by the position information encoding unit 8712. For example, the attribute information encoding unit 8713 refers to a node in which the parent node in the octree is the same as the target node among the surrounding nodes or adjacent nodes. Note that the method for determining the reference relationship is not limited to this.

[0470] Also, the encoding process of the position information or the attribute information may include at least one of a quantization process, a prediction process, and an arithmetic encoding process. In this case, reference means using a reference node to calculate a predicted value of the attribute information, or using the state of the reference node (for example, occupancy information indicating whether a point group is included in the reference node) to determine an encoding parameter. For example, the encoding parameter is a quantization parameter in the quantization process or a context in the arithmetic encoding, etc.

[0471] The normal vector generation unit 8716 calculates a normal vector for each divided data. Note that the input data does not necessarily have to be divided. In this case, the normal vector generation unit 8716 may calculate a normal vector for each point instead of a normal vector for each divided data. Or, the normal vector generation unit 8716 may calculate both a normal vector for each divided data and a normal vector for each point.

[0472] The additional information encoding unit 8714 generates encoded additional information by encoding the additional information included in the point cloud data, the additional information regarding data division generated at the time of division by the division unit 8711, and the normal vector generated by the normal vector generation unit 8716.

[0473] The multiplexing unit 8715 generates encoded data (encoded stream) by multiplexing a plurality of encoded position information, a plurality of encoded attribute information, and the encoded additional information, and sends out the generated encoded data. Also, the encoded additional information is used at the time of decoding.

[0474] The configuration of the three-dimensional data decoding apparatus according to the present embodiment will be described below. FIG. 79 is a block diagram showing the configuration of the three-dimensional data decoding apparatus 8720. The three-dimensional data decoding apparatus 8720 restores point cloud data by decoding encoded data (encoded stream) generated by encoding the point cloud data. This three-dimensional data decoding apparatus 8720 includes a demultiplexing unit 8721, a plurality of position information decoding units 8722, a plurality of attribute information decoding units 8723, an additional information decoding unit 8724, a combining unit 8725, a normal vector extraction unit 8726, a random access control unit 8727, and a selection unit 8728.

[0475] The demultiplexing unit 8721 generates a plurality of encoded position information, a plurality of encoded attribute information, and encoded additional information by demultiplexing the encoded data (encoded stream). The additional information decoding unit 8724 generates additional information by decoding the encoded additional information.

[0476] The normal vector extraction unit 8726 extracts a normal vector from the additional information. The random access control unit 8727 determines the divided data to be extracted, for example, based on the normal vector for each divided data. The selection unit 8728 extracts a plurality of divided data (a plurality of encoded position information and a plurality of encoded attribute information) determined by the random access control unit 8727 from the plurality of divided data (a plurality of encoded position information and a plurality of encoded attribute information). Note that the selection unit 8728 may extract one divided data.

[0477] The plurality of position information decoding units 8722 generate a plurality of divided position information by decoding the plurality of encoded position information extracted by the selection unit 8728. For example, the plurality of position information decoding units 8722 perform parallel processing on the plurality of encoded position information.

[0478] The plurality of attribute information decoding units 8723 generate a plurality of divided attribute information by decoding the plurality of encoded attribute information extracted by the selection unit 8728. For example, the plurality of attribute information decoding units 8723 perform parallel processing on the plurality of encoded attribute information.

[0479] The combining unit 8725 generates position information by combining a plurality of divided position information using additional information. The combining unit 8725 generates attribute information by combining a plurality of divided attribute information using additional information.

[0480] Next, a first example of generating and encoding a normal vector for each point will be described. FIG. 80 is a diagram showing an example of point cloud data. FIG. 81 is a diagram showing an example of a normal vector for each point. Encoding of the normal vector can be performed independently for each three-dimensional group. FIGS. 80 and 81 show the present three-dimensional point group and the normal vector of the three-dimensional point group. As shown in FIG. 81, there are a plurality of normal vectors extending in the upward, rightward, and forward directions. Here, the surface of the present invention is a plane, and the normal vectors of a plurality of points on a certain surface extend in the same direction. On the other hand, when the surface is round, the normal vectors extend in a plurality of directions according to the normal of the surface.

[0481] FIG. 82 is a diagram showing a syntax example of a normal vector in a bit stream. In the normal vector NormalVector[i][face] shown in FIG. 82, "i" represents the counter of each three-dimensional point group, and [face] represents the x, y, and z axes representing the three-dimensional point group. That is, NormalVector represents the magnitude of the normal vector of each axis.

[0482] FIG. 83 is a flowchart of three-dimensional data encoding processing. First, the three-dimensional data encoding device encodes position information (geometry) and attribute information for each point (S8701). For example, the three-dimensional data encoding device encodes the position information for each point. Also, when there is attribute information corresponding to a point, the three-dimensional data encoding device may encode the attribute information for each point.

[0483] Next, the three-dimensional data encoding device encodes the normal vector (x, y, z) for each point (S8702). The three-dimensional data encoding device may encode the normal vector for each point. Further, the three-dimensional data encoding device may encode, for example, difference information indicating the difference between the normal vector of the point to be processed and the normal vectors of other points. Thereby, the data amount can be reduced. Also, the three-dimensional data encoding device may encode the normal vector by including it in the position information or in the attribute information. Further, the three-dimensional data encoding device may encode the normal vector independently of the position information and the attribute information. Note that when there are a plurality of normal vectors for one point, the three-dimensional data encoding device may encode the plurality of normal vectors for each point.

[0484] FIG. 84 is a flowchart of three-dimensional data decoding processing. First, the three-dimensional data decoding device decodes the position information and the attribute information from the bit stream for each point (S8706). Next, the three-dimensional data decoding device decodes the normal vector from the bit stream for each point (S8707).

[0485] Note that the processing order shown in FIGS. 83 and 84 is an example, and the encoding order and the decoding order may be interchanged.

[0486] Also, the three-dimensional data encoding device may reduce the data amount by encoding the normal vector using the position information or the correlation of the position information. In that case, the three-dimensional data decoding device decodes the normal vector using the position information. By the above method, the normal vector for each point in the point cloud can be encoded and decoded.

[0487] Next, a second example of generating and encoding the normal vector for each point will be described. As another method of encoding the normal vector of each point, the normal vector is encoded as one of the attribute information. Hereinafter, an example of performing encoding using an attribute information encoding unit or an attribute information decoding unit as one of the attribute information will be described.

[0488] For example, a three-dimensional data encoding device encodes color information as first attribute information and a normal vector as second attribute information. FIG. 85 is a diagram showing a configuration example of a bit stream. For example, Attr(0) shown in FIG. 85 is encoded data of the first attribute information, and Attr(1) is encoded data of the second attribute information. Also, metadata related to encoding is stored in a parameter set (APS). The three-dimensional data decoding device decodes the encoded data by referring to the APS corresponding to the encoded data.

[0489] Note that the SPS stores identification information (attribute_type = Normal Vector) indicating that the second attribute information is a normal vector. Also, when the attribute information is a normal vector, information indicating that the normal vector is data having three elements for each point may be stored in the SPS or the like. Also, the SPS stores identification information (attribute_type = Color) indicating that the first attribute information is color information.

[0490] FIG. 86 is a diagram showing an example of point cloud information having position information, color information, and a normal vector. The three-dimensional data encoding device encodes the uncompressed point cloud data shown in FIG. 86.

[0491] The value range of the normal vector is from the value -1 to 1 in floating point. To facilitate representation, the three-dimensional data encoding device may convert the floating point to an integer according to the required accuracy. For example, the three-dimensional data encoding device may convert the floating point to a value from -127 to 128 using an 8-bit representation. That is, the three-dimensional data encoding device may convert the floating point to an integer or a positive integer value. Since the normal vector is treated as one attribute information, different quantization processes can be applied. For example, different quantization parameters can be used for each attribute information. Thereby, different accuracy levels can be realized. For example, the quantization parameters are stored in the APS.

[0492] Figure 87 is a flowchart of three-dimensional data encoding processing. First, the three-dimensional data encoding device encodes position information and attribute information (such as color information) for each point (S8711). Also, the three-dimensional data encoding device encodes the normal vector for each point as attribute information with attribute_type = "normal vector" using a predetermined method (S8712).

[0493] Figure 88 is a flowchart of three-dimensional data decoding processing. The three-dimensional data decoding device decodes position information and attribute information from the bitstream for each point (S8716). Also, the three-dimensional data decoding device decodes the normal vector for each point from the bitstream as attribute information with attribute_type = "normal vector" using a predetermined method (S8717).

[0494] Note that the processing order shown in FIGS. 87 and 88 is an example, and the encoding order and the decoding order may be interchanged.

[0495] Next, an example of generating a normal vector for each data unit including a plurality of points will be described. The three-dimensional data encoding device divides the point cloud data into a plurality of objects or a plurality of regions based on the position information and characteristics of the point cloud. The divided data is, for example, tiles or slices, or hierarchical data. The three-dimensional data encoding device generates a normal vector for this divided data unit, that is, a data unit including one or more points.

[0496] Here, visibility can be determined by the normal vector representation of the objects within the block. FIGS. 89 and 90 are diagrams for explaining this process. For example, as shown in FIG. 89, the three-dimensional data encoding device divides the normal vector direction at angles with an interval of 30° with respect to the horizontal axis and the vertical axis. As a simpler method, as shown in FIG. 90, the three-dimensional data encoding device may divide the normal vector into six directions of (0, 0), (0, 90), (0, -90), (90, 0), (-90, 0), (180, 180).

[0497] In addition, the three-dimensional data encoding device may calculate a valid normal vector using the median value, average value, or other more effective algorithms. Also, the three-dimensional data encoding device may use a representative value as the value of the valid normal vector, or may use other methods.

[0498] Also, the normal vector for each piece of divided data may directly indicate the original x, y, and z values, or may be quantized every 30 degrees as described above, or may be quantized into information every 90 degrees. Quantization can reduce the amount of information.

[0499] FIG. 91 is an example of point cloud data and shows an example of a face object. FIG. 92 is a diagram showing an example of the normal vector in this case. As shown in FIG. 92, the normal vectors of the face object shown in FIG. 91 point in the (0, 0) and (90, 0) directions. The three-dimensional data encoding device can use 1 bit for each direction to indicate whether there is a normal vector of the object in that direction.

[0500] Thus, there may be two or more normal vectors for one piece of divided data. In that case, multiple normal vectors may be shown for one unit of divided data.

[0501] For example, an example of the data including the face object shown in FIGS. 91 and 92 is an example of showing the normal vectors of the data in units of 90 degrees and six normal vectors for each face. In this example, two normal vectors in the (0, 0) and (90, 0) directions are the normal vectors of this divided data.

[0502] Also, as a method of indicating the normal vector, each of the six normal vectors may be represented by 1-bit information. FIG. 93 is a diagram showing an example of this normal vector information. When the divided data has the corresponding normal vector, the 1-bit information is set to the value 1, and when it does not have it, it is set to 0. Thereby, compared with the method of directly indicating the x, y, and z values, the amount of information can be reduced by quantizing the data.

[0503] The following describes a simpler method of expressing the normal vector. A six-sided cube is used to represent the normal vector and its feasibility (visibility) from a specific camera viewpoint. FIGS. 94 to 97 are diagrams for explaining this process. FIG. 94 shows an example of a six-sided cube. FIGS. 95, 96, and 97 show the front and back faces a and b, the left and right faces c and d, and the top and bottom faces e and f, respectively. Depending on the direction of the object according to the viewing angle, the normal vector faces at least one or three faces. Six flags of 1 bit each can be used to represent one of the six faces (abcdef) of the cube representing each system. For example, when viewed from the front (100000), when viewed from the side (001000), and when viewed from below (000001) are generated. In this representation, the magnitude is not important, only the direction is represented. There may also be an object for which three faces are specified. Face a is the opposite face of face b, face c is the opposite face of face d, and face e is the opposite face of face f. Therefore, it is impossible to see faces a and b at the same time. That is, the normal vector can be represented using three flags (ace).

[0504] Thus, when the camera viewpoint (camera angle) is known in advance, the information of the normal vector can be represented by 3 bits. FIG. 98 is a diagram showing the visibility when viewing the object of slice A or slice B from the direction of face c. Since slice A is visible from the direction of face c, it is represented as ace = (010). On the other hand, since slice B is hidden by slice A when viewed from the direction of face c, it is represented as ace = (000).

[0505] Next, a first method of encoding and decoding the normal vector for each brick will be described. FIG. 99 is a diagram showing an example of the configuration of the bit stream in this case. In the example shown in FIG. 99, the information of the normal vector is stored in the slice header of the position information in each slice. Note that the information of the normal vector may be stored in the header of the attribute information, or may be stored in metadata independent of the position information and the attribute information.

[0506] Figure 100 is a diagram showing a syntax example of a geometry slice header information of position information. The geometry slice header information of position information includes normal_vector_number, normal_vector_x, normal_vector_y, and normal_vector_z.

[0507] normal_vector_number indicates the number of normal vectors corresponding to the slice data. normal_vector_x, normal_vector_y, and normal_vector_z indicate the elements (x, y, z) of the normal vectors corresponding to the slice data, respectively.

[0508] In this example, the number of normal_vectors can be changed. The number of normal_vector_number indicates the number of normal_vectors shown.

[0509] If the information of the normal vectors is common for all slices, normal_vector_number may be stored in GPS or SPS that can store common information for multiple slices.

[0510] Also, the values of the normal vectors of x, y, and z may be quantized. For example, the three-dimensional data encoding device may quantize the values of the original normal vectors by shifting them by a common bit amount s (bit), and send out the information indicating the bit amount s and the information indicating the quantized normal vectors (normal_vector_x<<s, normal_vector_y<<s, normal_vector_z<<s). This can reduce the bit amount.

[0511] Figure 101 is a diagram showing another syntax example of the geometry slice header of position information. This example shows the normal vectors simplified (quantized) for the six-sided data for each divided data. For each face, it is shown whether there is a normal vector or not.

[0512] This slice header of the position information includes is_normal_vector. is_normal_vector is set to 1 if there is a normal vector corresponding to the slice data, and set to 0 if there is no normal vector. For example, the order of a plurality of surfaces is predetermined.

[0513] Note that the quantization accuracy and the number or order of the normal vectors are not limited to this. These may be fixed or variable.

[0514] Figure 102 is a flowchart of three-dimensional data encoding processing. First, the three-dimensional data encoding device generates a plurality of divided data by dividing the point cloud data (S8721). Next, the three-dimensional data encoding device encodes the position information and the attribute information for each divided data (S8722). Next, the three-dimensional data encoding device stores the normal vector for each divided data in the slice header (S8723).

[0515] Figure 103 is a flowchart of three-dimensional data decoding processing. First, the three-dimensional data decoding device decodes the position information and the attribute information for each divided data from the bit stream (S8726). Next, the three-dimensional data decoding device decodes the normal vector for each divided data from the slice header for each divided data (S8727). Next, the three-dimensional data decoding device combines the plurality of divided data (S8728).

[0516] Figure 104 is a flowchart of three-dimensional data decoding processing when decoding part of the data. First, the three-dimensional data decoding device decodes the normal vector for each divided data from the slice header for each divided data (S8731). Next, the three-dimensional data decoding device determines the divided data to be decoded based on the normal vector, and decodes the determined divided data (S8732). Next, the plurality of decoded divided data are combined (S8733).

[0517] Next, a second method for encoding and decoding the normal vector for each brick will be described. Another method for encoding the information of the normal vector is a method using metadata (for example, SEI: Supplemental Enhancement Information). FIG. 105 is a diagram showing a configuration example of a bit stream. As shown in FIG. 105, the SEI may be included in the bit stream, or may be generated as a separate file separately from the main encoded bit stream depending on how the SEI is implemented in both the encoding device and the decoding device.

[0518] FIG. 106 is a diagram showing a syntax example of slice information included in the SEI. The slice information includes number_of_slice, bounding_box_origin_x, bounding_box_origin_y, and bounding_box_origin_z, bounding_box_width, bounding_box_height, and bounding_box_depth, normalVector_QP, number_of_normal_vector, normalVector_x, normalVector_y, and normalVector_z.

[0519] number_of_slice indicates the number of divided data. bounding_box_origin_x, bounding_box_origin_y, and bounding_box_origin_z indicate the origin coordinates of the bounding box of the slice data. bounding_box_width, bounding_box_height, and bounding_box_depth indicate the width, height, and depth of the bounding box of the slice data, respectively.

[0520] When the normal_vector is quantized, normalVector_QP indicates the quantization scale information or bit shift information thereof. number_of_normal_vector indicates the number of normal vectors included in the slice data. normalVector_x, normalVector_y, and normalVector_z indicate the components of the elements (x, y, z) of the normal vector, respectively.

[0521] FIG. 107 is a diagram showing another example of slice information included in the SEI. The example shown in FIG. 107 is an example showing the normal vectors simplified (quantized) into six-sided data for each divided data. For each surface, it is shown whether there is a normal vector or not.

[0522] This slice information includes is_normal_vector. is_normal_vector is set to 1 if there is a normal vector corresponding to the slice data, and is set to 0 if there is no normal vector. For example, the order of a plurality of surfaces is predetermined.

[0523] Note that the slice information may include a flag indicating whether the slice information includes information on the bounding box (origin, width, height, and depth) for each slice. In this case, when the flag is on (for example, 1), the slice information includes information on the bounding box for each slice, and when the flag is off (for example, 0), the slice information does not include information on the bounding box for each slice. Also, the slice information may include a flag indicating whether the slice information includes information on the normal vector for each slice. In this case, when the flag is on (for example, 1), the slice information includes information on the normal vector for each slice, and when the flag is off (for example, 0), the slice information does not include information on the normal vector for each slice.

[0524] Next, random access and partial decoding will be described. The three-dimensional data decoding device independently decodes data for each slice using either or both of the information for each slice, for example, the bounding box information and the normal vector of the slice.

[0525] FIG. 108 is a flowchart of the three-dimensional data decoding process. First, the three-dimensional data decoding device determines the slice to be decoded by a predetermined method and the decoding order of the slices (S8741). Next, the three-dimensional data decoding device decodes a specific slice in the determined order (S8742).

[0526] FIG. 109 is a diagram showing an example of this partial decoding process. For example, the three-dimensional data decoding device receives the slice-divided encoded data shown in FIG. 109(a). As shown in FIG. 109(b), the three-dimensional data decoding device decodes the encoded data of some slices and does not decode the encoded data of other slices. Alternatively, as shown in FIG. 109(c), the three-dimensional data decoding device decodes by changing the order of the encoded data.

[0527] FIG. 110 is a diagram showing a configuration example of the three-dimensional data decoding device. As shown in FIG. 110, the three-dimensional data decoding device includes an attribute information decoding unit 8731 and a random access control unit 8732. The attribute information decoding unit 8731 extracts the information of the bounding box and the normal vector for each slice from the encoded data. The random access control unit 8732 determines the number and order of the slices to be decoded based on the information of the bounding box and the normal vector for each slice and the sensor information acquired from the outside, for example, the camera angle (camera orientation) and the camera position.

[0528] FIG. 111 and FIG. 112 are diagrams showing processing examples of the random access control unit 8732. As shown in FIG. 111, for example, the random access control unit 8732 may calculate a bounding box for each slice and distance information indicating the distance from the camera for each slice based on the camera position. Alternatively, as shown in FIG. 112, the random access control unit 8732 may derive, for each slice, visible information indicating whether an object is visible from the camera from the normal vector for each slice and the camera angle. Note that the random access control unit 8732 may calculate either the distance information or the visible information, or both.

[0529] Hereinafter, the visible information and the distance information will be described. FIG. 113 is a diagram showing an example of the relationship between the distance and the resolution. For example, what is visible from the camera is decoded (frustum culling). Further, the decoded resolution depends on the distance between the virtual camera and the point cloud data.

[0530] That is, the three-dimensional data decoding device determines whether a slice is visible from the camera based on the normal vector for each slice and the camera viewpoint (camera angle), and decodes the slices visible from the camera. Further, the three-dimensional data decoding device calculates the distance from the camera of the slice to be decoded, and when the distance from the camera is short, decodes high-resolution data, and when the distance from the camera is long, may decode low-resolution data.

[0531] In this case, the encoded data is hierarchically encoded, and the three-dimensional data decoding device can independently decode the low-resolution data. Further, when the three-dimensional data decoding device decodes high-resolution data, it further decodes the difference information between the low-resolution data and the high-resolution data, and generates the high-resolution data by adding the difference information to the low-resolution data. Note that when the encoded data is not hierarchically encoded, the three-dimensional data decoding device may not perform the above processing, or may determine whether to perform the above processing according to whether it is hierarchically encoded.

[0532] Next, the visibility determination using the normal vector will be described. FIG. 114 is a diagram showing an example of a brick and a normal vector. In the example shown in FIG. 114, two bricks (for example, slices) on the front surface facing the camera (frustum), that is, the bricks whose normal vectors are facing the camera, are decoded.

[0533] First, for each slice data, the three-dimensional data decoding device determines whether there is a normal vector having a normal vector opposite to the camera direction among one or more normal vectors included in the metadata. When there is a normal vector having a normal vector opposite to the camera direction in the slice data of the target slice, the three-dimensional data decoding device determines that the target slice is visible and determines the target slice as a decoding target.

[0534] Note that when there are other slices between the camera and the target slice, the three-dimensional data decoding device may determine that the target slice is invisible (not visible). Also, instead of determining whether the normal vector and the camera direction are completely opposite, the three-dimensional data decoding device may determine whether the relationship between the normal vector and the camera direction is within a predetermined angle range to determine whether it is visible.

[0535] Next, the process using LoD (Level of Detail) will be described. Hereinafter, an example of the decoding process according to layers with different resolutions will be described.

[0536] FIG. 115 is a diagram showing an example of the level (LoD). FIG. 116 is a diagram showing an example of an octree structure. Each brick is divided into layers to control the level of resolution to be decoded. For example, the level is the depth of division when dividing into an octree. As shown in FIG. 115, the number of voxels included in each level may be defined as 2(3×level). Note that another definition may be used for the division method or the number of voxels according to the level.

[0537] By using LoD, the three-dimensional data decoding device can achieve high-speed visibility determination and distance calculation. The decoding time affects real-time rendering. By using LoD, it becomes possible to display intermediate bricks, so that real-time rendering and smooth correspondence can be realized.

[0538] FIG. 117 is a flowchart of three-dimensional data decoding processing using LoD. First, the three-dimensional data decoding device determines the level to be decoded according to the purpose (S8751). Next, the three-dimensional data decoding device decodes the first level (level 0) (S8752). Next, the three-dimensional data decoding device determines whether decoding of all levels to be decoded has been completed (S8753). If decoding of all levels has not been completed (No in S8753), the three-dimensional data decoding device decodes the next level (S8754). At this time, the three-dimensional data decoding device may decode the next level using the data of the previous level. When decoding of all levels to be decoded has been completed (Yes in S8753), the three-dimensional data decoding device displays the decoded data (S8755).

[0539] In this way, the three-dimensional data decoding device decodes data up to the determined level and does not decode data after the determined level. Thereby, the processing amount related to decoding can be reduced and the processing speed can be improved. Also, the three-dimensional data decoding device displays the data up to the determined level and does not display the data after the determined level. Thereby, the processing amount related to display can be reduced and the processing speed can be improved. Note that the three-dimensional data decoding device may determine the level of the brick to be decoded based on, for example, the distance of the brick from the camera or whether the brick is visible from the camera.

[0540] Next, an implementation example of the process using LoD will be described. FIG. 118 is a flowchart of the three-dimensional data decoding process. First, the three-dimensional data decoding device acquires encoded data (S8761). For example, the encoded data is point cloud data encoded and compressed using an arbitrary encoding method. The encoded data may be in the form of a bit stream or in a file format.

[0541] Next, the three-dimensional data decoding device acquires the normal vector and position information of the block to be processed from the encoded data (S8762). For example, the three-dimensional data decoding device acquires the normal vector for each block and the position information of the block from the metadata (SEI or data header) included in the encoded data. Note that the three-dimensional data decoding device may determine the distance between the block and the camera from the position information of the block and the information on the camera position. Further, the three-dimensional data decoding device may determine the visibility of the block (whether the block faces the direction of the camera) from the normal vector and the camera direction.

[0542] Next, the three-dimensional data decoding device determines which block to decode and decodes the first level (level 0) of the determined block (S8763). FIG. 119 is a diagram showing an example of the block to be decoded. As shown in FIG. 119, the three-dimensional data decoding device decodes all visible blocks at the resolution of level 0.

[0543] Next, the three-dimensional data decoding device determines whether to decode the next level of each block according to the position information, and decodes the next level of the block determined to be decoded (S8764). Also, this process is repeated until the decoding process for all levels is completed (S8765). Specifically, the resolution of the block close to the position of the virtual camera is set high. For example, according to resources such as memory, levels for preferentially and gradually decoding the blocks close to the camera are added.

[0544] FIG. 120 is a diagram showing an example of the level of the block to be decoded. As shown in FIG. 120, the three-dimensional data decoding device decodes the blocks closer to the camera at a higher resolution and the blocks farther from the camera at a lower resolution according to the distance from the camera. Also, the three-dimensional data decoding device does not decode the blocks that cannot be seen.

[0545] When the decoding of all levels is completed (Yes in S8765), the three-dimensional data decoding device outputs the obtained three-dimensional point cloud (S8766).

[0546] So far, the method of calculating and encoding the normal vector and the bounding information for each slice data in the three-dimensional data encoding device, and calculating the visibility and distance information based on the information and the sensor input information in the three-dimensional data decoding device, and determining the slice to be decoded has been described. Hereinafter, an example of calculating and encoding the visibility and distance information corresponding to the camera direction in advance in the three-dimensional data encoding device for the data of each slice will be described.

[0547] FIG. 121 is a diagram showing an example of the syntax of the geometry slice header information of the position information. The geometry slice header of the position information includes number_of_angle, view_angle, and visibility.

[0548] number_of_angle indicates the number of camera angles (camera directions). view_angle indicates the camera angle, for example, indicates the vector of the camera angle. visibility indicates whether the slice is visible from the corresponding camera angle. Note that the number of view_angle may be variable or may be a predetermined fixed value. Also, when the number and value of view_angle are predetermined, view_angle may be omitted.

[0549] Also, although an example showing visibility according to the camera angle has been shown here, as another example, the three-dimensional data encoding device may calculate in advance the visibility according to the camera position or camera parameters, and store the calculated visibility in the encoded data.

[0550] FIG. 122 is a flowchart of three-dimensional data encoding processing. First, the three-dimensional data encoding device divides the point cloud data into divided data (for example, slices) (S8771). Next, the three-dimensional data encoding device encodes the position information and the attribute information for each divided data unit (S8772). Also, the three-dimensional data encoding device stores, in the metadata, visibility information according to the camera angle for each divided data (S8773).

[0551] FIG. 123 is a flowchart of three-dimensional data decoding processing. First, the three-dimensional data decoding device acquires the visibility information according to the camera angle from the metadata for each divided data (S8776). Next, the three-dimensional data decoding device determines the divided data that is visible from the desired camera angle based on the visibility information, and decodes the visible divided data (S8777).

[0552] FIGS. 124 and 125 are diagrams showing examples of point cloud data. In the figure, a, c, d, and e represent planes. Therefore, the three-dimensional data encoding device can perform slice division by utilizing the fact that the three-dimensional points of each slice have a normal vector in the same direction in slice division. The same method can also be applied to tile division.

[0553] FIGS. 126 to 129 are diagrams showing a configuration example of a system including a three-dimensional data encoding device, a three-dimensional data decoding device, and a display device.

[0554] In the example shown in FIG. 126, the three-dimensional data encoding device generates encoded data by encoding slice data, the normal vector for each slice, and bounding box information. The three-dimensional data decoding device identifies the data to be decoded from the encoded data and sensor information, and generates decoded slice data by decoding the identified data. The display device displays the decoded slice data. With this configuration, the three-dimensional data decoding device can flexibly determine the visible information and whether to decode it.

[0555] In the example shown in FIG. 127, the three-dimensional data encoding device generates encoded data by encoding slice data, the normal vector for each slice, and bounding box information. The three-dimensional data decoding device determines the data and order to be decoded from the encoded data and sensor information, and decodes the determined data in the determined order. With this configuration, the three-dimensional data decoding device can decode the data that it wants to display first (for example, 3, 4, 5) first, so the comfort of display can be improved.

[0556] In the example shown in FIG. 128, the three-dimensional data encoding device generates encoded data by encoding slice data and visible information for each camera angle. The three-dimensional data decoding device identifies the data to be decoded from the encoded data information and sensor information, and decodes the identified data. Note that the three-dimensional data decoding device may further determine the decoding order. With this configuration, since the three-dimensional data decoding device does not have to calculate the visible information, the processing amount of the three-dimensional data decoding device can be reduced.

[0557] In the example shown in FIG. 129, the three-dimensional data decoding device notifies the three-dimensional data encoding device of the camera angle or camera position of the three-dimensional data decoding device via communication or the like. The three-dimensional data encoding device calculates the visible information for each slice, determines the data and order to be encoded, and generates encoded data by encoding the determined data in the determined order. The three-dimensional data decoding device decodes the transmitted slice data as it is. With this configuration, by using an interactive configuration, the necessary parts are encoded and decoded, so that the processing amount and communication bandwidth can be reduced.

[0558] In addition, when the camera position or the camera angle changes, the three-dimensional data decoding device may re-determine the slice to be decoded when the amount of change exceeds a predetermined value. In that case, high-speed decoding and display can be achieved by decoding the differential data other than the already decoded data.

[0559] Hereinafter, a method of storing encoded data in a file format such as ISOBMFF will be described. FIG. 130 is a diagram showing a configuration example of a bit stream. FIG. 131 is a diagram showing a configuration example of a three-dimensional data encoding device. The three-dimensional data encoding device includes an encoding unit 8741 and a file conversion unit 8742. The encoding unit 8741 generates a bit stream including encoded data and control information by encoding point cloud data. The file conversion unit 8742 converts the bit stream into a file format.

[0560] FIG. 132 is a diagram showing a configuration example of a three-dimensional data decoding device. The three-dimensional data decoding device includes a file inverse conversion unit 8751 and a decoding unit 8752. The file inverse conversion unit 8751 converts the file format into a bit stream including encoded data and control information. The decoding unit 8752 generates point cloud data by decoding the bit stream.

[0561] FIG. 133 is a diagram showing the basic structure of ISOBMFF. FIG. 134 is a protocol stack diagram when storing NAL units common to the PCC codec in ISOBMFF. Here, what is stored in ISOBMFF is the NAL unit of the PCC codec.

[0562] The NAL unit includes a NAL unit for data and a NAL unit for metadata. The NAL unit for data includes position information slice data (Geometry Slice Data), attribute information slice data (Attribute Slice Data), and the like. The NAL unit for metadata includes SPS, GPS, APS, and SEI, and the like.

[0563] ISOBMFF (ISO based media file format) is a file format standard defined in ISO / IEC 14496-12, which defines a format that can multiplex and store various media such as video, audio, and text, and is a media-independent standard.

[0564] The basic unit in ISOBMFF is a box. A box is composed of type, length, and data, and a set of boxes with various types combined is a file. Mainly, a file is composed of boxes such as ftyp that indicates the file brand in 4CC, moov that stores metadata such as control information, and mdat that stores data.

[0565] The storage method for each media in ISOBMFF is separately defined. For example, the storage methods for AVC video and HEVC video are defined in ISO / IEC 14496-15. Also, in order to store and transmit PCC encoded data, it is conceivable to extend and use the functions of ISOBMFF.

[0566] When storing NAL units for metadata in ISOBMFF, SEI may be stored in the "mdat box" together with PCC data, or may be stored in the "track box" that describes control information related to the stream. Also, when data is packetized and transmitted, SEI may be stored in the packet header. By indicating SEI at the system layer, access to attribute information, tiles, and slice data becomes easier, and the access speed is improved.

[0567] Next, the method for generating a PCC random access table will be described. The three-dimensional data encoding device generates a random access table using metadata including bounding box information and normal vector information for each slice. Figure 135 is a diagram showing an example of converting a bitstream into a file format.

[0568] The three-dimensional data encoding device stores slice data in mdat of the file format respectively. The three-dimensional data encoding device calculates the memory position of the slice data as offset information (offsets 1 to 4 in FIG. 135) at the head of the file, and includes the calculated offset information in a random access table (PCC random access table).

[0569] FIG. 136 is a diagram showing a syntax example of slice information. FIGS. 137 to 139 are diagrams showing syntax examples of the PCC random access table.

[0570] The PCC random access table includes bounding box information (bounding_box_info), normal vector information (normal_vector_info), and offset information (offset) stored in slice information (slice_information).

[0571] The three-dimensional data decoding device analyzes the PCC random access table to identify the slice to be decoded. The three-dimensional data decoding device can access the desired data by obtaining offset information from the PCC random access table.

[0572] As described above, the three-dimensional data encoding device according to the present embodiment performs the processing shown in FIG. 140. The three-dimensional data encoding device generates a bit stream by encoding the position information and one or more attribute information of each of the plurality of three-dimensional points included in the point cloud data (S8781). In the encoding (S8781), the normal vector of each of the plurality of three-dimensional points is encoded as one piece of attribute information included in the one or more pieces of attribute information.

[0573] According to this, the three-dimensional data encoding device can process the normal vector in the same way as other attribute information by encoding the normal vector as attribute information. Therefore, the three-dimensional data encoding device can reduce the processing amount. That is, the three-dimensional data encoding device can encode the normal vector as attribute information without changing the definition of the attribute information or the like.

[0574] For example, in the encoding (S8781), the three-dimensional data encoding device converts the normal vector expressed in floating point numbers into an integer and then encodes it. According to this, the three-dimensional data encoding device can process the normal vector in the same way as other attribute information, for example, when other attribute information is expressed in integers.

[0575] For example, the bitstream includes control information (e.g., SPS) common to the position information and one or more pieces of attribute information, and the control information (e.g., SPS) includes information (e.g., attribute_type = Normal Vector) indicating that one piece of attribute information included in the one or more pieces of attribute information indicates a normal vector, or at least one of the information indicating that the normal vector is data having three elements for each point.

[0576] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor performs the above processing using the memory.

[0577] Further, the three-dimensional data decoding device according to the present embodiment performs the processing shown in FIG. 141. The three-dimensional data decoding device acquires (S8786) a bitstream generated by encoding the position information and one or more pieces of attribute information of each of a plurality of three-dimensional points included in the point cloud data, and the normal vector of each of the plurality of three-dimensional points is encoded as one piece of attribute information included in the one or more pieces of attribute information, and acquires the normal vector by decoding one piece of attribute information from the bitstream (S8787).

[0578] According to this, the three-dimensional data decoding device can process the normal vector in the same way as other attribute information by decoding the normal vector as attribute information. Therefore, the three-dimensional data decoding device can reduce the processing amount.

[0579] For example, in the acquisition of the normal vector (S8787), the three-dimensional data decoding device acquires the normal vector represented by an integer. According to this, the three-dimensional data decoding device can process the normal vector in the same way as other attribute information when, for example, other attribute information is represented by an integer.

[0580] For example, the bitstream includes control information (e.g., SPS) common to the position information and one or more pieces of attribute information, and the control information (e.g., SPS) includes at least one of information indicating that one piece of attribute information included in the one or more pieces of attribute information indicates a normal vector (e.g., attribute_type = Normal Vector), or information indicating that the normal vector is data having three elements for each point.

[0581] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor performs the above processing using the memory.

[0582] In addition, the three-dimensional data encoding device according to the present embodiment performs the processing shown in FIG. 142. The three-dimensional data encoding device divides the point cloud data into a plurality of divided data (e.g., bricks, slices, or tiles) (S8791), and generates a bitstream by encoding the plurality of divided data (S8792). The bitstream includes information indicating the normal vector of each of the plurality of divided data.

[0583] According to this, the three-dimensional data encoding device can reduce the processing amount and the encoding amount compared to the case of encoding the normal vector for each point by encoding the normal vector for each divided data. For example, each of the plurality of divided data is a random access unit.

[0584] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor performs the above processing using the memory.

[0585] In addition, the three-dimensional data decoding device according to the present embodiment performs the processing shown in FIG. 143. The three-dimensional data decoding device acquires a bitstream generated by encoding a plurality of divided data (for example, bricks, slices, or tiles) generated by dividing point cloud data (S8796), and acquires information indicating the normal vector of each of the plurality of divided data from the bitstream (S8797).

[0586] According to this, the three-dimensional data decoding device can reduce the processing amount compared to the case of decoding the normal vector for each point by decoding the normal vector for each divided data. For example, each of the plurality of divided data is a random access unit.

[0587] For example, the three-dimensional data decoding device further determines the divided data to be decoded from the plurality of divided data based on the normal vector, and decodes the divided data to be decoded.

[0588] For example, the three-dimensional data decoding device further determines the decoding order of the plurality of divided data based on the normal vector, and decodes the plurality of divided data in the determined decoding order.

[0589] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor performs the above processing using the memory.

[0590] (Embodiment 8) Next, the configuration of the three-dimensional data creation device 810 according to the present embodiment will be described. FIG. 144 is a block diagram showing a configuration example of the three-dimensional data creation device 810 according to the present embodiment. This three-dimensional data creation device 810 is mounted on a vehicle, for example. The three-dimensional data creation device 810 transmits and receives three-dimensional data to and from an external traffic monitoring cloud, a preceding vehicle, or a following vehicle, and creates and stores three-dimensional data.

[0591] The three-dimensional data creation device 810 includes a data reception unit 811, a communication unit 812, a reception control unit 813, a format conversion unit 814, a plurality of sensors 815, a three-dimensional data creation unit 816, a three-dimensional data synthesis unit 817, a three-dimensional data storage unit 818, a communication unit 819, a transmission control unit 820, a format conversion unit 821, and a data transmission unit 822.

[0592] The data reception unit 811 receives three-dimensional data 831 from the traffic monitoring cloud or the preceding vehicle. The three-dimensional data 831 includes information such as point cloud, visible light video, depth information, sensor position information, or speed information, for example, including areas that cannot be detected by the sensors 815 of the host vehicle.

[0593] The communication unit 812 communicates with the traffic monitoring cloud or the preceding vehicle and transmits a data transmission request or the like to the traffic monitoring cloud or the preceding vehicle.

[0594] The reception control unit 813 exchanges information such as the corresponding format with the communication destination via the communication unit 812 and establishes communication with the communication destination.

[0595] The format conversion unit 814 generates three-dimensional data 832 by performing format conversion or the like on the three-dimensional data 831 received by the data reception unit 811. Further, when the three-dimensional data 831 is compressed or encoded, the format conversion unit 814 performs decompression or decoding processing.

[0596] The plurality of sensors 815 is a group of sensors that acquire information outside the vehicle, such as LiDAR, a visible light camera, or an infrared camera, and generates sensor information 833. For example, when the sensor 815 is a laser sensor such as LiDAR, the sensor information 833 is three-dimensional data such as a point cloud (point group data). Note that the number of sensors 815 does not have to be plural.

[0597] The three-dimensional data creation unit 816 generates three-dimensional data 834 from the sensor information 833. The three-dimensional data 834 includes information such as, for example, point cloud, visible light video, depth information, sensor position information, or speed information.

[0598] The three-dimensional data synthesis unit 817 synthesizes the three-dimensional data 832 created by the traffic monitoring cloud or the vehicle ahead, etc., with the three-dimensional data 834 created based on the sensor information 833 of the host vehicle, thereby constructing three-dimensional data 835 that includes the space in front of the vehicle ahead that cannot be detected by the sensors 815 of the host vehicle.

[0599] The three-dimensional data storage unit 818 stores the generated three-dimensional data 835, etc.

[0600] The communication unit 819 communicates with the traffic monitoring cloud or the following vehicle, and transmits a data transmission request, etc. to the traffic monitoring cloud or the following vehicle.

[0601] The transmission control unit 820 exchanges information such as the corresponding format with the communication destination via the communication unit 819 to establish communication with the communication destination. Further, the transmission control unit 820 determines a transmission area, which is the space of the three-dimensional data to be transmitted, based on the three-dimensional data construction information of the three-dimensional data 832 generated by the three-dimensional data synthesis unit 817 and the data transmission request from the communication destination.

[0602] Specifically, the transmission control unit 820 determines a transmission area that includes the space in front of the host vehicle that cannot be detected by the sensors of the following vehicle in response to a data transmission request from the traffic monitoring cloud or the following vehicle. Further, the transmission control unit 820 determines the transmission area by judging whether there is an update of the space that can be transmitted or the transmitted space based on the three-dimensional data construction information. For example, the transmission control unit 820 determines the area specified in the data transmission request and where the corresponding three-dimensional data 835 exists as the transmission area. Then, the transmission control unit 820 notifies the format conversion unit 821 of the corresponding format of the communication destination and the transmission area.

[0603] The format conversion unit 821 generates three-dimensional data 837 by converting the three-dimensional data 836 in the transmission area among the three-dimensional data 835 stored in the three-dimensional data storage unit 818 into the format supported by the receiving side. Note that the format conversion unit 821 may reduce the data volume by compressing or encoding the three-dimensional data 837.

[0604] The data transmission unit 822 transmits the three-dimensional data 837 to the traffic monitoring cloud or the following vehicle. The three-dimensional data 837 includes information such as point cloud, visible light video, depth information, or sensor position information in front of the host vehicle, for example, including areas that are blind spots for the following vehicle.

[0605] Here, an example where format conversion and the like are performed by the format conversion units 814 and 821 has been described, but the format conversion may not be performed.

[0606] With such a configuration, the three-dimensional data creation device 810 acquires three-dimensional data 831 of an area that cannot be detected by the sensors 815 of the host vehicle from the outside, and generates three-dimensional data 835 by synthesizing the three-dimensional data 831 and three-dimensional data 834 based on the sensor information 833 detected by the sensors 815 of the host vehicle. Thereby, the three-dimensional data creation device 810 can generate three-dimensional data in a range that cannot be detected by the sensors 815 of the host vehicle.

[0607] In addition, the three-dimensional data creation device 810 can transmit three-dimensional data including the space in front of the host vehicle that cannot be detected by the sensors of the following vehicle to the traffic monitoring cloud or the following vehicle, etc., in response to a data transmission request from the traffic monitoring cloud or the following vehicle.

[0608] Next, the procedure for transmitting three-dimensional data to the following vehicle in the three-dimensional data creation device 810 will be described. FIG. 145 is a flowchart showing an example of the procedure for transmitting three-dimensional data to the traffic monitoring cloud or the following vehicle by the three-dimensional data creation device 810.

[0609] First, the three-dimensional data creation device 810 generates and updates three-dimensional data 835 of a space including the space on the road ahead of the host vehicle (S801). Specifically, the three-dimensional data creation device 810 synthesizes the three-dimensional data 831 created by a traffic monitoring cloud or a preceding vehicle, etc., with the three-dimensional data 834 created based on the sensor information 833 of the host vehicle, etc., to construct three-dimensional data 835 including the space in front of the preceding vehicle that cannot be detected by the sensor 815 of the host vehicle.

[0610] Next, the three-dimensional data creation device 810 determines whether the three-dimensional data 835 included in the transmitted space has changed (S802).

[0611] If a vehicle or a person enters the transmitted space from the outside and the three-dimensional data 835 included in the space changes (Yes in S802), the three-dimensional data creation device 810 transmits the three-dimensional data including the three-dimensional data 835 of the changed space to the traffic monitoring cloud or the following vehicle (S803).

[0612] Note that the three-dimensional data creation device 810 may transmit the three-dimensional data of the changed space in accordance with the transmission timing of the three-dimensional data transmitted at a predetermined interval, or may transmit it immediately after detecting the change. That is, the three-dimensional data creation device 810 may transmit the three-dimensional data of the changed space with priority over the three-dimensional data transmitted at a predetermined interval.

[0613] Also, the three-dimensional data creation device 810 may transmit all of the three-dimensional data of the changed space as the three-dimensional data of the changed space, or may transmit only the difference of the three-dimensional data (for example, information on three-dimensional points that have appeared or disappeared, or displacement information of three-dimensional points, etc.).

[0614] Further, the three-dimensional data creation device 810 may transmit metadata related to the danger avoidance operation of the host vehicle, such as an emergency braking warning, to the following vehicle prior to the three-dimensional data of the changed space. According to this, the following vehicle can recognize the emergency braking of the preceding vehicle, etc., earlier and can start a danger avoidance operation such as deceleration earlier.

[0615] If there is no change in the three-dimensional data 835 included in the transmitted space (No in S802), or after step S803, the three-dimensional data creation device 810 transmits the three-dimensional data included in a space of a predetermined shape at the forward distance L of the host vehicle to the traffic monitoring cloud or the following vehicle (S804).

[0616] Also, for example, the processes of steps S801 to S804 are repeatedly performed at predetermined time intervals.

[0617] Further, if there is no difference between the three-dimensional data 835 of the current transmission target space and the three-dimensional map, the three-dimensional data 837 of the space may not be transmitted by the three-dimensional data creation device 810.

[0618] In the present embodiment, the client device transmits sensor information obtained by a sensor to a server or another client device.

[0619] First, the configuration of the system according to the present embodiment will be described. FIG. 146 is a diagram showing the configuration of a three-dimensional map and sensor information transmission / reception system according to the present embodiment. This system includes a server 901 and client devices 902A and 902B. When the client devices 902A and 902B are not particularly distinguished, they are also referred to as the client device 902.

[0620] The client device 902 is, for example, an in-vehicle device mounted on a moving body such as a vehicle. The server 9 having a plurality of client devices 901 is, for example, a traffic monitoring cloud or the like and can communicate with the client devices 902.

[0621] The server 901 transmits a three-dimensional map composed of point clouds to the client device 902. Note that the configuration of the three-dimensional map is not limited to point clouds and may represent other three-dimensional data such as a mesh structure.

[0622] The client device 902 transmits the sensor information acquired by the client device 902 to the server 901. The sensor information includes, for example, at least one of LiDAR acquisition information, visible light image, infrared image, depth image, sensor position information, and speed information.

[0623] The data transmitted and received between the server 901 and the client device 902 may be compressed for data reduction, or may remain uncompressed to maintain the accuracy of the data. When compressing the data, for example, a three-dimensional compression method based on an octree structure can be used for the point cloud. Also, a two-dimensional image compression method can be used for the visible light image, infrared image, and depth image. The two-dimensional image compression method is, for example, MPEG-4 AVC or HEVC standardized by MPEG.

[0624] In addition, the server 901 transmits the three-dimensional map managed by the server 901 to the client device 902 in response to a transmission request for the three-dimensional map from the client device 902. Note that the server 901 may transmit the three-dimensional map without waiting for a transmission request for the three-dimensional map from the client device 902. For example, the server 901 may broadcast the three-dimensional map to one or more client devices 902 in a predetermined space. Also, the server 901 may transmit a three-dimensional map suitable for the position of the client device 902 to the client device 902 that has received the transmission request once every certain period of time. Further, the server 901 may transmit the three-dimensional map to the client device 902 each time the three-dimensional map managed by the server 901 is updated.

[0625] The client device 902 issues a transmission request for the three-dimensional map to the server 901. For example, when the client device 902 wants to perform self-position estimation during traveling, the client device 902 transmits a transmission request for the three-dimensional map to the server 901.

[0626] In addition, in the following cases, the client device 902 may send a request to the server 901 to transmit the three-dimensional map. When the three-dimensional map held by the client device 902 is old, the client device 902 may send a request to the server 901 to transmit the three-dimensional map. For example, when a certain period of time has elapsed since the client device 902 acquired the three-dimensional map, the client device 902 may send a request to the server 901 to transmit the three-dimensional map.

[0627] Before a certain time when the client device 902 exits from the space represented by the three-dimensional map held by the client device 902, the client device 902 may send a request to the server 901 to transmit the three-dimensional map. For example, when the client device 902 exists within a predetermined distance from the boundary of the space represented by the three-dimensional map held by the client device 902, the client device 902 may send a request to the server 901 to transmit the three-dimensional map. Further, when the movement path and movement speed of the client device 902 can be grasped, based on these, the time when the client device 902 exits from the space represented by the three-dimensional map held by the client device 902 may be predicted.

[0628] When the error at the time of alignment between the three-dimensional data created by the client device 902 from the sensor information and the three-dimensional map is a certain value or more, the client device 902 may send a request to the server 901 to transmit the three-dimensional map.

[0629] The client device 902 transmits sensor information to the server 901 in response to a transmission request for sensor information sent from the server 901. Note that the client device 902 may send the sensor information to the server 901 without waiting for a transmission request for sensor information from the server 901. For example, when the client device 902 once obtains a transmission request for sensor information from the server 901, it may periodically transmit the sensor information to the server 901 for a certain period of time. Also, when the error at the time of alignment between the three-dimensional data created by the client device 902 based on the sensor information and the three-dimensional map obtained from the server 901 is equal to or greater than a certain value, the client device 902 determines that there may be a change in the three-dimensional map around the client device 902, and may transmit the fact and the sensor information to the server 901.

[0630] The server 901 issues a transmission request for sensor information to the client device 902. For example, the server 901 receives position information of the client device 902 such as GPS from the client device 902. When the server 901 determines based on the position information of the client device 902 that the client device 902 is approaching a space with little information in the three-dimensional map managed by the server 901, the server 901 issues a transmission request for sensor information to the client device 902 to generate a new three-dimensional map. Also, when the server 901 wants to update the three-dimensional map, when it wants to check the road conditions such as during snow accumulation or a disaster, or when it wants to check the traffic jam situation, or an accident situation, etc., it may issue a transmission request for sensor information.

[0631] Also, the client device 902 may set the data volume of the sensor information to be transmitted to the server 901 according to the communication state or bandwidth at the time of receiving a transmission request for sensor information received from the server 901. Setting the data volume of the sensor information to be transmitted to the server 901 means, for example, increasing or decreasing the data itself, or appropriately selecting a compression method.

[0632] FIG. 147 is a block diagram showing a configuration example of the client device 902. The client device 902 receives a three-dimensional map composed of a point cloud or the like from the server 901, and estimates its own position of the client device 902 from the three-dimensional data created based on the sensor information of the client device 902. Further, the client device 902 transmits the acquired sensor information to the server 901.

[0633] The client device 902 includes a data reception unit 1011, a communication unit 1012, a reception control unit 1013, a format conversion unit 1014, a plurality of sensors 1015, a three-dimensional data creation unit 1016, a three-dimensional image processing unit 1017, a three-dimensional data storage unit 1018, a format conversion unit 1019, a communication unit 1020, a transmission control unit 1021, and a data transmission unit 1022.

[0634] The data reception unit 1011 receives the three-dimensional map 1031 from the server 901. The three-dimensional map 1031 is data including a point cloud such as WLD or SWLD. The three-dimensional map 1031 may include either compressed data or uncompressed data.

[0635] The communication unit 1012 communicates with the server 901 and transmits a data transmission request (for example, a transmission request for a three-dimensional map) to the server 901.

[0636] The reception control unit 1013 exchanges information such as a corresponding format with the communication destination via the communication unit 1012 and establishes communication with the communication destination.

[0637] The format conversion unit 1014 generates a three-dimensional map 1032 by performing format conversion or the like on the three-dimensional map 1031 received by the data reception unit 1011. Further, when the three-dimensional map 1031 is compressed or encoded, the format conversion unit 1014 performs decompression or decoding processing. Note that the format conversion unit 1014 does not perform decompression or decoding processing if the three-dimensional map 1031 is uncompressed data.

[0638] The plurality of sensors 1015 are a group of sensors that acquire information outside the vehicle on which the client device 902 is mounted, such as LiDAR, a visible light camera, an infrared camera, or a depth sensor, and generate sensor information 1033. For example, when the sensor 1015 is a laser sensor such as LiDAR, the sensor information 1033 is three-dimensional data such as point cloud (point group data). Note that the number of sensors 1015 does not have to be plural.

[0639] The three-dimensional data creation unit 1016 creates three-dimensional data 1034 around the host vehicle based on the sensor information 1033. For example, the three-dimensional data creation unit 1016 creates point cloud data with color information around the host vehicle using the information acquired by LiDAR and the visible light video obtained by the visible light camera.

[0640] The three-dimensional image processing unit 1017 performs self-position estimation processing of the host vehicle using the three-dimensional map 1032 such as the received point cloud and the three-dimensional data 1034 around the host vehicle generated from the sensor information 1033. Note that the three-dimensional image processing unit 1017 may create three-dimensional data 1035 around the host vehicle by synthesizing the three-dimensional map 1032 and the three-dimensional data 1034, and perform self-position estimation processing using the created three-dimensional data 1035.

[0641] The three-dimensional data storage unit 1018 stores the three-dimensional map 1032, the three-dimensional data 1034, the three-dimensional data 1035, etc.

[0642] The format conversion unit 1019 generates sensor information 1037 by converting the sensor information 1033 into a format supported by the receiving side. Note that the format conversion unit 1019 may reduce the data amount by compressing or encoding the sensor information 1037. Also, the format conversion unit 1019 may omit the process when format conversion is not necessary. Further, the format conversion unit 1019 may control the data amount to be transmitted according to the specification of the transmission range.

[0643] The communication unit 1020 communicates with the server 901 and receives from the server 901 a data transmission request (a request for transmitting sensor information) and the like.

[0644] The transmission control unit 1021 exchanges information such as a corresponding format with the communication destination via the communication unit 1020 and establishes communication.

[0645] The data transmission unit 1022 transmits the sensor information 1037 to the server 901. The sensor information 1037 includes information acquired by a plurality of sensors 1015 such as, for example, information acquired by LiDAR, a luminance image acquired by a visible light camera, an infrared image acquired by an infrared camera, a depth image acquired by a depth sensor, sensor position information, and speed information.

[0646] Next, the configuration of the server 901 will be described. FIG. 148 is a block diagram showing a configuration example of the server 901. The server 901 receives sensor information transmitted from the client device 902 and creates three-dimensional data based on the received sensor information. The server 901 updates a three-dimensional map managed by the serve...

Claims

1. A coding method executed by a coding device, comprising: coding a plurality of three-dimensional points, each located in one of a plurality of regions, for each region; generating connection information based on a relationship between a predetermined region among the plurality of regions and a plurality of other regions other than the predetermined region among the plurality of regions, the connection information including: (i) tile information indicating values uniquely assigned to each of the plurality of regions, and (ii) association information indicating that there is an association between the predetermined region and the other regions based on the tile information; generating a bit stream including the generated connection information and the coded plurality of three-dimensional points; wherein the relationship is a relationship based on visibility with respect to a viewpoint within the predetermined region. Three-dimensional data coding method.

2. wherein the relationship is a positional relationship between a predetermined region among the plurality of regions and the plurality of other regions other than the predetermined region among the plurality of regions. The three-dimensional data coding method according to claim 1.

3. In the generation of the connection information, generating the connection information including the association information indicating decoding of coded three-dimensional points located in other regions that are in contact with or overlap the predetermined region among the plurality of other regions. The three-dimensional data coding method according to claim 1 or 2.

4. In the generation of the connection information, generating the connection information including the association information indicating decoding of coded three-dimensional points located in other regions positioned in the direction as seen from the predetermined region, based on direction information indicating the direction from the predetermined region. The three-dimensional data coding method according to any one of claims 1 to 3.

5. In the generation of the connection information, generating the connection information including the association information indicating that, among the plurality of other regions, the closer the distance to the predetermined region, the earlier the order of the regions for decoding the coded three-dimensional points. The three-dimensional data coding method according to any one of claims 1 to 4.

6. In the generation of the connection information, determining, based on the relationship, to which of a plurality of predetermined groups each of the plurality of other regions belongs; generating the connection information including group information indicating the determined predetermined group. The three-dimensional data coding method according to any one of claims 1 to 5.

7. A decoding method executed by a decoding device, comprising: Obtain a bitstream including a plurality of three-dimensional points each located in one of a plurality of regions, where the plurality of three-dimensional points are encoded for each region. Obtain connection information generated based on the relationship between a predetermined region among the plurality of regions and a plurality of other regions other than the predetermined region among the plurality of regions, the connection information including: (i) tile information indicating values uniquely assigned to each of the plurality of regions, and (ii) association information indicating that there is an association between the predetermined region and the other regions based on the tile information. Decode the plurality of encoded three-dimensional points based on the obtained connection information. The relationship is a relationship based on visibility with respect to a viewpoint within the predetermined region. Three-dimensional data decoding method.

8. Selectively decode the plurality of encoded three-dimensional points for each region based on the obtained connection information. The three-dimensional data decoding method according to claim 7.

9. The relationship is a positional relationship between a predetermined region among the plurality of regions and the plurality of other regions other than the predetermined region among the plurality of regions. The three-dimensional data decoding method according to claim 7 or 8.

10. In obtaining the connection information, obtain the connection information including the association information indicating that, among the plurality of other regions, encoded three-dimensional points located in other regions that are in contact with or overlap the predetermined region are to be decoded. The three-dimensional data decoding method according to any one of claims 7 to 9.

11. In obtaining the connection information, obtain the connection information including the association information generated based on orientation information indicating an orientation from the predetermined region, the association information indicating that encoded three-dimensional points located in other regions located in the orientation as viewed from the predetermined region are to be decoded. The three-dimensional data decoding method according to any one of claims 7 to 10.

12. In obtaining the connection information, obtain the connection information including the association information indicating that, among the plurality of other regions, the closer the distance to the predetermined region, the earlier the order of the regions in which encoded three-dimensional points are to be decoded. The three-dimensional data decoding method according to any one of claims 7 to 11.

13. In obtaining the connection information, obtain the connection information including group information indicating to which of a plurality of predetermined groups each of the plurality of other regions belongs based on the relationship. The three-dimensional data decoding method according to any one of claims 7 to 12.

14. The associated information indicates the strength of the association between the predetermined region and each of the plurality of other regions. The three-dimensional data decoding method according to claim 7 or 8.

15. Based on the obtained connection information, determine the order of the regions to be decoded among the plurality of regions. The three-dimensional data decoding method according to claim 14.

16. The plurality of regions include a first region and a second region. Decode the plurality of encoded three-dimensional points in the predetermined region. Based on the obtained connection information, select a first region with a high priority for the predetermined region. Decode the plurality of encoded three-dimensional points in the first region preferentially over the plurality of encoded three-dimensional points in the second region. The three-dimensional data decoding method according to claim 7 or 8.

17. A processor and a memory, wherein the processor uses the memory to encode each of a plurality of three-dimensional points located in any of a plurality of regions for each region, generate connection information based on the relationship between a predetermined region among the plurality of regions and a plurality of other regions other than the predetermined region among the plurality of regions, the connection information including (i) tile information indicating values uniquely assigned to each of the plurality of regions, and (ii) associated information indicating that there is an association between the predetermined region and the other regions by the tile information, generate a bit stream including the generated connection information and the plurality of encoded three-dimensional points, wherein the relationship is a relationship based on visibility with respect to a viewpoint within the predetermined region Three-dimensional data encoding apparatus.

18. A processor and a memory, wherein the processor uses the memory to obtain a bit stream including a plurality of three-dimensional points each located in any of a plurality of regions and encoded for each region, obtain connection information generated based on the relationship between a predetermined region among the plurality of regions and a plurality of other regions other than the predetermined region among the plurality of regions, the connection information including (i) tile information indicating values uniquely assigned to each of the plurality of regions, and (ii) associated information indicating that there is an association between the predetermined region and the other regions by the tile information, decode the plurality of encoded three-dimensional points based on the obtained connection information. The relationship is a relationship based on visibility with reference to viewpoints within the predetermined region. Three-dimensional data decoding device.

Citation Information

Patent Citations

  • A system for encoding multiple videos acquired from moving objects in a scene by multiple fixed cameras

    JP2007519285A

  • Image processing device and image processing method

    WO2018025660A1

  • Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device

    WO2019146691A1

  • Map display device

    WO2014020663A1