Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device
Patent Information
- Application Number
- MYPI2021003337
- Authority / Receiving Office
- MY · MY
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-12-26
- Filing Date
- 2019-12-26
- Publication Date
- 2026-08-10
- Estimated Expiration
- 2039-12-26
AI Technical Summary
Current methods for encoding and decoding three-dimensional data, particularly point cloud data, face challenges in efficiently transmitting and storing attribute information, lacking a standardized format that supports multiple encoding methods, leading to inefficiencies in data compression and decoding processes.
A three-dimensional data encoding method that encodes attribute information using parameters, generating a bitstream with control information and type information, allowing for correct decoding of attribute information by identifying the type of information through identification tags, and supports coexistence of multiple encoding methods like PCC, enabling efficient data transmission and storage.
This approach enables correct and efficient decoding of attribute information, reduces data transmission bandwidth, and allows for the coexistence of multiple encoding methods within a standardized format, enhancing the encoding efficiency and compatibility of three-dimensional data systems.
Abstract
Description
Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device
[0001] This disclosure relates to a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, and a three-dimensional data decoding device.
[0002] In the future, devices and services utilizing three-dimensional data are expected to become widespread in a wide range of fields, including computer vision for autonomous operation of automobiles or robots, map information, monitoring, infrastructure inspection, and video distribution. Three-dimensional data can be acquired in various ways, such as using distance sensors like rangefinders, stereo cameras, or combinations of multiple monocular cameras.
[0003] One method of representing three-dimensional data is called a point cloud, which represents the shape of a three-dimensional structure using a cloud of points in three-dimensional space. In a point cloud, the position and color of the points are stored. Point clouds are expected to become the mainstream method of representing three-dimensional data, but point clouds are extremely large in size. Therefore, in the storage or transmission of three-dimensional data, data compression through encoding is essential, similar to two-dimensional moving images (for example, MPEG-4 AVC or HEVC, which are standardized by MPEG).
[0004] Furthermore, point cloud compression is partially supported by publicly available libraries (such as the Point Cloud Library) that handle point cloud-related processing.
[0005] Furthermore, there is a known technique for searching for and displaying facilities located around a vehicle using three-dimensional map data (see, for example, Patent Document 1).
[0006] International Publication No. 2014 / 020663
[0007] In the encoding and decoding of three-dimensional data, it is desirable to be able to correctly decode the attribute information of three-dimensional points.
[0008] The purpose of this disclosure is to provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device that can correctly decode the attribute information of three-dimensional points.
[0009] A three-dimensional data encoding method according to one aspect of the present disclosure encodes a plurality of attribute information having each of a plurality of three-dimensional points using parameters, generates a bitstream including the encoded plurality of attribute information, control information, and a plurality of first attribute control information, wherein the control information includes a plurality of type information corresponding to the plurality of attribute information, each indicating a different type of attribute information, and the plurality of first attribute control information includes first identification information indicating that each of the plurality of first attribute control information corresponds to the plurality of attribute information and is associated with one of the plurality of type information.
[0010] A three-dimensional data decoding method according to one aspect of the present disclosure decodes a plurality of attribute information having each of a plurality of three-dimensional points by acquiring a bitstream to acquire a plurality of encoded attribute information and parameters, and decoding the plurality of encoded attribute information using the parameters, wherein the bitstream includes control information and a plurality of first attribute control information, the control information includes a plurality of type information corresponding to the plurality of attribute information and each of which indicates a different type of attribute information, the plurality of first attribute control information each corresponds to the plurality of attribute information and each of the plurality of first attribute control information includes first identification information indicating that it is associated with one of the plurality of type information.
[0011] This disclosure provides a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device that can correctly decode attribute information of three-dimensional points.
[0012] Figure 1 is a diagram showing the configuration of a three-dimensional data encoding and decoding system according to Embodiment 1. Figure 2 is a diagram showing an example of the configuration of point cloud data according to Embodiment 1. Figure 3 is a diagram showing an example of the configuration of a data file describing point cloud data information according to Embodiment 1. Figure 4 is a diagram showing the types of point cloud data according to Embodiment 1. Figure 5 is a diagram showing the configuration of the first encoding unit according to Embodiment 1. Figure 6 is a block diagram of the first encoding unit according to Embodiment 1. Figure 7 is a diagram showing the configuration of the first decoding unit according to Embodiment 1. Figure 8 is a block diagram of the first decoding unit according to Embodiment 1. Figure 9 is a diagram showing the configuration of the second encoding unit according to Embodiment 1. Figure 10 is a block diagram of the second encoding unit according to Embodiment 1. Figure 11 is a diagram showing the configuration of the second decoding unit according to Embodiment 1. Figure 12 is a block diagram of the second decoding unit according to Embodiment 1. Figure 13 is a diagram showing the protocol stack related to PCC encoded data according to Embodiment 1. Figure 14 is a block diagram of the encoding unit according to Embodiment 1. Figure 15 is a block diagram of the decoding unit according to Embodiment 1. Figure 16 is a flowchart of the encoding process according to Embodiment 1. Figure 17 is a flowchart of the decoding process according to Embodiment 1. Figure 18 is a diagram showing the basic structure of ISOBMFF according to Embodiment 2. Figure 19 is a diagram showing the protocol stack according to Embodiment 2. Figure 20 is a diagram showing an example of storing the NAL unit according to Embodiment 2 in a file for codec 1. Figure 21 is a diagram showing an example of storing the NAL unit according to Embodiment 2 in a file for codec 2. Figure 22 is a diagram showing the configuration of the first multiplexing unit according to Embodiment 2. Figure 23 is a diagram showing the configuration of the first demultiplexing unit according to Embodiment 2. Figure 24 is a diagram showing the configuration of the second multiplexing unit according to Embodiment 2. Figure 25 is a diagram showing the configuration of the second demultiplexing unit according to Embodiment 2. Figure 26 is a flowchart of the processing by the first multiplexing unit according to Embodiment 2. Figure 27 is a flowchart of the processing by the second multiplexing unit according to Embodiment 2. Figure 28 is a flowchart of the processing by the first demultiplexing unit and the first decoding unit according to Embodiment 2.Figure 29 is a flowchart of the processing performed by the second demultiplexing unit and the second decoding unit according to Embodiment 2. Figure 30 is a diagram showing the configuration of the encoding unit and the third multiplexing unit according to Embodiment 3. Figure 31 is a diagram showing the configuration of the third demultiplexing unit and the decoding unit according to Embodiment 3. Figure 32 is a flowchart of the processing performed by the third multiplexing unit according to Embodiment 3. Figure 33 is a flowchart of the processing performed by the third demultiplexing unit and the decoding unit according to Embodiment 3. Figure 34 is a flowchart of the processing performed by the three-dimensional data storage device according to Embodiment 3. Figure 35 is a flowchart of the processing performed by the three-dimensional data acquisition device according to Embodiment 3. Figure 36 is a diagram showing the configuration of the encoding unit and the multiplexing unit according to Embodiment 4. Figure 37 is a diagram showing an example of the configuration of encoded data according to Embodiment 4. Figure 38 is a diagram showing an example of the configuration of encoded data and the NAL unit according to Embodiment 4. Figure 39 is a diagram showing an example of the semantics of pcc_nal_unit_type according to Embodiment 4. Figure 40 is a diagram showing an example of the transmission order of the NAL unit according to Embodiment 4. Figure 41 is a flowchart of the processing performed by the three-dimensional data encoding device according to Embodiment 4. Figure 42 is a flowchart of the processing performed by the three-dimensional data decoding device according to Embodiment 4. Figure 43 is a flowchart of the multiplexing process according to Embodiment 4. Figure 44 is a flowchart of the demultiplexing process according to Embodiment 4. Figure 45 is a flowchart of the processing performed by the three-dimensional data encoding device according to Embodiment 4. Figure 46 is a flowchart of the processing performed by the three-dimensional data decoding device according to Embodiment 4. Figure 47 is a block diagram of the division section according to Embodiment 5. Figure 48 is a diagram showing examples of slice and tile division according to Embodiment 5. Figure 49 is a diagram showing examples of slice and tile division patterns according to Embodiment 5. Figure 50 is a block diagram of the first encoding section according to Embodiment 6. Figure 51 is a block diagram of the first decoding section according to Embodiment 6. Figure 52 is a diagram showing examples of tile shapes according to Embodiment 6. Figure 53 is a diagram showing examples of tiles and slices according to Embodiment 6. Figure 54 is a block diagram of the division section according to Embodiment 6. Figure 55 is a diagram showing an example of a map of point cloud data viewed from above according to Embodiment 6.Figure 56 is a diagram showing an example of tile division according to Embodiment 6. Figure 57 is a diagram showing an example of tile division according to Embodiment 6. Figure 58 is a diagram showing an example of tile division according to Embodiment 6. Figure 59 is a diagram showing an example of tile data stored on a server according to Embodiment 6. Figure 60 is a diagram showing a system related to tile division according to Embodiment 6. Figure 61 is a diagram showing an example of slice division according to Embodiment 6. Figure 62 is a diagram showing an example of dependency relationships according to Embodiment 6. Figure 63 is a diagram showing an example of the data decoding order according to Embodiment 6. Figure 64 is a diagram showing an example of encoded tile data according to Embodiment 6. Figure 65 is a block diagram of the coupling part according to Embodiment 6. Figure 66 is a diagram showing an example of the configuration of encoded data and NAL unit according to Embodiment 6. Figure 67 is a flowchart of the encoding process according to Embodiment 6. Figure 68 is a flowchart of the decoding process according to Embodiment 6. Figure 69 is a diagram showing an example of the syntax of tile addition information according to Embodiment 6. Figure 70 is a block diagram of the encoding / decoding system according to Embodiment 6. Figure 71 is a diagram showing an example of the syntax of slice addition information according to Embodiment 6. Figure 72 is a flowchart of the encoding process according to Embodiment 6. Figure 73 is a flowchart of the decoding process according to Embodiment 6. Figure 74 is a flowchart of the encoding process according to Embodiment 6. Figure 75 is a flowchart of the decoding process according to Embodiment 6. Figure 76 is a flowchart of the reinitialization process of the CABAC encoding / decoding engine according to the CABAC initialization flag during encoding or decoding according to Embodiment 7. Figure 77 is a block diagram showing the configuration of the first encoding unit included in the three-dimensional data encoding device according to Embodiment 7. Figure 78 is a block diagram showing the configuration of the division unit according to Embodiment 7. Figure 79 is a block diagram showing the configuration of the location information encoding unit and attribute information encoding unit according to Embodiment 7. Figure 80 is a block diagram showing the configuration of the first decoding unit according to Embodiment 7. Figure 81 is a block diagram showing the configuration of the location information decoding unit and attribute information decoding unit according to Embodiment 7.Figure 82 is a flowchart showing an example of the process for initializing CABAC in encoding location information or attribute information according to Embodiment 7. Figure 83 is a diagram showing an example of the timing of CABAC initialization in point cloud data as a bitstream according to Embodiment 7. Figure 84 is a diagram showing the structure of encoded data and the method of storing encoded data in a NAL unit according to Embodiment 7. Figure 85 is a flowchart showing an example of the process for initializing CABAC in decoding location information or attribute information according to Embodiment 7. Figure 86 is a flowchart of the point cloud data encoding process according to Embodiment 7. Figure 87 is a flowchart showing an example of the process for updating additional information according to Embodiment 7. Figure 88 is a flowchart showing an example of the process for initializing CABAC according to Embodiment 7. Figure 89 is a flowchart showing the point cloud data decoding process according to Embodiment 7. Figure 90 is a flowchart showing an example of the process for initializing the CABAC decoding unit according to Embodiment 7. Figure 91 is a diagram showing examples of tiles and slices according to Embodiment 7. Figure 92 is a flowchart showing an example of the method for initializing CABAC and determining the context initial value according to Embodiment 7. Figure 93 is a diagram showing an example of dividing a map obtained from a top view using point cloud data obtained by LiDAR according to Embodiment 7 into tiles. Figure 94 is a flowchart showing another example of the CABAC initialization and context initial value determination method according to Embodiment 7. Figure 95 is a diagram for explaining the processing of the quantization unit and the dequantization unit according to Embodiment 8. Figure 96 is a diagram for explaining the default value of the quantization value and the quantization delta according to Embodiment 8. Figure 97 is a block diagram showing the configuration of the first encoding unit included in the three-dimensional data encoding device according to Embodiment 8. Figure 98 is a block diagram showing the configuration of the division unit according to Embodiment 8. Figure 99 is a block diagram showing the configuration of the location information encoding unit and the attribute information encoding unit according to Embodiment 8. Figure 100 is a block diagram showing the configuration of the first decoding unit according to Embodiment 8. Figure 101 is a block diagram showing the configuration of the location information decoding unit and the attribute information decoding unit according to Embodiment 8.Figure 102 is a flowchart illustrating an example of the process for determining quantization values in encoding location information or attribute information according to Embodiment 8. Figure 103 is a flowchart illustrating an example of the decoding process for location information and attribute information according to Embodiment 8. Figure 104 is a diagram illustrating a first example of the method for transmitting quantization parameters according to Embodiment 8. Figure 105 is a diagram illustrating a second example of the method for transmitting quantization parameters according to Embodiment 8. Figure 106 is a diagram illustrating a third example of the method for transmitting quantization parameters according to Embodiment 8. Figure 107 is a flowchart illustrating the encoding process for point cloud data according to Embodiment 8. Figure 108 is a flowchart illustrating an example of the process for determining QP values and updating additional information according to Embodiment 8. Figure 109 is a flowchart illustrating an example of the process for encoding the determined QP values according to Embodiment 8. Figure 110 is a flowchart illustrating the decoding process for point cloud data according to Embodiment 8. Figure 111 is a flowchart illustrating an example of the process for obtaining QP values and decoding the QP values of slices or tiles according to Embodiment 8. Figure 112 is a diagram illustrating an example of GPS syntax according to Embodiment 8. Figure 113 is a diagram showing an example of APS syntax according to Embodiment 8. Figure 114 is a diagram showing an example of location information header syntax according to Embodiment 8. Figure 115 is a diagram showing an example of attribute information header syntax according to Embodiment 8. Figure 116 is a diagram illustrating another example of the quantization parameter transmission method according to Embodiment 8. Figure 117 is a diagram illustrating another example of the quantization parameter transmission method according to Embodiment 8. Figure 118 is a diagram illustrating a ninth example of the quantization parameter transmission method according to Embodiment 8. Figure 119 is a diagram illustrating an example of QP value control according to Embodiment 8. Figure 120 is a flowchart illustrating an example of a method for determining QP values based on object quality according to Embodiment 8. Figure 121 is a flowchart illustrating an example of a method for determining QP values based on rate control according to Embodiment 8. Figure 122 is a flowchart of the encoding process according to Embodiment 8. Figure 123 is a flowchart of the decoding process according to Embodiment 8.Figure 124 is a diagram illustrating an example of a quantization parameter transmission method according to Embodiment 9. Figure 125 is a diagram showing a first example of the APS syntax and the attribute information header syntax according to Embodiment 9. Figure 126 is a diagram showing a second example of the APS syntax according to Embodiment 9. Figure 127 is a diagram showing a second example of the attribute information header syntax according to Embodiment 9. Figure 128 is a diagram showing the relationship between SPS, APS, and the attribute information header according to Embodiment 9. Figure 129 is a flowchart of the encoding process according to Embodiment 9. Figure 130 is a flowchart of the decoding process according to Embodiment 9.
[0013] A three-dimensional data encoding method according to one aspect of the present disclosure encodes a plurality of attribute information having each of a plurality of three-dimensional points using parameters, generates a bitstream including the encoded plurality of attribute information, control information, and a plurality of first attribute control information, wherein the control information includes a plurality of type information corresponding to the plurality of attribute information, each indicating a different type of attribute information, and the plurality of first attribute control information includes first identification information indicating that each of the plurality of first attribute control information corresponds to the plurality of attribute information and is associated with one of the plurality of type information.
[0014] According to this, since the first attribute control information generates a bitstream that includes first identification information for identifying the type of attribute information it corresponds to, the three-dimensional data decoding device that receives the bitstream can correctly and efficiently decode the attribute information of the three-dimensional point.
[0015] For example, the plurality of type information may be stored in the control information in a predetermined order, and the first identification information may indicate that the first attribute control information, which includes the first identification information, is associated with type information in one of the predetermined order.
[0016] According to this method, since the type information is displayed in a predetermined order without the need to add information indicating the type, the amount of data in the bitstream can be reduced, and the amount of bitstream transmission can be reduced.
[0017] For example, the bitstream may further include a plurality of second attribute control pieces corresponding to the plurality of attribute pieces, and each of the plurality of second attribute control pieces may include a reference value for a parameter used to encode the corresponding attribute piece.
[0018] According to this, since each of the multiple second attribute control information includes a reference value for the parameter, the attribute information corresponding to the second attribute control information can be encoded using the reference value.
[0019] For example, the first attribute control information may include difference information, which is the difference of the parameter from the reference value.
[0020] Therefore, encoding efficiency can be improved.
[0021] For example, the bitstream may further include a plurality of second attribute control pieces corresponding to the plurality of attribute pieces, and each of the plurality of second attribute control pieces may have second identification information indicating that it is associated with one of the plurality of type pieces.
[0022] According to this, a bitstream is generated that includes a second identification information for identifying the type of attribute information that the second attribute control information corresponds to, thereby enabling the generation of a bitstream that can correctly and efficiently decode the attribute information of a three-dimensional point.
[0023] For example, each of the plurality of first attribute control information has N fields in which N parameters (where N is 2 or more) are stored, and in a specific first attribute control information corresponding to a specific type of attribute among the plurality of first attribute control information, one of the N fields may contain a value indicating invalidity.
[0024] Therefore, a three-dimensional data decoding device that receives a bitstream can use the first identification information to identify the type of first attribute information, and can skip the decoding process in the case of specific first attribute control information, thereby enabling accurate and efficient decoding of the attribute information of three-dimensional points.
[0025] For example, in the encoding described above, the multiple attribute information may be quantized using the quantization parameter as the parameter.
[0026] According to this method, since parameters are represented using the difference from a reference value, the coding efficiency of quantization can be improved.
[0027] Furthermore, a three-dimensional data decoding method according to one aspect of the present disclosure decodes the multiple attribute information possessed by each of a plurality of three-dimensional points by acquiring a bitstream to acquire a plurality of encoded attribute information and parameters, and decoding the plurality of encoded attribute information using the parameters, wherein the bitstream includes control information and a plurality of first attribute control information, the control information includes a plurality of type information corresponding to the plurality of attribute information and each of which indicates a different type of attribute information, the plurality of first attribute control information each corresponds to the plurality of attribute information and includes first identification information indicating that each of the plurality of first attribute control information is associated with one of the plurality of type information.
[0028] According to this, the type of attribute information corresponding to the first attribute control information can be identified using the first identification information, thereby enabling accurate and efficient decoding of the attribute information of a three-dimensional point.
[0029] For example, the plurality of type information may be stored in the control information in a predetermined order, and the first identification information may indicate that the first attribute control information, which includes the first identification information, is associated with type information in one of the predetermined order.
[0030] According to this method, since the type information is displayed in a predetermined order without the need to add information indicating the type, the amount of data in the bitstream can be reduced, and the amount of bitstream transmission can be reduced.
[0031] For example, the bitstream may further include a plurality of second attribute control pieces corresponding to the plurality of attribute pieces, and each of the plurality of second attribute control pieces may include a reference value for a parameter used to encode the corresponding attribute piece.
[0032] According to this method, the attribute information corresponding to the second attribute control information can be decoded using a reference value, thus enabling accurate and efficient decoding of the attribute information of a three-dimensional point.
[0033] For example, the first attribute control information may include difference information, which is the difference of the parameter from the reference value.
[0034] According to this method, attribute information can be decoded using reference values and difference information, thus enabling accurate and efficient decoding of attribute information of three-dimensional points.
[0035] For example, the bitstream may further include a plurality of second attribute control pieces corresponding to the plurality of attribute pieces, and each of the plurality of second attribute control pieces may have second identification information indicating that it is associated with one of the plurality of type pieces.
[0036] According to this, by using the second identification information, the type of attribute information corresponding to the second attribute control information can be identified, thereby enabling accurate and efficient decoding of the attribute information of a three-dimensional point.
[0037] For example, each of the plurality of first attribute control information has a plurality of fields in which a plurality of parameters are stored, and in the decoding, parameters stored in a specific field of the plurality of fields of a specific first attribute control information corresponding to a specific type of attribute among the plurality of first attribute control information may be ignored.
[0038] According to this method, the type of first attribute information can be identified using the first identification information, and the decoding process can be omitted in the case of specific first attribute control information, thereby enabling accurate and efficient decoding of the attribute information of three-dimensional points.
[0039] For example, in the decoding process, the multiple encoded attribute information may be dequantized using the quantization parameter as the parameter.
[0040] According to this method, the attribute information of a three-dimensional point can be correctly decoded.
[0041] Furthermore, a three-dimensional data encoding device according to one aspect of the present disclosure comprises a processor and a memory, the processor using the memory to encode a plurality of attribute information having each of a plurality of three-dimensional points using parameters, and generates a bitstream including the encoded plurality of attribute information, control information, and a plurality of first attribute control information, the control information including a plurality of type information corresponding to the plurality of attribute information, each of which indicates a different type of attribute information, and the plurality of first attribute control information including (i) a first identification information indicating that it corresponds to each of the plurality of attribute information and (ii) an first identification information indicating that it is associated with any of the plurality of type information.
[0042] According to this, a bitstream is generated that includes first identification information for identifying the type of attribute information that the first attribute control information corresponds to, thereby enabling the generation of a bitstream that can correctly and efficiently decode the attribute information of a three-dimensional point.
[0043] Furthermore, a three-dimensional data decoding device according to one aspect of the present disclosure comprises a processor and a memory, wherein the processor uses the memory to acquire a bitstream to acquire a plurality of encoded attribute information and parameters, and decodes the plurality of attribute information having each of a plurality of three-dimensional points by decoding the encoded plurality of attribute information using the parameters, wherein the bitstream includes control information and a plurality of first attribute control information, the control information includes a plurality of type information corresponding to the plurality of attribute information and each indicating a different type of attribute information, and the plurality of first attribute control information includes (i) a first identification information indicating that it corresponds to each of the plurality of attribute information and (ii) an first identification information indicating that it is associated with any of the plurality of type information.
[0044] According to this, the type of attribute information corresponding to the first attribute control information can be identified using the first identification information, thereby enabling accurate and efficient decoding of the attribute information of a three-dimensional point.
[0045] These comprehensive or specific embodiments may be implemented as a system, method, integrated circuit, computer program, or recording medium such as a computer-readable CD-ROM, or as any combination of a system, method, integrated circuit, computer program, and recording medium.
[0046] The embodiments will be described in detail below with reference to the drawings. Note that the embodiments described below are all specific examples of this disclosure. The numerical values, shapes, materials, components, arrangement and connection configurations of components, steps, and the order of steps shown in the following embodiments are examples only and are not intended to limit this disclosure. Furthermore, among the components in the following embodiments, those not described in the independent claim representing the highest-level concept will be described as optional components.
[0047] (Embodiment 1) When using encoded point cloud data in an actual device or service, it is desirable to send and receive information as needed depending on the application in order to reduce network bandwidth. However, until now, such a function has not existed in the encoded structure of three-dimensional data, nor has there been an encoding method for that purpose.
[0048] This embodiment describes a three-dimensional data encoding method and a three-dimensional data encoding device for providing a function to send and receive information necessary for use in encoded data of a three-dimensional point cloud, a three-dimensional data decoding method and a three-dimensional data decoding device for decoding the encoded data, a three-dimensional data multiplexing method for multiplexing the encoded data, and a three-dimensional data transmission method for transmitting the encoded data.
[0049] In particular, while a first encoding method and a second encoding method are currently being considered as methods for encoding point cloud data, the structure of the encoded data and the method for storing the encoded data in a system format have not been defined. As a result, there is a problem in that MUX processing (multiplexing), transmission, or storage cannot be performed in the encoding unit.
[0050] Furthermore, there has been no method to date to support formats that combine two codecs, such as PCC (Point Cloud Compression), which uses a first encoding method and a second encoding method.
[0051] This embodiment describes the structure of PCC encoded data in which two codecs, a first encoding method and a second encoding method, coexist, and a method for storing the encoded data in a system format.
[0052] First, the configuration of the three-dimensional data (point cloud data) encoding and decoding system according to this embodiment will be described. Figure 1 is a diagram showing an example of the configuration of the three-dimensional data encoding and decoding system according to this embodiment. As shown in Figure 1, the three-dimensional data encoding and decoding system includes a three-dimensional data encoding system 4601, a three-dimensional data decoding system 4602, a sensor terminal 4603, and an external connection unit 4604.
[0053] The three-dimensional data encoding system 4601 generates encoded data or multiplexed data by encoding point cloud data, which is three-dimensional data. The three-dimensional data encoding system 4601 may be a three-dimensional data encoding device implemented by a single device, or it may be a system implemented by multiple devices. Furthermore, the three-dimensional data encoding device may include some of the multiple processing units included in the three-dimensional data encoding system 4601.
[0054] The three-dimensional data encoding system 4601 includes a point cloud data generation system 4611, a presentation unit 4612, an encoding unit 4613, a multiplexing unit 4614, an input / output unit 4615, and a control unit 4616. The point cloud data generation system 4611 includes a sensor information acquisition unit 4617 and a point cloud data generation unit 4618.
[0055] The sensor information acquisition unit 4617 acquires sensor information from the sensor terminal 4603 and outputs the sensor information to the point cloud data generation unit 4618. The point cloud data generation unit 4618 generates point cloud data from the sensor information and outputs the point cloud data to the encoding unit 4613.
[0056] The display unit 4612 presents sensor information or point cloud data to the user. For example, the display unit 4612 displays information or images based on sensor information or point cloud data.
[0057] The encoding unit 4613 encodes (compresses) the point cloud data and outputs the resulting encoded data, control information obtained during the encoding process, and other additional information to the multiplexing unit 4614. The additional information includes, for example, sensor information.
[0058] The multiplexing unit 4614 generates multiplexed data by multiplexing the encoded data input from the encoding unit 4613, control information, and additional information. The format of the multiplexed data is, for example, a file format for storage or a packet format for transmission.
[0059] The input / output unit 4615 (for example, the communication unit or interface) outputs the multiplexed data to the outside. Alternatively, the multiplexed data is stored in a storage unit such as internal memory. The control unit 4616 (or the application execution unit) controls each processing unit. In other words, the control unit 4616 performs control such as encoding and multiplexing.
[0060] The sensor information may also be input to the encoding unit 4613 or the multiplexing unit 4614. Furthermore, the input / output unit 4615 may output the point cloud data or encoded data directly to the outside.
[0061] The transmission signal (multiplexed data) output from the three-dimensional data encoding system 4601 is input to the three-dimensional data decoding system 4602 via the external connection unit 4604.
[0062] The three-dimensional data decoding system 4602 generates point cloud data, which is three-dimensional data, by decoding encoded data or multiplexed data. The three-dimensional data decoding system 4602 may be a three-dimensional data decoding device implemented by a single device, or it may be a system implemented by multiple devices. Furthermore, the three-dimensional data decoding device may include some of the multiple processing units included in the three-dimensional data decoding system 4602.
[0063] The three-dimensional data decoding system 4602 includes a sensor information acquisition unit 4621, an input / output unit 4622, a demultiplexing unit 4623, a decoding unit 4624, a presentation unit 4625, a user interface 4626, and a control unit 4627.
[0064] The sensor information acquisition unit 4621 acquires sensor information from the sensor terminal 4603.
[0065] The input / output unit 4622 acquires the transmission signal, decodes the multiplexed data (file format or packet) from the transmission signal, and outputs the multiplexed data to the demultiplexing unit 4623.
[0066] The demultiplexing unit 4623 acquires encoded data, control information, and additional information from the multiplexed data, and outputs the encoded data, control information, and additional information to the decoding unit 4624.
[0067] The decoding unit 4624 reconstructs the point cloud data by decoding the encoded data.
[0068] The presentation unit 4625 presents point cloud data to the user. For example, the presentation unit 4625 displays information or images based on the point cloud data. The user interface 4626 acquires instructions based on user operations. The control unit 4627 (or application execution unit) controls each processing unit. In other words, the control unit 4627 performs control such as demultiplexing, decoding, and presentation.
[0069] The input / output unit 4622 may acquire point cloud data or encoded data directly from an external source. The presentation unit 4625 may acquire additional information such as sensor information and present information based on that additional information. The presentation unit 4625 may also make presentations based on user instructions acquired through the user interface 4626.
[0070] The sensor terminal 4603 generates sensor information, which is information obtained from the sensor. The sensor terminal 4603 is a terminal equipped with a sensor or camera, and may be, for example, a mobile object such as an automobile, an aerial object such as an airplane, a mobile terminal, or a camera.
[0071] The sensor information that can be acquired by the sensor terminal 4603 includes, for example, (1) the distance between the sensor terminal 4603 and the object, or the reflectivity of the object, obtained from a LiDAR, millimeter-wave radar, or infrared sensor; and (2) the distance between the camera and the object, or the reflectivity of the object, obtained from multiple monocular camera images or stereo camera images. The sensor information may also include the sensor's attitude, orientation, gyroscope (angular velocity), position (GPS information or altitude), speed, or acceleration. The sensor information may also include temperature, atmospheric pressure, humidity, or magnetism.
[0072] The external connection unit 4604 is implemented by an integrated circuit (LSI or IC), an external storage unit, communication with a cloud server via the internet, or broadcasting, etc.
[0073] Next, we will explain point cloud data. Figure 2 shows the structure of point cloud data. Figure 3 shows an example of the structure of a data file containing information about point cloud data.
[0074] Point cloud data contains data for multiple points. Each point's data includes location information (three-dimensional coordinates) and attribute information related to that location. A collection of these points is called a point cloud. For example, a point cloud represents the three-dimensional shape of an object.
[0075] Position information, such as three-dimensional coordinates, is sometimes called geometry. Furthermore, the data for each point may include attribute information of multiple attribute types. Attribute types include, for example, color or reflectance.
[0076] One location information may be associated with one attribute information, or multiple attribute information of different attribute types may be associated with one location information. Furthermore, multiple attribute information of the same attribute type may be associated with one location information.
[0077] The example data file structure shown in Figure 3 represents a case where location information and attribute information correspond one-to-one, and it shows the location information and attribute information of the N points that make up the point cloud data.
[0078] Location information includes, for example, information for the three axes: x, y, and z. Attribute information includes, for example, RGB color information. A typical data file is a ply file.
[0079] Next, we will explain the types of point cloud data. Figure 4 is a diagram illustrating the types of point cloud data. As shown in Figure 4, point cloud data includes static objects and dynamic objects.
[0080] A static object is three-dimensional point cloud data at any given time (a specific moment). A dynamic object is three-dimensional point cloud data that changes over time. Hereafter, three-dimensional point cloud data at a given time will be referred to as a PCC frame, or simply a frame.
[0081] The object can be a point cloud with a somewhat limited area, like regular video data, or it can be a large-scale point cloud with no area limitations, like map information.
[0082] Furthermore, point cloud data of various densities may exist, including both sparse and dense point cloud data.
[0083] The details of each processing unit are described below. Sensor information is acquired by various methods, such as distance sensors like LIDAR or rangefinders, stereo cameras, or combinations of multiple monocular cameras. The point cloud data generation unit 4618 generates point cloud data based on the sensor information obtained by the sensor information acquisition unit 4617. The point cloud data generation unit 4618 generates position information as point cloud data and adds attribute information to the position information.
[0084] The point cloud data generation unit 4618 may process the point cloud data when generating position information or adding attribute information. For example, the point cloud data generation unit 4618 may reduce the amount of data by deleting point clouds with overlapping positions. The point cloud data generation unit 4618 may also transform the position information (such as shifting, rotating, or normalizing it) or render the attribute information.
[0085] In Figure 1, the point cloud data generation system 4611 is included in the three-dimensional data encoding system 4601, but it may also be provided independently outside of the three-dimensional data encoding system 4601.
[0086] The encoding unit 4613 generates encoded data by encoding the point cloud data based on a predetermined encoding method. There are two main types of encoding methods. The first is an encoding method using positional information, which will be referred to as the first encoding method hereafter. The second is an encoding method using a video codec, which will be referred to as the second encoding method hereafter.
[0087] The decoding unit 4624 decodes the point cloud data by decoding the encoded data based on a predetermined encoding method.
[0088] The multiplexing unit 4614 generates multiplexed data by multiplexing the encoded data using an existing multiplexing method. The generated multiplexed data is transmitted or stored. In addition to PCC encoded data, the multiplexing unit 4614 multiplexes other media such as video, audio, subtitles, applications, files, or reference time information. Furthermore, the multiplexing unit 4614 may also multiplex attribute information related to sensor information or point cloud data.
[0089] Multiplexing methods or file formats include ISOBMFF, ISOBMFF-based transmission methods such as MPEG-DASH, MMT, MPEG-2 TS Systems, and RMP.
[0090] The demultiplexing unit 4623 extracts PCC encoded data, other media, and time information from the multiplexed data.
[0091] The input / output unit 4615 transmits the multiplexed data using a method appropriate to the transmission medium or storage medium, such as broadcasting or communication. The input / output unit 4615 may communicate with other devices via the Internet, or with storage units such as cloud servers.
[0092] Communication protocols such as HTTP, FTP, TCP, or UDP can be used. A pull-type communication method or a push-type communication method may be used.
[0093] Either wired or wireless transmission may be used. Wired transmission methods include Ethernet®, USB, RS-232C, HDMI®, or coaxial cable. Wireless transmission methods include wireless LAN, Wi-Fi®, Bluetooth®, or millimeter wave.
[0094] Furthermore, broadcasting formats such as DVB-T2, DVB-S2, DVB-C2, ATSC3.0, or ISDB-S3 may be used.
[0095] Figure 5 shows the configuration of a first encoding unit 4630, which is an example of an encoding unit 4613 that performs encoding using the first encoding method. Figure 6 is a block diagram of the first encoding unit 4630. The first encoding unit 4630 generates encoded data (encoded stream) by encoding point cloud data using the first encoding method. This first encoding unit 4630 includes a location information encoding unit 4631, an attribute information encoding unit 4632, an additional information encoding unit 4633, and a multiplexing unit 4634.
[0096] The first encoding unit 4630 is characterized by performing encoding while being aware of the three-dimensional structure. Furthermore, the first encoding unit 4630 is characterized in that the attribute information encoding unit 4632 performs encoding using information obtained from the location information encoding unit 4631. The first encoding method is also called GPCC (Geometry-based PCC).
[0097] The point cloud data is PCC point cloud data such as a PLY file, or PCC point cloud data generated from sensor information, and includes position information, attribute information, and other additional information (MetaData). The position information is input to the position information encoding unit 4631, the attribute information is input to the attribute information encoding unit 4632, and the additional information is input to the additional information encoding unit 4633.
[0098] The location information encoding unit 4631 generates encoded location information (Compressed Geometry), which is encoded data, by encoding location information. For example, the location information encoding unit 4631 encodes location information using an N-tree structure such as an octree. Specifically, in an octree, the target space is divided into eight nodes (subspaces), and eight bits of information (occupancy code) are generated to indicate whether or not a point cloud is contained in each node. Furthermore, nodes containing point clouds are further divided into eight nodes, and eight bits of information are generated to indicate whether or not a point cloud is contained in each of these eight nodes. This process is repeated until the number of point clouds contained in a predetermined hierarchy or node falls below a threshold.
[0099] The attribute information encoding unit 4632 generates encoded attribute information (Compressed Attribute), which is encoded data, by encoding it using the configuration information generated by the location information encoding unit 4631. For example, the attribute information encoding unit 4632 determines the reference point (reference node) to be referenced in encoding the target point (target node) to be processed, based on the octave tree structure generated by the location information encoding unit 4631. For example, the attribute information encoding unit 4632 references a surrounding node or adjacent node whose parent node in the octave tree is the same as the target node. Note that the method for determining the reference relationship is not limited to this.
[0100] Furthermore, the attribute information encoding process may include at least one of the following: quantization, prediction, and arithmetic encoding. In this case, a reference means using a reference node to calculate the predicted value of the attribute information, or using the state of a reference node (for example, occupancy information indicating whether or not the reference node contains a point cloud) to determine the encoding parameters. For example, encoding parameters may be quantization parameters in the quantization process, or context in arithmetic encoding.
[0101] The additional information encoding unit 4633 generates encoded additional information (Compressed MetaData), which is encoded data, by encoding the compressible data among the additional information.
[0102] The multiplexing unit 4634 generates a compressed stream, which is encoded data, by multiplexing encoded position information, encoded attribute information, encoded additional information, and other additional information. The generated compressed stream is output to a processing unit of the system layer (not shown).
[0103] Next, we will describe a first decoding unit 4640, which is an example of a decoding unit 4624 that performs decoding of the first encoding method. Figure 7 is a diagram showing the configuration of the first decoding unit 4640. Figure 8 is a block diagram of the first decoding unit 4640. The first decoding unit 4640 generates point cloud data by decoding the encoded data (encoded stream) encoded by the first encoding method using the first encoding method. This first decoding unit 4640 includes a demultiplexing unit 4641, a location information decoding unit 4642, an attribute information decoding unit 4643, and an additional information decoding unit 4644.
[0104] A encoded stream, which is encoded data, is input to the first decoding unit 4640 from a processing unit of the system layer (not shown).
[0105] The demultiplexing unit 4641 separates encoded location information (Compressed Geometry), encoded attribute information (Compressed Attribute), encoded additional information (Compressed MetaData), and other additional information from the encoded data.
[0106] The position information decoding unit 4642 generates position information by decoding the encoded position information. For example, the position information decoding unit 4642 reconstructs the position information of a point cloud represented by three-dimensional coordinates from encoded position information represented by an N-tree structure such as an octree.
[0107] The attribute information decoding unit 4643 decodes the encoded attribute information based on the configuration information generated by the location information decoding unit 4642. For example, the attribute information decoding unit 4643 determines the reference point (reference node) to be referenced in the decoding of the target point (target node) to be processed, based on the octave tree structure obtained by the location information decoding unit 4642. For example, the attribute information decoding unit 4643 references a surrounding node or adjacent node whose parent node in the octave tree is the same as the target node. However, the method for determining the reference relationship is not limited to this.
[0108] Furthermore, the attribute information decoding process may include at least one of the following: inverse quantization, prediction, and arithmetic decoding. In this case, "reference" means using a reference node to calculate the predicted value of the attribute information, or using the state of the reference node (for example, occupancy information indicating whether or not the reference node contains a point cloud) to determine the decoding parameters. For example, decoding parameters may be quantization parameters in the inverse quantization process, or context in arithmetic decoding.
[0109] The additional information decoding unit 4644 generates additional information by decoding the encoded additional information. The first decoding unit 4640 uses the additional information necessary for decoding location information and attribute information during decoding and outputs the additional information necessary for the application to the outside.
[0110] Next, we will describe a second encoding unit 4650, which is an example of an encoding unit 4613 that performs encoding using the second encoding method. Figure 9 is a diagram showing the configuration of the second encoding unit 4650. Figure 10 is a block diagram of the second encoding unit 4650.
[0111] The second encoding unit 4650 generates encoded data (encoded stream) by encoding the point cloud data using a second encoding method. This second encoding unit 4650 includes an additional information generation unit 4651, a position image generation unit 4652, an attribute image generation unit 4653, a video encoding unit 4654, an additional information encoding unit 4655, and a multiplexing unit 4656.
[0112] The second encoding unit 4650 is characterized by generating a position image and an attribute image by projecting a three-dimensional structure onto a two-dimensional image, and encoding the generated position image and attribute image using an existing video encoding scheme. The second encoding method is also called VPCC (Video-based PCC).
[0113] The point cloud data is PCC point cloud data such as a PLY file, or PCC point cloud data generated from sensor information, and includes position information, attribute information, and other additional information (MetaData).
[0114] The additional information generation unit 4651 generates map information for multiple two-dimensional images by projecting a three-dimensional structure onto a two-dimensional image.
[0115] The location image generation unit 4652 generates a location image (Geometry Image) based on location information and map information generated by the additional information generation unit 4651. This location image is, for example, a depth image in which the distance is indicated as a pixel value. This depth image may be an image of multiple point clouds viewed from one viewpoint (an image of multiple point clouds projected onto a single two-dimensional plane), or multiple images of multiple point clouds viewed from multiple viewpoints, or a single image formed by integrating these multiple images.
[0116] The attribute image generation unit 4653 generates an attribute image based on attribute information and map information generated by the additional information generation unit 4651. This attribute image is, for example, an image in which attribute information (e.g., color (RGB)) is shown as pixel values. This image may be an image of multiple point clouds viewed from one viewpoint (an image of multiple point clouds projected onto a single two-dimensional plane), or multiple images of multiple point clouds viewed from multiple viewpoints, or a single image formed by integrating these multiple images.
[0117] The video encoding unit 4654 generates encoded data, namely a encoded geometry image and a encoded attribute image, by encoding the position image and attribute image using a video encoding scheme. Any known encoding method may be used as the video encoding scheme. For example, the video encoding scheme may be AVC or HEVC.
[0118] The additional information encoding unit 4655 generates encoded additional information (Compressed MetaData) by encoding additional information and map information included in the point cloud data.
[0119] The multiplexing unit 4656 generates a compressed stream, which is encoded data, by multiplexing the encoded position image, encoded attribute image, encoded additional information, and other additional information. The generated compressed stream is output to a processing unit of a system layer (not shown).
[0120] Next, we will describe a second decoding unit 4660, which is an example of a decoding unit 4624 that performs decoding of the second encoding method. Figure 11 is a diagram showing the configuration of the second decoding unit 4660. Figure 12 is a block diagram of the second decoding unit 4660. The second decoding unit 4660 generates point cloud data by decoding the encoded data (encoded stream) encoded by the second encoding method using the second encoding method. This second decoding unit 4660 includes a demultiplexing unit 4661, a video decoding unit 4662, an additional information decoding unit 4663, a location information generation unit 4664, and an attribute information generation unit 4665.
[0121] A encoded stream (Compressed Stream), which is encoded data, is input to the second decoding unit 4660 from a processing unit of the system layer (not shown).
[0122] The demultiplexing unit 4661 separates the encoded location image (Compressed Geometry Image), encoded attribute image (Compressed Attribute Image), encoded additional information (Compressed MetaData), and other additional information from the encoded data.
[0123] The video decoding unit 4662 generates a position image and an attribute image by decoding the encoded position image and the encoded attribute image using a video encoding scheme. Any known encoding scheme may be used as the video encoding scheme. For example, the video encoding scheme may be AVC or HEVC.
[0124] The additional information decoding unit 4663 generates additional information, including map information, by decoding the encoded additional information.
[0125] The location information generation unit 4664 generates location information using the location image and map information. The attribute information generation unit 4665 generates attribute information using the attribute image and map information.
[0126] The second decoding unit 4660 uses the additional information necessary for decoding during decoding and outputs the additional information necessary for the application to the outside.
[0127] The following describes the challenges in the PCC encoding scheme. Figure 13 is a diagram showing the protocol stack involved in PCC encoded data. Figure 13 shows an example in which data from other media such as video (e.g., HEVC) or audio is multiplexed onto PCC encoded data and transmitted or stored.
[0128] Multiplexing schemes and file formats have the function of multiplexing, transmitting, or storing various encoded data. In order to transmit or store encoded data, the encoded data must be converted into the format of the multiplexing scheme. For example, HEVC specifies a technique in which encoded data is stored in a data structure called a NAL unit, and the NAL unit is stored in ISOBMFF.
[0129] On the other hand, while a first encoding method (Codec1) and a second encoding method (Codec2) are currently being considered as methods for encoding point cloud data, the structure of the encoded data and the method for storing the encoded data in a system format have not been defined. As a result, there is a problem in that MUX processing (multiplexing), transmission, and storage cannot be performed in the encoding unit.
[0130] In the following text, unless a specific encoding method is mentioned, either the first encoding method or the second encoding method will be referred to.
[0131] The following describes the method for defining NAL units according to this embodiment. For example, in conventional codecs such as HEVC, one NAL unit for one format is defined for each codec. However, there has been no method to support a format in which two codecs (hereinafter referred to as PCC codecs), such as PCC, coexist, consisting of a first encoding method and a second encoding method.
[0132] First, we will describe the encoding unit 4670, which has the functions of both the first encoding unit 4630 and the second encoding unit 4650 described above, and the decoding unit 4680, which has the functions of both the first decoding unit 4640 and the second decoding unit 4660.
[0133] Figure 14 is a block diagram of the encoding unit 4670 according to this embodiment. This encoding unit 4670 includes the first encoding unit 4630 and the second encoding unit 4650 described above, and a multiplexing unit 4671. The multiplexing unit 4671 multiplexes the encoded data generated by the first encoding unit 4630 and the encoded data generated by the second encoding unit 4650, and outputs the obtained encoded data.
[0134] Figure 15 is a block diagram of the decoding unit 4680 according to this embodiment. This decoding unit 4680 includes the first decoding unit 4640 and the second decoding unit 4660 described above, and a demultiplexing unit 4681. The demultiplexing unit 4681 extracts encoded data using the first encoding method and encoded data using the second encoding method from the input encoded data. The demultiplexing unit 4681 outputs the encoded data using the first encoding method to the first decoding unit 4640 and outputs the encoded data using the second encoding method to the second decoding unit 4660.
[0135] With the above configuration, the encoding unit 4670 can encode point cloud data by selectively using the first encoding method and the second encoding method. Furthermore, the decoding unit 4680 can decode encoded data encoded using the first encoding method, encoded data encoded using the second encoding method, and encoded data encoded using both the first and second encoding methods.
[0136] For example, the encoding unit 4670 may switch the encoding method (first encoding method and second encoding method) on a point cloud data basis or on a frame basis. Alternatively, the encoding unit 4670 may switch the encoding method on an encodingable unit basis.
[0137] The encoding unit 4670 generates encoded data (encoded stream) that includes, for example, identification information for the PCC codec.
[0138] The demultiplexing unit 4681 included in the decoding unit 4680 identifies the data using, for example, the identification information of the PCC codec. If the data is encoded using the first encoding method, the demultiplexing unit 4681 outputs the data to the first decoding unit 4640, and if the data is encoded using the second encoding method, it outputs the data to the second decoding unit 4660.
[0139] Furthermore, the encoding unit 4670 may also send control information indicating whether both encoding methods were used or only one of the encoding methods was used, in addition to the identification information of the PCC codec.
[0140] Next, the encoding process according to this embodiment will be described. Figure 16 is a flowchart of the encoding process according to this embodiment. By using the identification information of the PCC codec, encoding processing that supports multiple codecs becomes possible.
[0141] First, the encoding unit 4670 encodes the PCC data using either the first encoding method, the second encoding method, or both of these codecs (S4681).
[0142] If the codec used is the second encoding method (the second encoding method in S4682), the encoding unit 4670 sets the pcc_codec_type included in the NAL unit header to a value indicating that the data included in the NAL unit's payload is data encoded using the second encoding method (S4683). Next, the encoding unit 4670 sets the identifier of the NAL unit for the second encoding method in pcc_nal_unit_type of the NAL unit header (S4684). Then, the encoding unit 4670 generates an NAL unit having the set NAL unit header and containing encoded data in the payload. Finally, the encoding unit 4670 transmits the generated NAL unit (S4685).
[0143] On the other hand, if the codec used is the first encoding method (first encoding method in S4682), the encoding unit 4670 sets the pcc_codec_type included in the NAL unit header to a value indicating that the data included in the NAL unit payload is data encoded using the first encoding method (S4686). Next, the encoding unit 4670 sets the identifier of the NAL unit for the first encoding method in the pcc_nal_unit_type included in the NAL unit header (S4687). Next, the encoding unit 4670 generates an NAL unit having the set NAL unit header and containing encoded data in the payload. Then, the encoding unit 4670 transmits the generated NAL unit (S4685).
[0144] Next, the decoding process according to this embodiment will be described. Figure 17 is a flowchart of the decoding process according to this embodiment. By using the identification information of the PCC codec, decoding processing that supports multiple codecs becomes possible.
[0145] First, the decoding unit 4680 receives the NAL unit (S4691). For example, this NAL unit is generated by the processing in the encoding unit 4670 described above.
[0146] Next, the decoding unit 4680 determines whether the pcc_codec_type included in the NAL unit header indicates the first encoding method or the second encoding method (S4692).
[0147] If pcc_codec_type indicates a second encoding method (second encoding method in S4692), the decoding unit 4680 determines that the data contained in the NAL unit payload is data encoded using the second encoding method (S4693). The second decoding unit 4660 then identifies the data, assuming that pcc_nal_unit_type contained in the NAL unit header is the identifier for the NAL unit used for the second encoding method (S4694). The decoding unit 4680 then decodes the PCC data using the decoding process of the second encoding method (S4695).
[0148] On the other hand, if pcc_codec_type indicates the first encoding method (first encoding method in S4692), the decoding unit 4680 determines that the data contained in the NAL unit payload is data encoded using the first encoding method (S4696). The decoding unit 4680 then identifies the data, assuming that pcc_nal_unit_type contained in the NAL unit header is the identifier for the NAL unit for the first encoding method (S4697). The decoding unit 4680 then decodes the PCC data using the decoding process of the first encoding method (S4698).
[0149] As described above, a three-dimensional data encoding device according to one aspect of the present disclosure generates an encoded stream by encoding three-dimensional data (e.g., point cloud data), and stores information indicating which encoding method was used for the encoding from among the first encoding method and the second encoding method (e.g., codec identification information) in the control information (e.g., parameter set) of the encoded stream.
[0150] According to this, when decoding an encoded stream generated by a three-dimensional data encoding device, the three-dimensional data decoding device can determine the encoding method used for encoding using information stored in the control information. Therefore, the three-dimensional data decoding device can correctly decode the encoded stream even when multiple encoding methods are used.
[0151] For example, the three-dimensional data includes location information. The three-dimensional data encoding device encodes the location information during the encoding process. During storage, the three-dimensional data encoding device stores information in the control information for the location information indicating which of the first and second encoding methods was used to encode the location information.
[0152] For example, the three-dimensional data includes location information and attribute information. The three-dimensional data encoding device encodes the location information and the attribute information in the encoding process. In the storage process, the three-dimensional data encoding device stores in the control information for the location information information information that indicates which of the first and second encoding methods was used to encode the location information, and in the control information for the attribute information information information that indicates which of the first and second encoding methods was used to encode the attribute information.
[0153] According to this method, different encoding methods can be used for location information and attribute information, thereby improving encoding efficiency.
[0154] For example, the three-dimensional data encoding method further stores the encoded stream in one or more units (e.g., NAL units).
[0155] For example, the unit has a format common to both the first and second encoding methods, and includes information indicating the type of data contained in the unit, which has an independent definition for both the first and second encoding methods (e.g., pcc_nal_unit_type).
[0156] For example, the unit has a format independent of the first encoding method and the second encoding method, and includes information indicating the type of data contained in the unit, which has a definition independent of the first encoding method and the second encoding method (for example, codec1_nal_unit_type or codec2_nal_unit_type).
[0157] For example, the unit has a format common to both the first encoding method and the second encoding method, and includes information indicating the type of data contained in the unit, which has a definition common to both the first encoding method and the second encoding method (e.g., pcc_nal_unit_type).
[0158] For example, a three-dimensional data encoding device comprises a processor and memory, and the processor uses the memory to perform the above processing.
[0159] Furthermore, the three-dimensional data decoding device according to this embodiment determines the encoding method used to encode the encoded stream based on information (for example, codec identification information) indicating which of the first and second encoding methods was used to encode the three-dimensional data, which is included in the control information (for example, parameter set) of the encoded stream generated by encoding the three-dimensional data, and decodes the encoded stream using the determined encoding method.
[0160] According to this, when decoding an encoded stream, the three-dimensional data decoding device can determine the encoding method used by using the information stored in the control information. Therefore, the three-dimensional data decoding device can correctly decode an encoded stream even when multiple encoding methods are used.
[0161] For example, the three-dimensional data includes location information, and the encoded stream includes encoded data of the location information. In the determination, the three-dimensional data decoding device determines the encoding method used to encode the location information based on information indicating which of the first and second encoding methods was used to encode the location information, which is included in the control information of the location information included in the encoded stream. In the decoding, the three-dimensional data decoding device decodes the encoded data of the location information using the determined encoding method used to encode the location information.
[0162] For example, the three-dimensional data includes location information and attribute information, and the encoded stream includes encoded data of the location information and encoded data of the attribute information. In the determination, the three-dimensional data decoding device determines the encoding method used for encoding the location information based on information indicating which of the first and second encoding methods was used for encoding the location information, which is included in the control information of the location information included in the encoded stream, and determines the encoding method used for encoding the attribute information based on information indicating which of the first and second encoding methods was used for encoding the attribute information, which is included in the control information of the attribute information included in the encoded stream. In the decoding, the three-dimensional data decoding device decodes the encoded data of the location information using the determined encoding method used for encoding the location information, and decodes the encoded data of the attribute information using the determined encoding method used for encoding the attribute information.
[0163] According to this method, different encoding methods can be used for location information and attribute information, thereby improving encoding efficiency.
[0164] For example, the encoded stream is stored in one or more units (e.g., NAL units), and the three-dimensional data decoding device further acquires the encoded stream from the one or more units.
[0165] For example, the unit has a format common to both the first and second encoding methods, and includes information indicating the type of data contained in the unit, which has an independent definition for both the first and second encoding methods (e.g., pcc_nal_unit_type).
[0166] For example, the unit has a format independent of the first encoding method and the second encoding method, and includes information indicating the type of data contained in the unit, which has a definition independent of the first encoding method and the second encoding method (for example, codec1_nal_unit_type or codec2_nal_unit_type).
[0167] For example, the unit has a format common to both the first encoding method and the second encoding method, and includes information indicating the type of data contained in the unit, which has a definition common to both the first encoding method and the second encoding method (e.g., pcc_nal_unit_type).
[0168] For example, a three-dimensional data decoding device comprises a processor and memory, and the processor uses the memory to perform the above processing.
[0169] (Embodiment 2) This embodiment describes a method for storing NAL units in an ISOBMFF file.
[0170] ISOBMFF (ISO-based media file format) is a file format standard defined in ISO / IEC 14496-12. ISOBMFF specifies a format that can store various media such as video, audio, and text in multiplexed format, and is a media-independent standard.
[0171] This section explains the basic structure (file) of ISOBMFF. The basic unit in ISOBMFF is a box. A box consists of type, length, and data, and a file is a collection of boxes of various types.
[0172] Figure 18 shows the basic structure (file) of ISOBMFF. An ISOBMFF file mainly contains boxes such as ftyp, which indicates the file brand using 4CC (4-character code), moov, which stores metadata such as control information, and mdat, which stores data.
[0173] The method for storing each type of media in an ISOBMFF file is specified separately; for example, the method for storing AVC video and HEVC video is specified in ISO / IEC 14496-15. Here, it is conceivable to extend the functionality of ISOBMFF to store or transmit PCC encoded data, but there is currently no provision for storing PCC encoded data in an ISOBMFF file. Therefore, this embodiment describes a method for storing PCC encoded data in an ISOBMFF file.
[0174] Figure 19 shows the protocol stack when a common NAL unit for PCC codecs is stored in an ISOBMFF file. Here, a common NAL unit for PCC codecs is stored in an ISOBMFF file. Although the NAL unit is common to PCC codecs, multiple PCC codecs are stored in the NAL unit, so it is desirable to define a storage method (Carriage of Codec1, Carriage of Codec2) according to each codec.
[0175] Next, we will explain how to store a common PCC NAL unit that supports multiple PCC codecs in an ISOBMFF file. Figure 20 shows an example of storing a common PCC NAL unit in an ISOBMFF file using the storage method for codec 1 (Carriage of Codec 1). Figure 21 shows an example of storing a common PCC NAL unit in an ISOBMFF file using the storage method for codec 2 (Carriage of Codec 2).
[0176] Here, ftyp is important information for identifying the file format, and a different identifier is defined for each codec for ftyp. When PCC encoded data encoded with the first encoding method (encoding scheme) is stored in the file, ftyp is set to pcc1. When PCC encoded data encoded with the second encoding method is stored in the file, ftyp is set to pcc2.
[0177] Here, pcc1 indicates that PCC codec 1 (first encoding method) is used. pcc2 indicates that PCC codec 2 (second encoding method) is used. In other words, pcc1 and pcc2 indicate that the data is PCC (encoded data of three-dimensional data (point cloud data)) and also indicate PCC codecs (first encoding method and second encoding method).
[0178] The following describes how to store NAL units in an ISOBMFF file. The multiplexing unit analyzes the NAL unit header, and if pcc_codec_type = Codec1, it writes pcc1 to ftyp in ISOBMFF.
[0179] Furthermore, the multiplexing unit analyzes the NAL unit header and, if pcc_codec_type = Codec2, writes pcc2 to ftyp in ISOBMFF.
[0180] Furthermore, if pcc_nal_unit_type is metadata, the multiplexing unit stores the NAL unit in a predetermined manner, for example, in moov or mdat. If pcc_nal_unit_type is data, the multiplexing unit stores the NAL unit in a predetermined manner, for example, in moov or mdat.
[0181] For example, the multiplexing unit may store the NAL unit size in the NAL unit, similar to HEVC.
[0182] This storage method allows the demultiplexing unit (system layer) to analyze the ftyp contained in the file, thereby determining whether the PCC encoded data was encoded using the first or second encoding method. Furthermore, as described above, by determining whether the PCC encoded data was encoded using the first or second encoding method, it is possible to extract encoded data encoded using one of the two encoding methods from data containing a mixture of encoded data encoded using both methods. This reduces the amount of data transmitted when transmitting encoded data. In addition, this storage method allows the use of a common data format without setting different data (file) formats for the first and second encoding methods.
[0183] Furthermore, if the system layer metadata, such as ftyp in ISOBMFF, includes codec identification information, the multiplexing unit may store the NAL unit with pcc_nal_unit_type removed in the ISOBMFF file.
[0184] Next, the configuration and operation of the multiplexing unit of the three-dimensional data encoding system (three-dimensional data encoding device) according to this embodiment, and the demultiplexing unit of the three-dimensional data decoding system (three-dimensional data decoding device) according to this embodiment will be described.
[0185] Figure 22 shows the configuration of the first multiplexing unit 4710. The first multiplexing unit 4710 includes a file conversion unit 4711 that generates multiplexed data (file) by storing the encoded data and control information (NAL unit) generated by the first encoding unit 4630 into an ISOBMFF file. This first multiplexing unit 4710 is included, for example, in the multiplexing unit 4614 shown in Figure 1.
[0186] Figure 23 shows the configuration of the first demultiplexing unit 4720. The first demultiplexing unit 4720 includes a file inverse conversion unit 4721 that acquires encoded data and control information (NAL unit) from the multiplexed data (file) and outputs the acquired encoded data and control information to the first decoding unit 4640. This first demultiplexing unit 4720 is included, for example, in the demultiplexing unit 4623 shown in Figure 1.
[0187] Figure 24 shows the configuration of the second multiplexing unit 4730. The second multiplexing unit 4730 includes a file conversion unit 4731 that generates multiplexed data (file) by storing the encoded data and control information (NAL unit) generated by the second encoding unit 4650 into an ISOBMFF file. This second multiplexing unit 4730 is included, for example, in the multiplexing unit 4614 shown in Figure 1.
[0188] Figure 25 shows the configuration of the second demultiplexing unit 4740. The second demultiplexing unit 4740 includes a file inverse conversion unit 4741 that acquires encoded data and control information (NAL unit) from the multiplexed data (file) and outputs the acquired encoded data and control information to the second decoding unit 4660. This second demultiplexing unit 4740 is included, for example, in the demultiplexing unit 4623 shown in Figure 1.
[0189] Figure 26 is a flowchart of the multiplexing process performed by the first multiplexing unit 4710. First, the first multiplexing unit 4710 analyzes the pcc_codec_type included in the NAL unit header to determine whether the codec being used is the first encoding method or the second encoding method (S4701).
[0190] If pcc_codec_type indicates a second encoding method (second encoding method in S4702), the first multiplexing unit 4710 does not process the NAL unit (S4703).
[0191] On the other hand, if pcc_codec_type indicates a second encoding method (the first encoding method in S4702), the first multiplexing unit 4710 writes pcc1 to ftyp (S4704). In other words, the first multiplexing unit 4710 writes information to ftyp indicating that data encoded with the first encoding method is stored in the file.
[0192] Next, the first multiplexing unit 4710 analyzes the pcc_nal_unit_type included in the NAL unit header and stores the data in a box (moov or mdat, etc.) in a predetermined manner according to the data type indicated by pcc_nal_unit_type (S4705). Then, the first multiplexing unit 4710 creates an ISOBMFF file containing the ftyp and the box (S4706).
[0193] Figure 27 is a flowchart of the multiplexing process performed by the second multiplexing unit 4730. First, the second multiplexing unit 4730 analyzes the pcc_codec_type included in the NAL unit header to determine whether the codec being used is the first encoding method or the second encoding method (S4711).
[0194] If pcc_unit_type indicates a second encoding method (the second encoding method in S4712), the second multiplexing unit 4730 writes pcc2 to ftyp (S4713). In other words, the second multiplexing unit 4730 writes information to ftyp indicating that data encoded with the second encoding method is stored in the file.
[0195] Next, the second multiplexing unit 4730 analyzes the pcc_nal_unit_type included in the NAL unit header and stores the data in a box (moov or mdat, etc.) in a predetermined manner according to the data type indicated by pcc_nal_unit_type (S4714). Then, the second multiplexing unit 4730 creates an ISOBMFF file containing the ftyp and the box (S4715).
[0196] On the other hand, if pcc_unit_type indicates the first encoding method (first encoding method in S4712), the second multiplexing unit 4730 does not process the NAL unit (S4716).
[0197] The above process illustrates an example of encoding PCC data using either the first encoding method or the second encoding method. The first multiplexing unit 4710 and the second multiplexing unit 4730 store the desired NAL units in a file by identifying the codec type of the NAL units. If PCC codec identification information is included in addition to the NAL unit header, the first multiplexing unit 4710 and the second multiplexing unit 4730 may, in steps S4701 and S4711, use the PCC codec identification information included in addition to the NAL unit header to identify the codec type (first encoding method or second encoding method).
[0198] Furthermore, the first multiplexing unit 4710 and the second multiplexing unit 4730 may, in steps S4706 and S4714, remove pcc_nal_unit_type from the NAL unit header before storing the data in the file.
[0199] Figure 28 is a flowchart showing the processing performed by the first demultiplexing unit 4720 and the first decoding unit 4640. First, the first demultiplexing unit 4720 analyzes the ftyp contained in the ISOBMFF file (S4721). If the codec indicated by ftyp is the second encoding method (pcc2) (second encoding method in S4722), the first demultiplexing unit 4720 determines that the data contained in the NAL unit payload is data encoded using the second encoding method (S4723). The first demultiplexing unit 4720 also transmits the result of this determination to the first decoding unit 4640. The first decoding unit 4640 does not process the NAL unit (S4724).
[0200] On the other hand, if the codec indicated by ftyp is the first encoding method (pcc1) (first encoding method in S4722), the first demultiplexing unit 4720 determines that the data contained in the NAL unit's payload is data encoded using the first encoding method (S4725). The first demultiplexing unit 4720 also transmits the result of this determination to the first decoding unit 4640.
[0201] The first decoding unit 4640 identifies the data by assuming that pcc_nal_unit_type included in the NAL unit header is the identifier for the NAL unit for the first encoding method (S4726). Then, the first decoding unit 4640 decodes the PCC data using the decoding process of the first encoding method (S4727).
[0202] Figure 29 is a flowchart showing the processing performed by the second demultiplexing unit 4740 and the second decoding unit 4660. First, the second demultiplexing unit 4740 analyzes the ftyp contained in the ISOBMFF file (S4731). If the codec indicated by ftyp is the second encoding method (pcc2) (second encoding method in S4732), the second demultiplexing unit 4740 determines that the data contained in the NAL unit payload is data encoded using the second encoding method (S4733). The second demultiplexing unit 4740 also transmits the result of this determination to the second decoding unit 4660.
[0203] The second decoding unit 4660 identifies the data as having pcc_nal_unit_type included in the NAL unit header as the identifier for the NAL unit for the second encoding method (S4734). Then, the second decoding unit 4660 decodes the PCC data using the decoding process of the second encoding method (S4735).
[0204] On the other hand, if the codec indicated by ftyp is the first encoding method (pcc1) (first encoding method in S4732), the second demultiplexing unit 4740 determines that the data contained in the payload of the NAL unit is data encoded using the first encoding method (S4736). The second demultiplexing unit 4740 also transmits the result of this determination to the second decoding unit 4660. The second decoding unit 4660 does not process the NAL unit (S4737).
[0205] Thus, for example, by identifying the codec type of the NAL unit in the first demultiplexing unit 4720 or the second demultiplexing unit 4740, the codec type can be identified at an early stage. Furthermore, the desired NAL unit can be input to the first decoding unit 4640 or the second decoding unit 4660, and unnecessary NAL units can be removed. In this case, the process of analyzing the codec identification information in the first decoding unit 4640 or the second decoding unit 4660 may become unnecessary. However, the first decoding unit 4640 or the second decoding unit 4660 may again refer to the NAL unit type and perform the process of analyzing the codec identification information.
[0206] Furthermore, if the first multiplexing unit 4710 or the second multiplexing unit 4730 has removed pcc_nal_unit_type from the NAL unit header, the first demultiplexing unit 4720 or the second demultiplexing unit 4740 may add pcc_nal_unit_type to the NAL unit and then output it to the first decoding unit 4640 or the second decoding unit 4660.
[0207] (Embodiment 3) In this embodiment, a multiplexing unit and a demultiplexing unit corresponding to the encoding unit 4670 and decoding unit 4680 that correspond to multiple codecs as described in Embodiment 1 will be described. Figure 30 is a diagram showing the configuration of the encoding unit 4670 and the third multiplexing unit 4750 according to this embodiment.
[0208] The encoding unit 4670 encodes the point cloud data using either the first encoding method or the second encoding method, or both. The encoding unit 4670 may switch the encoding method (the first encoding method and the second encoding method) on a point cloud data basis or on a frame basis. Alternatively, the encoding unit 4670 may switch the encoding method on an encodeable basis.
[0209] The encoding unit 4670 generates encoded data (encoded stream) that includes identification information for the PCC codec.
[0210] The third multiplexing unit 4750 includes a file conversion unit 4751. The file conversion unit 4751 converts the NAL units output from the encoding unit 4670 into a PCC data file. The file conversion unit 4751 analyzes the codec identification information contained in the NAL unit header and determines whether the PCC encoded data is data encoded using a first encoding method, data encoded using a second encoding method, or data encoded using both methods. The file conversion unit 4751 writes a brand name that can identify the codec in ftyp. For example, if it indicates that the data is encoded using both methods, pcc3 is written in ftyp.
[0211] Furthermore, if the encoding unit 4670 contains identification information for the PCC codec in addition to the NAL unit, the file conversion unit 4751 may use this identification information to determine the PCC codec (encoding method).
[0212] Figure 31 shows the configuration of the third demultiplexing unit 4760 and decoding unit 4680 according to this embodiment.
[0213] The third demultiplexing unit 4760 includes a file inverse conversion unit 4761. The file inverse conversion unit 4761 analyzes the ftyp contained in the file and determines whether the PCC encoded data is data encoded by the first encoding method, data encoded by the second encoding method, or data encoded by both methods.
[0214] If the PCC encoded data is encoded using either one of the encoding methods, the data is input to the corresponding decoding unit of the first decoding unit 4640 and the second decoding unit 4660, and the data is not input to the other decoding unit. If the PCC encoded data is encoded using both encoding methods, the data is input to the decoding unit 4680 corresponding to both methods.
[0215] The decoding unit 4680 decodes the PCC encoded data using either the first encoding method or the second encoding method, or both.
[0216] Figure 32 is a flowchart showing the processing performed by the third multiplexing unit 4750 according to this embodiment.
[0217] First, the third multiplexing unit 4750 analyzes the pcc_codec_type included in the NAL unit header to determine whether the codec being used is the first encoding method, the second encoding method, or both the first and second encoding methods (S4741).
[0218] If the second encoding method is used (Yes in S4742 and the second encoding method in S4743), the third multiplexing unit 4750 writes pcc2 to ftyp (S4744). In other words, the third multiplexing unit 4750 writes information to ftyp indicating that data encoded using the second encoding method is stored in the file.
[0219] Next, the third multiplexing unit 4750 analyzes the pcc_nal_unit_type included in the NAL unit header and stores the data in a box (moov or mdat, etc.) in a predetermined manner according to the data type indicated by pcc_unit_type (S4745). Then, the third multiplexing unit 4750 creates an ISOBMFF file containing the ftyp and the box (S4746).
[0220] On the other hand, if the first encoding method is used (Yes in S4742 and the first encoding method in S4743), the third multiplexing unit 4750 writes pcc1 to ftyp (S4747). In other words, the third multiplexing unit 4750 writes information to ftyp indicating that data encoded using the first encoding method is stored in the file.
[0221] Next, the third multiplexing unit 4750 analyzes the pcc_nal_unit_type included in the NAL unit header and stores the data in a box (moov or mdat, etc.) in a predetermined manner according to the data type indicated by pcc_unit_type (S4748). Then, the third multiplexing unit 4750 creates an ISOBMFF file containing the ftyp and the box (S4746).
[0222] On the other hand, if both the first encoding method and the second encoding method are used (No in S4742), the third multiplexing unit 4750 writes pcc3 to ftyp (S4749). In other words, the third multiplexing unit 4750 writes information to ftyp indicating that data encoded using both encoding methods is stored in the file.
[0223] Next, the third multiplexing unit 4750 analyzes the pcc_nal_unit_type included in the NAL unit header and stores the data in a box (moov or mdat, etc.) in a predetermined manner according to the data type indicated by pcc_unit_type (S4750). Then, the third multiplexing unit 4750 creates an ISOBMFF file containing the ftyp and the box (S4746).
[0224] Figure 33 is a flowchart showing the processing performed by the third demultiplexing unit 4760 and the decoding unit 4680. First, the third demultiplexing unit 4760 analyzes the ftyp contained in the ISOBMFF file (S4761). If the codec indicated by ftyp is the second encoding method (pcc2) (Yes in S4762 and the second encoding method in S4763), the third demultiplexing unit 4760 determines that the data contained in the NAL unit payload is data encoded using the second encoding method (S4764). The third demultiplexing unit 4760 also transmits the result of this determination to the decoding unit 4680.
[0225] The decoding unit 4680 identifies the data as having pcc_nal_unit_type included in the NAL unit header as the identifier for the NAL unit for the second encoding method (S4765). Then, the decoding unit 4680 decodes the PCC data using the decoding process of the second encoding method (S4766).
[0226] On the other hand, if the codec indicated by ftyp is the first encoding method (pcc1) (Yes in S4762 and the first encoding method in S4763), the third demultiplexing unit 4760 determines that the data contained in the NAL unit's payload is data encoded using the first encoding method (S4767). The third demultiplexing unit 4760 also transmits the result of this determination to the decoding unit 4680.
[0227] The decoding unit 4680 identifies the data by assuming that pcc_nal_unit_type included in the NAL unit header is the identifier for the NAL unit for the first encoding method (S4768). Then, the decoding unit 4680 decodes the PCC data using the decoding process of the first encoding method (S4769).
[0228] On the other hand, if ftyp indicates that both encoding methods are used (pcc3) (No in S4762), the third demultiplexing unit 4760 determines that the data included in the NAL unit's payload is data encoded using both the first and second encoding methods (S4770). The third demultiplexing unit 4760 also transmits the result of this determination to the decoding unit 4680.
[0229] The decoding unit 4680 identifies the data by assuming that pcc_nal_unit_type included in the NAL unit header is the identifier for the NAL unit for the codec described in pcc_codec_type (S4771). Then, the decoding unit 4680 decodes the PCC data using the decoding processes of both encoding methods (S4772). In other words, the decoding unit 4680 decodes the data encoded by the first encoding method using the decoding process of the first encoding method, and decodes the data encoded by the second encoding method using the decoding process of the second encoding method.
[0230] The following describes modifications of this embodiment. The following types may be indicated in the identification information as the brand type shown in ftyp. In addition, combinations of the following types may be indicated in the identification information.
[0231] The identification information may indicate whether the object of the original data before PCC encoding is a point cloud with a limited area, or a large point cloud with no area limitations, such as map information.
[0232] The identification information may indicate whether the original data before PCC encoding is a static or dynamic object.
[0233] As described above, the identification information may indicate whether the PCC encoded data is data encoded using the first encoding method or data encoded using the second encoding method.
[0234] The identification information may indicate the algorithm used in PCC coding. Here, the algorithm is, for example, an coding method that can be used in the first coding method or the second coding method.
[0235] The identification information may indicate differences in how the PCC encoded data is stored in ISOBMFF files. For example, the identification information may indicate whether the storage method used is a storage method for storage or a storage method for real-time transmission such as dynamic streaming.
[0236] Furthermore, while embodiments 2 and 3 describe examples in which ISOBMFF is used as the file format, other methods may also be used. For example, the same method as in this embodiment may be used when storing PCC encoded data in MPEG-2 TS Systems, MPEG-DASH, MMT, or RMP.
[0237] Furthermore, while the above example shows how to store metadata such as identification information in ftyp, this metadata may also be stored in a location other than ftyp. For example, this metadata may be stored in moov.
[0238] As described above, the three-dimensional data storage device (or three-dimensional data multiplexer, or three-dimensional data encoding device) performs the processing shown in Figure 34.
[0239] First, the three-dimensional data storage device (for example, including a first multiplexing unit 4710, a second multiplexing unit 4730, or a third multiplexing unit 4750) acquires one or more units (e.g., NAL units) in which encoded streams of point cloud data are stored (S4781). Next, the three-dimensional data storage device stores one or more units in a file (e.g., an ISOBMFF file) (S4782). In addition, during the storage (S4782), the three-dimensional data storage device stores information (e.g., pcc1, pcc2, or pcc3) indicating that the data stored in the file is data of encoded point cloud data in the control information of the file (e.g., ftyp).
[0240] According to this, a device that processes files generated by the three-dimensional data storage device can refer to the file's control information to quickly determine whether the data stored in the file is encoded point cloud data. Therefore, it is possible to reduce the processing load of the device or speed up processing.
[0241] For example, the information further indicates the encoding method used to encode the point cloud data from among the first encoding method and the second encoding method. Note that the fact that the data stored in the file is data in which the point cloud data has been encoded, and the encoding method used to encode the point cloud data from among the first encoding method and the second encoding method, may be indicated by a single piece of information or by different pieces of information.
[0242] According to this, a device that processes files generated by the three-dimensional data storage device can refer to the file's control information to quickly determine the codec used for the data stored in the file. Therefore, it is possible to reduce the processing load of the device or speed up processing.
[0243] For example, the first encoding method is a method (GPCC) that encodes positional information represented by an N (where N is an integer of 2 or more) subtree of the position of point cloud data, and encodes attribute information using the positional information, while the second encoding method is a method (VPCC) that generates a two-dimensional image from the point cloud data and encodes the two-dimensional image using a video encoding method.
[0244] For example, the aforementioned file conforms to ISOBMFF (ISO-based media file format).
[0245] For example, a three-dimensional data storage device comprises a processor and memory, and the processor uses the memory to perform the above processing.
[0246] Furthermore, as described above, the three-dimensional data acquisition device (or three-dimensional data demultiplexing device, or three-dimensional data decoding device) performs the processing shown in Figure 35.
[0247] The three-dimensional data acquisition device (for example, including a first demultiplexing unit 4720, a second demultiplexing unit 4740, or a third demultiplexing unit 4760) acquires a file (for example, an ISOBMFF file) that stores one or more units (for example, NAL units) in which encoded streams of point cloud data are stored (S4791). Next, the three-dimensional data acquisition device acquires one or more units from the file (S4792). The file control information (for example, ftyp) also includes information (for example, pcc1, pcc2, or pcc3) indicating that the data stored in the file is data in which point cloud data has been encoded.
[0248] For example, the three-dimensional data acquisition device refers to the aforementioned information to determine whether the data stored in the file is encoded point cloud data. If the three-dimensional data acquisition device determines that the data stored in the file is encoded point cloud data, it generates point cloud data by decoding the encoded point cloud data contained in one or more units. Alternatively, if the three-dimensional data acquisition device determines that the data stored in the file is encoded point cloud data, it outputs (notifies) a subsequent processing unit (for example, the first decoding unit 4640, the second decoding unit 4660, or the decoding unit 4680) information indicating that the data contained in one or more units is encoded point cloud data.
[0249] According to this, the three-dimensional data acquisition device can refer to the file's control information to quickly determine whether the data stored in the file is encoded point cloud data. Therefore, it is possible to reduce the processing load or speed up processing for the three-dimensional data acquisition device or subsequent devices.
[0250] For example, the information further indicates the encoding method used for the encoding, among the first and second encoding methods. Note that the fact that the data stored in the file is data in which point cloud data has been encoded, and the encoding method used for encoding the point cloud data, among the first and second encoding methods, may be indicated by a single piece of information or by different pieces of information.
[0251] According to this, the three-dimensional data acquisition device can refer to the file's control information to quickly determine the codec used for the data stored in the file. Therefore, it is possible to reduce the processing load or speed up processing for the three-dimensional data acquisition device or subsequent devices.
[0252] For example, the three-dimensional data acquisition device acquires data encoded using one of the encoding methods from encoded point cloud data, which includes data encoded using the first encoding method and data encoded using the second encoding method, based on the aforementioned information.
[0253] For example, the first encoding method is a method (GPCC) that encodes positional information represented by an N (where N is an integer of 2 or more) subtree of the position of point cloud data, and encodes attribute information using the positional information, while the second encoding method is a method (VPCC) that generates a two-dimensional image from the point cloud data and encodes the two-dimensional image using a video encoding method.
[0254] For example, the aforementioned file conforms to ISOBMFF (ISO-based media file format).
[0255] For example, a three-dimensional data acquisition device comprises a processor and memory, and the processor uses the memory to perform the above processing.
[0256] (Embodiment 4) In this embodiment, the types of encoded data (location information (Geometry), attribute information (Attribute), additional information (Metadata)) generated by the first encoding unit 4630 or the second encoding unit 4650 described above, the method for generating additional information (metadata), and the multiplexing process in the multiplexing unit will be described. Note that additional information (metadata) may also be referred to as parameter set or control information.
[0257] In this embodiment, we will explain using the dynamic object (three-dimensional point cloud data that changes over time) described in Figure 4 as an example, but the same method may be used for static objects (three-dimensional point cloud data at any given time).
[0258] Figure 36 shows the configuration of the encoding unit 4801 and the multiplexing unit 4802 included in the three-dimensional data encoding device according to this embodiment. The encoding unit 4801 corresponds, for example, to the first encoding unit 4630 or the second encoding unit 4650 described above. The multiplexing unit 4802 corresponds to the multiplexing unit 4634 or 46456 described above.
[0259] The encoding unit 4801 encodes point cloud data from multiple PCC (Point Cloud Compression) frames and generates multiple encoded data (Multiple Compressed Data) containing location information, attribute information, and additional information.
[0260] The multiplexing unit 4802 converts data of multiple data types (location information, attribute information, and additional information) into NAL units, thereby transforming the data into a data configuration that takes into account data access by the decoding device.
[0261] Figure 37 shows an example of the structure of encoded data generated by the encoding unit 4801. The arrows in the figure indicate dependencies related to the decoding of encoded data, with the source of the arrow depending on the data at the end of the arrow. In other words, the decoding device decodes the data at the end of the arrow and uses that decoded data to decode the source of the arrow. To put it another way, dependency means that the dependent data is referenced (used) in the processing of the dependent data (encoding or decoding, etc.).
[0262] First, the process for generating encoded location data will be explained. The encoding unit 4801 generates encoded location data (Compressed Geometry Data) for each frame by encoding the location information of each frame. The encoded location data is represented by G(i), where i represents the frame number or the time of the frame.
[0263] Furthermore, the encoding unit 4801 generates a position parameter set (GPS(i)) corresponding to each frame. The position parameter set includes parameters that can be used to decode the encoded position data. Also, the encoded position data for each frame depends on the corresponding position parameter set.
[0264] Furthermore, encoded position data consisting of multiple frames is defined as a position sequence (Geometry Sequence). The encoding unit 4801 generates a position sequence parameter set (Geometry Sequence PS: also written as Position SPS) that stores parameters commonly used for decoding multiple frames within the position sequence. The position sequence depends on the Position SPS.
[0265] Next, the process for generating encoded attribute data will be explained. The encoding unit 4801 generates encoded attribute data for each frame by encoding the attribute information of each frame. The encoded attribute data is represented by A(i). Figure 37 shows an example where attribute X and attribute Y exist, with the encoded attribute data for attribute X represented by AX(i) and the encoded attribute data for attribute Y represented by AY(i).
[0266] Furthermore, the encoding unit 4801 generates an attribute parameter set (APS(i)) corresponding to each frame. The attribute parameter set for attribute X is represented by AXPS(i), and the attribute parameter set for attribute Y is represented by AYPS(i). The attribute parameter set includes parameters that can be used to decode the encoded attribute information. The encoded attribute data depends on the corresponding attribute parameter set.
[0267] Furthermore, encoded attribute data consisting of multiple frames is defined as an attribute sequence. The encoding unit 4801 generates an attribute sequence parameter set (Attribute Sequence PS, also written as attribute SPS) that stores parameters commonly used for decoding multiple frames within the attribute sequence. The attribute sequence depends on the attribute SPS.
[0268] Furthermore, in the first encoding method, the encoded attribute data depends on the encoded position data.
[0269] Figure 37 also shows an example where there are two types of attribute information (attribute X and attribute Y). When there are two types of attribute information, for example, two encoding units generate the respective data and metadata. Also, for example, an attribute sequence is defined for each type of attribute information, and an attribute SPS is generated for each type of attribute information.
[0270] Note that Figure 37 shows an example where there is one type of positional information and two types of attribute information, but the system is not limited to this; there may be one type of attribute information or three or more types. In this case as well, encoded data can be generated using the same method. Furthermore, in the case of point cloud data that does not have attribute information, attribute information is not required. In that case, the encoding unit 4801 does not need to generate a parameter set related to attribute information.
[0271] Next, the process of generating additional information (metadata) will be described. The encoding unit 4801 generates a PCC stream PS (also written as Stream PS), which is a parameter set for the entire PCC stream. The encoding unit 4801 stores in Stream PS parameters that can be used in common for decoding one or more position sequences and one or more attribute sequences. For example, Stream PS includes identification information indicating the codec of the point cloud data, and information indicating the algorithm used for encoding. The position sequences and attribute sequences depend on Stream PS.
[0272] Next, the Access Unit and GOF will be described. In this embodiment, the concepts of Access Unit (AU) and GOF (Group of Frame) are newly introduced.
[0273] An access unit is the basic unit for accessing data during decryption, and consists of one or more data points and one or more metadata points. For example, an access unit consists of location information at the same time and one or more attribute information points. A GOF (Group of Four) is a random access unit and consists of one or more access units.
[0274] The encoding unit 4801 generates an access unit header (AU Header) as identification information indicating the beginning of an access unit. The encoding unit 4801 stores parameters related to the access unit in the access unit header. For example, the access unit header includes the structure or information of the encoded data contained in the access unit. The access unit header also includes parameters commonly used in the data contained in the access unit, such as parameters related to decoding the encoded data.
[0275] The encoding unit 4801 may generate an access unit delimiter that does not include parameters related to the access unit, instead of an access unit header. This access unit delimiter is used as identification information to indicate the beginning of the access unit. The decoding device identifies the beginning of the access unit by detecting the access unit header or the access unit delimiter.
[0276] Next, the generation of identification information for the beginning of the GOF will be explained. The encoding unit 4801 generates a GOF header as identification information indicating the beginning of the GOF. The encoding unit 4801 stores parameters related to the GOF in the GOF header. For example, the GOF header includes the structure or information of the encoded data contained in the GOF. The GOF header also includes parameters commonly used in the data contained in the GOF, such as parameters related to decoding the encoded data.
[0277] The encoding unit 4801 may generate a GOF delimiter that does not include parameters related to the GOF, instead of a GOF header. This GOF delimiter is used as identification information to indicate the beginning of the GOF. The decoding device identifies the beginning of the GOF by detecting the GOF header or the GOF delimiter.
[0278] In PCC encoded data, for example, an access unit is defined as a PCC frame. The decoding device accesses the PCC frame based on the identification information at the beginning of the access unit.
[0279] Furthermore, for example, a GOF is defined as a single random access unit. The decryption device accesses the random access unit based on the identification information at the beginning of the GOF. For example, if PCC frames are independent of each other and can be decrypted individually, then a PCC frame may be defined as a random access unit.
[0280] Furthermore, two or more PCC frames may be assigned to a single access unit, and multiple random access units may be assigned to a single GOF.
[0281] Furthermore, the encoding unit 4801 may define and generate parameter sets or metadata other than those described above. For example, the encoding unit 4801 may generate SEI (Supplemental Enhancement Information) which stores parameters that may not necessarily be used during decoding (optional parameters).
[0282] Next, we will explain the structure of the encoded data and how to store the encoded data in the NAL unit.
[0283] For example, a data format is defined for each type of encoded data. Figure 38 shows an example of encoded data and a NAL unit.
[0284] For example, as shown in Figure 38, encoded data includes a header and a payload. The encoded data may also include length information indicating the length (data volume) of the encoded data, header, or payload. Furthermore, the encoded data does not necessarily have to include a header.
[0285] The header includes, for example, identification information to identify the data. This identification information may indicate, for example, the data type or frame number.
[0286] The header contains, for example, identification information indicating a reference relationship. This identification information is stored in the header when there is a dependency between data, and it is information used to reference the referenced data from the source. For example, the header of the referenced data contains identification information to identify that data. The header of the referenced data contains identification information indicating the referenced data.
[0287] Furthermore, if the referenced or source can be identified or derived from other information, the identifying information for identifying the data or identifying information indicating the reference relationship may be omitted.
[0288] The multiplexing unit 4802 stores the encoded data in the payload of the NAL unit. The NAL unit header contains pcc_nal_unit_type, which is identification information for the encoded data. Figure 39 shows an example of the semantics of pcc_nal_unit_type.
[0289] As shown in Figure 39, when pcc_codec_type is codec 1 (Codec1: first encoding method), the values of pcc_nal_unit_type 0 to 10 are the encoded position data (Geometry), encoded attribute X data (AttributeX), encoded attribute Y data (AttributeY), position PS (Geom.PS), attribute XPS (AttriX.PS), attribute YPS (AttriX.PS), position SPS (Geometry Sequence PS), attribute XSPS (AttributeX Sequence PS), attribute YSPS (AttributeY Sequence PS), AU header (AU Header), GOF header (GOF It is assigned to the Header. Values 11 and above are also assigned to the backup of Codec 1.
[0290] If pcc_codec_type is Codec 2 (the second encoding method), then values of pcc_nal_unit_type from 0 to 2 are assigned to the codec's Data A, Metadata A, and Metadata B. Values 3 and above are assigned to the backup of Codec 2.
[0291] Next, we will explain the data transmission order. The following describes the constraints on the transmission order of the NAL unit.
[0292] The multiplexing unit 4802 sends out NAL units in groups of GOF or AU units. The multiplexing unit 4802 places a GOF header at the beginning of each GOF and an AU header at the beginning of each AU.
[0293] The multiplexing unit 4802 may provide a sequence parameter set (SPS) for each AU so that the decoding device can decode from the next AU even if data is lost due to packet loss or other reasons.
[0294] If there are dependencies in the encoded data related to decoding, the decoding device decodes the referenced data first, and then decodes the source data. In order to enable decoding in the order in which the data was received without rearranging the data in the decoding device, the multiplexing unit 4802 sends the referenced data first.
[0295] Figure 40 shows an example of the transmission order of the NAL unit. Figure 40 shows three examples: location information priority, parameter priority, and data integration.
[0296] The location-prioritized transmission order is an example where location information and attribute information are transmitted together. In this transmission order, the transmission of location information is completed earlier than the transmission of attribute information.
[0297] For example, by using this transmission order, a decoding device that does not decode attribute information may be able to create a period of time where it does not process attribute information by ignoring the decoding of attribute information. Also, for example, a decoding device that wants to decode location information quickly may be able to decode the location information faster by obtaining the encoded location information data earlier.
[0298] Note that in Figure 40, the attributes XSPS and YSPS are combined and labeled as attribute SPS, but attributes XSPS and YSPS may also be placed separately.
[0299] In a parameter set priority transmission order, the parameter set is sent first, followed by the data.
[0300] As long as the constraints on the NAL unit transmission order are followed as described above, the multiplexing unit 4802 may transmit the NAL units in any order. For example, sequence identification information may be defined, and the multiplexing unit 4802 may have the function of transmitting NAL units in multiple patterns of order. For example, the sequence identification information of the NAL units may be stored in the stream PS.
[0301] The three-dimensional data decoding device may perform decoding based on sequence identification information. The three-dimensional data decoding device may instruct the three-dimensional data encoding device to send a desired transmission order, and the three-dimensional data encoding device (multiplexing unit 4802) may control the transmission order according to the instructed transmission order.
[0302] Furthermore, the multiplexing unit 4802 may generate encoded data that merges multiple functions, as long as it adheres to the constraints of the transmission order, such as the transmission order of data integration. For example, as shown in Figure 40, the GOF header and the AU header may be integrated, or the AXPS and AYPS may be integrated. In this case, pcc_nal_unit_type is defined as an identifier indicating that the data has multiple functions.
[0303] The following describes modifications of this embodiment. PS has levels, such as frame-level PS, sequence-level PS, and PCC sequence-level PS. If the PCC sequence level is considered a higher level and the frame level a lower level, the following method may be used to store the parameters.
[0304] The default PS value is shown in the higher-level PS. If the value of a lower-level PS differs from the value of a higher-level PS, the PS value is shown in the lower-level PS. Alternatively, the PS value is not listed in the higher-level PS, but is listed in the lower-level PS. Alternatively, information indicating whether the PS value is shown in the lower-level PS, the higher-level PS, or both is shown in either the lower-level PS or the higher-level PS, or both. Alternatively, the lower-level PS may be merged with the higher-level PS. Alternatively, if the lower-level PS and the higher-level PS overlap, the multiplexing unit 4802 may omit sending one of them.
[0305] The encoding unit 4801 or the multiplexing unit 4802 may divide the data into slices or tiles and send out the divided data. The divided data includes information for identifying the divided data, and the parameters used for decoding the divided data are included in the parameter set. In this case, pcc_nal_unit_type is defined as an identifier indicating that it is data that stores data or parameters related to tiles or slices.
[0306] The following describes the processing related to sequence identification information. Figure 41 is a flowchart of the processing performed by the three-dimensional data encoding device (encoding unit 4801 and multiplexing unit 4802) related to the transmission order of the NAL unit.
[0307] First, the three-dimensional data encoding device determines the transmission order of the NAL units (position information priority or parameter set priority) (S4801). For example, the three-dimensional data encoding device determines the transmission order based on a specification from the user or an external device (e.g., a three-dimensional data decoding device).
[0308] If the determined transmission order prioritizes location information (location information priority in S4802), the three-dimensional data encoding device sets the sequence identification information included in the stream PS to prioritize location information (S4803). In other words, in this case, the sequence identification information indicates that the NAL units will be transmitted in the order prioritizing location information. The three-dimensional data encoding device then transmits the NAL units in the order prioritizing location information (S4804).
[0309] On the other hand, if the determined transmission order is parameter set priority (parameter set priority in S4802), the three-dimensional data encoding device sets the sequence identification information included in the stream PS to parameter set priority (S4805). In other words, in this case, the sequence identification information indicates that the NAL units will be transmitted in the parameter set priority order. The three-dimensional data encoding device then transmits the NAL units in the parameter set priority order (S4806).
[0310] Figure 42 is a flowchart of the processing performed by the three-dimensional data decoding device regarding the transmission order of the NAL unit. First, the three-dimensional data decoding device analyzes the sequence identification information contained in the stream PS (S4811).
[0311] If the transmission order indicated by the sequence identification information is based on positional information priority (positional information priority in S4812), the three-dimensional data decoding device decodes the NAL units assuming that the transmission order of the NAL units is based on positional information priority (S4813).
[0312] On the other hand, if the transmission order indicated by the sequence identification information is parameter set priority (parameter set priority in S4812), the three-dimensional data decoding device decodes the NAL units assuming that the transmission order of the NAL units is parameter set priority (S4814).
[0313] For example, if the three-dimensional data decoding device does not decode attribute information, in step S4813 it may not acquire all NAL units, but instead acquire NAL units related to location information and decode the location information from the acquired NAL units.
[0314] Next, the process related to the generation of AUs and GOFs will be explained. Figure 43 is a flowchart of the processing by the three-dimensional data encoding device (multiplexing unit 4802) related to the generation of AUs and GOFs in the multiplexing of NAL units.
[0315] First, the three-dimensional data encoding device determines the type of encoded data (S4821). Specifically, the three-dimensional data encoding device determines whether the encoded data to be processed is data starting with AU, data starting with GOF, or other data.
[0316] If the encoded data is the first data of the GOF (GOF first in S4822), the three-dimensional data encoding device places the GOF header and AU header at the beginning of the encoded data belonging to the GOF and generates a NAL unit (S4823).
[0317] If the encoded data is data that starts with AU (AU at the beginning in S4822), the three-dimensional data encoding device places the AU header at the beginning of the encoded data belonging to AU and generates a NAL unit (S4824).
[0318] If the encoded data is neither GOF-first nor AU-first (i.e., not GOF-first or AU-first in S4822), the three-dimensional data encoding device generates a NAL unit by placing the encoded data after the AU header of the AU to which the encoded data belongs (S4825).
[0319] Next, the processing related to accessing the AU and GOF will be explained. Figure 44 is a flowchart of the processing of the three-dimensional data decoding device related to accessing the AU and GOF in the demultiplexing of the NAL unit.
[0320] First, the three-dimensional data decoding device determines the type of encoded data contained in the NAL unit by analyzing the nal_unit_type contained in the NAL unit (S4831). Specifically, the three-dimensional data decoding device determines whether the encoded data contained in the NAL unit is data at the beginning of the AU, data at the beginning of the GOF, or other data.
[0321] If the encoded data contained in the NAL unit is the data at the beginning of the GOF (the beginning of the GOF in S4832), the three-dimensional data decoding device determines that the NAL unit is the starting position for random access, accesses the NAL unit, and starts the decoding process (S4833).
[0322] On the other hand, if the encoded data contained in the NAL unit is the data at the beginning of the AU (AU at the beginning in S4832), the three-dimensional data decoding device determines that the NAL unit is at the beginning of the AU, accesses the data contained in the NAL unit, and decodes the AU (S4834).
[0323] On the other hand, if the encoded data contained in the NAL unit is neither GOF-first nor AU-first (i.e., in S4832, neither GOF-first nor AU-first), the three-dimensional data decoding device does not process the NAL unit.
[0324] As described above, the three-dimensional data encoding device performs the processing shown in Figure 45. The three-dimensional data encoding device encodes time-series three-dimensional data (for example, point cloud data of a dynamic object). The three-dimensional data includes position information and attribute information for each time step.
[0325] First, the three-dimensional data encoding device encodes location information (S4841). Next, the three-dimensional data encoding device encodes the attribute information of the data to be processed by referring to the location information at the same time as the attribute information of the data to be processed (S4842). Here, as shown in Figure 37, the location information and attribute information at the same time constitute an access unit (AU). In other words, the three-dimensional data encoding device encodes the attribute information of the data to be processed by referring to the location information included in the same access unit as the attribute information of the data to be processed.
[0326] According to this, the three-dimensional data encoding device can simplify the control of references in encoding using an access unit. Therefore, the three-dimensional data encoding device can reduce the processing load of the encoding process.
[0327] For example, a three-dimensional data encoding device generates a bitstream that includes encoded location information (encoded location data), encoded attribute information (encoded attribute data), and information indicating the location information of the reference point for the attribute information to be processed.
[0328] For example, a bitstream includes a position parameter set (position PS) containing control information for the position information at each time point, and an attribute parameter set (attribute PS) containing control information for the attribute information at each time point.
[0329] For example, a bitstream includes a position sequence parameter set (position SPS) containing control information common to position information at multiple time points, and an attribute sequence parameter set (attribute SPS) containing control information common to attribute information at multiple time points.
[0330] For example, a bitstream includes a stream parameter set (stream PS) that contains control information common to the position information and attribute information of multiple time points.
[0331] For example, a bitstream includes an access unit header (AU header) that contains common control information within the access unit.
[0332] For example, a three-dimensional data encoding device encodes a GOF (Group of Frames), which consists of one or more access units, in a way that allows for independent decoding. In other words, a GOF is a random access unit.
[0333] For example, a bitstream includes a GOF header containing common control information within the GOF.
[0334] For example, a three-dimensional data encoding device comprises a processor and memory, and the processor uses the memory to perform the above processing.
[0335] Furthermore, as described above, the three-dimensional data decoding device performs the processing shown in Figure 46. The three-dimensional data decoding device decodes time-series three-dimensional data (for example, point cloud data of a dynamic object). The three-dimensional data includes position information and attribute information for each time point. Position information and attribute information at the same time point constitute an access unit (AU).
[0336] First, the three-dimensional data decoding device decodes the location information from the bitstream (S4851). In other words, the three-dimensional data decoding device generates location information by decoding the encoded location information (encoded location data) contained in the bitstream.
[0337] Next, the three-dimensional data decoding device decodes the attribute information to be processed from the bitstream, referring to the position information at the same time as the attribute information to be processed (S4852). In other words, the three-dimensional data decoding device generates attribute information by decoding the encoded attribute information (encoded attribute data) contained in the bitstream. At this time, the three-dimensional data decoding device refers to the decoded position information contained in the same access unit as the attribute information.
[0338] According to this, the three-dimensional data decoding device can simplify the control of references during decoding using an access unit. Therefore, this three-dimensional data decoding method can reduce the processing load of the decoding process.
[0339] For example, a three-dimensional data decoding device obtains information from a bitstream indicating the location of the reference point for the attribute information to be processed, and decodes the attribute information to be processed by referring to the reference point location indicated by the obtained information.
[0340] For example, a bitstream includes a position parameter set (position PS) containing control information for the position information at each time point, and an attribute parameter set (attribute PS) containing control information for the attribute information at each time point. In other words, the three-dimensional data decoding device uses the control information contained in the position parameter set for the time point to be processed to decode the position information for the time point to be processed, and uses the control information contained in the attribute parameter set for the time point to be processed to decode the attribute information for the time point to be processed.
[0341] For example, a bitstream includes a position sequence parameter set (position SPS) containing control information common to position information at multiple times, and an attribute sequence parameter set (attribute SPS) containing control information common to attribute information at multiple times. In other words, the three-dimensional data decoding device uses the control information contained in the position sequence parameter set to decode the position information at multiple times, and uses the control information contained in the attribute sequence parameter set to decode the attribute information at multiple times.
[0342] For example, a bitstream includes a stream parameter set (stream PS) containing control information common to the position information and attribute information of multiple time points. In other words, a three-dimensional data decoding device uses the control information contained in the stream parameter set to decode the position information and attribute information of multiple time points.
[0343] For example, a bitstream includes an access unit header (AU header) containing common control information within the access unit. In other words, the 3D data decoding device uses the control information contained in the access unit header to decode the location information and attribute information contained in the access unit.
[0344] For example, a three-dimensional data decoding device independently decodes a GOF (Group of Frames), which consists of one or more access units. In other words, a GOF is a random access unit.
[0345] For example, a bitstream includes a GOF header containing common control information within the GOF. In other words, a three-dimensional data decoding device uses the control information contained in the GOF header to decode the position information and attribute information contained in the GOF.
[0346] For example, a three-dimensional data decoding device comprises a processor and memory, and the processor uses the memory to perform the above processing.
[0347] (Embodiment 5) Next, the configuration of the division section 4911 will be described. Figure 47 is a block diagram of the division section 4911. The division section 4911 includes a slice division section 4931, a geometry tile division section 4932, and an attribute information tile division section 4933.
[0348] The slice division unit 4931 generates multiple slice position information by dividing position information (Position (Geometry)) into slices. The slice division unit 4931 also generates multiple slice attribute information by dividing attribute information (Attribute) into slices. Furthermore, the slice division unit 4931 outputs slice additional information (SliceMetaData) which includes information related to slice division and information generated during slice division.
[0349] The location information tile division unit 4932 generates multiple divided location information (multiple tile location information) by dividing multiple slice location information into tiles. The location information tile division unit 4932 also outputs location tile additional information (Geometry Tile MetaData) which includes information related to the tile division of the location information and information generated in the tile division of the location information.
[0350] The attribute information tile division unit 4933 generates multiple divided attribute information (multiple tile attribute information) by dividing multiple slice attribute information into tiles. The attribute information tile division unit 4933 also outputs attribute tile additional information (Attribute Tile MetaData) which includes information related to the tile division of the attribute information and information generated in the tile division of the attribute information.
[0351] The number of slices or tiles to be divided must be one or more. In other words, it is not necessary to divide the slices or tiles.
[0352] Furthermore, while this example shows tile division after slicing, slicing may also be performed after tile division. In addition to slicing and tiling, new division types may be defined, and division may be performed using three or more division types.
[0353] The following describes methods for dividing point cloud data. Figure 48 shows examples of slicing and tiling.
[0354] First, the method of slicing will be explained. The slicing unit 4911 divides the three-dimensional point cloud data into arbitrary point clouds in slice units. In slicing, the slicing unit 4911 does not separate the position information and attribute information that constitute the points, but rather separates the position information and attribute information together. That is, the slicing unit 4911 performs slicing so that the position information and attribute information at any given point belong to the same slice. Note that the number of divisions and the division method can be any method as long as these are followed. Also, the smallest unit of division is a point. For example, the number of divisions for position information and attribute information is the same. For example, the three-dimensional point corresponding to the position information after slicing and the three-dimensional point corresponding to the attribute information are included in the same slice.
[0355] Furthermore, the division unit 4911 generates slice supplemental information, which is additional information relating to the number of divisions and the division method, during slice division. The slice supplemental information is the same for both positional information and attribute information. For example, the slice supplemental information includes information indicating the reference coordinate position, size, or side length of the bounding box after division. The slice supplemental information also includes information indicating the number of divisions and the division type.
[0356] Next, the method of tile division will be described. The division unit 4911 divides the sliced data into slice position information (G slice) and slice attribute information (A slice), and divides the slice position information and the slice attribute information into tile units respectively.
[0357] Note that in FIG. 48, an example of division in a quadtree structure is shown, but the number of divisions and the division method may be any method.
[0358] Also, the division unit 4911 may divide the position information and the attribute information by different division methods, or may divide them by the same division method. Also, the division unit 4911 may divide a plurality of slices into tiles by different division methods, or may divide them into tiles by the same division method.
[0359] Further, the division unit 4911 generates tile addition information related to the number of divisions and the division method at the time of tile division. The tile addition information (position tile addition information and attribute tile addition information) is independent of the position information and the attribute information. For example, the tile addition information includes information indicating the reference coordinate position, size, or side length of the bounding box after division. Also, the tile addition information includes information indicating the number of divisions, the division type, and the like.
[0360] Next, an example of a method for dividing point cloud data into slices or tiles will be described. The division unit 4911 may use a predetermined method as the method of slice or tile division, or may adaptively switch the method used according to the point cloud data.
[0361] At the time of slice division, the division unit 4911 divides the three-dimensional space at once for the position information and the attribute information. For example, the division unit 4911 determines the shape of the object and divides the three-dimensional space into slices according to the shape of the object. For example, the division unit 4911 extracts an object such as a tree or a building and performs division in object units. For example, the division unit 4911 performs slice division so that the whole of one or a plurality of objects is included in one slice. Or, the division unit 4911 divides one object into a plurality of slices.
[0362] In this case, the encoding device may change the encoding method for each slice, for example. For example, the encoding device may use a high-quality compression method for a specific object or a specific part of an object. In this case, the encoding device may store information indicating the encoding method for each slice in additional information (metadata).
[0363] Further, the dividing unit 4911 may perform slice division so that each slice corresponds to a predetermined coordinate space based on map information or position information.
[0364] At the time of tile division, the dividing unit 4911 divides the position information and the attribute information independently. For example, the dividing unit 4911 divides a slice into tiles according to the data amount or the processing amount. For example, the dividing unit 4911 determines whether the data amount of a slice (for example, the number of three-dimensional points included in the slice) is more than a predetermined threshold. When the data amount of the slice is more than the threshold, the dividing unit 4911 divides the slice into tiles. When the data amount of the slice is less than the threshold, the dividing unit 4911 does not divide the slice into tiles.
[0365] For example, the dividing unit 4911 divides a slice into tiles so that the processing amount or the processing time in the decoding device is within a certain range (not more than a predetermined value). Thereby, the processing amount per tile in the decoding device becomes constant, and parallel processing in the decoding device becomes easy.
[0366] Further, when the processing amounts of the position information and the attribute information are different, for example, when the processing amount of the position information is more than the processing amount of the attribute information, the dividing unit 4911 makes the number of divisions of the position information larger than the number of divisions of the attribute information.
[0367] Further, for example, when, depending on the content, in the decoding device, the position information may be decoded and displayed quickly and the attribute information may be decoded and displayed slowly later, the dividing unit 4911 may also make the number of divisions of the position information larger than the number of divisions of the attribute information. Thereby, the decoding device can increase the parallelism of the position information, so that the processing of the position information can be made faster than the processing of the attribute information.
[0368] Furthermore, the decoding device does not necessarily need to process the sliced or tiled data in parallel; it may decide whether or not to process them in parallel depending on the number or capacity of the decoding processing units.
[0369] By dividing the data in the manner described above, adaptive encoding can be achieved according to the content or object. Furthermore, parallel processing can be implemented in the decoding process. This improves the flexibility of the point cloud coding system or point cloud decoding system.
[0370] Figure 49 shows examples of slice and tile division patterns. In the figure, DU stands for Data Unit, representing data for a tile or slice. Each DU also includes a slice index and a tile index. The number in the upper right corner of the DU indicates the slice index, and the number in the lower left corner indicates the tile index.
[0371] In Pattern 1, the number of divisions and the division method are the same for G slices and A slices in slice partitioning. In tile partitioning, the number of divisions and the division method for G slices are different from those for A slices. Also, the same number of divisions and division method are used between multiple G slices. The same number of divisions and division method are used between multiple A slices.
[0372] In Pattern 2, the number of divisions and the division method are the same for G slices and A slices in slice division. In tile division, the number of divisions and the division method for G slices are different from those for A slices. Also, the number of divisions and the division method differ between multiple G slices. The number of divisions and the division method differ between multiple A slices.
[0373] (Embodiment 6) The following describes an example in which slicing is performed after tile division. In autonomous applications such as autonomous driving of vehicles, point cloud data is required not for the entire area, but for the area around the vehicle or the area in the direction of the vehicle's movement. Here, tiles and slices can be used to selectively decode the original point cloud data. By dividing the three-dimensional point cloud data into tiles and then further dividing it into slices, encoding efficiency can be improved or parallel processing can be achieved. When the data is divided, additional information (metadata) is generated, and the generated additional information is sent to the multiplexing unit.
[0374] Figure 50 is a block diagram showing the configuration of a first encoding unit 5010 included in the three-dimensional data encoding device according to this embodiment. The first encoding unit 5010 generates encoded data (encoded stream) by encoding point cloud data using a first encoding method (GPCC (Geometry-based PCC)). This first encoding unit 5010 includes a division unit 5011, a plurality of location information encoding units 5012, a plurality of attribute information encoding units 5013, an additional information encoding unit 5014, and a multiplexing unit 5015.
[0375] The division unit 5011 generates multiple division data by dividing the point cloud data. Specifically, the division unit 5011 generates multiple division data by dividing the space of the point cloud data into multiple subspaces. Here, a subspace is either a tile or a slice, or a combination of a tile and a slice. More specifically, the point cloud data includes location information, attribute information, and additional information. The division unit 5011 divides the location information into multiple division location information and divides the attribute information into multiple division attribute information. The division unit 5011 also generates additional information related to the division.
[0376] For example, the division unit 5011 first divides the point cloud into tiles. Next, the division unit 5011 further divides the obtained tiles into slices.
[0377] Multiple location information encoding units 5012 generate multiple encoded location information by encoding multiple divided location information. For example, multiple location information encoding units 5012 process multiple divided location information in parallel.
[0378] Multiple attribute information encoding units 5013 generate multiple encoded attribute information by encoding multiple divided attribute information. For example, multiple attribute information encoding units 5013 process multiple divided attribute information in parallel.
[0379] The additional information encoding unit 5014 generates encoded additional information by encoding the additional information contained in the point cloud data and the additional information related to data division generated by the division unit 5011 during division.
[0380] The multiplexing unit 5015 generates encoded data (encoded stream) by multiplexing multiple encoded position information, multiple encoded attribute information, and encoded additional information, and transmits the generated encoded data. The encoded additional information is used during decoding.
[0381] In Figure 50, an example is shown where there are two location information encoding units 5012 and two attribute information encoding units 5013. However, the number of location information encoding units 5012 and attribute information encoding units 5013 may be one or three or more. Furthermore, multiple divided data may be processed in parallel within the same chip, such as multiple cores in a CPU, or in parallel across the cores of multiple chips, or in parallel across multiple cores of multiple chips.
[0382] Next, the decoding process will be described. Figure 51 is a block diagram showing the configuration of the first decoding unit 5020. The first decoding unit 5020 restores the point cloud data by decoding the encoded data (encoded stream) generated when the point cloud data is encoded using the first encoding method (GPCC). This first decoding unit 5020 includes a demultiplexing unit 5021, a plurality of location information decoding units 5022, a plurality of attribute information decoding units 5023, an additional information decoding unit 5024, and a coupling unit 5025.
[0383] The demultiplexing unit 5021 generates multiple encoded position information, multiple encoded attribute information, and encoded additional information by demultiplexing the encoded data (encoded stream).
[0384] Multiple location information decoding units 5022 generate multiple segmented location information by decoding multiple encoded location information. For example, multiple location information decoding units 5022 process multiple encoded location information in parallel.
[0385] The multiple attribute information decoding unit 5023 generates multiple segmented attribute information by decoding multiple encoded attribute information. For example, the multiple attribute information decoding unit 5023 processes multiple encoded attribute information in parallel.
[0386] Multiple additional information decoding units 5024 generate additional information by decoding encoded additional information.
[0387] The merging unit 5025 generates position information by combining multiple division position information using additional information. The merging unit 5025 generates attribute information by combining multiple division attribute information using additional information. For example, the merging unit 5025 first generates point cloud data corresponding to tiles by combining decoded point cloud data for slices using slice additional information. Next, the merging unit 5025 restores the original point cloud data by combining point cloud data corresponding to tiles using tile additional information.
[0388] In Figure 50, an example is shown where there are two location information decoding units 5022 and two attribute information decoding units 5023. However, the number of location information decoding units 5022 and attribute information decoding units 5023 may be one or three or more. Furthermore, multiple divided data may be processed in parallel within the same chip, such as multiple cores in a CPU, or in parallel across the cores of multiple chips, or across multiple cores of multiple chips.
[0389] Next, we will explain how to segment point cloud data. In autonomous applications such as self-driving vehicles, point cloud data is not needed for the entire area, but rather for the area around the vehicle or the area in the direction of the vehicle's movement.
[0390] Figure 52 shows an example of tile shapes. As shown in Figure 52, various shapes such as circles, rectangles, or ellipses may be used as tile shapes.
[0391] FIG. 53 is a diagram showing examples of tiles and slices. The configuration of the slices may be different between the tiles. For example, the configuration or the configuration of the tiles or slices may be optimized based on the data volume. Or the configuration of the tiles or slices may be optimized based on the decoding speed.
[0392] Also, tile division may be performed based on position information. In this case, the attribute information is divided in the same way as the corresponding position information.
[0393] Also, in the slice division after tile division, the position information and the attribute information may be divided into slices by different methods. For example, the method of slice division in each tile may be selected according to a request from an application. Based on the request from the application, different methods of slice division or tile division may be used.
[0394] For example, the dividing unit 5011 divides the three-dimensional point cloud data into one or more tiles based on position information such as map information in the two-dimensional shape seen from above. Then, the dividing unit 5011 divides each tile into one or more slices.
[0395] Note that the dividing unit 5011 may divide the position information (Geometry) and the attribute information (Attribute) into slices in the same way.
[0396] Note that the position information and the attribute information may each be of one type or two or more types. Also, in the case of point cloud data without attribute information, the attribute information may not be present.
[0397] FIG. 54 is a block diagram of the dividing unit 5011. The dividing unit 5011 includes a tile dividing unit 5031 (Tile Divider), a position information slice dividing unit 5032 (Geometry Slice Divider), and an attribute information slice dividing unit 5033 (Attribute Slice Divider).
[0398] The tile division unit 5031 generates multiple tile position information by dividing position information (Position (Geometry)) into tiles. The tile division unit 5031 also generates multiple tile attribute information by dividing attribute information (Attribute) into tiles. Furthermore, the tile division unit 5031 outputs tile additional information (TileMetaData) which includes information related to tile division and information generated during tile division.
[0399] The location information slice division unit 5032 generates multiple divided location information (multiple slice location information) by dividing multiple tile location information into slices. The location information slice division unit 5032 also outputs location slice additional information (Geometry Slice MetaData) which includes information related to the slice division of the location information and information generated in the slice division of the location information.
[0400] The attribute information slice division unit 5033 generates multiple divided attribute information (multiple slice attribute information) by dividing multiple tile attribute information into slices. The attribute information slice division unit 5033 also outputs attribute slice additional information (Attribute Slice MetaData) which includes information related to the slice division of attribute information and information generated in the slice division of attribute information.
[0401] Next, we will describe examples of tile shapes. The entire three-dimensional map is divided into multiple tiles. The data from multiple tiles is selectively sent to the three-dimensional data decoding device. Alternatively, the data from multiple tiles is sent to the three-dimensional data decoding device in order of importance. Depending on the situation, the tile shape may be selected from multiple shapes.
[0402] Figure 55 shows an example of a map viewed from above using point cloud data obtained by LiDAR. The example shown in Figure 55 is point cloud data of a highway, including overpasses.
[0403] Figure 56 shows an example of dividing the point cloud data shown in Figure 55 into square tiles. Such division into squares can be easily performed on the map server. In addition, the tile height is set low for normal roads. In the case of overpasses, the tile height is set higher than for normal roads so that the tiles cover the overpass.
[0404] Figure 57 shows an example of dividing the point cloud data shown in Figure 55 into circular tiles. In this case, adjacent tiles may overlap in a plan view. The three-dimensional data encoding device transmits point cloud data of a cylindrical area (a circle in a top view) around the vehicle to the vehicle when the vehicle requires point cloud data of the surrounding area.
[0405] Also, similar to the example in Figure 56, the tile height is set lower for normal roads. In overpasses, the tile height is set higher than for normal roads so that the tiles cover the overpass.
[0406] The 3D data encoding device may change the height of the tiles according to, for example, the shape or height of a road or building. Alternatively, the 3D data encoding device may change the height of the tiles according to location information or area information. Furthermore, the 3D data encoding device may change the height of each tile individually. Or, the 3D data encoding device may change the height of each tile in sections containing multiple tiles. In other words, the 3D data encoding device may make the height of multiple tiles within a section the same. Also, tiles of different heights may overlap in a top view.
[0407] Figure 58 shows examples of tile divisions using tiles of various shapes, sizes, or heights. The tile shapes can be any shape, any size, or any combination thereof.
[0408] For example, in addition to the examples of dividing with square tiles without overlapping and dividing with overlapping circular tiles as described above, the three-dimensional data encoding device may also divide with overlapping square tiles. Furthermore, the shape of the tiles does not have to be square or circular; polygons with three or more vertices may be used, or shapes without vertices may be used.
[0409] Furthermore, there may be two or more types of tile shapes, and tiles of different shapes may overlap. Also, there may be one or more types of tile shapes, and within the same shape being divided, shapes of different sizes may be combined, and these may overlap.
[0410] For example, in areas without objects such as roads, larger tiles are used than in areas where objects exist. Furthermore, the 3D data encoding device may adaptively change the shape or size of the tiles depending on the objects.
[0411] Furthermore, for example, a three-dimensional data encoding device may set the tiles in the direction of travel to a larger size because it is highly likely that it will need to read tiles far ahead of the vehicle, which is the direction of travel of the vehicle. Conversely, since it is unlikely that the vehicle will move to the side, the tiles to the side may be set to a smaller size than the tiles in the direction of travel.
[0412] Figure 59 shows an example of tile data stored on the server. For example, point cloud data is pre-divided into tiles and encoded, and the resulting encoded data is stored on the server. The user retrieves the desired tile data from the server when needed. Alternatively, the server (three-dimensional data encoding device) may perform tile division and encoding to include the data desired by the user, according to the user's instructions.
[0413] For example, if the moving object (vehicle) is moving at a high speed, it is conceivable that a wider range of point cloud data will be required. Therefore, the server may determine the shape and size of the tiles and perform tiling based on the pre-estimated speed of the vehicle (e.g., the legal speed limit on the road, the speed of the vehicle that can be estimated from the width and shape of the road, or statistical speed). Alternatively, as shown in Figure 59, the server may pre-encode tiles of multiple shapes or sizes and store the obtained data. The moving object may acquire tile data of an appropriate shape and size according to the direction and speed of the moving object.
[0414] Figure 60 shows an example of a system for tile division. As shown in Figure 60, the shape and area of the tiles may be determined based on the position of an antenna (base station), which is a communication means for transmitting point cloud data, or the communication area supported by the antenna. Alternatively, if point cloud data is generated by a sensor such as a camera, the shape and area of the tiles may be determined based on the position of the sensor or the target range (detection range) of the sensor.
[0415] One tile may be assigned to one antenna or sensor, or one tile may be assigned to multiple antennas or sensors. Multiple tiles may be assigned to one antenna or sensor. The antenna or sensor may be fixed or movable.
[0416] For example, encoded data divided into tiles may be managed by a server connected to an antenna or sensor for the area assigned to each tile. The server may manage the encoded data for its own area and the tile information for adjacent areas. Multiple encoded data for multiple tiles may be managed in a centralized management server (cloud) that manages multiple servers corresponding to each tile. Alternatively, there may be no servers corresponding to each tile, and the antenna or sensor may be directly connected to the centralized management server.
[0417] The target range of the antenna or sensor may vary depending on the radio wave power, equipment differences, and installation conditions, and the shape and size of the tiles may also be changed accordingly. Based on the target range of the antenna or sensor, slices or PCC frames may be assigned instead of tiles.
[0418] Next, we will explain a technique for dividing tiles into slices. By assigning similar objects to the same slice, encoding efficiency can be improved.
[0419] For example, a three-dimensional data encoding device may use the features of the point cloud data to recognize objects (roads, buildings, trees, etc.) and perform slicing by clustering the point cloud for each object.
[0420] Alternatively, the three-dimensional data encoding device may perform slicing by grouping objects with the same attributes and assigning slices to each group. Here, attributes refer to information about movement, for example, and objects are grouped by classifying them into dynamic information such as pedestrians and cars, semi-dynamic information such as accidents and traffic jams, semi-static information such as traffic regulations and road construction, and static information such as road surfaces and structures.
[0421] Note that data may be duplicated across multiple slices. For example, when slicing into multiple object groups, any object may belong to one object group or to two or more object groups.
[0422] Figure 61 shows an example of this slicing method. For example, in the example shown in Figure 61, the tile is a rectangular prism. However, the tile may be cylindrical or of other shapes.
[0423] The point cloud contained in a tile is grouped into object groups, such as roads, buildings, and trees. Then, the tile is sliced so that each object group is contained in a single slice. Each slice is then encoded individually.
[0424] Next, the method for encoding the divided data will be described. The three-dimensional data encoding device (first encoding unit 5010) encodes each of the divided data. When encoding attribute information, the three-dimensional data encoding device generates dependency information as additional information that indicates which configuration information (location information, additional information, or other attribute information) was used as the basis for encoding. In other words, the dependency information indicates, for example, the configuration information of the reference (dependent). In this case, the three-dimensional data encoding device generates the dependency information based on the configuration information corresponding to the division shape of the attribute information. Note that the three-dimensional data encoding device may generate dependency information based on configuration information corresponding to multiple division shapes.
[0425] Dependency information may be generated by a three-dimensional data encoding device and sent to a three-dimensional data decoding device. Alternatively, the three-dimensional data decoding device may generate the dependency information, and the three-dimensional data encoding device may not send it. Furthermore, the dependencies used by the three-dimensional data encoding device may be predetermined, and the three-dimensional data encoding device may not send the dependency information.
[0426] Figure 62 shows an example of the dependencies between data. In the figure, the tip of the arrow indicates the dependent data, and the base of the arrow indicates the dependent data. The three-dimensional data decoding device decodes the data in the order of dependent data to dependent data. Also, in the figure, data shown with solid lines are data that is actually transmitted, and data shown with dotted lines are data that is not transmitted.
[0427] In the same figure, G indicates location information and A indicates attribute information. Gt1 indicates location information for tile number 1, and Gt2 indicates location information for tile number 2. Gt1s1 indicates location information for tile number 1 and slice number 1, Gt1s2 indicates location information for tile number 1 and slice number 2, Gt2s1 indicates location information for tile number 2 and slice number 1, and Gt2s2 indicates location information for tile number 2 and slice number 2. Similarly, At1 indicates attribute information for tile number 1, and At2 indicates attribute information for tile number 2. At1s1 indicates attribute information for tile number 1 and slice number 1, At1s2 indicates attribute information for tile number 1 and slice number 2, At2s1 indicates attribute information for tile number 2 and slice number 1, and At2s2 indicates attribute information for tile number 2 and slice number 2.
[0428] Mtile indicates tile addition information, MGslice indicates position slice addition information, and MAslice indicates attribute slice addition information. Dt1s1 indicates dependency information of attribute information At1s1, and Dt2s1 indicates dependency information of attribute information At2s1.
[0429] Depending on the application, different tile or slice partitioning structures may be used.
[0430] Furthermore, the three-dimensional data encoding device may rearrange the data in the order of decoding so that the three-dimensional data decoding device does not need to rearrange the data. Alternatively, the data may be rearranged in the three-dimensional data decoding device, or both the three-dimensional data encoding device and the three-dimensional data decoding device may rearrange the data.
[0431] Figure 63 shows an example of the data decoding order. In the example in Figure 63, decoding is performed sequentially from left to right. The 3D data decoding device decodes dependent data first among dependent data. For example, the 3D data encoding device pre-arranges and sends the data in this order. Any order is acceptable as long as the dependent data comes first. The 3D data encoding device may also send additional information and dependency information before the data.
[0432] Furthermore, the three-dimensional data decoding device may selectively decode tiles based on requests from the application and information obtained from the NAL unit header. Figure 64 shows an example of encoded tile data. For example, the decoding order of tiles is arbitrary; that is, there may be no dependencies between tiles.
[0433] Next, the configuration of the coupling unit 5025 included in the first decoding unit 5020 will be described. Figure 65 is a block diagram showing the configuration of the coupling unit 5025. The coupling unit 5025 includes a geometry slice combiner 5041, an attribute slice combiner 5042, and a tile combiner.
[0434] The position information slice joining unit 5041 generates multiple tile position information by joining multiple divided position information using position slice additional information. The attribute information slice joining unit 5042 generates multiple tile attribute information by joining multiple divided attribute information using attribute slice additional information.
[0435] The tile joining unit 5043 generates position information by combining multiple tile position information using tile addition information. The tile joining unit 5043 also generates attribute information by combining multiple tile attribute information using tile addition information.
[0436] The number of slices or tiles to be divided must be one or more. In other words, the slices or tiles do not need to be divided at all.
[0437] Next, the structure of the sliced or tiled encoded data and the method of storing the encoded data in the NAL unit (multiplexing method) will be described. Figure 66 is a diagram showing the structure of the encoded data and the method of storing the encoded data in the NAL unit.
[0438] The encoded data (splitting position information and splitting attribute information) is stored in the NAL unit's payload.
[0439] Encoded data includes a header and a payload. The header includes identification information to identify the data contained in the payload. This identification information includes, for example, the type of slice or tile division (slice_type, tile_type), index information to identify the slice or tile (slice_idx, tile_idx), location information of the data (slice or tile), or the address of the data (address). Index information to identify a slice is also written as slice index (SliceIndex). Index information to identify a tile is also written as tile index (TileIndex). The type of division can be, for example, a method based on the object shape as described above, a method based on map information or location information, or a method based on the amount of data or processing amount.
[0440] Furthermore, the header of the encoded data includes identification information indicating dependencies. In other words, if there are dependencies between data, the header includes identification information for referencing the dependent data from the dependent data source. For example, the header of the dependent data includes identification information to identify that data. The header of the dependent data includes identification information indicating the dependent data. Note that if the identification information for identifying the data, additional information related to slicing or tiling, and identification information indicating dependencies can be identified or derived from other information, this information may be omitted.
[0441] Next, the flow of the point cloud data encoding and decoding processes according to this embodiment will be described. Figure 67 is a flowchart of the point cloud data encoding process according to this embodiment.
[0442] First, the three-dimensional data encoding device determines the division method to be used (S5011). This division method includes whether or not to perform tile division or slice division. The division method may also include the number of divisions if tile division or slice division is performed, and the type of division. The type of division refers to methods based on object shape, methods based on map information or location information, or methods based on data volume or processing volume, as described above. The division method may also be predetermined.
[0443] If tile division is performed (Yes in S5012), the three-dimensional data encoding device generates multiple tile position information and multiple tile attribute information by dividing the position information and attribute information together (S5013). The three-dimensional data encoding device also generates tile addition information related to tile division. The three-dimensional data encoding device may divide the position information and attribute information independently.
[0444] If slicing is performed (Yes in S5014), the three-dimensional data encoding device generates multiple slicing position information and multiple slicing attribute information by independently slicing multiple tile position information and multiple tile attribute information (or position information and attribute information) (S5015). The three-dimensional data encoding device also generates position slice addition information and attribute slice addition information related to the slicing. The three-dimensional data encoding device may also slice the tile position information and tile attribute information together.
[0445] Next, the three-dimensional data encoding device generates multiple encoded location information and multiple encoded attribute information by encoding each of the multiple division location information and multiple division attribute information (S5016). The three-dimensional data encoding device also generates dependency information.
[0446] Next, the three-dimensional data encoding device generates encoded data (encoded stream) by NAL unitizing (multiplexing) multiple encoded position information, multiple encoded attribute information, and additional information (S5017). The three-dimensional data encoding device also transmits the generated encoded data.
[0447] Figure 68 is a flowchart of the point cloud data decoding process according to this embodiment. First, the three-dimensional data decoding device determines the division method by analyzing the additional information related to the division method (tile additional information, position slice additional information, and attribute slice additional information) included in the encoded data (encoded stream) (S5021). This division method includes whether or not to perform tile division and whether or not to perform slice division. The division method may also include the number of divisions and the type of division when tile division or slice division is performed.
[0448] Next, the three-dimensional data decoding device generates partitioning location information and partitioning attribute information by decoding multiple encoded position information and multiple encoded attribute information contained in the encoded data using dependency information contained in the encoded data (S5022).
[0449] If the additional information indicates that slice division has been performed (Yes in S5023), the three-dimensional data decoding device generates multiple tile position information and multiple tile attribute information by combining multiple division position information and multiple division attribute information in their respective methods based on the position slice additional information and attribute slice additional information (S5024). The three-dimensional data decoding device may combine multiple division position information and multiple division attribute information in the same method.
[0450] If the additional information indicates that tile division has been performed (Yes in S5025), the three-dimensional data decoding device generates position information and attribute information by combining multiple tile position information and multiple tile attribute information (multiple division position information and multiple division attribute information) in the same way based on the tile additional information (S5026). The three-dimensional data decoding device may combine the multiple tile position information and multiple tile attribute information in different ways.
[0451] Next, we will explain the tile appending information. The three-dimensional data encoding device generates tile appending information, which is metadata about the tile division method, and transmits the generated tile appending information to the three-dimensional data decoding device.
[0452] Figure 69 shows an example of the syntax for tile data (TileMetaData). As shown in Figure 69, for example, tile data includes division method information (type_of_divide), shape information (topview_shape), overlap flag (tile_overlap_flag), overlap information (type_of_overlap), height information (tile_height), number of tiles (tile_number), and tile position information (global_position, relative_position).
[0453] The division method information (type_of_divide) indicates how the tiles are divided. For example, the division method information indicates whether the tiles are divided based on map information, i.e., based on a top view (top_view), or otherwise (other).
[0454] Shape information (topview_shape) is included in the tile information when, for example, the tile division method is based on a top view. Shape information indicates the shape of the tile when viewed from above. For example, this shape includes squares and circles. This shape may also include polygons other than ellipses, rectangles, or quadrilaterals, or other shapes. Furthermore, shape information is not limited to the shape of the tile when viewed from above, but may also indicate the three-dimensional shape of the tile (for example, cubes and cylinders).
[0455] The overlap flag (tile_overlap_flag) indicates whether tiles overlap or not. For example, the overlap flag is included in the tile information when the tile division method is based on a top view. In this case, the overlap flag indicates whether tiles overlap in a top view. The overlap flag may also indicate whether tiles overlap in three-dimensional space.
[0456] Duplicate information (type_of_overlap) is included in the tile information when tiles overlap, for example. Duplicate information indicates how the tiles overlap, etc. For example, it indicates the size of the overlapping area.
[0457] The height information (tile_height) indicates the height of the tile. The height information may also include information indicating the shape of the tile. For example, if the shape of the tile when viewed from above is rectangular, this information may indicate the lengths of the sides (vertical and horizontal lengths) of that rectangle. Alternatively, if the shape of the tile when viewed from above is circular, this information may indicate the diameter or radius of that circle.
[0458] Furthermore, the height information may indicate the height of each tile, or it may indicate a common height for multiple tiles. Alternatively, multiple height types for roads and overpasses may be predefined, and the height information may indicate the height of each height type and the height type of each tile. Or, the height of each height type may be predefined, and the height information may indicate the height type of each tile. In other words, the height of each height type does not necessarily have to be indicated by the height information.
[0459] The tile number (tile_number) indicates the number of tiles. Note that tile information may also include information indicating the spacing between tiles.
[0460] Tile position information (global_position, relative_position) is information used to identify the position of each tile. For example, tile position information indicates the absolute or relative coordinates of each tile.
[0461] Some or all of the above information may be provided for each tile, or for multiple tiles (for example, for each frame or for multiple frames).
[0462] The three-dimensional data encoding device may include the tile addition information in the SEI (Supplemental Enhancement Information) and send it. Alternatively, the three-dimensional data encoding device may store the tile addition information in an existing parameter set (PPS, GPS, or APS, etc.) and send it.
[0463] For example, if the tile information changes from frame to frame, the tile information may be stored in a parameter set for each frame (such as GPS or APS). If the tile information does not change within a sequence, the tile information may be stored in a parameter set for each sequence (such as location SPS or attribute SPS). Furthermore, if the same tile division information is used for both location information and attribute information, the tile information may be stored in the parameter set of the PCC stream (stream PS).
[0464] Furthermore, tile information may be stored in any of the parameter sets described above, or in multiple parameter sets. Additionally, tile information may be stored in the header of the encoded data. Furthermore, tile information may be stored in the header of the NAL unit.
[0465] Furthermore, all or part of the tile addition information may be stored in one of the headers of the division location information and the division attribute information, but not in the other. For example, if the same tile addition information is used for both location information and attribute information, the tile addition information may be included in one of the headers of the location information or attribute information. For example, if attribute information depends on location information, the location information is processed first. Therefore, the header of the location information may contain this tile addition information, while the header of the attribute information may not. In this case, the three-dimensional data decoding device will determine, for example, that the attribute information of the dependency belongs to the same tile as the tile of the location information to which it depends.
[0466] The 3D data decoding device reconstructs the tiled point cloud data based on the tile information. If there is duplicate point cloud data, the 3D data decoding device identifies the multiple duplicate point cloud data, selects one, or merges the multiple point cloud data.
[0467] Furthermore, the three-dimensional data decoding device may perform decoding using tile-added information. For example, if multiple tiles overlap, the three-dimensional data decoding device may decode each tile, perform processing using the decoded data (e.g., smoothing or filtering), and generate point cloud data. This may enable highly accurate decoding.
[0468] Figure 70 shows an example of a system configuration including a three-dimensional data encoding device and a three-dimensional data decoding device. The tile division unit 5051 divides point cloud data, including position information and attribute information, into first tiles and second tiles. The tile division unit 5051 also sends tile addition information related to tile division to the decoding unit 5053 and the tile joining unit 5054.
[0469] The encoding unit 5052 generates encoded data by encoding the first tile and the second tile.
[0470] The decoding unit 5053 reconstructs the first and second tiles by decoding the encoded data generated by the encoding unit 5052. The tile joining unit 5054 reconstructs the point cloud data (position information and attribute information) by joining the first and second tiles using the tile addition information.
[0471] Next, we will explain slice appending information. The three-dimensional data encoding device generates slice appending information, which is metadata about the slice division method, and transmits the generated slice appending information to the three-dimensional data decoding device.
[0472] Figure 71 shows an example of the syntax for slice data (SliceMetaData). As shown in Figure 71, for example, slice data includes division method information (type_of_divide), duplicate flag (slice_overlap_flag), duplicate information (type_of_overlap), slice number (slice_number), slice position information (global_position, relative_position), and slice size information (slice_bounding_box_size).
[0473] The division method information (type_of_divide) indicates the method of dividing the slice. For example, the division method information indicates whether the slice division method is based on object information (object) as shown in Figure 61. Note that the slice addition information may also include information indicating the method of object division. For example, this information indicates whether one object is divided into multiple slices or assigned to one slice. This information may also indicate the number of divisions when one object is divided into multiple slices.
[0474] The duplicate flag (slice_overlap_flag) indicates whether or not the slices are duplicates. Duplicate information (type_of_overlap) is included in the slice append information, for example, when the slices are duplicates. Duplicate information indicates how the slices are duplicated, etc. For example, it indicates the size of the overlapping area.
[0475] The slice number (slice_number) indicates the number of slices.
[0476] Slice position information (global_position, relative_position) and slice size information (slice_bounding_box_size) are information about the region of the slice. Slice position information is information used to identify the position of each slice. For example, slice position information indicates the absolute or relative coordinates of each slice. Slice size information (slice_bounding_box_size) indicates the size of each slice. For example, slice size information indicates the size of the bounding box of each slice.
[0477] The three-dimensional data encoding device may include slice addition information in the SEI and send it out. Alternatively, the three-dimensional data encoding device may store the slice addition information in an existing parameter set (PPS, GPS, or APS, etc.) and send it out.
[0478] For example, if slice addition information changes from frame to frame, the slice addition information may be stored in a parameter set for each frame (such as GPS or APS). If the slice addition information does not change within a sequence, the slice addition information may be stored in a parameter set for each sequence (location SPS or attribute SPS). Furthermore, if the same slice division information is used for both location information and attribute information, the slice addition information may be stored in the parameter set of the PCC stream (stream PS).
[0479] Furthermore, slice addition information may be stored in any of the parameter sets described above, or in multiple parameter sets. Also, slice addition information may be stored in the header of the encoded data. Furthermore, slice addition information may be stored in the header of the NAL unit.
[0480] Furthermore, all or part of the slice addition information may be stored in one of the headers of the division location information and the division attribute information, but not in the other. For example, if the same slice addition information is used for both location information and attribute information, the slice addition information may be included in one of the headers of the location information or attribute information. For example, if attribute information depends on location information, the location information is processed first. Therefore, the header of the location information may contain this slice addition information, while the header of the attribute information may not. In this case, the three-dimensional data decoding device will determine, for example, that the attribute information that depends on the location information belongs to the same slice as the slice of the location information it depends on.
[0481] The 3D data decoding device reconstructs the sliced point cloud data based on the slice addition information. If there is duplicate point cloud data, the 3D data decoding device identifies the multiple duplicate point cloud data, selects one, or merges the multiple point cloud data.
[0482] Furthermore, the three-dimensional data decoding device may perform decoding using slice-added information. For example, if multiple slices overlap, the three-dimensional data decoding device may decode each slice, perform processing (e.g., smoothing or filtering) using the decoded data, and generate point cloud data. This may enable highly accurate decoding.
[0483] Figure 72 is a flowchart of the three-dimensional data encoding process, including the generation of tile-added information, using the three-dimensional data encoding device according to this embodiment.
[0484] First, the three-dimensional data encoding device determines the method of dividing the tiles (S5031). Specifically, the three-dimensional data encoding device determines whether to use a division method based on a top view (top_view) or another method (other) for dividing the tiles. The three-dimensional data encoding device also determines the shape of the tiles when using the division method based on a top view. Furthermore, the three-dimensional data encoding device determines whether or not a tile overlaps with other tiles.
[0485] If the tile division method determined in step S5031 is a division method based on a top view (Yes in S5032), the three-dimensional data encoding device records in the tile addition information that the tile division method is a division method based on a top view (top_view) (S5033).
[0486] On the other hand, if the tile division method determined in step S5031 is other than the division method based on the top view (No in S5032), the three-dimensional data encoding device indicates in the tile addition information that the tile division method is other than the division method based on the top view (top_view) (S5034).
[0487] Furthermore, if the shape of the tile viewed from above, as determined in step S5031, is a square (square in S5035), the three-dimensional data encoding device records that the shape of the tile viewed from above is a square in the tile addition information (S5036). On the other hand, if the shape of the tile viewed from above, as determined in step S5031, is a circle (circle in S5035), the three-dimensional data encoding device records that the shape of the tile viewed from above is a circle in the tile addition information (S5037).
[0488] Next, the three-dimensional data encoding device determines whether a tile overlaps with another tile (S5038). If a tile overlaps with another tile (Yes in S5038), the three-dimensional data encoding device records that the tile overlaps in the tile addition information (S5039). On the other hand, if a tile does not overlap with another tile (No in S5038), the three-dimensional data encoding device records that the tile does not overlap in the tile addition information (S5040).
[0489] Next, the three-dimensional data encoding device divides the tiles based on the tile division method determined in step S5031, encodes each tile, and sends out the generated encoded data and tile-additional information (S5041).
[0490] Figure 73 is a flowchart of the three-dimensional data decoding process using tile-added information by the three-dimensional data decoding device according to this embodiment.
[0491] First, the three-dimensional data decoding device analyzes the tile addition information contained in the bitstream (S5051).
[0492] If the tile information indicates that a tile does not overlap with other tiles (No in S5052), the 3D data decoding device generates point cloud data for each tile by decoding each tile (S5053). Next, the 3D data decoding device reconstructs point cloud data from the point cloud data of each tile based on the tile division method and tile shape indicated in the tile information (S5054).
[0493] On the other hand, if the tile addition information indicates that a tile overlaps with other tiles (Yes in S5052), the 3D data decoding device generates point cloud data for each tile by decoding each tile. The 3D data decoding device also identifies the overlapping portion of the tiles based on the tile addition information (S5055). The 3D data decoding device may use multiple pieces of overlapping information to perform the decoding process for the overlapping portion. Next, the 3D data decoding device reconstructs point cloud data from the point cloud data of each tile based on the tile division method, tile shape, and overlapping information indicated in the tile addition information (S5056).
[0494] The following describes variations related to slicing. The three-dimensional data encoding device may transmit additional information indicating the type of object (road, building, tree, etc.) or attributes (dynamic information, static information, etc.). Alternatively, encoding parameters may be predetermined according to the object, and the three-dimensional data encoding device may notify the three-dimensional data decoding device of the encoding parameters by sending the type of object or attributes.
[0495] The following methods may be used for the encoding order and transmission order of slice data. For example, the 3D data encoding device may encode slice data in order from data that is easy to recognize or cluster. Alternatively, the 3D data encoding device may encode slice data in order from slice data that has been clustered first. The 3D data encoding device may also transmit the encoded slice data in order. Alternatively, the 3D data encoding device may transmit slice data in order of the decoding priority in the application. For example, if the decoding priority of dynamic information is high, the 3D data encoding device may transmit slice data in order from slices grouped by dynamic information.
[0496] Furthermore, if the order of encoded data differs from the order of decoding priority, the three-dimensional data encoding device may rearrange the encoded data before sending it out. Also, when storing encoded data, the three-dimensional data encoding device may rearrange the encoded data before storing it.
[0497] The application (3D data decoding device) requests the server (3D data encoding device) to send slices containing the desired data. The server sends the slice data required by the application, and does not need to send unnecessary slice data.
[0498] The application requests the server to send tiles containing the desired data. The server sends the tile data that the application needs, and does not need to send any tile data that is not needed.
[0499] As described above, the three-dimensional data encoding device according to this embodiment performs the processing shown in Figure 74. First, the three-dimensional data encoding device generates multiple encoded data by encoding multiple subspaces (e.g., tiles) obtained by dividing the target space containing multiple three-dimensional points (S5061). The three-dimensional data encoding device generates a bitstream that includes the multiple encoded data and first information (e.g., topview_shape) indicating the shape of the multiple subspaces (S5062).
[0500] According to this, a three-dimensional data encoding device can improve encoding efficiency because it can select any shape from multiple types of subspace shapes.
[0501] For example, the shape is the two-dimensional or three-dimensional shape of the plurality of subspaces. For example, the shape is the shape of the plurality of subspaces viewed from above. In other words, the first information indicates the shape of the subspaces viewed from a specific direction (e.g., from above). To put it another way, the first information indicates the shape of the subspaces viewed from above. For example, the shape is a rectangle or a circle.
[0502] For example, the bitstream includes second information (e.g., tile_overlap_flag) indicating whether the multiple sub-sections overlap or not.
[0503] According to this, a three-dimensional data encoding device can duplicate subspaces, thus enabling the generation of subspaces without complicating their shapes.
[0504] For example, the bitstream includes third information (e.g., type_of_divide) indicating whether the method of dividing the plurality of sub-sections is a top-view method.
[0505] For example, the bitstream includes a fourth piece of information (e.g., tile_height) indicating at least one of the height, width, depth, and radius of the plurality of subsections.
[0506] For example, the bitstream includes a fifth piece of information (e.g., global_position or relative_position) indicating the position of each of the plurality of sub-sections.
[0507] For example, the bitstream includes a sixth piece of information (e.g., tile_number) indicating the number of sub-sections.
[0508] For example, the bitstream includes a seventh piece of information indicating the interval between the plurality of sub-intervals.
[0509] For example, a three-dimensional data encoding device comprises a processor and memory, and the processor uses the memory to perform the above processing.
[0510] Furthermore, the three-dimensional data decoding device according to this embodiment performs the processing shown in Figure 75. First, the three-dimensional data decoding device reconstructs the multiple subspaces by decoding multiple encoded data generated by encoding multiple subspaces (e.g., tiles) which are obtained by dividing the target space containing multiple three-dimensional points included in the bitstream (S5071). The three-dimensional data decoding device reconstructs the target space by combining the multiple subspaces using first information (e.g., topview_shape) which indicates the shape of the multiple subspaces included in the bitstream (S5072). For example, the three-dimensional data decoding device can grasp the position and range of each subspace within the target space by recognizing the shape of the multiple subspaces using the first information. The three-dimensional data decoding device can combine the multiple subspaces based on the grasped position and range of the multiple subspaces. As a result, the three-dimensional data decoding device can correctly combine the multiple subspaces.
[0511] For example, the shape is the shape in two dimensions or three dimensions of the plurality of subspaces. For example, the shape is a rectangle or a circle.
[0512] For example, the bitstream includes second information (e.g., tile_overlap_flag) indicating whether the multiple sub-sections overlap. In restoring the target space, the three-dimensional data decoding device further uses the second information to combine the multiple sub-spaces. For example, the three-dimensional data decoding device uses the second information to determine whether the sub-spaces overlap. If the sub-spaces overlap, the three-dimensional data decoding device identifies the overlapping region and performs a predetermined action on the identified overlapping region.
[0513] For example, the bitstream includes third information (e.g., type_of_divide) indicating whether the method of dividing the plurality of sub-sections is a top-view method. If the third information indicates that the method of dividing the plurality of sub-sections is a top-view method, the three-dimensional data decoder combines the plurality of sub-spaces using the first information.
[0514] For example, the bitstream includes fourth information (e.g., tile_height) indicating at least one of the height, width, depth, and radius of the plurality of sub-sections. In reconstructing the target space, the three-dimensional data decoder further uses the fourth information to combine the plurality of sub-spaces. For example, by using the fourth information to recognize the height of the plurality of sub-spaces, the three-dimensional data decoder can grasp the position and range of each sub-space within the target space. Based on the grasped position and range of the plurality of sub-spaces, the three-dimensional data decoder can combine the plurality of sub-spaces.
[0515] For example, the bitstream includes a fifth piece of information (e.g., global_position or relative_position) indicating the position of each of the multiple sub-sections. In reconstructing the target space, the three-dimensional data decoder further uses the fifth piece of information to combine the multiple sub-spaces. For example, the three-dimensional data decoder can use the fifth piece of information to recognize the positions of the multiple sub-spaces and thereby grasp the position of each sub-space within the target space. Based on the grasped positions of the multiple sub-spaces, the three-dimensional data decoder can combine the multiple sub-spaces.
[0516] For example, the bitstream includes a sixth piece of information (e.g., tile_number) indicating the number of sub-sections. The three-dimensional data decoder further uses the sixth piece of information to combine the sub-spaces in the reconstruction of the target space.
[0517] For example, the bitstream includes seventh information indicating the intervals between the multiple sub-intervals. In reconstructing the target space, the three-dimensional data decoder further uses the seventh information to combine the multiple sub-spaces. For example, by using the seventh information to recognize the intervals between the multiple sub-spaces, the three-dimensional data decoder can grasp the position and range of each sub-space within the target space. Based on the grasped positions and ranges of the multiple sub-spaces, the three-dimensional data decoder can combine the multiple sub-spaces.
[0518] For example, a three-dimensional data decoding device comprises a processor and memory, and the processor uses the memory to perform the above processing.
[0519] (Embodiment 7) To divide point cloud data into tiles and slices and efficiently encode or decode the divided data, it is necessary to control the encoding and decoding sides appropriately. By making the encoding and decoding of the divided data independent of the divided data, the divided data can be processed in parallel on each thread / core using a multithreaded or multicore processor, improving performance.
[0520] There are various methods for dividing point cloud data into tiles and slices. For example, some methods divide the data based on the attributes of the object in the point cloud data, such as a road surface, or on the characteristics of the point cloud data, such as color information, such as green.
[0521] CABAC stands for Context-Based Adaptive Binary Arithmetic Coding. It is an encoding method that improves the accuracy of probabilities by sequentially updating the context (a model that estimates the probability of occurrence of input binary symbols) based on encoded information, thereby achieving highly compressible arithmetic coding (entropy coding).
[0522] To process partitioned data, such as tiles or slices, in parallel, each partitioned data must be able to be encoded or decoded independently. However, to keep the CABAC independent between partitioned data, the CABAC needs to be initialized at the beginning of each partitioned data during encoding and decoding, but there is no mechanism to do this.
[0523] The CABACABAC initialization flag is used to initialize CABAC during CABAC coding and decoding.
[0524] Figure 76 is a flowchart showing the process of initializing CABACCABAC in coding or decoding, according to the CABAC initialization flag.
[0525] The three-dimensional data encoding device or the three-dimensional data decoding device determines whether the CABAC initialization flag is 1 during encoding or decoding (S5201).
[0526] If the CABAC initialization flag is 1 (Yes in S5201), the three-dimensional data encoding device or three-dimensional data decoding device initializes the CABAC encoding / decoding unit to the default state (S5202) and continues encoding or decoding.
[0527] If the CABAC initialization flag is not 1 (No in S5201), the three-dimensional data encoding device or three-dimensional data decoding device continues encoding or decoding without initialization.
[0528] In other words, when initializing CABAC, CABAC_init_flag=1 is set, and the encoding or decoding unit of CABAC is initialized or reinitialized. When initializing, the initial value (default state) of the context used for CABAC processing is set.
[0529] The encoding process will now be explained. Figure 77 is a block diagram showing the configuration of the first encoding unit 5200 included in the three-dimensional data encoding device according to this embodiment. Figure 78 is a block diagram showing the configuration of the division unit 5201 according to this embodiment. Figure 79 is a block diagram showing the configuration of the position information encoding unit 5202 and the attribute information encoding unit 5203 according to this embodiment.
[0530] The first encoding unit 5200 generates encoded data (encoded stream) by encoding point cloud data using a first encoding method (GPCC (Geometry-based PCC)). This first encoding unit 5200 includes a division unit 5201, a plurality of location information encoding units 5202, a plurality of attribute information encoding units 5203, an additional information encoding unit 5204, and a multiplexing unit 5205.
[0531] The division unit 5201 generates multiple division data by dividing the point cloud data. Specifically, the division unit 5201 generates multiple division data by dividing the space of the point cloud data into multiple subspaces. Here, a subspace is either a tile or a slice, or a combination of a tile and a slice. More specifically, the point cloud data includes location information, attribute information, and additional information. The division unit 5201 divides the location information into multiple division location information and the attribute information into multiple division attribute information. The division unit 5201 also generates additional information related to the division.
[0532] As shown in Figure 78, the division unit 5201 includes a tile division unit 5211 and a slice division unit 5212. For example, the tile division unit 5211 divides the point cloud into tiles. The tile division unit 5211 may determine the quantization value to be used for each divided tile as tile additional information.
[0533] The slice division unit 5212 further divides the tile obtained by the tile division unit 5211 into slices. The slice division unit 5212 may determine the quantization value to be used for each divided slice as slice additional information.
[0534] Multiple location information encoding units 5202 generate multiple encoded location information by encoding multiple divided location information. For example, multiple location information encoding units 5202 process multiple divided location information in parallel.
[0535] As shown in Figure 79, the position information encoding unit 5202 includes a CABAC initialization unit 5221 and an entropy encoding unit 5222. The CABAC initialization unit 5221 initializes or reinitializes CABAC according to the CABAC initialization flag. The entropy encoding unit 5222 encodes the division position information using CABAC.
[0536] Multiple attribute information encoding units 5203 generate multiple encoded attribute information by encoding multiple divided attribute information. For example, multiple attribute information encoding units 5203 process multiple divided attribute information in parallel.
[0537] As shown in Figure 79, the attribute information encoding unit 5203 includes a CABAC initialization unit 5231 and an entropy encoding unit 5232. The CABAC initialization unit 5221 initializes or reinitializes CABAC according to the CABAC initialization flag. The entropy encoding unit 5232 encodes the segmented attribute information using CABAC.
[0538] The additional information encoding unit 5204 generates encoded additional information by encoding the additional information contained in the point cloud data and the additional information related to data division generated by the division unit 5201 during division.
[0539] The multiplexing unit 5205 generates encoded data (encoded stream) by multiplexing multiple encoded position information, multiple encoded attribute information, and encoded additional information, and transmits the generated encoded data. The encoded additional information is used during decoding.
[0540] In Figure 77, an example is shown where there are two location information encoding units 5202 and two attribute information encoding units 5203. However, the number of location information encoding units 5202 and attribute information encoding units 5203 may be one or three or more. Furthermore, multiple divided data may be processed in parallel within the same chip, such as multiple cores in a CPU, or in parallel across the cores of multiple chips, or across multiple cores of multiple chips.
[0541] Next, the decoding process will be explained. Figure 80 is a block diagram showing the configuration of the first decoding unit 5240. Figure 81 is a block diagram showing the configuration of the location information decoding unit 5242 and the attribute information decoding unit 5243.
[0542] The first decoding unit 5240 reconstructs the point cloud data by decoding the encoded data (encoded stream) generated when the point cloud data is encoded using a first encoding method (GPCC). This first decoding unit 5240 includes a demultiplexing unit 5241, a plurality of location information decoding units 5242, a plurality of attribute information decoding units 5243, an additional information decoding unit 5244, and a coupling unit 5245.
[0543] The demultiplexing unit 5241 generates multiple encoded position information, multiple encoded attribute information, and encoded additional information by demultiplexing the encoded data (encoded stream).
[0544] Multiple location information decoding units 5242 generate multiple quantized location information by decoding multiple encoded location information. For example, multiple location information decoding units 5242 process multiple encoded location information in parallel.
[0545] As shown in Figure 81, the location information decoding unit 5242 includes a CABAC initialization unit 5251 and an entropy decoding unit 5252. The CABAC initialization unit 5251 initializes or reinitializes CABAC according to the CABAC initialization flag. The entropy decoding unit 5252 decodes the location information using CABAC.
[0546] The multiple attribute information decoding unit 5243 generates multiple segmented attribute information by decoding multiple encoded attribute information. For example, the multiple attribute information decoding unit 5243 processes multiple encoded attribute information in parallel.
[0547] As shown in Figure 81, the attribute information decoding unit 5243 includes a CABAC initialization unit 5261 and an entropy decoding unit 5262. The CABAC initialization unit 5261 initializes or reinitializes CABAC according to the CABAC initialization flag. The entropy decoding unit 5262 decodes the attribute information using CABAC.
[0548] Multiple additional information decoding units 5244 generate additional information by decoding encoded additional information.
[0549] The merging unit 5245 generates position information by combining multiple division position information using additional information. The merging unit 5245 generates attribute information by combining multiple division attribute information using additional information. For example, the merging unit 5245 first generates point cloud data corresponding to a tile by combining decoded point cloud data for a slice using slice additional information. Next, the merging unit 5245 restores the original point cloud data by combining point cloud data corresponding to a tile using tile additional information.
[0550] In Figure 80, an example is shown where there are two location information decoding units 5242 and two attribute information decoding units 5243. However, the number of location information decoding units 5242 and attribute information decoding units 5243 may be one or three or more. Furthermore, multiple divided data may be processed in parallel within the same chip, such as multiple cores in a CPU, or in parallel across the cores of multiple chips, or across multiple cores of multiple chips.
[0551] Figure 82 is a flowchart showing an example of the process related to the initialization of CABAC in encoding location information or attribute information.
[0552] First, the three-dimensional data encoding device determines, based on predetermined conditions, whether or not to perform CABAC initialization for each slice by encoding the position information of that slice (S5201).
[0553] If the three-dimensional data encoding device determines that it should perform CABAC initialization (Yes in S5202), it determines the initial context value to be used for encoding the location information (S5203). The initial context value is set to an initial value that takes the encoding characteristics into consideration. The initial value may be a predetermined value, or it may be determined adaptively according to the characteristics of the data in the slice.
[0554] Next, the three-dimensional data encoding device sets the CABAC initialization flag for location information to 1 and sets the context initial value (S5204). When CABAC initialization is performed, the initialization process is executed using the context initial value during the encoding of location information.
[0555] On the other hand, if the three-dimensional data encoding device determines that CABAC initialization is not required (No in S5202), it sets the CABAC initialization flag for the location information to 0 (S5205).
[0556] Next, the three-dimensional data encoding device determines, for each slice, whether or not to perform CABAC initialization by encoding the attribute information of that slice, based on predetermined conditions (S5206).
[0557] If the three-dimensional data encoding device determines that it should perform CABAC initialization (Yes in S5207), it determines the initial context value to be used for encoding attribute information (S5208). The initial context value is set to an initial value that takes encoding characteristics into consideration. The initial value may be a predetermined value, or it may be determined adaptively according to the characteristics of the data in the slice.
[0558] Next, the three-dimensional data encoding device sets the CABAC initialization flag for attribute information to 1 and sets the context initial value (S5209). When CABAC initialization is performed, the initialization process is executed using the context initial value during the encoding of attribute information.
[0559] On the other hand, if the three-dimensional data encoding device determines that CABAC initialization is not required (No in S5207), it sets the CABAC initialization flag in the attribute information to 0 (S5210).
[0560] In the flowchart shown in Figure 82, the processing order of location information and attribute information may be reversed or in parallel.
[0561] Although the flowchart in Figure 82 uses slice-level processing as an example, processing at the tile level or other data units can be done in the same way. In other words, the slices in the flowchart of Figure 82 can be interpreted as tiles or other data units.
[0562] Furthermore, the specified conditions may be the same for location information and attribute information, or they may be different.
[0563] Figure 83 shows an example of the timing of CABAC initialization in point cloud data converted into a bitstream.
[0564] Point cloud data includes location information and attribute information (zero or more). That is, point cloud data may have no attribute information, or it may have multiple attribute information.
[0565] For example, a single three-dimensional point may have attribute information such as color information, color information and reflection information, or one or more pieces of color information associated with one or more viewpoints.
[0566] The method described in this embodiment can be applied to any of the configurations.
[0567] Next, we will explain the conditions for determining whether CABAC is initialized.
[0568] CABAC may be initialized in the encoding of location information or attribute information if the following conditions are met.
[0569] For example, CABAC may be initialized with the first data of location information or attribute information (or each attribute if there are multiple). For example, CABAC may be initialized with the first data of a PCC frame that can be decrypted independently. In other words, as shown in Figure 83(a), if a PCC frame can be decrypted on a frame-by-frame basis, CABAC may be initialized with the first data of the PCC frame.
[0570] Furthermore, for example, as shown in Figure 83(b), if the inter-frame conjecture is used between PCC frames, or if a frame cannot be decoded individually, the CABAC may be initialized with the first data of a random access unit (e.g., GOF).
[0571] Alternatively, as shown in Figure 83(c), CABAC may be initialized at the beginning of a slice data set divided into one or more parts, at the beginning of a tile data set divided into one or more parts, or at the beginning of other divided data.
[0572] Figure 83(c) shows an example of a tile, but the same applies to a slice. Initialization may or may not be required at the beginning of a tile or slice.
[0573] Figure 84 shows the structure of the encoded data and the method of storing the encoded data in the NAL unit.
[0574] Initialization information may be stored in the header of the encoded data, or in the metadata. Alternatively, initialization information may be stored in both the header and the metadata. Initialization information may be, for example, caba_init_flag, CABAC initial value, or an index on a table that can identify the initial value.
[0575] In this embodiment, the part described as being stored in metadata may be interpreted as being stored in the header of the encoded data, and vice versa.
[0576] Initialization information may be stored in the header of the encoded data, for example, in the first NAL unit of the encoded data. The location information contains initialization information for encoding the location information, and the attribute information contains initialization information for encoding the attribute information.
[0577] The cabac_init_flag for encoding attribute information and the cabac_init_flag for encoding location information may be the same value or different values. If they are the same value, the cabac_init_flag for location information and attribute information may be common. If they are different values, the cabac_init_flag for location information and attribute information will each represent different values.
[0578] Initialization information may be stored in metadata common to both location information and attribute information, in at least one of the individual metadata for location information and / or the individual metadata for attribute information, or in both the common metadata and the individual metadata. Alternatively, a flag may be used to indicate whether or not the information is contained in the individual metadata for location information, the individual metadata for attribute information, or the common metadata.
[0579] Figure 85 is a flowchart showing an example of the process related to the initialization of CABAC in decoding location information or attribute information.
[0580] The three-dimensional data decoding device analyzes the encoded data and obtains the CABAC initialization flag for location information, the CABAC initialization flag for attribute information, and the context initial value (S5211).
[0581] Next, the three-dimensional data decoding device determines whether the CABAC initialization flag for the position information is 1 or not (S5212).
[0582] If the CABAC initialization flag for location information is 1 (Yes in S5212), the three-dimensional data decoding device initializes the CABAC decoding of location information coding using the initial context value of location information coding (S5213).
[0583] On the other hand, if the CABAC initialization flag for location information is 0 (No in S5212), the three-dimensional data decoding device does not initialize CABAC decoding in location information coding (S5214).
[0584] Next, the three-dimensional data decoding device determines whether the CABAC initialization flag in the attribute information is 1 or not (S5215).
[0585] If the CABAC initialization flag for attribute information is 1 (Yes in S5215), the three-dimensional data decoding device initializes the CABAC decoding of attribute information encoding using the context initialization value of attribute information encoding (S5216).
[0586] On the other hand, if the CABAC initialization flag for attribute information is 0 (No in S5215), the three-dimensional data decoding device does not initialize CABAC decoding in attribute information encoding (S5217).
[0587] In the flowchart shown in Figure 85, the processing order of location information and attribute information may be reversed or in parallel.
[0588] The flowchart in Figure 85 is applicable to both slice partitioning and tile partitioning.
[0589] Next, the flow of the point cloud data encoding and decoding processes according to this embodiment will be described. Figure 86 is a flowchart of the point cloud data encoding process according to this embodiment.
[0590] First, the three-dimensional data encoding device determines the division method to be used (S5221). This division method includes whether or not to perform tiling division or slicing division. The division method may also include the number of divisions if tiling or slicing division is performed, and the type of division. The type of division refers to methods based on object shape, methods based on map information or location information, or methods based on data volume or processing volume, as described above. The division method may also be predetermined.
[0591] If tile division is performed (Yes in S5222), the three-dimensional data encoding device generates multiple tile position information and multiple tile attribute information by dividing the position information and attribute information into tile units (S5223). The three-dimensional data encoding device also generates tile addition information related to the tile division.
[0592] If slice division is performed (Yes in S5224), the three-dimensional data encoding device generates multiple divided position information and multiple divided attribute information by dividing multiple tile position information and multiple tile attribute information (or position information and attribute information) (S5225). The three-dimensional data encoding device also generates position slice addition information and attribute slice addition information related to slice division.
[0593] Next, the three-dimensional data encoding device generates multiple encoded location information and multiple encoded attribute information by encoding each of the multiple division location information and multiple division attribute information (S5226). The three-dimensional data encoding device also generates dependency information.
[0594] Next, the three-dimensional data encoding device generates encoded data (encoded stream) by NAL unitizing (multiplexing) multiple encoded position information, multiple encoded attribute information, and additional information (S5227). The three-dimensional data encoding device also transmits the generated encoded data.
[0595] Figure 87 is a flowchart showing an example of a process in which the value of the CABAC initialization flag is determined and additional information is updated during tile division (S5222) or slice division (S5225).
[0596] In steps S5222 and S5225, the positional information and attribute information of the tiles and / or slices may be divided independently and individually in their respective ways, or they may be divided together in common. This generates additional information divided for each tile and / or slice.
[0597] At this point, the three-dimensional data encoding device decides whether to set the CABAC initialization flag to 1 or 0 (S5231).
[0598] Then, the three-dimensional data encoding device updates the additional information to include the determined CABAC initialization flag (S5232).
[0599] Figure 88 is a flowchart showing an example of the CABAC initialization process in the encoding (S5226) process.
[0600] The three-dimensional data encoding device determines whether the CABAC initialization flag is 1 or not (S5241).
[0601] If the CABAC initialization flag is 1 (Yes in S5241), the three-dimensional data encoding device reinitializes the CABAC encoding unit to its default state (S5242).
[0602] The three-dimensional data encoding device then continues the encoding process until the conditions for stopping the encoding process are met, for example, until there is no more data to be encoded (S5243).
[0603] Figure 89 is a flowchart of the point cloud data decoding process according to this embodiment. First, the three-dimensional data decoding device determines the division method by analyzing the additional information related to the division method (tile additional information, position slice additional information, and attribute slice additional information) included in the encoded data (encoded stream) (S5251). This division method includes whether or not to perform tile division and whether or not to perform slice division. The division method may also include the number of divisions and the type of division when tile division or slice division is performed.
[0604] Next, the three-dimensional data decoding device generates partitioning location information and partitioning attribute information by decoding multiple encoded position information and multiple encoded attribute information contained in the encoded data using dependency information contained in the encoded data (S5252).
[0605] If the additional information indicates that slice division has been performed (Yes in S5253), the three-dimensional data decoding device generates multiple tile position information and multiple tile attribute information by combining multiple division position information and multiple division attribute information based on the position slice additional information and attribute slice additional information (S5254).
[0606] If the additional information indicates that tile division has been performed (Yes in S5255), the three-dimensional data decoding device generates position information and attribute information by combining multiple tile position information and multiple tile attribute information (multiple division position information and multiple division attribute information) based on the tile additional information (S5256).
[0607] Figure 90 is a flowchart showing an example of the process for initializing the CABAC decoding unit when combining information divided by slice (S5254) or combining information divided by tile (S5256).
[0608] The location and attribute information of slices or tiles may be combined using their respective methods, or they may be combined using the same method.
[0609] The three-dimensional data decoding device decodes the CABAC initialization flag from the additional information of the encoded stream (S5261).
[0610] Next, the three-dimensional data decoding device determines whether the CABAC initialization flag is 1 or not (S5262).
[0611] If the CABAC initialization flag is 1 (Yes in S5262), the three-dimensional data decoding device reinitializes the CABAC decoding unit to its default state (S5263).
[0612] On the other hand, if the CABAC initialization flag is not 1 (No in S5262), the three-dimensional data decoding device proceeds to step S5264 without reinitializing the CABAC decoding unit.
[0613] The three-dimensional data decoding device then continues the decoding process until the conditions for stopping the decoding process are met, for example, until there is no more data to be decoded (S5264).
[0614] Next, we will explain the other conditions for determining CABAC initialization.
[0615] Whether or not to initialize the encoding of location information or attribute information may be determined by considering the encoding efficiency of data units such as tiles or slices. In this case, CABAC may be initialized in the first data of a tile or slice that satisfies predetermined conditions.
[0616] Next, we will explain the criteria for determining CABAC initialization in location information encoding.
[0617] The three-dimensional data encoding device may, for example, determine the density of the point cloud data for each slice, that is, the number of points per unit area belonging to the slice, compare the data density of the slice with that of other slices, and if the change in data density does not exceed predetermined conditions, it may decide that not initializing CABAC is more efficient for encoding and therefore decide not to initialize CABAC. On the other hand, if the change in data density does not satisfy predetermined conditions, the three-dimensional data encoding device may decide that initializing CABAC is more efficient for encoding and therefore decide to initialize CABAC.
[0618] Here, "other slices" may refer to, for example, the slice immediately preceding the decoding order, or to spatially adjacent slices. Furthermore, the three-dimensional data encoding device may determine whether or not to initialize CABAC based on whether the data density of the slice in question meets a predetermined data density, without comparing it with the data density of other slices.
[0619] When the three-dimensional data encoding device determines that CABAC initialization is required, it determines the context initial value to be used for encoding location information. The context initial value is set to an initial value with good encoding characteristics according to the data density. The three-dimensional data encoding device may also maintain an initial value table for data density in advance and select the optimal initial value from the table.
[0620] Furthermore, the three-dimensional data encoding device may determine whether or not to initialize CABAC based not only on the example of slice density, but also on the number of points, the distribution of points, the bias of points, etc. Alternatively, the three-dimensional data encoding device may determine whether or not to initialize CABAC based on the feature quantities obtained from the point information, the number of feature points, or the recognized object. In that case, the determination criteria may be stored in memory in advance as a table associated with the feature quantities obtained from the point information, the number of feature points, or the object recognized based on the point information.
[0621] The three-dimensional data encoding device may, for example, determine objects in the location information of map information and decide whether or not to initialize CABAC based on the objects based on the location information, or it may decide whether or not to initialize CABAC based on information or features obtained by projecting the three-dimensional data onto two dimensions.
[0622] Next, we will explain the criteria for determining CABAC initialization in attribute information encoding.
[0623] The three-dimensional data encoding device may, for example, compare the color characteristics of the previous slice with those of the current slice, and if the change in color characteristics satisfies predetermined conditions, it may determine that not initializing CABAC is more efficient and decide not to initialize CABAC. On the other hand, if the change in color characteristics does not satisfy predetermined conditions, the three-dimensional data encoding device may determine that initializing CABAC is more efficient and decide to initialize. Color characteristics include, for example, luminance, chromaticity, saturation, their histograms, and color continuity.
[0624] Here, "other slices" may refer to, for example, the slice immediately preceding the decoding order, or to spatially adjacent slices. Furthermore, the three-dimensional data encoding device may determine whether or not to initialize CABAC based on whether the data density of the slice in question meets a predetermined data density, without comparing it with the data density of other slices.
[0625] When the three-dimensional data encoding device determines that CABAC initialization is required, it determines the context initial value to be used for encoding attribute information. The context initial value is set to an initial value with good encoding characteristics according to the data density. The three-dimensional data encoding device may also maintain an initial value table for data density in advance and select the optimal initial value from the table.
[0626] The three-dimensional data encoding device may determine whether or not to initialize CABAC based on the information derived from the reflectance if the attribute information is reflectance.
[0627] When a three-dimensional point has multiple attribute pieces of information, the three-dimensional data encoding device may independently determine initialization information based on each piece of attribute information, or it may determine initialization information for multiple attribute pieces of information based on one of the attribute pieces of information, or it may use multiple attribute pieces of information to determine initialization information for multiple attribute pieces of information.
[0628] While an example was given where location information initialization information is determined based on location information and attribute information initialization information is determined based on attribute information, the location information and attribute information initialization information may also be determined based on location information, or based on attribute information, or based on both types of information.
[0629] The three-dimensional data encoding device may determine the initialization information based on the results of a prior simulation of encoding efficiency, for example, by setting cabac_init_flag to on or off, or by using one or more initial values from an initial value table.
[0630] If a three-dimensional data encoding device determines a method for dividing data into slices or tiles based on location information or attribute information, it may also determine initialization information based on the same information used to determine the division method.
[0631] Figure 91 shows examples of tiles and slices.
[0632] For example, a slice in a tile containing part of the PCC data is identified as shown in the legend. The CABAC initialization flag can be used to determine whether context reinitialization is necessary in consecutive slices. For example, in Figure 91, if a tile contains slice data divided by object (moving object, walkway, building, tree, and other objects), the CABAC initialization flag for the moving object, walkway, and tree slices is set to 1, and the CABAC initialization flag for the building and other slices is set to 0. This is because, for example, if both the walkway and the building are dense permanent structures and may have similar coding efficiency, coding efficiency may be improved by not reinitializing CABAC between the walkway and building slices. On the other hand, if the density and coding efficiency of the building and the tree may differ significantly, coding efficiency may be improved by initializing CABAC between the building and tree slices.
[0633] Figure 92 is a flowchart showing an example of how to initialize CABAC and determine the initial context values.
[0634] First, the three-dimensional data encoding device divides the point cloud data into slices based on the objects determined from the position information (S5271).
[0635] Next, the three-dimensional data encoding device determines, for each slice, whether or not to perform CABAC initialization of the location information encoding and attribute information encoding based on the data density of the objects in that slice (S5272). In other words, the three-dimensional data encoding device determines CABAC initialization information (CABAC initialization flag) for the location information encoding and attribute information encoding based on the location information. The three-dimensional data encoding device determines an initialization with good encoding efficiency, for example, based on the point cloud data density. Note that the CABAC initialization information may be shown in a common cabac_init_flag for both location information and attribute information.
[0636] Next, if the three-dimensional data encoding device determines that it should initialize CABAC (Yes in S5273), it determines the initial context value for encoding the location information (S5274).
[0637] Next, the three-dimensional data encoding device determines the initial context value for encoding attribute information (S5275).
[0638] Next, the three-dimensional data encoding device sets the CABAC initialization flag for location information to 1, setting the context initial value for location information, and also sets the CABAC initialization flag for attribute information to 1, setting the context initial value for attribute information (S5276). When performing CABAC initialization, the three-dimensional data encoding device performs initialization processing using the context initial value for both the encoding of location information and the encoding of attribute information.
[0639] On the other hand, if the three-dimensional data encoding device determines that CABAC initialization is not required (No in S5273), it sets the CABAC initialization flag for the location information to 0 and the CABAC initialization flag for the attribute information to 0 (S5277).
[0640] Figure 93 shows an example of dividing a top-view map of point cloud data obtained by LiDAR into tiles. Figure 94 is a flowchart showing another example of how to initialize CABAC and determine the initial context values.
[0641] The three-dimensional data encoding device divides point cloud data into one or more tiles in large-scale map data based on location information using a two-dimensional division method in a top-down view (S5281). The three-dimensional data encoding device may divide the data into square regions, for example, as shown in Figure 93. Alternatively, the three-dimensional data encoding device may divide the point cloud data into tiles of various shapes and sizes. The tile division may be performed using one or more predetermined methods, or it may be performed adaptively.
[0642] Next, the three-dimensional data encoding device determines the objects within each tile and decides whether or not to initialize CABAC by encoding the position information or attribute information of the tile (S5282). In the case of slicing, the three-dimensional data encoding device recognizes objects (trees, people, moving objects, buildings) and performs slicing and determines the initial values according to the object.
[0643] If the three-dimensional data encoding device determines that it should perform CABAC initialization (Yes in S5283), it determines the initial context value for encoding the location information (S5284).
[0644] Next, the three-dimensional data encoding device determines the initial context value for encoding attribute information (S5285).
[0645] In steps S5284 and S5285, initial values for tiles with specific coding characteristics may be stored as initial values and used as initial values for tiles with the same coding characteristics.
[0646] Next, the three-dimensional data encoding device sets the CABAC initialization flag for location information to 1, setting the context initial value for location information, and also sets the CABAC initialization flag for attribute information to 1, setting the context initial value for attribute information (S5286). When performing CABAC initialization, the three-dimensional data encoding device performs initialization processing using the context initial value for both the encoding of location information and the encoding of attribute information.
[0647] On the other hand, if the three-dimensional data encoding device determines that CABAC initialization is not required (No in S5283), it sets the CABAC initialization flag for the location information to 0 and the CABAC initialization flag for the attribute information to 0 (S5287).
[0648] (Embodiment 8) The quantization parameters will be described below.
[0649] Slicing and tiling are used to divide point cloud data based on its characteristics and location. However, due to hardware limitations and real-time processing requirements, the required quality of each divided point cloud data may differ. For example, when dividing and encoding the data into slices for each object, slice data containing plants, which are less important, can be quantized to reduce their resolution (quality). On the other hand, important slice data can be given higher resolution (quality) by setting a lower quantization value. Quantization parameters are used to enable this control over quantization values.
[0650] Here, the data to be quantized, the scale used for quantization, and the quantized data calculated by quantization are expressed by (Equation G1) and (Equation G2) below.
[0651] Quantized data = data / scale (Equation G1)
[0652] Data = Quantized Data * Scale (Equation G2)
[0653] Figure 95 is a diagram illustrating the processing of the quantization unit 5323, which quantizes data, and the inverse quantization unit 5333, which inversely quantizes the quantized data.
[0654] The quantization unit 5323 calculates quantized data by quantizing the data using a scale, that is, by performing a process using equation G1.
[0655] The inverse quantization unit 5333 calculates the inversely quantized data by performing a process using equation G2, that is, by inversely quantizing the quantized data using a scale.
[0656] Furthermore, the scale and the quantization value (QP (Quantization Parameter) value) are expressed by the following (Equation G3).
[0657] Quantized value (QP value) = log(scale) (Equation G3)
[0658] Quantized value (QP value) = default value (reference value) + quantized delta (difference information) (Equation G4)
[0659] These parameters are collectively called quantization parameters.
[0660] For example, as shown in Figure 96, the quantized value is a value based on the default value and is calculated by adding the quantization delta to the default value. If the quantized value is smaller than the default value, the quantization delta will be a negative value. If the quantized value is larger than the default value, the quantization delta will be a positive value. If the quantized value is equal to the default value, the quantization delta will be 0. If the quantization delta is 0, the quantization delta may be omitted.
[0661] The encoding process will now be explained. Figure 97 is a block diagram showing the configuration of the first encoding unit 5300 included in the three-dimensional data encoding device according to this embodiment. Figure 98 is a block diagram showing the configuration of the division unit 5301 according to this embodiment. Figure 99 is a block diagram showing the configuration of the position information encoding unit 5302 and the attribute information encoding unit 5303 according to this embodiment.
[0662] The first encoding unit 5300 generates encoded data (encoded stream) by encoding point cloud data using a first encoding method (GPCC (Geometry-based PCC)). This first encoding unit 5300 includes a division unit 5301, a plurality of location information encoding units 5302, a plurality of attribute information encoding units 5303, an additional information encoding unit 5304, and a multiplexing unit 5305.
[0663] The division unit 5301 generates multiple division data by dividing the point cloud data. Specifically, the division unit 5301 generates multiple division data by dividing the space of the point cloud data into multiple subspaces. Here, a subspace is either a tile or a slice, or a combination of a tile and a slice. More specifically, the point cloud data includes location information, attribute information, and additional information. The division unit 5301 divides the location information into multiple division location information and the attribute information into multiple division attribute information. The division unit 5301 also generates additional information related to the division.
[0664] As shown in Figure 98, the division unit 5301 includes a tile division unit 5311 and a slice division unit 5312. For example, the tile division unit 5311 divides the point cloud into tiles. The tile division unit 5311 may determine the quantization value to be used for each divided tile as tile additional information.
[0665] The slice division unit 5312 further divides the tile obtained by the tile division unit 5311 into slices. The slice division unit 5312 may determine the quantization value to be used for each divided slice as slice additional information.
[0666] Multiple location information encoding units 5302 generate multiple encoded location information by encoding multiple divided location information. For example, multiple location information encoding units 5302 process multiple divided location information in parallel.
[0667] As shown in Figure 99, the position information encoding unit 5302 includes a quantization value calculation unit 5321 and an entropy encoding unit 5322. The quantization value calculation unit 5321 obtains the quantization value (quantization parameter) of the partitioned position information to be encoded. The entropy encoding unit 5322 calculates quantized position information by quantizing the partitioned position information using the quantization value (quantization parameter) obtained by the quantization value calculation unit 5321.
[0668] Multiple attribute information encoding units 5303 generate multiple encoded attribute information by encoding multiple divided attribute information. For example, multiple attribute information encoding units 5303 process multiple divided attribute information in parallel.
[0669] As shown in Figure 99, the attribute information encoding unit 5303 includes a quantization value calculation unit 5331 and an entropy encoding unit 5332. The quantization value calculation unit 5331 obtains the quantization values (quantization parameters) of the partitioned attribute information to be encoded. The entropy encoding unit 5332 calculates quantized attribute information by quantizing the partitioned attribute information using the quantization values (quantization parameters) obtained by the quantization value calculation unit 5331.
[0670] The additional information encoding unit 5304 generates encoded additional information by encoding the additional information contained in the point cloud data and the additional information related to data division generated by the division unit 5301 during division.
[0671] The multiplexing unit 5305 generates encoded data (encoded stream) by multiplexing multiple encoded position information, multiple encoded attribute information, and encoded additional information, and transmits the generated encoded data. The encoded additional information is used during decoding.
[0672] In Figure 97, an example is shown where there are two location information encoding units 5302 and two attribute information encoding units 5303. However, the number of location information encoding units 5302 and attribute information encoding units 5303 may be one or three or more. Furthermore, multiple divided data may be processed in parallel within the same chip, such as multiple cores in a CPU, or in parallel across the cores of multiple chips, or in parallel across multiple cores of multiple chips.
[0673] Next, the decoding process will be explained. Figure 100 is a block diagram showing the configuration of the first decoding unit 5340. Figure 101 is a block diagram showing the configuration of the location information decoding unit 5342 and the attribute information decoding unit 5343.
[0674] The first decoding unit 5340 reconstructs the point cloud data by decoding the encoded data (encoded stream) generated when the point cloud data is encoded using a first encoding method (GPCC). This first decoding unit 5340 includes a demultiplexing unit 5341, a plurality of position information decoding units 5342, a plurality of attribute information decoding units 5343, an additional information decoding unit 5344, and a coupling unit 5345.
[0675] The demultiplexing unit 5341 generates multiple encoded position information, multiple encoded attribute information, and encoded additional information by demultiplexing the encoded data (encoded stream).
[0676] Multiple position information decoding units 5342 generate multiple quantized position information by decoding multiple encoded position information. For example, multiple position information decoding units 5342 process multiple encoded position information in parallel.
[0677] As shown in Figure 101, the position information decoding unit 5342 includes a quantization value calculation unit 5351 and an entropy decoding unit 5352. The quantization value calculation unit 5351 acquires the quantized value of the quantized position information. The entropy decoding unit 5352 calculates the position information by dequantizing the quantized position information using the quantization value acquired by the quantization value calculation unit 5351.
[0678] The multiple attribute information decoding unit 5343 generates multiple segmented attribute information by decoding multiple encoded attribute information. For example, the multiple attribute information decoding unit 5343 processes multiple encoded attribute information in parallel.
[0679] As shown in Figure 101, the attribute information decoding unit 5343 includes a quantization value calculation unit 5361 and an entropy decoding unit 5362. The quantization value calculation unit 5361 obtains the quantization value of the quantized attribute information. The entropy decoding unit 5362 calculates the attribute information by dequantizing the quantized attribute information using the quantization value obtained by the quantization value calculation unit 5361.
[0680] Multiple additional information decoding units 5344 generate additional information by decoding encoded additional information.
[0681] The merging unit 5345 generates position information by combining multiple division position information using additional information. The merging unit 5345 generates attribute information by combining multiple division attribute information using additional information. For example, the merging unit 5345 first generates point cloud data corresponding to a tile by combining decoded point cloud data for a slice using slice additional information. Next, the merging unit 5345 restores the original point cloud data by combining point cloud data corresponding to a tile using tile additional information.
[0682] In Figure 100, an example is shown where there are two location information decoding units 5342 and two attribute information decoding units 5343. However, the number of location information decoding units 5342 and attribute information decoding units 5343 may be one or three or more. Furthermore, multiple divided data may be processed in parallel within the same chip, such as multiple cores in a CPU, or in parallel across the cores of multiple chips, or across multiple cores of multiple chips.
[0683] [Method for Determining Quantization Parameters] Figure 102 is a flowchart showing an example of the process for determining quantization parameters (QP values) in the encoding of geometry or attribute information.
[0684] The QP value is determined, for example, for each data unit of positional information or attribute information that constitutes a PCC frame, taking encoding efficiency into consideration. If the data unit is a divided tile unit or a divided slice unit, the QP value is determined for the divided data unit, taking encoding efficiency for that data unit into consideration. Alternatively, the QP value may be determined for the data unit before division.
[0685] As shown in Figure 102, the three-dimensional data encoding device determines the QP value used for encoding the location information (S5301). The three-dimensional data encoding device may determine the QP value for each of the divided slices based on a predetermined method. Specifically, the three-dimensional data encoding device determines the QP value based on the characteristics or quality of the location information data. For example, the three-dimensional data encoding device may determine the density of the point cloud data for each data unit, that is, the number of points per unit area belonging to the slice, and determine the value corresponding to the density of the point cloud data as the QP value. Alternatively, the three-dimensional data encoding device may determine the corresponding value as the QP value based on the number of points in the point cloud data, the distribution of points, the bias of points, or the feature quantities obtained from the point information, the number of feature points, or the recognized object. Furthermore, the three-dimensional data encoding device may determine the object in the location information of the map and determine the QP value based on the object based on the location information, or it may determine the QP value based on information or features obtained by projecting the three-dimensional point cloud into two dimensions. The corresponding QP values may be stored in memory as a table that is pre-associated with the density, number of points, distribution of points, or bias of the point cloud data. Alternatively, the corresponding QP values may be stored in memory as a table that is pre-associated with the feature quantities or number of feature points obtained from the point information, or with the objects recognized based on the point information. Furthermore, the corresponding QP values may be determined based on the results of simulations of coding rates and other factors using various QP values when encoding the positional information of the point cloud data.
[0686] Next, the three-dimensional data encoding device determines the reference value (default value) and difference information (quantization delta) of the QP value of the position information (S5302). Specifically, the three-dimensional data encoding device determines the reference value and difference information to be transmitted using the determined QP value and a predetermined method, and sets (adds) the determined reference value and difference information to at least one of the additional information and the data header.
[0687] Next, the three-dimensional data encoding device determines the QP value to be used for encoding the attribute information (S5303). The three-dimensional data encoding device may determine the QP value for each of the divided slices based on a predetermined method. Specifically, the three-dimensional data encoding device determines the QP value based on the characteristics or quality of the attribute information data. For example, the three-dimensional data encoding device may determine the QP value for each data unit based on the characteristics of the attribute information. Color characteristics include, for example, luminance, chromaticity, saturation, their histograms, and color continuity. If the attribute information is reflectance, the determination may be made according to the information based on reflectance. For example, if the three-dimensional data encoding device detects a face as an object from point cloud data, it may determine a good quality QP value for the point cloud data constituting the object in which the face was detected. In this way, the three-dimensional data encoding device may determine the QP value for the point cloud data constituting the object according to the type of object.
[0688] Furthermore, if a three-dimensional point has multiple attribute pieces of information, the three-dimensional data encoding device may independently determine the QP value based on each piece of attribute information, or it may determine the QP value of multiple attribute pieces of information based on one of the attribute pieces of information, or it may use multiple attribute pieces of information to determine the QP value of the multiple attribute pieces of information.
[0689] Next, the three-dimensional data encoding device determines the reference value (default value) and difference information (quantization delta) of the QP value of the attribute information (S5304). Specifically, the three-dimensional data encoding device determines the reference value and difference information to be transmitted using the determined QP value and a predetermined method, and sets (adds) the determined reference value and difference information to at least one of the additional information and the data header.
[0690] Then, the three-dimensional data encoding device quantizes and encodes the location information and attribute information based on the determined QP values of the location information and attribute information (S5305).
[0691] While the example described shows that the QP value of location information is determined based on the location information and the QP value of attribute information is determined based on the attribute information, this is not an exhaustive example. For instance, the QP values of location information and attribute information may be determined based on location information, or based on attribute information, or based on both location information and attribute information.
[0692] Furthermore, the QP values for location information and attribute information may be adjusted considering the balance between the quality of location information and the quality of attribute information in the point cloud data. For example, the QP values for location information and attribute information may be determined such that the quality of location information is set high and the quality of attribute information is set lower than the quality of location information. For example, the QP value for attribute information may be determined to satisfy the constraint that it is greater than or equal to the QP value of location information.
[0693] Furthermore, the QP value may be adjusted so that the encoded data falls within a predetermined rate range. For example, if the encoding amount of the previous data unit is likely to exceed a predetermined rate, that is, if the difference to the predetermined rate is less than the first difference, the QP value may be adjusted so that the encoding quality of the data unit is less than the first difference. On the other hand, if the difference to the predetermined rate is greater than a second difference which is greater than the first difference, and there is a sufficiently large difference, the QP value may be adjusted so that the encoding quality of the data unit is improved. Adjustments between data units may be, for example, between PCC frames, between tiles, or between slices. The QP value of attribute information may be adjusted based on the encoding rate of the location information.
[0694] In the flowchart shown in Figure 102, the processing order of location information and attribute information may be reversed or in parallel.
[0695] Although the flowchart in Figure 102 uses slice-level processing as an example, processing at the tile level or other data units can be done in the same way. In other words, the slices in the flowchart of Figure 102 can be interpreted as tiles or other data units.
[0696] Figure 103 is a flowchart showing an example of the decoding process for location information and attribute information.
[0697] As shown in Figure 103, the three-dimensional data decoding device acquires reference values and difference information indicating the QP value of the location information, and reference values and difference information indicating the QP value of the attribute information (S5311). Specifically, the three-dimensional data decoding device analyzes either or both of the transmitted metadata and the header of the encoded data to acquire reference values and difference information for deriving the QP value.
[0698] Next, the three-dimensional data decoding device uses the acquired reference value and difference information to derive the QP value based on a predetermined method (S5312).
[0699] The three-dimensional data decoding device then acquires quantized position information and dequantizes the position information by dequantizing the quantized position information using the derived QP value (S5313).
[0700] Next, the three-dimensional data decoding device acquires quantized attribute information and dequantizes the attribute information by dequantizing the quantized attribute information using the derived QP value (S5314).
[0701] Next, we will explain the method for transmitting quantization parameters.
[0702] Figure 104 is a diagram illustrating a first example of a method for transmitting quantization parameters. Figure 104(a) shows an example of the relationship between QP values.
[0703] In Figure 104, Q G and Q A These represent the absolute values of the QP values used for encoding location information and the absolute values of the QP values used for encoding attribute information, respectively. G This is an example of a first quantization parameter used to quantize the positional information of multiple three-dimensional points. Also, Δ(Q) A Q G ) is Q A Q used in the derivation G This shows the difference information, which indicates the difference from Q. A QG and Δ(Q A , Q G ) are used for derivation. Thus, the QP value is transmitted by being divided into a reference value (absolute value) and difference information (relative value). Further, in decoding, a desired QP value is derived from the transmitted reference value and difference information.
[0704] For example, in (a) of FIG. 104, the absolute value Q G and the difference information Δ(Q A , Q G ) are transmitted, and in decoding, as shown in the following (Equation G5), Q A is derived by adding Δ(Q A , Q A , Q G ) to Q
[0705] Q A = Q G + Δ(Q A , Q G ) (Equation G5)
[0706] A method for transmitting the QP value when slice-dividing point cloud data composed of position information and attribute information using (b) and (c) of FIG. 104 will be described. (b) of FIG. 104 is a diagram showing a first example of the relationship between the reference value and difference information of each QP value. (c) of FIG. 104 is a diagram showing a first example of the transmission order of the QP value, position information, and attribute information.
[0707] The QP value is roughly divided into a QP value per frame of the PCC (frame QP) and a QP value per data unit (data QP) for each position information and each attribute information. The QP value per data unit is the QP value used for encoding determined in step S5301 of FIG. 102.
[0708] Here, Q G , which is the QP value used for encoding the position information in the PCC frame unit, is used as a reference value, and the QP value per data unit is generated and transmitted as difference information indicating the difference from Q G . Q G : QP value for encoding position information in the PCC frame... transmitted as a reference value "1." using GPS A : QP value for encoding attribute information in the PCC frame... using APS for QG The difference information "2." is sent out, indicating the difference from Q. Gs1 Q Gs2 : QP value of the encoding of location information in slice data... Using the header of the location information encoding data, Q G The difference information "3." and "5." indicating the difference from Q are sent out. As1 Q As2 : QP value of attribute information encoding in slice data... Using the header of the attribute information encoding data, Q A The difference information "4." and "6." indicating the difference from are sent.
[0709] The information used to derive the frame QP is contained in the frame metadata (GPS, APS), and the information used to derive the data QP is contained in the data metadata (header of the encoded data).
[0710] In this way, data QP is generated and sent as difference information indicating the difference from frame QP. Therefore, the amount of data in data QP can be reduced.
[0711] The first decoding unit 5340 refers to the metadata indicated by the arrow in Figure 104(c) for each encoded data and obtains the reference value and difference information corresponding to the encoded data. Then, the first decoding unit 5340 derives the QP value corresponding to the encoded data to be decoded based on the obtained reference value and difference information.
[0712] The first decoding unit 5340 obtains, for example, the reference information "1." and the difference information "2." and "6." indicated by the arrows in Figure 104(c) from the metadata or header, and adds the difference information "2." and "6." to the reference information "1." as shown in (Equation G6) below, A s2 Derive the QP value.
[0713] Q AS2 = Q G +Δ(Q) A Q G ) + Δ(Q As2 Q A ) (Formula G6)
[0714] Point cloud data includes location information and attribute information (zero or more). That is, point cloud data may have no attribute information, or it may have multiple attribute information.
[0715] For example, a single three-dimensional point may have attribute information such as color information, color information and reflection information, or one or more pieces of color information associated with one or more viewpoints.
[0716] Here, an example of having two pieces of color information and reflection information will be explained using Figure 105. Figure 105 is a diagram illustrating a second example of a method for transmitting quantization parameters. Figure 105(a) shows a second example of the relationship between the reference value of each QP value and the difference information. Figure 105(b) shows a second example of the transmission order of QP values, position information, and attribute information.
[0717] Q G This is an example of the first quantization parameter, similar to Figure 104.
[0718] Each of the two color information sets is represented by luminance (lumen) Y and chrominance (chromen) Cb and Cr. The QP value used to encode the luminance Y1 of the first color is Q. Y1 Q is the reference value. G And the difference shown is Δ(Q Y1 Q G It is derived using ). Brightness Y1 is an example of the first brightness, Q Y1 This is an example of a second quantization parameter used to quantize the luminance Y1 as the first luminance. Δ(Q) Y1 Q G ) is the difference information "2.".
[0719] Furthermore, QP value is used to encode the color differences Cb1 and Cr1 of the first color. Cb1 Q Cr1 These represent QY1 and Δ(Q), respectively, which indicate the difference between them. Cb1 Q Y1 ), Δ(Q Cr1 Q Y1 ) is derived using ). Color differences Cb1 and Cr1 are examples of the first color difference, and Q Cb1 Q Cr1This is an example of a third quantization parameter used to quantize the color differences Cb1 and Cr1 as the first color differences. Δ(Q Cb1 Q Y1 ) is the difference information "3.", and Δ(Q Cr1 Q Y1 ) is the difference information "4.". Δ(Q Cb1 Q Y1 ) and Δ(Q Cr1 Q Y1 These are examples of the first difference, respectively.
[0720] Q Cb1 and Q Cr1 The same value may be used for each, or a common value may be used. If a common value is used, Q Cb1 and Q Cr1 Since only one of them needs to be used, the other is not necessary.
[0721] Furthermore, Q is the QP value used to encode the luminance Y1D of the first color in the slice data. Y1D Q Y1 And the difference shown is Δ(Q Y1D Q Y1 ) is derived using . The luminance Y1D of the first color in the slice data is an example of the first luminance of one or more three-dimensional points contained in the subspace, Q Y1D This is an example of a fifth quantization parameter used to quantize luminance Y1D. Δ(Q) Y1D Q Y1 ) is difference information "10.", and is an example of the second difference.
[0722] Similarly, Q is the QP value used to encode the color differences Cb1D and Cr1D of the first color in the slice data. Cb1D Q Cr1D These are, respectively, Q Cb1 Q Cr1 And the difference shown is Δ(Q Cb1D Q Cb1 ), Δ(Q Cr1D Q Cr1 ) is derived using . The color differences Cb1D and Cr1D of the first color in the slice data are examples of the first color differences of one or more three-dimensional points contained in the subspace, Q Cb1D QCr1D is an example of the sixth quantization parameter used to quantize the color difference Cb1D and Cr1D. Δ(Q Cb1D , Q Cb1 ) is the difference information "11.", and Δ(Q Cr1D , Q Cr1 ) is the difference information "12.". Δ(Q Cb1D , Q Cb1 ) and Δ(Q Cr1D , Q Cr1 ) are an example of the third difference.
[0723] Since the relationship of the QP value in the first color also holds for the second color, the explanation is omitted.
[0724] Q R , which is the QP value used for encoding the reflectance R, is derived using the reference value Q G and Δ(Q R , Q G ) indicating the difference therefrom. Q R is an example of the fourth quantization parameter used to quantize the reflectance R. Δ(Q R , Q G ) is the difference information "8.".
[0725] Also, Q RD , which is the QP value used for encoding the reflectance RD in the slice data, is derived using Q R and Δ(Q RD , Q R ) indicating the difference therefrom. Δ(Q RD , Q R ) is the difference information "16.".
[0726] Thus, the difference information "9." to "16." indicates the difference information between the data QP and the frame QP.
[0727] Note that, for example, when the values of the data QP and the frame QP are the same, the difference information may be set to 0, or it may be regarded as 0 by not transmitting the difference information.
[0728] The first decoding unit 5340, for example, when decoding the color difference Cr2 of the second color, obtains the reference information "1." and difference information "5.", "7.", and "15." indicated by the arrows in Figure 105(b) from metadata or a header, and derives the QP value of the color difference Cr2 by adding the difference information "5.", "7.", and "15." to the reference information "1." as shown in (Equation G7) below.
[0729] Q Cr2D = Q G +Δ(Q) Y2 Q G ) + Δ(Q Cr2 Q Y2 ) + Δ(Q Cr2D Q Cr2 ) (Formula G7)
[0730] Next, an example of dividing location information and attribute information into two tiles, and then into two slices, will be explained using Figure 106. Figure 106 is a diagram illustrating a third example of a method for transmitting quantization parameters. Figure 106(a) shows a third example of the relationship between the reference value of each QP value and the difference information. Figure 106(b) shows a third example of the transmission order of QP values, location information, and attribute information. Figure 106(c) is a diagram illustrating the intermediate generated values of the difference information in the third example.
[0731] When dividing into multiple tiles and then further dividing into multiple slices, as shown in Figure 106(c), the QP value (Q) for each tile after division into tiles is obtained. At1 ) and difference information Δ(Q At1 Q A ) is generated as an intermediate generated value. Then, after dividing into slices, the QP value (Q) for each slice is generated. At1s1 Q At1s2 ) and difference information (Δ(Q At1s1 Q At1 ), Δ(Q At1s2 Q At1 )) is generated.
[0732] In this case, for example, the difference information "4." in Figure 106(a) is derived by the following (Equation G8).
[0733] Δ(Q)At1s1 Q A ) = Δ(Q At1 Q A ) + Δ(Q At1s1 Q At1 ) (Formula G8)
[0734] The first decoding unit 5340, for example, receives attribute information A of slice 1 in tile 2. t2s1 When decrypting, the reference information "1." and difference information "2." and "8." indicated by the arrows in Figure 106(b) are obtained from the metadata or header, and the difference information "2." and "8." are added to the reference information "1." as shown in (Equation G9) below to obtain attribute information A t2s1 Derive the QP value.
[0735] Q At2s1 = Q G +Δ(Q) At2s1 Q A ) + Δ(Q A Q G ) (Formula G9)
[0736] Next, the flow of the point cloud data encoding and decoding processes according to this embodiment will be described. Figure 107 is a flowchart of the point cloud data encoding process according to this embodiment.
[0737] First, the three-dimensional data encoding device determines the division method to be used (S5321). This division method includes whether or not to perform tiling division or slicing division. The division method may also include the number of divisions if tiling or slicing division is performed, and the type of division. The type of division refers to methods based on object shape, methods based on map information or location information, or methods based on data volume or processing volume, as described above. The division method may also be predetermined.
[0738] If tile division is performed (Yes in S5322), the three-dimensional data encoding device generates multiple tile position information and multiple tile attribute information by dividing the position information and attribute information into tile units (S5323). The three-dimensional data encoding device also generates tile addition information related to the tile division.
[0739] If slice division is performed (Yes in S5324), the three-dimensional data encoding device generates multiple divided position information and multiple divided attribute information by dividing multiple tile position information and multiple tile attribute information (or position information and attribute information) (S5325). The three-dimensional data encoding device also generates position slice addition information and attribute slice addition information related to slice division.
[0740] Next, the three-dimensional data encoding device generates multiple encoded location information and multiple encoded attribute information by encoding each of the multiple division location information and multiple division attribute information (S5326). The three-dimensional data encoding device also generates dependency information.
[0741] Next, the three-dimensional data encoding device generates encoded data (encoded stream) by NAL unitizing (multiplexing) multiple encoded position information, multiple encoded attribute information, and additional information (S5327). The three-dimensional data encoding device also transmits the generated encoded data.
[0742] Figure 108 is a flowchart showing an example of a process for determining the QP value and updating additional information during tile division (S5323) or slice division (S5325).
[0743] In steps S5323 and S5325, the positional information and attribute information of the tiles and / or slices may be divided independently and individually in their respective ways, or they may be divided together in common. This generates additional information divided for each tile and / or slice.
[0744] At this time, the three-dimensional data encoding device determines the reference value and difference information of the QP value for each divided tile and / or slice (S5331). Specifically, the three-dimensional data encoding device determines the reference value and difference information as illustrated in Figures 104 to 106.
[0745] Then, the three-dimensional data encoding device updates the additional information to include the determined reference value and difference information (S5332).
[0746] Figure 109 is a flowchart showing an example of the process of encoding the determined QP value in the encoding process (S5326).
[0747] The three-dimensional data encoding device encodes the QP value determined in step S5331 (S5341). Specifically, the three-dimensional data encoding device encodes the reference value and difference information of the QP value included in the updated additional information.
[0748] The three-dimensional data encoding device then continues the encoding process until the conditions for stopping the encoding process are met, for example, until there is no more data to be encoded (S5342).
[0749] Figure 110 is a flowchart of the point cloud data decoding process according to this embodiment. First, the three-dimensional data decoding device determines the division method by analyzing the additional information related to the division method (tile additional information, position slice additional information, and attribute slice additional information) included in the encoded data (encoded stream) (S5351). This division method includes whether or not to perform tile division and whether or not to perform slice division. The division method may also include the number of divisions and the type of division when tile division or slice division is performed.
[0750] Next, the three-dimensional data decoding device generates partitioning location information and partitioning attribute information by decoding multiple encoded position information and multiple encoded attribute information contained in the encoded data using dependency information contained in the encoded data (S5352).
[0751] If the additional information indicates that slice division has been performed (Yes in S5353), the three-dimensional data decoding device generates multiple tile position information and multiple tile attribute information by combining multiple division position information and multiple division attribute information based on the position slice additional information and attribute slice additional information (S5354).
[0752] If the additional information indicates that tile division has been performed (Yes in S5355), the three-dimensional data decoding device generates position information and attribute information by combining multiple tile position information and multiple tile attribute information (multiple division position information and multiple division attribute information) based on the tile additional information (S5356).
[0753] Figure 111 is a flowchart showing an example of a process for obtaining QP values and decoding the QP values of slices or tiles when combining information divided by slice (S5354) or information divided by tile (S5356).
[0754] The location and attribute information of slices or tiles may be combined using their respective methods, or they may be combined using the same method.
[0755] The three-dimensional data decoding device decodes the reference value and difference information from the additional information of the encoded stream (S5361).
[0756] Next, the three-dimensional data decoding device calculates quantization values using the decoded reference values and difference information, and updates the QP values used for inverse quantization to the calculated QP values (S5362). This makes it possible to derive QP values for inverse quantization of the quantization attribute information for each tile or slice.
[0757] The three-dimensional data decoding device then continues the decoding process until the conditions for stopping the decoding process are met, for example, until there is no more data to be decoded (S5363).
[0758] Figure 112 shows an example of GPS syntax. Figure 113 shows an example of APS syntax. Figure 114 shows an example of location information header syntax. Figure 115 shows an example of attribute information header syntax.
[0759] As shown in Figure 112, for example, GPS, which is additional location information, includes QP_value, which represents the absolute value that serves as the basis for deriving the QP value. QP_value is, for example, the Q shown in Figures 104 to 106. G It corresponds to this.
[0760] Furthermore, as shown in Figure 113, for example, if the APS, which is additional attribute information, has multiple color information from multiple viewpoints for a three-dimensional point, a default viewpoint may be defined, and the 0th item may always contain information for the default viewpoint. For example, when a three-dimensional data encoding device decodes or displays a single color information, it only needs to decode or display the 0th attribute information.
[0761] APS includes QP_delta_Attribute_to_Geometry. QP_delta_Attribute_to_Geometry indicates the difference information from the reference value (QP_value) listed in GPS. This difference information is, for example, the difference information from luminance if the attribute information is color information.
[0762] Furthermore, GPS may include a flag in the Geometry_header (location information header) indicating whether or not there is difference information for calculating the QP value. Similarly, APS may include a flag in the Attribute_header (attribute information header) indicating whether or not there is difference information for calculating the QP value. The flag may also indicate in the attribute information whether or not there is difference information between the data QP and the frame QP for calculating the data QP.
[0763] Thus, the encoded stream may include identification information (flags) indicating that, when the first color among the attribute information is represented by a first luminance and a first color difference, quantization using a second quantization parameter for quantizing the first luminance, and quantization using a third quantization parameter for quantizing the first color difference, quantization using a fifth and sixth quantization parameter, was performed.
[0764] Furthermore, as shown in Figure 114, the location information header may include QP_delta_data_to_frame, which indicates the difference information from the reference value (QP_value) listed in the GPS. Alternatively, the location information header may be divided into information for each tile and / or slice, with the corresponding QP value shown for each tile and / or slice.
[0765] Furthermore, as shown in Figure 115, the attribute information header may include QP_delta_data_to_frame, which indicates the difference information with the QP value described in APS.
[0766] In Figures 104 to 106, the reference value for QP was explained as the QP value of the position information in the PCC frame, but other values may also be used as the reference value.
[0767] Figure 116 illustrates another example of a method for transmitting quantization parameters.
[0768] Figures 116(a) and (b) show a fourth example in which a common reference value Q is set for the QP values of location information and attribute information in the PCC frame. In the fourth example, the reference value Q is stored in the GPS, and the QP value of the location information (Q) is calculated from the reference value Q. G The difference information of ) is stored in the GPS, and the QP value of the attribute information (Q Y and Q R The difference information of ) is stored in APS. Note that the reference value Q may be stored in SPS.
[0769] Figures 116(c) and (d) show a fifth example in which reference values are set independently for each location and attribute information. In the fifth example, the GPS and APS store the reference QP values (absolute values) for the location and attribute information, respectively. That is, the location information is stored with the reference value Q G The color information of the attribute information is set to the reference value Q. Y The reflectance of the attribute information is set to the reference value Q. R These are set accordingly. In this way, a reference value for the QP value may be set for each of the location information and multiple types of attribute information. Note that the fifth example may be combined with other examples. That is, the Q in the first example A Q in the second example Y1 Q Y2 Q R This may be the reference value for the QP value.
[0770] Figures 116(e) and (f) show a sixth example in which a common reference value Q is set for multiple PCC frames when there are multiple PCC frames. In the sixth example, the reference value Q is stored in SPS or GSPS, and the difference information between the QP value of the position information of each PCC frame and the reference value is stored in GPS. Note that, for example, within the range of a random access unit, such as GOF, the first frame of the random access unit is used as the reference value, and the difference information between PCC frames (e.g., Δ(Q)) is used. G(1) Q G(0) You may send out )).
[0771] Even if a tile or slice is further divided, the difference information between the QP value of the division unit and the current value is stored in the data header in the same manner and sent.
[0772] Figure 117 illustrates another example of a method for transmitting quantization parameters.
[0773] Figures 117(a) and (b) show the common reference value Q for positional and attribute information in the PCC frame. G A seventh example of setting the reference value Q is shown. In the seventh example, the reference value Q is set. G The data is stored in the GPS, and the difference information between it and the location information or attribute information is stored in the respective data headers. Reference value Q G It may be stored in SPS.
[0774] Furthermore, Figures 117(c) and (d) show an eighth example in which the QP value of attribute information is shown as the difference information with the QP value of position information belonging to the same slice and tile. In the eighth example, the reference value Q G This may be stored in SPS.
[0775] Figure 118 is a diagram illustrating a ninth example of a method for transmitting quantization parameters.
[0776] Figures 118(a) and (b) are a ninth example showing the difference information between the QP value of the location information and the common QP value of the attribute information, respectively, after determining the common QP value of the attribute information.
[0777] Figure 119 is a diagram illustrating an example of controlling the QP value.
[0778] Lower quantization parameter values improve quality, but require more bits, thus reducing encoding efficiency.
[0779] For example, when dividing three-dimensional point cloud data into tiles and encoding them, if the point cloud data contained in a tile represents a major road, it is encoded using a predefined attribute information QP value. On the other hand, since the surrounding tiles do not contain important information, it may be possible to degrade the data quality and improve encoding efficiency by setting the difference information of the QP values to a positive value.
[0780] Furthermore, when dividing the 3D point cloud data, which is divided into tiles, into slices and encoding them, sidewalks, trees, and buildings are set to negative QP values because they are important for position estimation (localization and mapping) in autonomous driving, while moving objects and others are set to positive QP values because they are of lower importance.
[0781] Figure 119(b) shows an example of deriving difference information when quantized delta values are set in advance based on the objects contained in the tiles or slices. For example, if the divided data is slice data of a "building" contained in a tile that is a "main road", the quantized delta value of the tile that is a "main road" is 0 and the quantized delta value of the slice data that is a "building" is -5, and the difference information is derived as -5.
[0782] Figure 120 is a flowchart showing an example of a method for determining the QP value based on the quality of an object.
[0783] The three-dimensional data encoding device divides the point cloud data into one or more tiles based on the map information and determines the objects contained in each of the one or more tiles (S5371). Specifically, the three-dimensional data encoding device performs object recognition processing to recognize what the objects are, for example, using a learning model obtained through machine learning.
[0784] Next, the three-dimensional data encoding device determines whether or not to encode the tiles to be processed with high quality (S5372). High-quality encoding means, for example, encoding at a bit rate higher than a predetermined rate.
[0785] Next, if the three-dimensional data encoding device encodes the tiles to be processed with high quality (Yes in S5372), it sets the QP value of the tiles to achieve high encoding quality (S5373).
[0786] On the other hand, if the three-dimensional data encoding device does not encode the tiles to be processed with high quality (No in S5372), it sets the QP value of the tiles to be encoded in a way that results in lower encoding quality (S5374).
[0787] After step S5373 or step S5374, the three-dimensional data encoding device determines the objects within the tile and divides them into one or more slices (S5375).
[0788] Next, the three-dimensional data encoding device determines whether or not to encode the slice to be processed with high quality (S5376).
[0789] Next, if the three-dimensional data encoding device encodes the slice to be processed with high quality (Yes in S5376), it sets the QP value of the slice to achieve high encoding quality (S5377).
[0790] On the other hand, if the three-dimensional data encoding device does not encode the slice to be processed with high quality (No in S5376), it sets the QP value of the slice to a lower encoding quality (S5378).
[0791] Next, the three-dimensional data encoding device determines the reference value and difference information to be transmitted in a predetermined manner based on the set QP value, and stores the determined reference value and difference information in at least one of the additional information and the data header (S5379).
[0792] Next, the three-dimensional data encoding device quantizes and encodes the position information and attribute information based on the determined QP value (S5380).
[0793] Figure 121 is a flowchart showing an example of a method for determining the QP value based on rate control.
[0794] The three-dimensional data encoding device encodes the point cloud data sequentially (S5381).
[0795] Next, the three-dimensional data encoding device determines the rate control status related to the encoding process from the amount of encoding data and the amount occupied by the encoding buffer, and determines the quality of the next encoding (S5382).
[0796] Next, the three-dimensional data encoding device determines whether or not to improve the encoding quality (S5383).
[0797] Next, if the three-dimensional data encoding device wants to improve the encoding quality (Yes in S5383), it sets the QP value of the tiles to improve the encoding quality (S5384).
[0798] On the other hand, if the three-dimensional data encoding device does not want to increase the encoding quality (No in S5383), it sets the QP value of the tiles to a lower encoding quality (S5385).
[0799] Next, the three-dimensional data encoding device determines the reference value and difference information to be transmitted in a predetermined manner based on the set QP value, and stores the determined reference value and difference information in at least one of the additional information and the data header (S5386).
[0800] Next, the three-dimensional data encoding device quantizes and encodes the position information and attribute information based on the determined QP value (S5387).
[0801] As described above, the three-dimensional data encoding device according to this embodiment performs the processing shown in Figure 122. First, the three-dimensional data encoding device quantizes the position information of each of the plurality of three-dimensional points using a first quantization parameter (S5391). The three-dimensional data encoding device quantizes the first luminance and the first color difference, which represent the first color among the attribute information of each of the plurality of three-dimensional points, using a second quantization parameter, and quantizes the first color difference using a third quantization parameter (S5392). The three-dimensional data encoding device generates a bitstream including the quantized position information, the quantized first luminance, the quantized first color difference, the first quantization parameter, the second quantization parameter, and the first difference between the second quantization parameter and the third quantization parameter (S5393).
[0802] According to this method, the third quantization parameter is represented in the bitstream by the first difference from the second quantization parameter, thereby improving encoding efficiency.
[0803] For example, the three-dimensional data encoding device further quantizes the reflectance of the attribute information of each of the plurality of three-dimensional points using a fourth quantization parameter. In addition, the generation process generates a bitstream that further includes the quantized reflectance and the fourth quantization parameter.
[0804] For example, in quantization using the second quantization parameter, when quantizing the first luminance of one or more three-dimensional points in each of the multiple subspaces obtained by dividing the target space containing the multiple three-dimensional points, the fifth quantization parameter is further used to quantize the first luminance of one or more three-dimensional points in the subspace. In quantization using the third quantization parameter, when quantizing the first color difference of one or more three-dimensional points, the sixth quantization parameter is further used to quantize the first color difference of one or more three-dimensional points. The generation further generates a bitstream that includes the second difference between the second quantization parameter and the fifth quantization parameter, and the third difference between the third quantization parameter and the sixth quantization parameter.
[0805] According to this, in the bitstream, the fifth quantization parameter is represented by the second difference from the second quantization parameter, and the sixth quantization parameter is represented by the third difference from the third quantization parameter, thereby improving encoding efficiency.
[0806] For example, in the generation process, when quantization is performed using the fifth and sixth quantization parameters in addition to the second and third quantization parameters, a bitstream is generated that further includes identification information indicating that quantization was performed using the fifth and sixth quantization parameters.
[0807] According to this, a three-dimensional data decoding device that has acquired a bitstream can determine that it has been quantized using the fifth and sixth quantization parameters using identification information, thereby reducing the processing load of the decoding process.
[0808] For example, the three-dimensional data encoding device further quantizes the second luminance and second color difference, which represent the second color among the attribute information of each of the plurality of three-dimensional points, using a seventh quantization parameter and a seventh quantization parameter to quantize the second luminance and the eighth quantization parameter. The generation further generates a bitstream that includes the quantized second luminance, the quantized second color difference, the seventh quantization parameter, and a fourth difference between the seventh quantization parameter and the eighth quantization parameter.
[0809] According to this method, the eighth quantization parameter is represented in the bitstream by the fourth difference from the seventh quantization parameter, thereby improving encoding efficiency. Furthermore, two types of color information can be included in the attribute information of a three-dimensional point.
[0810] For example, a three-dimensional data encoding device comprises a processor and memory, and the processor uses the memory to perform the above processing.
[0811] Furthermore, the three-dimensional data decoding device according to this embodiment performs the processing shown in Figure 123. First, the three-dimensional data decoding device acquires a bitstream to obtain quantized position information, quantized first luminance, quantized first color difference, a first quan...
Claims
1. A three-dimensional data encoding method comprising: encoding a plurality of pieces of attribute information possessed by each of a plurality of three-dimensional points using parameters; generating a bit stream including the encoded plurality of pieces of attribute information, control information, and a plurality of pieces of first attribute control information; the control information including a plurality of pieces of type information corresponding to the plurality of pieces of attribute information and each piece of type information indicating a different type of attribute information; the plurality of pieces of first attribute control information corresponding to the plurality of pieces of attribute information, respectively; and each piece of first attribute control information including first identification information indicating that it is associated with one of the plurality of type information.
2. The three-dimensional data encoding method according to claim 1, wherein the plurality of type information are stored in a predetermined order in the control information, and the first identification information indicates that the first attribute control information including the first identification information is associated with type information in one of the predetermined orders.
3. The three-dimensional data encoding method according to claim 1 or 2, wherein the bit stream further includes a plurality of pieces of second attribute control information corresponding to the plurality of pieces of attribute information, and each of the plurality of pieces of second attribute control information includes a reference value of a parameter used in encoding the corresponding piece of attribute information.
4. The three-dimensional data encoding method according to claim 3, wherein the first attribute control information includes difference information that is a difference between the parameter and the reference value.
5. A three-dimensional data encoding method according to claim 1 or 2, wherein the bit stream further includes a plurality of second attribute control information corresponding to the plurality of attribute information, and each of the plurality of second attribute control information has second identification information indicating that it corresponds to one of the plurality of type information.
6. A three-dimensional data encoding method according to any one of claims 1 to 5, wherein each of the plurality of first attribute control information has N fields in which N parameters (N is 2 or more) are stored, and in a specific first attribute control information among the plurality of first attribute control information that corresponds to a specific type of attribute, one of the N fields contains a value indicating that it is invalid.
7. The three-dimensional data encoding method according to any one of claims 1 to 6, wherein the encoding quantizes the plurality of pieces of attribute information using a quantization parameter as the parameter.
8. A three-dimensional data decoding method comprising: acquiring a bitstream to acquire encoded multiple pieces of attribute information and parameters; decoding the encoded multiple pieces of attribute information using the parameters to decode the multiple pieces of attribute information possessed by each of multiple three-dimensional points; the bitstream including control information and multiple pieces of first attribute control information; the control information including multiple pieces of type information corresponding to the multiple pieces of attribute information and each piece of type information indicating a different type of attribute information; the multiple pieces of first attribute control information corresponding to the multiple pieces of attribute information, respectively; and each piece of first attribute control information including first identification information indicating that it is associated with one of the multiple pieces of type information.
9. A three-dimensional data decoding method according to claim 8, wherein the plurality of type information are stored in a predetermined order in the control information, and the first identification information indicates that the first attribute control information including the first identification information is associated with type information in one of the predetermined orders.
10. A three-dimensional data decoding method according to claim 8 or 9, wherein the bit stream further includes a plurality of second attribute control information corresponding to the plurality of attribute information, and each of the plurality of second attribute control information includes a reference value of a parameter used in encoding the corresponding attribute information.
11. The three-dimensional data decoding method according to claim 10, wherein the first attribute control information includes difference information that is a difference between the parameter and the reference value.
12. A three-dimensional data decoding method as described in claim 8 or 9, wherein the bit stream further includes a plurality of second attribute control information corresponding to the plurality of attribute information, and each of the plurality of second attribute control information has second identification information indicating that it corresponds to one of the plurality of type information.
13. A three-dimensional data decoding method according to any one of claims 8 to 12, wherein each of the plurality of first attribute control information has a plurality of fields in which a plurality of parameters are stored, and wherein the decoding ignores parameters stored in a specific field of the plurality of fields of specific first attribute control information corresponding to a specific type of attribute among the plurality of first attribute control information.
14. A three-dimensional data decoding method according to any one of claims 8 to 13, wherein the decoding involves dequantizing the encoded plurality of pieces of attribute information using a quantization parameter as the parameter.
15. A three-dimensional data encoding device comprising: a processor; and a memory, wherein the processor uses the memory to encode a plurality of pieces of attribute information possessed by each of a plurality of three-dimensional points using parameters; and generates a bit stream including the encoded plurality of pieces of attribute information, control information, and a plurality of pieces of first attribute control information, wherein the control information includes a plurality of pieces of type information corresponding to the plurality of pieces of attribute information and each piece of type information indicating a different type of attribute information, and wherein the plurality of pieces of first attribute control information (i) respectively correspond to the plurality of pieces of attribute information, and (ii) include first identification information indicating that it is associated with one of the plurality of type information.
16. A three-dimensional data decoding device comprising: a processor; and a memory, wherein the processor uses the memory to acquire a plurality of pieces of attribute information and parameters encoded by acquiring a bit stream; and decodes the plurality of pieces of attribute information possessed by each of a plurality of three-dimensional points by decoding the encoded plurality of pieces of attribute information using the parameters; wherein the bit stream includes control information and a plurality of pieces of first attribute control information; the control information includes a plurality of pieces of type information corresponding to the plurality of pieces of attribute information and each piece of type information indicating a different type of attribute information; and the plurality of pieces of first attribute control information (i) respectively correspond to the plurality of pieces of attribute information, and (ii) include first identification information indicating that it is associated with one of the plurality of type information.