Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, three-dimensional data decoding device, and program

By dividing three-dimensional data into sub-spaces and applying specific movement amounts, the method addresses inefficiencies in encoding and decoding, enhancing storage and transmission efficiency.

JP2026086825APending Publication Date: 2026-05-26PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
Filing Date
2026-02-19
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing three-dimensional data encoding and decoding methods are inefficient, particularly in handling large point cloud data, which hinders effective storage and transmission.

Method used

The method involves dividing three-dimensional data into sub-spaces and applying common and individual movement amounts to improve encoding efficiency, utilizing techniques such as octree division and bitstream encoding to reduce positional information.

Benefits of technology

This approach enhances encoding efficiency by reducing the amount of positional information, enabling effective storage and transmission of three-dimensional data with improved decoding accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026086825000001_ABST
    Figure 2026086825000001_ABST
Patent Text Reader

Abstract

This provides a three-dimensional data encoding method that can improve encoding efficiency. [Solution] A three-dimensional data encoding method for encoding three-dimensional data, wherein each of the three-dimensional data is generated into a plurality of sub-three-dimensional data containing positional information of a part of the three-dimensional data, a common displacement amount is calculated for the plurality of sub-three-dimensional data, a plurality of individual displacement amounts corresponding to each of the plurality of sub-three-dimensional data is calculated, each of the plurality of sub-three-dimensional data is moved using the common displacement amount and the corresponding individual displacement amount, each of the plurality of sub-three-dimensional data moved using the common displacement amount and the corresponding individual displacement amount is encoded, wherein the common displacement amount is a movement of the same distance for the plurality of sub-three-dimensional data and is a movement in the direction from the first point to the second point in three-dimensional space, and the individual displacement amounts are different distances for the plurality of sub-three-dimensional data and are movements in the direction from the second point to the third point in three-dimensional space.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates to a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, a three-dimensional data decoding device, and a program. [Background technology]

[0002] In the future, devices and services utilizing three-dimensional data are expected to become widespread in a wide range of fields, including computer vision for autonomous operation of automobiles or robots, map information, monitoring, infrastructure inspection, and video distribution. Three-dimensional data can be acquired in various ways, such as using distance sensors like rangefinders, stereo cameras, or combinations of multiple monocular cameras.

[0003] One method of representing three-dimensional data is called a point cloud, which represents the shape of a three-dimensional structure using a cloud of points in three-dimensional space. In a point cloud, the position and color of the points are stored. Point clouds are expected to become the mainstream method of representing three-dimensional data, but point clouds are extremely large in size. Therefore, in the storage or transmission of three-dimensional data, data compression through encoding is essential, just as with two-dimensional moving images (for example, MPEG-4 AVC or HEVC, which are standardized by MPEG).

[0004] Furthermore, point cloud compression is partially supported by publicly available libraries (such as the Point Cloud Library) that handle point cloud-related processing.

[0005] Furthermore, there is a known technique for searching for and displaying facilities located around a vehicle using three-dimensional map data (see, for example, Patent Document 1). [Prior art documents] [Patent Documents]

[0006] [Patent Document 1] International Publication No. 2014 / 020663

Summary of the Invention

Problems to be Solved by the Invention

[0007] In the encoding process and decoding process of three-dimensional data, it is desired to improve the encoding efficiency.

[0008] The present disclosure aims to provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding apparatus, or a three-dimensional data decoding apparatus that can improve the encoding efficiency.

Means for Solving the Problems

[0009] A three-dimensional data encoding method according to an aspect of the present disclosure is a three-dimensional data encoding method executed by a three-dimensional data encoding apparatus, which generates a plurality of sub-three-dimensional data each including position information of a part of the three-dimensional data, calculates a common movement amount for the plurality of sub-three-dimensional data, calculates a plurality of individual movement amounts respectively corresponding to the plurality of sub-three-dimensional data, moves each of the plurality of sub-three-dimensional data using the common movement amount and the corresponding individual movement amount, encodes each of the plurality of sub-three-dimensional data moved using the common movement amount and the corresponding individual movement amount, the common movement amount is a movement of the same distance in the plurality of sub-three-dimensional data and is a movement in a direction from a first point to a second point in a three-dimensional space, and the individual movement amount is a different distance in the plurality of sub-three-dimensional data and is a movement in a direction from the second point to a third point in the three-dimensional space.

[0010] A three-dimensional data decoding method according to an aspect of the present disclosure is a three-dimensional data decoding method executed by a three-dimensional data decoding apparatus, including: obtaining common movement information, a plurality of individual movement information, and a plurality of sub-three-dimensional data each including position information of a part of the three-dimensional data; decoding the plurality of sub-three-dimensional data to obtain a plurality of decoded position information; calculating a common movement amount using the common movement information; calculating an individual movement amount corresponding to each of the plurality of sub-three-dimensional data using each of the plurality of individual movement information; and moving each of the plurality of sub-three-dimensional data using the common movement amount and the corresponding individual movement amount.

[0011] A three-dimensional data encoding method according to an aspect of the present disclosure is a three-dimensional data encoding method for encoding three-dimensional data, including: generating a plurality of sub-three-dimensional data each including position information of a part of the three-dimensional data; calculating a common movement amount for the plurality of sub-three-dimensional data; calculating a plurality of individual movement amounts corresponding to each of the plurality of sub-three-dimensional data; moving each of the plurality of sub-three-dimensional data using the common movement amount and the corresponding individual movement amount; and encoding each of the plurality of sub-three-dimensional data moved using the common movement amount and the corresponding individual movement amount.

[0012] A three-dimensional data decoding method according to an aspect of the present disclosure includes: obtaining common movement information, a plurality of individual movement information, and a plurality of sub-three-dimensional data each including position information of a part of the three-dimensional data; decoding the plurality of sub-three-dimensional data to obtain a plurality of decoded position information; calculating a common movement amount using the common movement information; calculating an individual movement amount corresponding to each of the plurality of sub-three-dimensional data using each of the plurality of individual movement information; moving each of the plurality of sub-three-dimensional data using the common movement amount and the corresponding individual movement amount; wherein the common movement amount is a movement of the same distance in the plurality of sub-three-dimensional data and is a movement in a direction from a first point to a second point in three-dimensional space, and the individual movement amount is a different distance in the plurality of sub-three-dimensional data and is a movement in a direction from the second point to a third point in the three-dimensional space.

[0013] A three-dimensional data encoding method according to one aspect of the present disclosure is a three-dimensional data encoding method for encoding point cloud data indicating a plurality of three-dimensional positions in a three-dimensional space, wherein the three-dimensional space is divided into a plurality of subspaces, thereby dividing the point cloud data into a plurality of sub-point cloud data, a common displacement amount for the plurality of sub-point cloud data is calculated, the common displacement amount is a displacement of the same distance between the plurality of sub-point cloud data and indicates a displacement toward a predetermined point in the three-dimensional space, the plurality of sub-point cloud data is moved by the common displacement amount, a plurality of individual displacement amounts corresponding to each of the plurality of sub-point cloud data moved by the common displacement amount is calculated, the plurality of individual displacement amounts are a displacement of different distances between the plurality of sub-point cloud data moved by the common displacement amount and indicates a displacement toward the predetermined point, for each of the plurality of sub-point cloud data moved by the common displacement amount, the sub-point cloud data is moved by the individual displacement amount corresponding to the sub-point cloud data, and the plurality of sub-point cloud data moved by the corresponding individual displacement amount are encoded.

[0014] A three-dimensional data decoding method according to one aspect of the present disclosure involves decoding from a bitstream a plurality of sub-point cloud data obtained by dividing a three-dimensional space into a plurality of sub-spaces, wherein each sub-point cloud data is moved by a common amount and a corresponding individual amount, along with common movement information for calculating the common amount and a plurality of individual movement information for calculating the respective individual amounts obtained by moving the plurality of sub-point cloud data. The common amount represents movement of the same distance between the plurality of sub-point cloud data and movement toward a predetermined point in the three-dimensional space, and the plurality of individual amounts obtained by moving the plurality of sub-point cloud data represent movement of different distances between the plurality of sub-point cloud data moved by the common amount and movement toward the predetermined point. The method then recovers the point cloud data by moving each of the plurality of sub-point cloud data by an amount equal to the sum of the common amount and the individual amount obtained by moving the sub-point cloud data.

[0015] A three-dimensional data encoding method according to one aspect of the present disclosure is a three-dimensional data encoding method for encoding point cloud data indicating a plurality of three-dimensional positions in a three-dimensional space, wherein the point cloud data is moved by a first displacement amount, the three-dimensional space is divided into a plurality of sub-spaces to divide the point cloud data into a plurality of sub-point cloud data, each of the plurality of sub-point cloud data included in the point cloud data after being moved by the first displacement amount is moved by a second displacement amount based on the position in the sub-space containing the sub-point cloud data, and a bitstream is generated by encoding the plurality of sub-point cloud data after the movement, wherein the bitstream includes first displacement information for calculating the first displacement amount and a plurality of second displacement amounts obtained by moving the plurality of sub-point cloud data.

[0016] A three-dimensional data decoding method according to one aspect of the present disclosure involves decoding from a bitstream a plurality of sub-point cloud data obtained by dividing a three-dimensional space into a plurality of sub-spaces, each of which is a plurality of sub-point cloud data obtained by dividing point cloud data indicating a plurality of three-dimensional positions, wherein each sub-point cloud data is moved by a first movement amount and a corresponding second movement amount, along with first movement information for calculating the first movement amount and a plurality of second movement information for calculating the plurality of second movement amounts obtained by moving the plurality of sub-point cloud data, and then restoring the point cloud data by moving each of the plurality of sub-point cloud data by a movement amount obtained by adding the first movement amount and the corresponding second movement amount. [Effects of the Invention]

[0017] This disclosure provides a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device that can improve encoding efficiency. [Brief explanation of the drawing]

[0018] [Figure 1] Figure 1 shows the configuration of a three-dimensional data encoding and decoding system according to Embodiment 1. [Figure 2] Figure 2 shows an example of the configuration of point cloud data according to Embodiment 1. [Figure 3] Figure 3 shows an example of the configuration of a data file containing point cloud data information according to Embodiment 1. [Figure 4] Figure 4 is a diagram showing the types of point cloud data according to Embodiment 1. [Figure 5] Figure 5 shows the configuration of the first encoding unit according to Embodiment 1. [Figure 6] Figure 6 is a block diagram of the first encoding unit according to Embodiment 1. [Figure 7] Figure 7 shows the configuration of the first decoding unit according to Embodiment 1. [Figure 8] Figure 8 is a block diagram of the first decoding unit according to Embodiment 1. [Figure 9]Figure 9 shows the configuration of the second encoding unit according to Embodiment 1. [Figure 10] Figure 10 is a block diagram of the second encoding unit according to Embodiment 1. [Figure 11] Figure 11 shows the configuration of the second decoding unit according to Embodiment 1. [Figure 12] Figure 12 is a block diagram of the second decoding unit according to Embodiment 1. [Figure 13] Figure 13 is a diagram showing the protocol stack related to PCC encoded data according to Embodiment 1. [Figure 14] Figure 14 shows the basic structure of ISOBMFF according to Embodiment 2. [Figure 15] Figure 15 is a diagram showing the protocol stack according to Embodiment 2. [Figure 16] Figure 16 shows an example of storing the NAL unit according to Embodiment 2 in a file for codec 1. [Figure 17] Figure 17 shows an example of storing the NAL unit according to Embodiment 2 in a file for codec 2. [Figure 18] Figure 18 shows the configuration of the first multiplexing unit according to Embodiment 2. [Figure 19] Figure 19 shows the configuration of the first demultiplexing unit according to Embodiment 2. [Figure 20] Figure 20 shows the configuration of the second multiplexing unit according to Embodiment 2. [Figure 21] Figure 21 is a diagram showing the configuration of the second demultiplexing unit according to Embodiment 2. [Figure 22] Figure 22 is a flowchart of the processing performed by the first multiplexing unit according to Embodiment 2. [Figure 23] Figure 23 is a flowchart of the processing performed by the second multiplexing unit according to Embodiment 2. [Figure 24] Figure 24 is a flowchart of the processing performed by the first demultiplexing unit and the first decoding unit according to Embodiment 2. [Figure 25]Figure 25 is a flowchart of the processing performed by the second demultiplexing unit and the second decoding unit according to Embodiment 2. [Figure 26] Figure 26 shows the configuration of the encoding unit and the third multiplexing unit according to Embodiment 3. [Figure 27] Figure 27 shows the configuration of the third demultiplexing unit and decoding unit according to Embodiment 3. [Figure 28] Figure 28 is a flowchart of the processing performed by the third multiplexing unit according to Embodiment 3. [Figure 29] Figure 29 is a flowchart of the processing performed by the third demultiplexing unit and decoding unit according to Embodiment 3. [Figure 30] Figure 30 is a flowchart of the processing performed by the three-dimensional data storage device according to Embodiment 3. [Figure 31] Figure 31 is a flowchart of the processing performed by the three-dimensional data acquisition device according to Embodiment 3. [Figure 32] Figure 32 shows the configuration of the encoding unit and the multiplexing unit according to Embodiment 4. [Figure 33] Figure 33 shows an example of the structure of encoded data according to Embodiment 4. [Figure 34] Figure 34 shows an example of the configuration of encoded data and NAL unit according to Embodiment 4. [Figure 35] Figure 35 shows an example of the semantics of pcc_nal_unit_type according to Embodiment 4. [Figure 36] Figure 36 shows an example of the transmission sequence of the NAL unit according to Embodiment 4. [Figure 37] Figure 37 shows an example of slice and tile division according to Embodiment 5. [Figure 38] Figure 38 shows an example of a slice and tile division pattern according to Embodiment 5. [Figure 39] Figure 39 is a block diagram of the first encoding unit according to Embodiment 6. [Figure 40]Figure 40 is a block diagram of the first decoding unit according to Embodiment 6. [Figure 41] Figure 41 shows an example of the tile shape according to Embodiment 6. [Figure 42] Figure 42 shows examples of tiles and slices according to Embodiment 6. [Figure 43] Figure 43 is a block diagram of the divided section according to Embodiment 6. [Figure 44] Figure 44 shows an example of a map of point cloud data viewed from above according to Embodiment 6. [Figure 45] Figure 45 shows an example of tile division according to Embodiment 6. [Figure 46] Figure 46 shows an example of tile division according to Embodiment 6. [Figure 47] Figure 47 shows an example of tile division according to Embodiment 6. [Figure 48] Figure 48 shows an example of tile data stored on the server according to Embodiment 6. [Figure 49] Figure 49 is a diagram showing a system for tile division according to Embodiment 6. [Figure 50] Figure 50 shows an example of slice division according to Embodiment 6. [Figure 51] Figure 51 is a diagram showing an example of a dependency relationship according to Embodiment 6. [Figure 52] Figure 52 shows an example of the data decoding order according to Embodiment 6. [Figure 53] Figure 53 shows an example of encoded tile data according to Embodiment 6. [Figure 54] Figure 54 is a block diagram of the joint according to Embodiment 6. [Figure 55] Figure 55 shows an example of the configuration of encoded data and a NAL unit according to Embodiment 6. [Figure 56] Figure 56 is a flowchart of the encoding process according to Embodiment 6. [Figure 57]Figure 57 is a flowchart of the decoding process according to Embodiment 6. [Figure 58] Figure 58 shows an example of the syntax for tile addition information according to Embodiment 6. [Figure 59] Figure 59 is a block diagram of the coding and decoding system according to Embodiment 6. [Figure 60] Figure 60 shows an example of the syntax for slice addition information according to Embodiment 6. [Figure 61] Figure 61 is a flowchart of the encoding process according to Embodiment 6. [Figure 62] Figure 62 is a flowchart of the decoding process according to Embodiment 6. [Figure 63] Figure 63 is a flowchart of the encoding process according to Embodiment 6. [Figure 64] Figure 64 is a flowchart of the decoding process according to Embodiment 6. [Figure 65] Figure 65 shows an example of a division method according to Embodiment 7. [Figure 66] Figure 66 shows an example of point cloud data division according to Embodiment 7. [Figure 67] Figure 67 shows an example of the syntax for tile addition information according to Embodiment 7. [Figure 68] Figure 68 shows an example of index information according to Embodiment 7. [Figure 69] Figure 69 is a diagram showing an example of a dependency relationship according to Embodiment 7. [Figure 70] Figure 70 shows an example of transmission data according to Embodiment 7. [Figure 71] Figure 71 shows an example of the configuration of the NAL unit according to Embodiment 7. [Figure 72] Figure 72 shows an example of a dependency relationship according to Embodiment 7. [Figure 73] Figure 73 shows an example of the data decoding order according to Embodiment 7. [Figure 74]Figure 74 shows an example of a dependency relationship according to Embodiment 7. [Figure 75] Figure 75 shows an example of the data decoding order according to Embodiment 7. [Figure 76] Figure 76 is a flowchart of the encoding process according to Embodiment 7. [Figure 77] Figure 77 is a flowchart of the decoding process according to Embodiment 7. [Figure 78] Figure 78 is a flowchart of the encoding process according to Embodiment 7. [Figure 79] Figure 79 is a flowchart of the encoding process according to Embodiment 7. [Figure 80] Figure 80 shows examples of transmitted and received data according to Embodiment 7. [Figure 81] Figure 81 is a flowchart of the decoding process according to Embodiment 7. [Figure 82] Figure 82 shows examples of transmitted and received data according to Embodiment 7. [Figure 83] Figure 83 is a flowchart of the decoding process according to Embodiment 7. [Figure 84] Figure 84 is a flowchart of the encoding process according to Embodiment 7. [Figure 85] Figure 85 shows an example of index information according to Embodiment 7. [Figure 86] Figure 86 is a diagram showing an example of a dependency relationship according to Embodiment 7. [Figure 87] Figure 87 shows an example of transmission data according to Embodiment 7. [Figure 88] Figure 88 shows examples of transmitted and received data according to Embodiment 7. [Figure 89] Figure 89 is a flowchart of the decoding process according to Embodiment 7. [Figure 90] Figure 90 is a flowchart of the encoding process according to Embodiment 7. [Figure 91]Figure 91 is a flowchart of the decoding process according to Embodiment 7. [Figure 92] Figure 92 is a block diagram showing an example of the configuration of a three-dimensional data encoding device according to Embodiment 8. [Figure 93] Figure 93 is a diagram illustrating the schematic of the encoding method using the three-dimensional data encoding device according to Embodiment 8. [Figure 94] Figure 94 is a diagram illustrating a first example of position shift according to Embodiment 8. [Figure 95] Figure 95 is a diagram illustrating a second example of position shift according to Embodiment 8. [Figure 96] Figure 96 is a flowchart showing an example of an encoding method according to Embodiment 8. [Figure 97] Figure 97 is a flowchart showing an example of a decoding method according to Embodiment 8. [Figure 98] Figure 98 is a diagram illustrating a third example of position shift according to Embodiment 8. [Figure 99] Figure 99 is a flowchart showing an example of an encoding method according to Embodiment 8. [Figure 100] Figure 100 is a flowchart showing an example of a decoding method according to Embodiment 8. [Figure 101] Figure 101 is a diagram illustrating a fourth example of position shift according to Embodiment 8. [Figure 102] Figure 102 is a flowchart showing an example of an encoding method according to Embodiment 8. [Figure 103] Figure 103 is a diagram illustrating a fifth example of position shift according to Embodiment 8. [Figure 104] Figure 104 is a diagram illustrating the encoding method according to Embodiment 8. [Figure 105] Figure 105 shows an example of GPS syntax according to Embodiment 8. [Figure 106] Figure 106 shows an example of the syntax of the location information header according to Embodiment 8. [Figure 107] Figure 107 is a flowchart showing an example of an encoding method for switching the processing according to Embodiment 8. [Figure 108] Figure 108 is a flowchart showing an example of a decoding method for switching the processing according to Embodiment 8. [Figure 109] Figure 109 is a diagram showing an example of the data structure of a bitstream according to Embodiment 8. [Figure 110] Figure 110 shows an example where the divided data of Figure 109 according to Embodiment 8 is used as a frame. [Figure 111] Figure 111 shows another example of a divided region according to Embodiment 8. [Figure 112] Figure 112 shows another example of a divided region according to Embodiment 8. [Figure 113] Figure 113 shows another example of a divided region according to Embodiment 8. [Figure 114] Figure 114 shows another example of a divided region according to Embodiment 8. [Figure 115] Figure 115 shows another example of the data configuration according to Embodiment 8. [Figure 116] Figure 116 shows another example of the data configuration according to Embodiment 8. [Figure 117] Figure 117 shows another example of the data configuration according to Embodiment 8. [Figure 118] Figure 118 shows another example of the data configuration according to Embodiment 8. [Figure 119] Figure 119 shows another example of the data configuration according to Embodiment 8. [Figure 120] Figure 120 is a flowchart of the encoding process according to Embodiment 8. [Figure 121] Figure 121 is a flowchart of the decoding process according to Embodiment 8. [Modes for carrying out the invention]

[0019] A three-dimensional data encoding method according to one aspect of the present disclosure is a three-dimensional data encoding method for encoding point cloud data indicating a plurality of three-dimensional positions in a three-dimensional space, wherein the point cloud data is moved by a first displacement amount, the three-dimensional space is divided into a plurality of sub-spaces to divide the point cloud data into a plurality of sub-point cloud data, each of the plurality of sub-point cloud data included in the point cloud data after being moved by the first displacement amount is moved by a second displacement amount based on the position in the sub-space containing the sub-point cloud data, and a bitstream is generated by encoding the plurality of sub-point cloud data after the movement, wherein the bitstream includes first displacement information for calculating the first displacement amount and a plurality of second displacement amounts obtained by moving the plurality of sub-point cloud data.

[0020] According to this method, since the divided sub-point cloud data is moved before encoding, the amount of positional information in each sub-point cloud data can be reduced, thereby improving encoding efficiency.

[0021] For example, the plurality of subspaces may be of equal size to one another, and each of the plurality of second movement information may include the number of the plurality of subspaces and first identification information for identifying the corresponding subspace.

[0022] Therefore, the amount of information in the second movement information can be reduced, and encoding efficiency can be improved.

[0023] For example, the first identification information may be in Morton order corresponding to each of the plurality of subspaces.

[0024] For example, each of the multiple subspaces may be a space obtained by dividing a single three-dimensional space using an octree, and the bitstream may include a second identification information indicating that the multiple subspaces are spaces obtained by dividing the space using an octree, and depth information indicating the depth of the octree.

[0025] Therefore, by dividing the point cloud data in three-dimensional space using an octave tree, the amount of positional information in each sub-point cloud data can be reduced, thereby improving encoding efficiency.

[0026] For example, the division may be performed after moving the point cloud data by the first displacement amount.

[0027] Furthermore, a three-dimensional data decoding method according to one aspect of the present disclosure may also be used to reconstruct the point cloud data by decoding from a bitstream a plurality of sub-point cloud data obtained by dividing a three-dimensional space into a plurality of sub-spaces, wherein each of the sub-point cloud data is moved by a first movement amount and a corresponding second movement amount, a first movement information for calculating the first movement amount, and a plurality of second movement information for calculating each of the plurality of second movement amounts obtained by moving the plurality of sub-point cloud data, and then moving each of the plurality of sub-point cloud data by a movement amount obtained by adding the first movement amount and the corresponding second movement amount.

[0028] According to this, point cloud data can be correctly decoded using a bitstream with improved encoding efficiency.

[0029] For example, the plurality of subspaces may be of equal size to one another, and each of the plurality of second movement information may include the number of the plurality of subspaces and first identification information for identifying the corresponding subspace.

[0030] For example, the first identification information may be in Morton order corresponding to each of the plurality of subspaces.

[0031] For example, each of the multiple subspaces may be a space obtained by dividing a single three-dimensional space using an octree, and the bitstream may include a second identification information indicating that the multiple subspaces are spaces obtained by dividing the space using an octree, and depth information indicating the depth of the octree.

[0032] Furthermore, a three-dimensional data encoding device according to one aspect of the present disclosure is a three-dimensional data encoding device for encoding point cloud data indicating a plurality of three-dimensional positions in a three-dimensional space, comprising a processor and a memory, wherein the processor uses the memory to move the point cloud data by a first displacement amount, divides the three-dimensional space into a plurality of subspaces, thereby dividing the point cloud data into a plurality of sub-point cloud data, moves each of the plurality of sub-point cloud data included in the point cloud data after moving by the first displacement amount by a second displacement amount based on the position in the subspace containing the sub-point cloud data, and generates a bitstream by encoding the plurality of sub-point cloud data after the movement, wherein the bitstream includes first displacement information for calculating the first displacement amount and a plurality of second displacement information for calculating a plurality of second displacement amounts obtained by moving the plurality of sub-point cloud data.

[0033] According to this method, since the divided sub-point cloud data is moved before encoding, the amount of positional information in each sub-point cloud data can be reduced, thereby improving encoding efficiency.

[0034] Furthermore, a three-dimensional data decoding device according to one aspect of the present disclosure comprises a processor and a memory, wherein the processor uses the memory to decode from a bitstream a plurality of sub-point cloud data obtained by dividing a three-dimensional space into a plurality of sub-spaces, wherein each sub-point cloud data is moved by a first movement amount and a corresponding second movement amount, a first movement information for calculating the first movement amount, and a plurality of second movement information for calculating the plurality of second movement amounts obtained by moving the plurality of sub-point cloud data, and restores the point cloud data by moving each of the plurality of sub-point cloud data by a movement amount obtained by adding the first movement amount and the corresponding second movement amount.

[0035] According to this, point cloud data can be correctly decoded using a bitstream with improved encoding efficiency.

[0036] These comprehensive or specific embodiments may be implemented as a system, method, integrated circuit, computer program, or recording medium such as a computer-readable CD-ROM, or as any combination of a system, method, integrated circuit, computer program, and recording medium.

[0037] The embodiments will be described in detail below with reference to the drawings. Note that the embodiments described below are all specific examples of this disclosure. The numerical values, shapes, materials, components, arrangement and connection configurations of components, steps, and the order of steps shown in the following embodiments are examples only and are not intended to limit this disclosure. Furthermore, among the components in the following embodiments, those not described in the independent claim representing the highest-level concept will be described as optional components.

[0038] (Embodiment 1) When using encoded point cloud data in actual devices or services, it is desirable to send and receive necessary information depending on the application in order to reduce network bandwidth. However, until now, such functionality has not existed in the encoded structure of three-dimensional data, nor has there been an encoding method for that purpose.

[0039] This embodiment describes a three-dimensional data encoding method and a three-dimensional data encoding device for providing a function to send and receive information necessary for use in encoded data of a three-dimensional point cloud, a three-dimensional data decoding method and a three-dimensional data decoding device for decoding the encoded data, a three-dimensional data multiplexing method for multiplexing the encoded data, and a three-dimensional data transmission method for transmitting the encoded data.

[0040] In particular, while two encoding methods (encoding schemes) for point cloud data are currently being considered, the structure of the encoded data and the method for storing the encoded data in a system format have not been defined. As a result, there is a problem in that MUX processing (multiplexing), transmission, or storage cannot be performed in the encoding unit.

[0041] Furthermore, there has been no existing method to support formats like PCC (Point Cloud Compression) that use a mixture of two codecs, a first encoding method and a second encoding method.

[0042] This embodiment describes the structure of PCC encoded data in which two codecs, a first encoding method and a second encoding method, coexist, and a method for storing the encoded data in a system format.

[0043] First, the configuration of the three-dimensional data (point cloud data) encoding and decoding system according to this embodiment will be described. Figure 1 is a diagram showing an example of the configuration of the three-dimensional data encoding and decoding system according to this embodiment. As shown in Figure 1, the three-dimensional data encoding and decoding system includes a three-dimensional data encoding system 4601, a three-dimensional data decoding system 4602, a sensor terminal 4603, and an external connection unit 4604.

[0044] The three-dimensional data encoding system 4601 generates encoded data or multiplexed data by encoding point cloud data, which is three-dimensional data. The three-dimensional data encoding system 4601 may be a three-dimensional data encoding device implemented by a single device, or it may be a system implemented by multiple devices. Furthermore, the three-dimensional data encoding device may include some of the multiple processing units included in the three-dimensional data encoding system 4601.

[0045] The three-dimensional data encoding system 4601 includes a point cloud data generation system 4611, a presentation unit 4612, an encoding unit 4613, a multiplexing unit 4614, an input / output unit 4615, and a control unit 4616. The point cloud data generation system 4611 includes a sensor information acquisition unit 4617 and a point cloud data generation unit 4618.

[0046] The sensor information acquisition unit 4617 acquires sensor information from the sensor terminal 4603 and outputs the sensor information to the point cloud data generation unit 4618. The point cloud data generation unit 4618 generates point cloud data from the sensor information and outputs the point cloud data to the encoding unit 4613.

[0047] The display unit 4612 presents sensor information or point cloud data to the user. For example, the display unit 4612 displays information or images based on sensor information or point cloud data.

[0048] The encoding unit 4613 encodes (compresses) the point cloud data and outputs the resulting encoded data, control information obtained during the encoding process, and other additional information to the multiplexing unit 4614. The additional information includes, for example, sensor information.

[0049] The multiplexing unit 4614 generates multiplexed data by multiplexing the encoded data input from the encoding unit 4613, control information, and additional information. The format of the multiplexed data is, for example, a file format for storage or a packet format for transmission.

[0050] The input / output unit 4615 (for example, the communication unit or interface) outputs the multiplexed data to the outside. Alternatively, the multiplexed data is stored in a storage unit such as internal memory. The control unit 4616 (or application execution unit) controls each processing unit. In other words, the control unit 4616 performs control such as encoding and multiplexing.

[0051] The sensor information may also be input to the encoding unit 4613 or the multiplexing unit 4614. Furthermore, the input / output unit 4615 may output the point cloud data or encoded data directly to the outside.

[0052] The transmission signal (multiplexed data) output from the three-dimensional data encoding system 4601 is input to the three-dimensional data decoding system 4602 via the external connection unit 4604.

[0053] The three-dimensional data decoding system 4602 generates point cloud data, which is three-dimensional data, by decoding encoded data or multiplexed data. The three-dimensional data decoding system 4602 may be a three-dimensional data decoding device implemented by a single device, or it may be a system implemented by multiple devices. Furthermore, the three-dimensional data decoding device may include some of the multiple processing units included in the three-dimensional data decoding system 4602.

[0054] The three-dimensional data decoding system 4602 includes a sensor information acquisition unit 4621, an input / output unit 4622, a demultiplexing unit 4623, a decoding unit 4624, a presentation unit 4625, a user interface 4626, and a control unit 4627.

[0055] The sensor information acquisition unit 4621 acquires sensor information from the sensor terminal 4603.

[0056] The input / output unit 4622 acquires the transmission signal, decodes the multiplexed data (file format or packet) from the transmission signal, and outputs the multiplexed data to the demultiplexing unit 4623.

[0057] The demultiplexing unit 4623 acquires encoded data, control information, and additional information from the multiplexed data, and outputs the encoded data, control information, and additional information to the decoding unit 4624.

[0058] The decoding unit 4624 reconstructs the point cloud data by decoding the encoded data.

[0059] The presentation unit 4625 presents point cloud data to the user. For example, the presentation unit 4625 displays information or images based on the point cloud data. The user interface 4626 acquires instructions based on user operations. The control unit 4627 (or application execution unit) controls each processing unit. In other words, the control unit 4627 performs control such as demultiplexing, decoding, and presentation.

[0060] The input / output unit 4622 may acquire point cloud data or encoded data directly from an external source. The presentation unit 4625 may acquire additional information such as sensor information and present information based on that additional information. The presentation unit 4625 may also make presentations based on user instructions acquired through the user interface 4626.

[0061] The sensor terminal 4603 generates sensor information, which is information obtained from the sensor. The sensor terminal 4603 is a terminal equipped with a sensor or camera, and may be, for example, a mobile object such as an automobile, an aerial object such as an airplane, a mobile terminal, or a camera.

[0062] The sensor information that can be acquired by the sensor terminal 4603 includes, for example, (1) the distance between the sensor terminal 4603 and the object, or the reflectivity of the object, obtained from a LiDAR, millimeter-wave radar, or infrared sensor, and (2) the distance between the camera and the object, or the reflectivity of the object, obtained from multiple monocular camera images or stereo camera images. The sensor information may also include the sensor's attitude, orientation, gyroscope (angular velocity), position (GPS information or altitude), speed, or acceleration. The sensor information may also include temperature, atmospheric pressure, humidity, or magnetism.

[0063] The external connection unit 4604 is implemented by an integrated circuit (LSI or IC), an external storage unit, communication with a cloud server via the internet, or broadcasting, etc.

[0064] Next, we will explain point cloud data. Figure 2 shows the structure of point cloud data. Figure 3 shows an example of the structure of a data file containing information about point cloud data.

[0065] Point cloud data contains data for multiple points. Each point's data includes location information (three-dimensional coordinates) and attribute information related to that location. A collection of these points is called a point cloud. For example, a point cloud represents the three-dimensional shape of an object.

[0066] Position information, such as three-dimensional coordinates, is sometimes referred to as geometry. Furthermore, the data for each point may include attribute information of multiple attribute types. Attribute types include, for example, color or reflectance.

[0067] One location information may be associated with one attribute information, or multiple attribute information of different attribute types may be associated with one location information. Furthermore, multiple attribute information of the same attribute type may be associated with one location information.

[0068] The example data file structure shown in Figure 3 represents a case where location information and attribute information correspond one-to-one, and it shows the location information and attribute information of the N points that make up the point cloud data.

[0069] Location information includes, for example, information for the three axes: x, y, and z. Attribute information includes, for example, RGB color information. A typical data file is a ply file.

[0070] Next, we will explain the types of point cloud data. Figure 4 is a diagram illustrating the types of point cloud data. As shown in Figure 4, point cloud data includes static objects and dynamic objects.

[0071] A static object is three-dimensional point cloud data for any given time (a specific moment). A dynamic object is three-dimensional point cloud data that changes over time. Hereafter, three-dimensional point cloud data for a given time will be referred to as a PCC frame, or simply a frame.

[0072] The object can be a point cloud with a somewhat limited area, like regular video data, or it can be a large-scale point cloud with no area limitations, like map information.

[0073] Furthermore, point cloud data of various densities may exist, including both sparse and dense point cloud data.

[0074] The details of each processing unit are described below. Sensor information is acquired by various methods, such as distance sensors like LIDAR or rangefinders, stereo cameras, or combinations of multiple monocular cameras. The point cloud data generation unit 4618 generates point cloud data based on the sensor information obtained by the sensor information acquisition unit 4617. The point cloud data generation unit 4618 generates position information as point cloud data and adds attribute information to the position information.

[0075] The point cloud data generation unit 4618 may process the point cloud data when generating position information or adding attribute information. For example, the point cloud data generation unit 4618 may reduce the amount of data by deleting point clouds with overlapping positions. The point cloud data generation unit 4618 may also transform the position information (such as position shifting, rotation, or normalization) or render the attribute information.

[0076] In Figure 1, the point cloud data generation system 4611 is included in the three-dimensional data encoding system 4601, but it may also be provided independently outside of the three-dimensional data encoding system 4601.

[0077] The encoding unit 4613 generates encoded data by encoding the point cloud data based on a predetermined encoding method. There are two main types of encoding methods. The first is an encoding method using positional information, which will be referred to as the first encoding method hereafter. The second is an encoding method using a video codec, which will be referred to as the second encoding method hereafter.

[0078] The decoding unit 4624 decodes the point cloud data by decoding the encoded data based on a predetermined encoding method.

[0079] The multiplexing unit 4614 generates multiplexed data by multiplexing the encoded data using an existing multiplexing method. The generated multiplexed data is transmitted or stored. In addition to PCC encoded data, the multiplexing unit 4614 multiplexes other media such as video, audio, subtitles, applications, files, or reference time information. Furthermore, the multiplexing unit 4614 may also multiplex attribute information related to sensor information or point cloud data.

[0080] Multiplexing methods or file formats include ISOBMFF, ISOBMFF-based transmission methods such as MPEG-DASH, MMT, MPEG-2 TS Systems, and RMP.

[0081] The demultiplexing unit 4623 extracts PCC encoded data, other media, and time information from the multiplexed data.

[0082] The input / output unit 4615 transmits the multiplexed data using a method appropriate to the transmission medium or storage medium, such as broadcasting or communication. The input / output unit 4615 may communicate with other devices via the internet or with storage units such as cloud servers.

[0083] Communication protocols such as HTTP, FTP, TCP, or UDP can be used. Either a pull-type or push-type communication method may be employed.

[0084] Either wired or wireless transmission may be used. Wired transmission methods include Ethernet®, USB, RS-232C, HDMI®, or coaxial cable. Wireless transmission methods include wireless LAN, Wi-Fi®, Bluetooth®, or millimeter wave.

[0085] Furthermore, broadcasting formats such as DVB-T2, DVB-S2, DVB-C2, ATSC3.0, or ISDB-S3 may be used.

[0086] Figure 5 shows the configuration of a first encoding unit 4630, which is an example of an encoding unit 4613 that performs encoding using the first encoding method. Figure 6 is a block diagram of the first encoding unit 4630. The first encoding unit 4630 generates encoded data (encoded stream) by encoding point cloud data using the first encoding method. This first encoding unit 4630 includes a location information encoding unit 4631, an attribute information encoding unit 4632, an additional information encoding unit 4633, and a multiplexing unit 4634.

[0087] The first encoding unit 4630 is characterized by performing encoding while being aware of the three-dimensional structure. Furthermore, the first encoding unit 4630 is characterized by the attribute information encoding unit 4632 performing encoding using information obtained from the location information encoding unit 4631. The first encoding method is also called GPCC (Geometry-based PCC).

[0088] The point cloud data is PCC point cloud data such as a PLY file, or PCC point cloud data generated from sensor information, and includes position information, attribute information, and other additional information (metadata). The position information is input to the position information encoding unit 4631, the attribute information is input to the attribute information encoding unit 4632, and the additional information is input to the additional information encoding unit 4633.

[0089] The location information encoding unit 4631 generates encoded location information (Compressed Geometry), which is encoded data, by encoding location information. For example, the location information encoding unit 4631 encodes location information using an N-tree structure such as an octree. Specifically, in an octree, the target space is divided into 8 nodes (subspaces), and 8 bits of information (occupancy code) are generated to indicate whether or not a point cloud is contained in each node. Furthermore, nodes containing point clouds are further divided into 8 nodes, and 8 bits of information are generated to indicate whether or not a point cloud is contained in each of these 8 nodes. This process is repeated until the number of point clouds contained in a predetermined hierarchy or node falls below a threshold.

[0090] The attribute information encoding unit 4632 generates encoded attribute information (Compressed Attribute), which is encoded data, by encoding it using the configuration information generated by the location information encoding unit 4631. For example, the attribute information encoding unit 4632 determines the reference point (reference node) to be referenced in encoding the target point (target node) to be processed, based on the octave tree structure generated by the location information encoding unit 4631. For example, the attribute information encoding unit 4632 references a surrounding node or adjacent node whose parent node in the octave tree is the same as the target node. Note that the method for determining the reference relationship is not limited to this.

[0091] Furthermore, the attribute information encoding process may include at least one of the following: quantization, prediction, and arithmetic encoding. In this case, a reference means using a reference node to calculate the predicted value of the attribute information, or using the state of a reference node (for example, occupancy information indicating whether or not the reference node contains a point cloud) to determine the encoding parameters. For example, encoding parameters may be quantization parameters in the quantization process, or context in arithmetic encoding.

[0092] The additional information encoding unit 4633 generates encoded data, or compressed additional information (Compressed MetaData), by encoding the compressible data from the additional information.

[0093] The multiplexing unit 4634 generates a compressed stream, which is encoded data, by multiplexing encoded position information, encoded attribute information, encoded additional information, and other additional information. The generated compressed stream is output to a processing unit of the system layer (not shown).

[0094] Next, we will describe a first decoding unit 4640, which is an example of a decoding unit 4624 that performs decoding of the first encoding method. Figure 7 is a diagram showing the configuration of the first decoding unit 4640. Figure 8 is a block diagram of the first decoding unit 4640. The first decoding unit 4640 generates point cloud data by decoding the encoded data (encoded stream) encoded by the first encoding method using the first encoding method. This first decoding unit 4640 includes a demultiplexing unit 4641, a location information decoding unit 4642, an attribute information decoding unit 4643, and an additional information decoding unit 4644.

[0095] A compressed stream, which is encoded data, is input to the first decoding unit 4640 from a processing unit of the system layer (not shown).

[0096] The demultiplexing unit 4641 separates encoded location information (Compressed Geometry), encoded attribute information (Compressed Attribute), encoded additional information (Compressed MetaData), and other additional information from the encoded data.

[0097] The location information decoding unit 4642 generates location information by decoding the encoded location information. For example, the location information decoding unit 4642 reconstructs the location information of a point cloud represented by three-dimensional coordinates from encoded location information represented by an N-tree structure such as an octree.

[0098] The attribute information decoding unit 4643 decodes the encoded attribute information based on the configuration information generated by the location information decoding unit 4642. For example, the attribute information decoding unit 4643 determines the reference point (reference node) to be referenced in the decoding of the target point (target node) to be processed, based on the octave tree structure obtained by the location information decoding unit 4642. For example, the attribute information decoding unit 4643 references a surrounding node or adjacent node whose parent node in the octave tree is the same as the target node. Note that the method for determining the reference relationship is not limited to this.

[0099] Furthermore, the attribute information decoding process may include at least one of the following: inverse quantization, prediction, and arithmetic decoding. In this case, "reference" means using a reference node to calculate the predicted value of the attribute information, or using the state of the reference node (for example, occupancy information indicating whether or not the reference node contains a point cloud) to determine the decoding parameters. For example, decoding parameters may be quantization parameters in the inverse quantization process, or context in arithmetic decoding.

[0100] The additional information decoding unit 4644 generates additional information by decoding the encoded additional information. The first decoding unit 4640 uses the additional information necessary for decoding location information and attribute information during decoding and outputs the additional information necessary for the application to the outside.

[0101] Next, we will describe a second encoding unit 4650, which is an example of an encoding unit 4613 that performs encoding using the second encoding method. Figure 9 is a diagram showing the configuration of the second encoding unit 4650. Figure 10 is a block diagram of the second encoding unit 4650.

[0102] The second encoding unit 4650 generates encoded data (encoded stream) by encoding the point cloud data using a second encoding method. This second encoding unit 4650 includes an additional information generation unit 4651, a position image generation unit 4652, an attribute image generation unit 4653, a video encoding unit 4654, an additional information encoding unit 4655, and a multiplexing unit 4656.

[0103] The second encoding unit 4650 generates a position image and an attribute image by projecting a three-dimensional structure onto a two-dimensional image, and then encodes the generated position image and attribute image using an existing video encoding scheme. The second encoding method is also called VPCC (Video based PCC).

[0104] The point cloud data is PCC point cloud data such as a PLY file, or PCC point cloud data generated from sensor information, and includes position information, attribute information, and other additional information (metadata).

[0105] The additional information generation unit 4651 generates map information for multiple two-dimensional images by projecting a three-dimensional structure onto a two-dimensional image.

[0106] The position image generation unit 4652 generates a position image (geometry image) based on position information and map information generated by the additional information generation unit 4651. This position image is, for example, a depth image in which the distance is indicated as a pixel value. This depth image may be an image of multiple point clouds viewed from one viewpoint (an image of multiple point clouds projected onto a single two-dimensional plane), or multiple images of multiple point clouds viewed from multiple viewpoints, or a single image formed by integrating these multiple images.

[0107] The attribute image generation unit 4653 generates an attribute image based on attribute information and map information generated by the additional information generation unit 4651. This attribute image is, for example, an image in which attribute information (e.g., color (RGB)) is shown as pixel values. This image may be an image of multiple point clouds viewed from one viewpoint (an image of multiple point clouds projected onto a single two-dimensional plane), or multiple images of multiple point clouds viewed from multiple viewpoints, or a single image formed by integrating these multiple images.

[0108] The video encoding unit 4654 generates encoded data, namely a compressed geometric image and a compressed attribute image, by encoding the position image and attribute image using a video encoding scheme. Any known encoding scheme may be used as the video encoding scheme. For example, the video encoding scheme may be AVC or HEVC.

[0109] The additional information encoding unit 4655 generates encoded additional information (Compressed MetaData) by encoding additional information and map information included in the point cloud data.

[0110] The multiplexing unit 4656 generates a compressed stream, which is encoded data, by multiplexing the encoded position image, encoded attribute image, encoded additional information, and other additional information. The generated compressed stream is output to a processing unit of the system layer (not shown).

[0111] Next, we will describe a second decoding unit 4660, which is an example of a decoding unit 4624 that performs decoding of the second encoding method. Figure 11 is a diagram showing the configuration of the second decoding unit 4660. Figure 12 is a block diagram of the second decoding unit 4660. The second decoding unit 4660 generates point cloud data by decoding the encoded data (encoded stream) encoded by the second encoding method using the second encoding method. This second decoding unit 4660 includes a demultiplexing unit 4661, a video decoding unit 4662, an additional information decoding unit 4663, a location information generation unit 4664, and an attribute information generation unit 4665.

[0112] A compressed stream, which is encoded data, is input to the second decoding unit 4660 from a processing unit of the system layer (not shown).

[0113] The demultiplexing unit 4661 separates the encoded location image (Compressed Geometry Image), encoded attribute image (Compressed Attribute Image), encoded additional information (Compressed MetaData), and other additional information from the encoded data.

[0114] The video decoding unit 4662 generates a position image and an attribute image by decoding the encoded position image and the encoded attribute image using a video encoding scheme. Any known encoding scheme may be used as the video encoding scheme. For example, the video encoding scheme may be AVC or HEVC.

[0115] The additional information decoding unit 4663 generates additional information, including map information, by decoding the encoded additional information.

[0116] The location information generation unit 4664 generates location information using the location image and map information. The attribute information generation unit 4665 generates attribute information using the attribute image and map information.

[0117] The second decoding unit 4660 uses the additional information necessary for decoding during the decoding process and outputs the additional information necessary for the application to the outside.

[0118] The following describes the challenges in the PCC encoding scheme. Figure 13 is a diagram showing the protocol stack involved in PCC encoded data. Figure 13 shows an example in which data from other media, such as video (e.g., HEVC) or audio, is multiplexed onto PCC encoded data and then transmitted or stored.

[0119] Multiplexing schemes and file formats have the function of multiplexing, transmitting, or storing various encoded data. In order to transmit or store encoded data, the encoded data must be converted into the format of the multiplexing scheme. For example, HEVC specifies a technique in which encoded data is stored in a data structure called a NAL unit, and the NAL unit is stored in ISOBMFF.

[0120] On the other hand, while two encoding methods are currently being considered for encoding point cloud data, the structure of the encoded data and the method for storing the encoded data in a system format have not been defined. As a result, there is a problem in that MUX processing (multiplexing), transmission, and storage cannot be performed in the encoding unit.

[0121] In the following text, unless a specific encoding method is mentioned, either the first encoding method or the second encoding method will be referred to.

[0122] (Embodiment 2) This embodiment describes a method for storing NAL units in an ISOBMFF file.

[0123] ISOBMFF (ISO based media file format) is a file format standard defined in ISO / IEC 14496-12. ISOBMFF specifies a format that can store various media such as video, audio, and text in multiplexed format, and is a media-independent standard.

[0124] This section explains the basic structure (file) of ISOBMFF. The basic unit in ISOBMFF is a box. A box consists of type, length, and data, and a file is a collection of boxes of various types.

[0125] Figure 14 shows the basic structure (file) of ISOBMFF. An ISOBMFF file mainly contains boxes such as ftyp, which indicates the file brand using 4CC (4-character code), moov, which stores metadata such as control information, and mdat, which stores data.

[0126] The method for storing each type of media in an ISOBMFF file is specified separately; for example, the method for storing AVC video and HEVC video is specified in ISO / IEC 14496-15. Here, it is conceivable to extend the functionality of ISOBMFF to store or transmit PCC encoded data, but there is currently no provision for storing PCC encoded data in an ISOBMFF file. Therefore, this embodiment describes a method for storing PCC encoded data in an ISOBMFF file.

[0127] Figure 15 shows the protocol stack when a common NAL unit for PCC codecs is stored in an ISOBMFF file. Here, a common NAL unit for PCC codecs is stored in an ISOBMFF file. Although the NAL unit is common to all PCC codecs, multiple PCC codecs are stored in the NAL unit, so it is desirable to define a storage method (Carriage of Codec1, Carriage of Codec2) according to each codec.

[0128] Next, we will explain how to store a common PCC NAL unit that supports multiple PCC codecs in an ISOBMFF file. Figure 16 shows an example of storing a common PCC NAL unit in an ISOBMFF file using the codec 1 storage method (Carriage of Codec1). Figure 17 shows an example of storing a common PCC NAL unit in an ISOBMFF file using the codec 2 storage method (Carriage of Codec2).

[0129] Here, ftyp is important information for identifying the file format, and a different identifier is defined for each codec for ftyp. If PCC encoded data encoded with the first encoding method (encoding scheme) is stored in the file, ftyp=pcc1 is set. If PCC encoded data encoded with the second encoding method is stored in the file, ftyp=pcc2 is set.

[0130] Here, pcc1 indicates that PCC codec 1 (first encoding method) is used. pcc2 indicates that PCC codec 2 (second encoding method) is used. In other words, pcc1 and pcc2 indicate that the data is PCC (encoded data of three-dimensional data (point cloud data)) and also indicate PCC codecs (first encoding method and second encoding method).

[0131] The following describes how to store NAL units in an ISOBMFF file. The multiplexing unit parses the NAL unit header and, if pcc_codec_type=Codec1, writes pcc1 to the ftyp field of the ISOBMFF file.

[0132] Furthermore, the multiplexing unit analyzes the NAL unit header and, if pcc_codec_type=Codec2, writes pcc2 to ftyp in ISOBMFF.

[0133] Furthermore, if pcc_nal_unit_type is metadata, the multiplexing unit stores the NAL unit in a predetermined manner, for example, in a moov or mdat file. If pcc_nal_unit_type is data, the multiplexing unit stores the NAL unit in a predetermined manner, for example, in a moov or mdat file.

[0134] For example, the multiplexing unit may store the NAL unit size in the NAL unit, similar to HEVC.

[0135] This storage method allows the demultiplexing unit (system layer) to analyze the ftyp contained in the file, thereby determining whether the PCC encoded data was encoded using the first or second encoding method. Furthermore, as described above, by determining whether the PCC encoded data was encoded using the first or second encoding method, it is possible to extract encoded data encoded using one of the two encoding methods from data containing a mixture of encoded data encoded using both methods. This reduces the amount of data transmitted when transmitting encoded data. In addition, this storage method allows the use of a common data format without setting different data (file) formats for the first and second encoding methods.

[0136] Furthermore, if the system layer metadata, such as ftyp in ISOBMFF, includes codec identification information, the multiplexer may store the NAL unit with pcc_nal_unit_type removed in the ISOBMFF file.

[0137] Next, the configuration and operation of the multiplexing unit of the three-dimensional data encoding system (three-dimensional data encoding device) according to this embodiment, and the demultiplexing unit of the three-dimensional data decoding system (three-dimensional data decoding device) according to this embodiment will be described.

[0138] Figure 18 shows the configuration of the first multiplexing unit 4710. The first multiplexing unit 4710 includes a file conversion unit 4711 that generates multiplexed data (file) by storing the encoded data and control information (NAL unit) generated by the first encoding unit 4630 into an ISOBMFF file. This first multiplexing unit 4710 is included, for example, in the multiplexing unit 4614 shown in Figure 1.

[0139] Figure 19 shows the configuration of the first demultiplexing unit 4720. The first demultiplexing unit 4720 includes a file inverse conversion unit 4721 that acquires encoded data and control information (NAL unit) from the multiplexed data (file) and outputs the acquired encoded data and control information to the first decoding unit 4640. This first demultiplexing unit 4720 is included, for example, in the demultiplexing unit 4623 shown in Figure 1.

[0140] Figure 20 shows the configuration of the second multiplexing unit 4730. The second multiplexing unit 4730 includes a file conversion unit 4731 that generates multiplexed data (file) by storing the encoded data and control information (NAL unit) generated by the second encoding unit 4650 into an ISOBMFF file. This second multiplexing unit 4730 is included, for example, in the multiplexing unit 4614 shown in Figure 1.

[0141] Figure 21 shows the configuration of the second demultiplexing unit 4740. The second demultiplexing unit 4740 includes a file inverse conversion unit 4741 that acquires encoded data and control information (NAL units) from the multiplexed data (file) and outputs the acquired encoded data and control information to the second decoding unit 4660. This second demultiplexing unit 4740 is included, for example, in the demultiplexing unit 4623 shown in Figure 1.

[0142] Figure 22 is a flowchart of the multiplexing process performed by the first multiplexing unit 4710. First, the first multiplexing unit 4710 analyzes the pcc_codec_type included in the NAL unit header to determine whether the codec being used is the first encoding method or the second encoding method (S4701).

[0143] If pcc_codec_type indicates a second encoding method (second encoding method in S4702), the first multiplexing unit 4710 does not process the NAL unit (S4703).

[0144] On the other hand, if pcc_codec_type indicates a second encoding method (the first encoding method in S4702), the first multiplexing unit 4710 writes pcc1 to ftyp (S4704). In other words, the first multiplexing unit 4710 writes information to ftyp indicating that data encoded with the first encoding method is stored in the file.

[0145] Next, the first multiplexing unit 4710 analyzes the pcc_nal_unit_type included in the NAL unit header and stores the data in a box (moov or mdat, etc.) in a predetermined manner according to the data type indicated by pcc_nal_unit_type (S4705). Then, the first multiplexing unit 4710 creates an ISOBMFF file containing the ftyp and the box (S4706).

[0146] Figure 23 is a flowchart of the multiplexing process performed by the second multiplexing unit 4730. First, the second multiplexing unit 4730 analyzes the pcc_codec_type included in the NAL unit header to determine whether the codec being used is the first encoding method or the second encoding method (S4711).

[0147] If pcc_unit_type indicates a second encoding method (second encoding method in S4712), the second multiplexing unit 4730 writes pcc2 to ftyp (S4713). In other words, the second multiplexing unit 4730 writes information to ftyp indicating that data encoded with the second encoding method is stored in the file.

[0148] Next, the second multiplexing unit 4730 analyzes the pcc_nal_unit_type included in the NAL unit header and stores the data in a box (moov or mdat, etc.) in a predetermined manner according to the data type indicated by pcc_nal_unit_type (S4714). Then, the second multiplexing unit 4730 creates an ISOBMFF file containing the ftyp and the box (S4715).

[0149] On the other hand, if pcc_unit_type indicates the first encoding method (first encoding method in S4712), the second multiplexing unit 4730 does not process the NAL unit (S4716).

[0150] The above process illustrates an example of encoding PCC data using either the first encoding method or the second encoding method. The first multiplexing unit 4710 and the second multiplexing unit 4730 store the desired NAL units in a file by identifying the codec type of the NAL units. If the PCC codec identification information is included in addition to the NAL unit header, the first multiplexing unit 4710 and the second multiplexing unit 4730 may use the PCC codec identification information included in addition to the NAL unit header to identify the codec type (first encoding method or second encoding method) in steps S4701 and S4711.

[0151] Furthermore, the first multiplexing unit 4710 and the second multiplexing unit 4730 may, in steps S4706 and S4714, remove the pcc_nal_unit_type from the NAL unit header before storing the data in the file.

[0152] Figure 24 is a flowchart showing the processing performed by the first demultiplexing unit 4720 and the first decoding unit 4640. First, the first demultiplexing unit 4720 analyzes the ftyp contained in the ISOBMFF file (S4721). If the codec indicated by ftyp is the second encoding method (pcc2) (second encoding method in S4722), the first demultiplexing unit 4720 determines that the data contained in the NAL unit payload is data encoded using the second encoding method (S4723). The first demultiplexing unit 4720 also transmits the result of this determination to the first decoding unit 4640. The first decoding unit 4640 does not process the NAL unit (S4724).

[0153] On the other hand, if the codec indicated by ftyp is the first encoding method (pcc1) (first encoding method in S4722), the first demultiplexer 4720 determines that the data contained in the NAL unit's payload is data encoded using the first encoding method (S4725). The first demultiplexer 4720 also transmits the result of this determination to the first decoding unit 4640.

[0154] The first decoding unit 4640 identifies the data by assuming that the pcc_nal_unit_type included in the NAL unit header is the identifier for the NAL unit for the first encoding method (S4726). Then, the first decoding unit 4640 decodes the PCC data using the decoding process of the first encoding method (S4727).

[0155] Figure 25 is a flowchart showing the processing performed by the second demultiplexing unit 4740 and the second decoding unit 4660. First, the second demultiplexing unit 4740 analyzes the ftyp contained in the ISOBMFF file (S4731). If the codec indicated by ftyp is the second encoding method (pcc2) (second encoding method in S4732), the second demultiplexing unit 4740 determines that the data contained in the NAL unit payload is data encoded using the second encoding method (S4733). The second demultiplexing unit 4740 then transmits the result of this determination to the second decoding unit 4660.

[0156] The second decoding unit 4660 identifies the data by assuming that the pcc_nal_unit_type included in the NAL unit header is the identifier for the NAL unit for the second encoding method (S4734). Then, the second decoding unit 4660 decodes the PCC data using the decoding process of the second encoding method (S4735).

[0157] On the other hand, if the codec indicated by ftyp is the first encoding method (pcc1) (first encoding method in S4732), the second demultiplexer 4740 determines that the data contained in the payload of the NAL unit is data encoded using the first encoding method (S4736). The second demultiplexer 4740 also transmits the result of this determination to the second decoding unit 4660. The second decoding unit 4660 does not process the NAL unit (S4737).

[0158] Thus, for example, by identifying the codec type of the NAL unit in the first demultiplexing unit 4720 or the second demultiplexing unit 4740, the codec type can be identified at an early stage. Furthermore, the desired NAL units can be input to the first decoding unit 4640 or the second decoding unit 4660, and unnecessary NAL units can be removed. In this case, the process of analyzing the codec identification information in the first decoding unit 4640 or the second decoding unit 4660 may become unnecessary. However, the first decoding unit 4640 or the second decoding unit 4660 may again refer to the NAL unit type and perform the process of analyzing the codec identification information.

[0159] Furthermore, if the first multiplexing unit 4710 or the second multiplexing unit 4730 has removed pcc_nal_unit_type from the NAL unit header, the first demultiplexing unit 4720 or the second demultiplexing unit 4740 may add pcc_nal_unit_type to the NAL unit before outputting it to the first decoding unit 4640 or the second decoding unit 4660.

[0160] (Embodiment 3) In this embodiment, a multiplexing unit and a demultiplexing unit corresponding to the encoding unit 4670 and decoding unit 4680 that support multiple codecs, as described in Embodiment 1, will be described. Figure 26 is a diagram showing the configuration of the encoding unit 4670 and the third multiplexing unit 4750 according to this embodiment.

[0161] The encoding unit 4670 encodes the point cloud data using either the first encoding method or the second encoding method, or both. The encoding unit 4670 may switch the encoding method (the first encoding method and the second encoding method) on a point cloud data basis or on a frame basis. Alternatively, the encoding unit 4670 may switch the encoding method on an encodeable basis.

[0162] The encoding unit 4670 generates encoded data (encoded stream) that includes identification information for the PCC codec.

[0163] The third multiplexing unit 4750 includes a file conversion unit 4751. The file conversion unit 4751 converts the NAL units output from the encoding unit 4670 into a PCC data file. The file conversion unit 4751 analyzes the codec identification information contained in the NAL unit header and determines whether the PCC encoded data is data encoded using a first encoding method, data encoded using a second encoding method, or data encoded using both methods. The file conversion unit 4751 writes a brand name that can identify the codec in ftyp. For example, if it indicates that the data is encoded using both methods, pcc3 is written in ftyp.

[0164] Furthermore, if the encoding unit 4670 contains identification information for the PCC codec in addition to the NAL unit, the file conversion unit 4751 may use this identification information to determine the PCC codec (encoding method).

[0165] Figure 27 shows the configuration of the third demultiplexing unit 4760 and decoding unit 4680 according to this embodiment.

[0166] The third demultiplexing unit 4760 includes a file inverse conversion unit 4761. The file inverse conversion unit 4761 analyzes the ftyp contained in the file and determines whether the PCC encoded data is data encoded using the first encoding method, data encoded using the second encoding method, or data encoded using both methods.

[0167] If the PCC encoded data is encoded using either one encoding method, the data is input to the corresponding decoding unit of the first decoding unit 4640 and the second decoding unit 4660, and the data is not input to the other decoding unit. If the PCC encoded data is encoded using both encoding methods, the data is input to the decoding unit 4680 corresponding to both methods.

[0168] The decoding unit 4680 decodes the PCC encoded data using either the first encoding method or the second encoding method, or both.

[0169] Figure 28 is a flowchart showing the processing performed by the third multiplexing unit 4750 according to this embodiment.

[0170] First, the third multiplexing unit 4750 analyzes the pcc_codec_type included in the NAL unit header to determine whether the codec being used is the first encoding method, the second encoding method, or both the first and second encoding methods (S4741).

[0171] If the second encoding method is used (Yes in S4742 and the second encoding method in S4743), the third multiplexing unit 4750 writes PCC2 to ftyp (S4744). In other words, the third multiplexing unit 4750 writes information to ftyp indicating that data encoded using the second encoding method is stored in the file.

[0172] Next, the third multiplexing unit 4750 analyzes the pcc_nal_unit_type included in the NAL unit header and stores the data in a box (moov or mdat, etc.) in a predetermined manner according to the data type indicated by pcc_unit_type (S4745). Then, the third multiplexing unit 4750 creates an ISOBMFF file containing the ftyp and the box (S4746).

[0173] On the other hand, if the first encoding method is used (Yes in S4742 and the first encoding method in S4743), the third multiplexing unit 4750 writes PCC1 to ftyp (S4747). In other words, the third multiplexing unit 4750 writes information to ftyp indicating that data encoded using the first encoding method is stored in the file.

[0174] Next, the third multiplexing unit 4750 analyzes the pcc_nal_unit_type included in the NAL unit header and stores the data in a box (moov or mdat, etc.) in a predetermined manner according to the data type indicated by pcc_unit_type (S4748). Then, the third multiplexing unit 4750 creates an ISOBMFF file containing the ftyp and the box (S4746).

[0175] On the other hand, if both the first encoding method and the second encoding method are used (No in S4742), the third multiplexing unit 4750 writes PCC3 to ftyp (S4749). In other words, the third multiplexing unit 4750 writes information to ftyp indicating that data encoded using both encoding methods is stored in the file.

[0176] Next, the third multiplexing unit 4750 analyzes the pcc_nal_unit_type included in the NAL unit header and stores the data in a box (moov or mdat, etc.) in a predetermined manner according to the data type indicated by pcc_unit_type (S4750). Then, the third multiplexing unit 4750 creates an ISOBMFF file containing the ftyp and the box (S4746).

[0177] Figure 29 is a flowchart showing the processing performed by the third demultiplexing unit 4760 and the decoding unit 4680. First, the third demultiplexing unit 4760 analyzes the ftyp contained in the ISOBMFF file (S4761). If the codec indicated by ftyp is the second encoding method (pcc2) (Yes in S4762 and the second encoding method in S4763), the third demultiplexing unit 4760 determines that the data contained in the NAL unit payload is data encoded using the second encoding method (S4764). The third demultiplexing unit 4760 then transmits the result of this determination to the decoding unit 4680.

[0178] The decoding unit 4680 identifies the data by assuming that the pcc_nal_unit_type included in the NAL unit header is the identifier for the NAL unit for the second encoding method (S4765). Then, the decoding unit 4680 decodes the PCC data using the decoding process of the second encoding method (S4766).

[0179] On the other hand, if the codec indicated by ftyp is the first encoding method (pcc1) (Yes in S4762 and the first encoding method in S4763), the third demultiplexing unit 4760 determines that the data contained in the NAL unit's payload is data encoded using the first encoding method (S4767). The third demultiplexing unit 4760 also transmits the result of this determination to the decoding unit 4680.

[0180] The decoding unit 4680 identifies the data by assuming that the pcc_nal_unit_type included in the NAL unit header is the identifier for the NAL unit for the first encoding method (S4768). Then, the decoding unit 4680 decodes the PCC data using the decoding process of the first encoding method (S4769).

[0181] On the other hand, if ftyp indicates that both encoding methods are used (pcc3) (No in S4762), the third demultiplexer 4760 determines that the data included in the NAL unit's payload is data encoded using both the first and second encoding methods (S4770). The third demultiplexer 4760 also transmits the result of this determination to the decoding unit 4680.

[0182] The decoding unit 4680 identifies the data by assuming that the pcc_nal_unit_type included in the NAL unit header is the identifier for the NAL unit for the codec described in pcc_codec_type (S4771). Then, the decoding unit 4680 decodes the PCC data using the decoding processes of both encoding methods (S4772). In other words, the decoding unit 4680 decodes the data encoded by the first encoding method using the decoding process of the first encoding method, and decodes the data encoded by the second encoding method using the decoding process of the second encoding method.

[0183] The following describes modifications of this embodiment. The following types may be indicated in the identification information as the brand type shown in ftyp. In addition, combinations of the following types may be indicated in the identification information.

[0184] The identification information may indicate whether the object of the original data before PCC encoding is a point cloud with a limited area, or a large point cloud with no area limitations, such as map information.

[0185] The identification information may indicate whether the original data before PCC encoding is a static or dynamic object.

[0186] As described above, the identification information may indicate whether the PCC encoded data is data encoded using the first encoding method or data encoded using the second encoding method.

[0187] The identification information may indicate the algorithm used in PCC coding. Here, the algorithm is, for example, an coding method that can be used in the first coding method or the second coding method.

[0188] The identification information may indicate differences in how PCC encoded data is stored in ISOBMFF files. For example, the identification information may indicate whether the storage method used is for storage or for real-time transmission, such as dynamic streaming.

[0189] Furthermore, while Embodiments 2 and 3 describe examples where ISOBMFF is used as the file format, other methods may also be used. For example, the same method as in these embodiments may be used when storing PCC encoded data in MPEG-2 TS Systems, MPEG-DASH, MMT, or RMP.

[0190] Furthermore, while the above example shows how to store metadata such as identification information in ftyp, this metadata may also be stored in a location other than ftyp. For example, this metadata may be stored in moov.

[0191] As described above, the three-dimensional data storage device (or three-dimensional data multiplexer, or three-dimensional data encoding device) performs the processing shown in Figure 30.

[0192] First, the three-dimensional data storage device (for example, including a first multiplexing unit 4710, a second multiplexing unit 4730, or a third multiplexing unit 4750) acquires one or more units (e.g., NAL units) in which encoded streams of point cloud data are stored (S4781). Next, the three-dimensional data storage device stores one or more units in a file (e.g., an ISOBMFF file) (S4782). In addition, during the storage (S4782), the three-dimensional data storage device stores information (e.g., pcc1, pcc2, or pcc3) indicating that the data stored in the file is data of encoded point cloud data in the file's control information (e.g., ftyp).

[0193] According to this, in a device that processes a file generated by the three-dimensional data storage device, by referring to the control information of the file, it is possible to early determine whether the data stored in the file is encoded data of point cloud data. Therefore, it is possible to reduce the processing amount or speed up the processing of the device.

[0194] For example, the information further indicates the encoding method used for encoding the point cloud data among the first encoding method and the second encoding method. Note that the fact that the data stored in the file is data obtained by encoding point cloud data and the encoding method used for encoding the point cloud data among the first encoding method and the second encoding method may be indicated by single information or different information.

[0195] According to this, in a device that processes a file generated by the three-dimensional data storage device, by referring to the control information of the file, it is possible to early determine the codec used for the data stored in the file. Therefore, it is possible to reduce the processing amount or speed up the processing of the device.

[0196] For example, the first encoding method is a method (GPCC) of encoding position information representing the position of point cloud data in an N-ary tree (N is an integer of 2 or more) and encoding attribute information using the position information, and the second encoding method is a method (VPCC) of generating a two-dimensional image from point cloud data and encoding the two-dimensional image using a video encoding method.

[0197] For example, the file complies with ISOBMFF (ISO based media file format).

[0198] For example, the three-dimensional data storage device includes a processor and a memory, and the processor performs the above processing using the memory.

[0199] Furthermore, as described above, the three-dimensional data acquisition device (or three-dimensional data demultiplexing device, or three-dimensional data decoding device) performs the processing shown in Figure 31.

[0200] The three-dimensional data acquisition device (for example, including a first demultiplexing unit 4720, a second demultiplexing unit 4740, or a third demultiplexing unit 4760) acquires a file (for example, an ISOBMFF file) that stores one or more units (for example, NAL units) in which encoded streams of point cloud data are stored (S4791). Next, the three-dimensional data acquisition device acquires one or more units from the file (S4792). The file control information (for example, ftyp) also includes information (for example, pcc1, pcc2, or pcc3) indicating that the data stored in the file is data in which point cloud data has been encoded.

[0201] For example, the three-dimensional data acquisition device refers to the aforementioned information to determine whether the data stored in the file is encoded point cloud data. If the three-dimensional data acquisition device determines that the data stored in the file is encoded point cloud data, it generates point cloud data by decoding the encoded point cloud data contained in one or more units. Alternatively, if the three-dimensional data acquisition device determines that the data stored in the file is encoded point cloud data, it outputs (notifies) information indicating that the data contained in one or more units is encoded point cloud data to a subsequent processing unit (for example, the first decoding unit 4640, the second decoding unit 4660, or the decoding unit 4680).

[0202] According to this, the three-dimensional data acquisition device can refer to the file's control information to quickly determine whether the data stored in the file is encoded point cloud data. Therefore, it is possible to reduce the processing load or speed up processing for the three-dimensional data acquisition device or subsequent devices.

[0203] For example, the information further indicates the encoding method used for the encoding, among the first and second encoding methods. Note that the fact that the data stored in the file is data in which point cloud data has been encoded, and the encoding method used for encoding the point cloud data, among the first and second encoding methods, may be indicated by a single piece of information or by different pieces of information.

[0204] According to this, the three-dimensional data acquisition device can refer to the file's control information to quickly determine the codec used for the data stored in the file. Therefore, it is possible to reduce the processing load or speed up processing for the three-dimensional data acquisition device or subsequent devices.

[0205] For example, a three-dimensional data acquisition device acquires data encoded using one of two encoding methods from encoded point cloud data, which includes data encoded using a first encoding method and data encoded using a second encoding method, based on the aforementioned information.

[0206] For example, the first encoding method is a method (GPCC) that encodes positional information represented by an N (where N is an integer greater than or equal to 2) subtree representing the positions of point cloud data, and then encodes attribute information using the positional information. The second encoding method is a method (VPCC) that generates a two-dimensional image from the point cloud data and then encodes the two-dimensional image using a video encoding method.

[0207] For example, the aforementioned file conforms to ISOBMFF (ISO based media file format).

[0208] For example, a three-dimensional data acquisition device comprises a processor and memory, and the processor uses the memory to perform the above processing.

[0209] (Embodiment 4) This embodiment describes the types of encoded data (geometry, attribute, and metadata) generated by the first encoding unit 4630 or the second encoding unit 4650 described above, the method for generating metadata, and the multiplexing process in the multiplexing unit. Note that metadata may also be referred to as parameter sets or control information.

[0210] In this embodiment, we will explain using the dynamic object (three-dimensional point cloud data that changes over time) described in Figure 4 as an example, but the same method may be used for static objects (three-dimensional point cloud data at any given time).

[0211] Figure 32 shows the configuration of the encoding unit 4801 and the multiplexing unit 4802 included in the three-dimensional data encoding device according to this embodiment. The encoding unit 4801 corresponds, for example, to the first encoding unit 4630 or the second encoding unit 4650 described above. The multiplexing unit 4802 corresponds to the multiplexing unit 4634 or 4656 described above.

[0212] The encoding unit 4801 encodes point cloud data from multiple PCC (Point Cloud Compression) frames and generates encoded data (Multiple Compressed Data) containing multiple location information, attribute information, and additional information.

[0213] The multiplexing unit 4802 converts data of multiple data types (location information, attribute information, and additional information) into NAL units, thereby transforming the data into a data configuration that takes into account data access by the decoding device.

[0214] FIG. 33 is a diagram showing a configuration example of encoded data generated by the encoding unit 4801. The arrows in the figure indicate the dependency relationships related to the decoding of the encoded data, and the source of the arrow depends on the data at the destination of the arrow. That is, the decoding device decodes the data at the destination of the arrow and uses the decoded data to decode the data at the source of the arrow. In other words, to depend means that the data at the destination is referenced (used) in the processing (such as encoding or decoding) of the data at the source of the dependency.

[0215] First, the generation process of the encoded data for the position information will be described. The encoding unit 4801 generates encoded position data (Compressed Geometry Data) for each frame by encoding the position information of each frame. Also, the encoded position data is represented by G(i). Here, i indicates the frame number, or the time of the frame, etc.

[0216] In addition, the encoding unit 4801 generates a position parameter set (GPS(i)) corresponding to each frame. The position parameter set includes parameters that can be used for decoding the encoded position data. Also, the encoded position data for each frame depends on the corresponding position parameter set.

[0217] Also, the encoded position data consisting of multiple frames is defined as a position sequence (Geometry Sequence). The encoding unit 4801 generates a position sequence parameter set (Geometry Sequence PS: also denoted as position SPS) that stores parameters commonly used for the decoding process for multiple frames within the position sequence. The position sequence depends on the position SPS.

[0218] Next, the process for generating encoded attribute data will be explained. The encoding unit 4801 generates compressed attribute data for each frame by encoding the attribute information of each frame. The compressed attribute data is represented by A(i). Figure 33 shows an example where attribute X and attribute Y exist, with the compressed attribute data for attribute X represented by AX(i) and the compressed attribute data for attribute Y represented by AY(i).

[0219] Furthermore, the encoding unit 4801 generates an attribute parameter set (APS(i)) corresponding to each frame. The attribute parameter set for attribute X is represented by AXPS(i), and the attribute parameter set for attribute Y is represented by AYPS(i). The attribute parameter set includes parameters that can be used to decode the encoded attribute information. The encoded attribute data depends on the corresponding attribute parameter set.

[0220] Furthermore, encoded attribute data consisting of multiple frames is sequenced into an attribute sequence (Attribute The attribute sequence is defined as an attribute sequence. The encoding unit 4801 generates an attribute sequence parameter set (Attribute Sequence PS: also written as attribute SPS) which stores parameters commonly used for decoding multiple frames within the attribute sequence. The attribute sequence depends on the attribute SPS.

[0221] Furthermore, in the first encoding method, the encoded attribute data depends on the encoded position data.

[0222] Figure 33 also shows an example where there are two types of attribute information (attribute X and attribute Y). When there are two types of attribute information, for example, two encoding units generate the respective data and metadata. Also, for example, an attribute sequence is defined for each type of attribute information, and an attribute SPS is generated for each type of attribute information.

[0223] Note that Figure 33 shows an example where there is one type of positional information and two types of attribute information, but the example is not limited to this; there may be one type of attribute information or three or more types. In this case as well, encoded data can be generated in the same way. Furthermore, in the case of point cloud data that does not have attribute information, attribute information is not required. In that case, the encoding unit 4801 does not need to generate a parameter set related to attribute information.

[0224] Next, the process of generating additional information (metadata) will be described. The encoding unit 4801 generates a PCC stream PS (also written as Stream PS), which is a parameter set for the entire PCC stream. The encoding unit 4801 stores in Stream PS parameters that can be used in common for decoding one or more location sequences and one or more attribute sequences. For example, Stream PS includes identification information indicating the codec of the point cloud data, and information indicating the algorithm used for encoding. The location sequences and attribute sequences depend on Stream PS.

[0225] Next, we will explain the Access Unit and GOF. In this embodiment, we introduce the new concepts of Access Unit (AU) and GOF (Group of Frame).

[0226] An access unit is the basic unit for accessing data during decryption, and consists of one or more data points and one or more metadata points. For example, an access unit consists of location information at the same time and one or more attribute information points. A GOF (Group of Four) is a random access unit and consists of one or more access units.

[0227] The encoding unit 4801 generates an access unit header (AU Header) as identification information indicating the beginning of an access unit. The encoding unit 4801 stores parameters related to the access unit in the access unit header. For example, the access unit header includes the structure or information of the encoded data contained in the access unit. The access unit header also includes parameters commonly used in the data contained in the access unit, such as parameters related to the decoding of the encoded data.

[0228] The encoding unit 4801 may generate an access unit delimiter that does not include parameters related to the access unit, instead of an access unit header. This access unit delimiter is used as identification information to indicate the beginning of the access unit. The decoding device identifies the beginning of the access unit by detecting the access unit header or the access unit delimiter.

[0229] Next, the generation of identification information at the beginning of the GOF will be explained. The encoding unit 4801 generates a GOF header as identification information indicating the beginning of the GOF. The encoding unit 4801 stores parameters related to the GOF in the GOF header. For example, the GOF header includes the structure or information of the encoded data included in the GOF. The GOF header also includes parameters commonly used in the data included in the GOF, such as parameters related to decoding the encoded data.

[0230] The encoding unit 4801 may generate a GOF delimiter that does not include parameters related to the GOF, instead of a GOF header. This GOF delimiter is used as identification information to indicate the beginning of the GOF. The decoding device identifies the beginning of the GOF by detecting either the GOF header or the GOF delimiter.

[0231] In PCC encoded data, for example, an access unit is defined as a PCC frame. The decoder accesses the PCC frame based on the identification information at the beginning of the access unit.

[0232] Furthermore, for example, a GOF (Group of Frames) is defined as a single random access unit. The decryption device accesses the random access unit based on the identification information at the beginning of the GOF. For example, if PCC frames are independent of each other and can be decrypted individually, then a PCC frame may be defined as a random access unit.

[0233] Furthermore, two or more PCC frames may be assigned to a single access unit, and multiple random access units may be assigned to a single GOF.

[0234] Furthermore, the encoding unit 4801 may define and generate parameter sets or metadata other than those described above. For example, the encoding unit 4801 may generate SEI (Supplemental Enhancement Information) that stores parameters that may not necessarily be used during decoding (optional parameters).

[0235] Next, we will explain the structure of the encoded data and how to store the encoded data in the NAL unit.

[0236] For example, a data format is defined for each type of encoded data. Figure 34 shows an example of encoded data and a NAL unit.

[0237] For example, as shown in Figure 34, encoded data includes a header and a payload. The encoded data may also include length information indicating the length (data volume) of the encoded data, header, or payload. Furthermore, the encoded data does not necessarily have to include a header.

[0238] The header includes, for example, identification information to identify the data. This identification information may indicate, for example, the data type or frame number.

[0239] The header contains, for example, identification information indicating a reference relationship. This identification information is stored in the header when there is a dependency between data, and it is information used to reference the referenced data from the source. For example, the header of the referenced data contains identification information to identify that data. The header of the referenced data contains identification information indicating the referenced data.

[0240] Furthermore, if the referenced or source can be identified or derived from other information, the identifying information for identifying the data or identifying information indicating the reference relationship may be omitted.

[0241] The multiplexing unit 4802 stores the encoded data in the payload of the NAL unit. The NAL unit header contains pcc_nal_unit_type, which is identification information for the encoded data. Figure 35 shows an example of the semantics of pcc_nal_unit_type.

[0242] As shown in Figure 35, when pcc_codec_type is Codec 1 (Codec1: First encoding method), values ​​of pcc_nal_unit_type from 0 to 10 are assigned to the encoded position data (Geometry), encoded attribute X data (AttributeX), encoded attribute Y data (AttributeY), position PS (Geom.PS), attribute XPS (AttrX.PS), attribute YPS (AttrX.PS), position SPS (Geometry Sequence PS), attribute XSPS (AttributeX Sequence PS), attribute YSPS (AttributeY Sequence PS), AU header (AU Header), and GOF header (GOF Header) in Codec 1. Values ​​11 and above are assigned to the reserves of Codec 1.

[0243] If pcc_codec_type is Codec 2 (the second encoding method), values ​​of pcc_nal_unit_type from 0 to 2 are assigned to the codec's Data A, Metadata A, and Metadata B. Values ​​3 and above are assigned to the backup of Codec 2.

[0244] Next, we will explain the data transmission order. Below, we will explain the constraints on the transmission order of the NAL unit.

[0245] The multiplexing unit 4802 sends out NAL units in groups of GOF or AU units. The multiplexing unit 4802 places a GOF header at the beginning of a GOF and an AU header at the beginning of an AU.

[0246] The multiplexing unit 4802 may provide a sequence parameter set (SPS) for each AU so that the decoding device can decode from the next AU even if data is lost due to packet loss or other reasons.

[0247] If there are dependencies in the encoded data related to decoding, the decoding device decodes the referenced data first, and then decodes the source data. In order for the decoding device to decode the data in the order it was received without rearranging it, the multiplexing unit 4802 sends the referenced data first.

[0248] Figure 36 shows an example of the transmission order of the NAL unit. Figure 36 shows three examples: location information priority, parameter priority, and data integration.

[0249] The location-prioritized transmission order is an example where location information and attribute information are transmitted together. In this transmission order, the transmission of location information is completed earlier than the transmission of attribute information.

[0250] For example, by using this transmission order, a decoding device that does not decode attribute information may be able to create a period of time where it does not process attribute information by ignoring the decoding of attribute information. Also, for example, a decoding device that wants to decode location information quickly may be able to decode the location information faster by obtaining the encoded location information data earlier.

[0251] Note that in Figure 36, the attributes XSPS and YSPS are combined and labeled as attribute SPS, but it is also acceptable to place attributes XSPS and YSPS separately.

[0252] In a parameter set priority transmission order, the parameter set is sent first, followed by the data.

[0253] As long as the constraints on the NAL unit transmission order are followed as described above, the multiplexing unit 4802 may transmit the NAL units in any order. For example, sequence identification information may be defined, and the multiplexing unit 4802 may have the function of transmitting NAL units in multiple patterns of order. For example, the sequence identification information of the NAL units may be stored in the stream PS.

[0254] The three-dimensional data decoding device may perform decoding based on sequence identification information. The three-dimensional data decoding device may instruct the three-dimensional data encoding device to send a desired transmission order, and the three-dimensional data encoding device (multiplexing unit 4802) may control the transmission order according to the instructed transmission order.

[0255] Furthermore, the multiplexing unit 4802 may generate encoded data that merges multiple functions, as long as it adheres to the constraints of the transmission order, such as the transmission order of data integration. For example, as shown in Figure 36, the GOF header and the AU header may be merged, or AXPS and AYPS may be merged. In this case, pcc_nal_unit_type is defined as an identifier indicating that the data has multiple functions.

[0256] The following describes modifications of this embodiment. PS has levels, such as frame-level PS, sequence-level PS, and PCC sequence-level PS. If we consider the PCC sequence level as the higher level and the frame level as the lower level, the following method may be used to store the parameters.

[0257] The default PS value is shown in the higher-level PS. If the value of the lower-level PS differs from the value of the higher-level PS, the PS value is shown in the lower-level PS. Alternatively, the PS value is not listed in the higher-level PS, but is listed in the lower-level PS. Alternatively, information on whether the PS value is shown in the lower-level PS, the higher-level PS, or both is shown in either the lower-level PS or the higher-level PS, or both. Alternatively, the lower-level PS may be merged with the higher-level PS. Alternatively, if the lower-level PS and the higher-level PS overlap, the multiplexing unit 4802 may omit sending one of them.

[0258] The encoding unit 4801 or the multiplexing unit 4802 may divide the data into slices or tiles and send out the divided data. The divided data includes information for identifying the divided data, and the parameter set includes parameters used for decoding the divided data. In this case, pcc_nal_unit_type is defined as an identifier indicating that it is data that stores data or parameters related to tiles or slices.

[0259] (Embodiment 5) The following describes methods for dividing point cloud data. Figure 37 shows examples of slicing and tiling.

[0260] First, let's explain the slicing method. The 3D data encoding device divides the 3D point cloud data into arbitrary point clouds in slice units. In slicing, the 3D data encoding device does not separate the position information and attribute information that constitute the points, but rather separates the position information and attribute information together. That is, the 3D data encoding device performs slicing so that the position information and attribute information of any given point belong to the same slice. Note that any number of divisions and division method is acceptable as long as these are followed. Also, the smallest unit of division is a point. For example, the number of divisions for position information and attribute information is the same. For example, the 3D point corresponding to the position information and the 3D point corresponding to the attribute information after slicing will be included in the same slice.

[0261] Furthermore, the three-dimensional data encoding device generates slice supplementary information, which is additional information related to the number of divisions and the division method, during slice division. The slice supplementary information is the same for both positional information and attribute information. For example, the slice supplementary information includes information indicating the reference coordinate position, size, or side length of the bounding box after division. It also includes information indicating the number of divisions and the division type.

[0262] Next, the tile division method will be explained. The three-dimensional data encoding device divides the sliced ​​data into slice position information (G slice) and slice attribute information (A slice), and then divides the slice position information and slice attribute information into tile units.

[0263] Although Figure 37 shows an example of partitioning using an octave tree structure, the number of partitions and the partitioning method can be any method.

[0264] Furthermore, the three-dimensional data encoding device may divide the position information and attribute information using different division methods, or it may divide them using the same division method. Also, the three-dimensional data encoding device may divide multiple slices into tiles using different division methods, or it may divide them into tiles using the same division method.

[0265] Furthermore, the three-dimensional data encoding device generates tile-additional information related to the number of divisions and the division method during tile division. The tile-additional information (positional tile-additional information and attribute tile-additional information) is independent of the positional information and attribute information. For example, the tile-additional information includes information indicating the reference coordinate position, size, or side length of the bounding box after division. The tile-additional information also includes information indicating the number of divisions and the division type.

[0266] Next, we will describe an example of how to divide point cloud data into slices or tiles. The three-dimensional data encoding device may use a predetermined method for slicing or tiling, or it may adaptively switch between methods depending on the point cloud data.

[0267] During slicing, the 3D data encoding device divides the 3D space collectively based on location information and attribute information. For example, the 3D data encoding device determines the shape of an object and divides the 3D space into slices according to the object's shape. For example, the 3D data encoding device extracts objects such as trees or buildings and divides them on an object-by-object basis. For example, the 3D data encoding device performs slicing so that the entirety of one or more objects is included in one slice. Alternatively, the 3D data encoding device divides a single object into multiple slices.

[0268] In this case, the encoding device may, for example, change the encoding method for each slice. For example, the encoding device may use a high-quality compression method for a specific object or a specific part of an object. In this case, the encoding device may store information indicating the encoding method for each slice in additional information (metadata).

[0269] Furthermore, the three-dimensional data encoding device may perform slicing based on map information or location information so that each slice corresponds to a predetermined coordinate space.

[0270] When dividing a slice into tiles, the 3D data encoding device divides the positional information and attribute information independently. For example, the 3D data encoding device divides a slice into tiles according to the amount of data or processing load. For example, the 3D data encoding device determines whether the amount of data in a slice (e.g., the number of 3D points in the slice) is greater than a predetermined threshold. If the amount of data in a slice is greater than the threshold, the 3D data encoding device divides the slice into tiles. If the amount of data in a slice is less than the threshold, the 3D data encoding device does not divide the slice into tiles.

[0271] For example, a three-dimensional data encoding device divides slices into tiles so that the processing load or processing time at the decoding device is within a certain range (below a predetermined value). This ensures that the processing load per tile at the decoding device is constant, facilitating distributed processing at the decoding device.

[0272] Furthermore, if the processing load differs between location information and attribute information, for example, if the processing load for location information is greater than that for attribute information, the three-dimensional data encoding device will increase the number of divisions for location information to be greater than the number of divisions for attribute information.

[0273] Furthermore, for example, if the content allows the decoding device to decode and display location information quickly and attribute information slowly later, the three-dimensional data encoding device may divide the location information into more segments than the attribute information. This allows the decoding device to process location information in parallel, thus speeding up the processing of location information compared to the processing of attribute information.

[0274] Furthermore, the decoding device does not necessarily need to process the sliced ​​or tiled data in parallel; it may decide whether or not to process them in parallel depending on the number or capacity of the decoding processing units.

[0275] By dividing the data in the manner described above, adaptive encoding can be achieved according to the content or object. Furthermore, parallel processing can be implemented in the decoding process. This improves the flexibility of the point cloud coding system or point cloud decoding system.

[0276] Figure 38 shows examples of slice and tile division patterns. In the figure, DU stands for Data Unit, representing data for a tile or slice. Each DU also includes a Slice Index and a Tile Index. The number in the upper right corner of the DU indicates the Slice Index, and the number in the lower left corner indicates the Tile Index.

[0277] In Pattern 1, the number of divisions and the division method are the same for G slices and A slices in slice partitioning. In tile partitioning, the number of divisions and the division method for G slices are different from those for A slices. Also, the same number of divisions and division method are used between multiple G slices. The same number of divisions and division method are used between multiple A slices.

[0278] In Pattern 2, the number of divisions and the division method are the same for G slices and A slices in slice partitioning. In tile partitioning, the number of divisions and the division method for G slices are different from those for A slices. Also, the number of divisions and the division method differ between multiple G slices. The number of divisions and the division method differ between multiple A slices.

[0279] (Embodiment 6) The following describes an example of performing slicing after tiling. In autonomous applications such as self-driving vehicles, point cloud data is needed not for the entire area, but for the area around the vehicle or the area in the direction of the vehicle's movement. Here, tiles and slices can be used to selectively decode the original point cloud data. By dividing the three-dimensional point cloud data into tiles and then further dividing it into slices, encoding efficiency can be improved or parallel processing can be achieved. When the data is divided, additional information (metadata) is generated, and this generated additional information is sent to the multiplexing unit.

[0280] Figure 39 is a block diagram showing the configuration of the first encoding unit 5010 included in the three-dimensional data encoding device according to this embodiment. The first encoding unit 5010 generates encoded data (encoded stream) by encoding point cloud data using a first encoding method (GPCC (Geometry based PCC)). This first encoding unit 5010 includes a division unit 5011, a plurality of position information encoding units 5012, a plurality of attribute information encoding units 5013, an additional information encoding unit 5014, and a multiplexing unit 5015.

[0281] The division unit 5011 generates multiple divided data by dividing the point cloud data. Specifically, the division unit 5011 generates multiple divided data by dividing the space of the point cloud data into multiple subspaces. Here, a subspace is either a tile or a slice, or a combination of a tile and a slice. More specifically, the point cloud data includes location information, attribute information, and additional information. The division unit 5011 divides the location information into multiple divided location information and the attribute information into multiple divided attribute information. The division unit 5011 also generates additional information related to the division.

[0282] For example, the division unit 5011 first divides the point cloud into tiles. Next, the division unit 5011 further divides the resulting tiles into slices.

[0283] Multiple location information encoding units 5012 generate multiple encoded location information by encoding multiple divided location information. For example, multiple location information encoding units 5012 process multiple divided location information in parallel.

[0284] The multiple attribute information encoding unit 5013 generates multiple encoded attribute information by encoding multiple divided attribute information. For example, the multiple attribute information encoding unit 5013 processes multiple divided attribute information in parallel.

[0285] The additional information encoding unit 5014 generates encoded additional information by encoding the additional information contained in the point cloud data and the additional information related to data division generated by the division unit 5011 during division.

[0286] The multiplexing unit 5015 generates encoded data (encoded stream) by multiplexing multiple encoded position information, multiple encoded attribute information, and encoded additional information, and transmits the generated encoded data. The encoded additional information is used during decoding.

[0287] In Figure 39, an example is shown where there are two location information encoding units 5012 and two attribute information encoding units 5013. However, the number of location information encoding units 5012 and attribute information encoding units 5013 may be one or three or more. Furthermore, multiple divided data may be processed in parallel within the same chip, such as multiple cores in a CPU, or in parallel across the cores of multiple chips, or in parallel across multiple cores of multiple chips.

[0288] Next, the decoding process will be explained. Figure 40 is a block diagram showing the configuration of the first decoding unit 5020. The first decoding unit 5020 restores the point cloud data by decoding the encoded data (encoded stream) generated when the point cloud data is encoded using the first encoding method (GPCC). This first decoding unit 5020 includes a demultiplexing unit 5021, multiple location information decoding units 5022, multiple attribute information decoding units 5023, an additional information decoding unit 5024, and a coupling unit 5025.

[0289] The demultiplexing unit 5021 generates multiple encoded position information, multiple encoded attribute information, and encoded additional information by demultiplexing the encoded data (encoded stream).

[0290] Multiple location information decoding units 5022 generate multiple segmented location information by decoding multiple encoded location information. For example, multiple location information decoding units 5022 process multiple encoded location information in parallel.

[0291] The multiple attribute information decoding unit 5023 generates multiple segmented attribute information by decoding multiple encoded attribute information. For example, the multiple attribute information decoding unit 5023 processes multiple encoded attribute information in parallel.

[0292] Multiple additional information decoding units 5024 generate additional information by decoding encoded additional information.

[0293] The merging unit 5025 generates position information by combining multiple division position information using additional information. The merging unit 5025 generates attribute information by combining multiple division attribute information using additional information. For example, the merging unit 5025 first generates point cloud data corresponding to tiles by combining decoded point cloud data for slices using slice additional information. Next, the merging unit 5025 restores the original point cloud data by combining point cloud data corresponding to tiles using tile additional information.

[0294] In Figure 39, an example is shown where there are two location information decoding units 5022 and two attribute information decoding units 5023. However, the number of location information decoding units 5022 and attribute information decoding units 5023 may be one or three or more. Furthermore, multiple divided data may be processed in parallel within the same chip, such as multiple cores in a CPU, or in parallel across the cores of multiple chips, or across multiple cores of multiple chips.

[0295] Next, we will explain how to segment point cloud data. In autonomous applications such as self-driving vehicles, point cloud data is not needed for the entire area, but rather for the area around the vehicle or the area in the direction of the vehicle's movement.

[0296] Figure 41 shows examples of tile shapes. As shown in Figure 41, various shapes such as circles, rectangles, or ellipses may be used as tile shapes.

[0297] Figure 42 shows examples of tiles and slices. The configuration of slices may differ between tiles. For example, the configuration of a tile or slice may be optimized based on the amount of data. Alternatively, the configuration of a tile or slice may be optimized based on the decoding speed.

[0298] Alternatively, tile division may be performed based on location information. In this case, attribute information is divided in the same way as the corresponding location information.

[0299] Furthermore, in the slicing division after tile division, the slices may be divided using a method different from that used for positional information and attribute information. For example, the slicing method for each tile may be selected according to the requirements of the application. Based on the requirements of the application, different slicing methods or tile division methods may be used.

[0300] For example, the division unit 5011 divides the three-dimensional point cloud data into one or more tiles based on location information such as map information, in a two-dimensional shape viewed from above. Then, the division unit 5011 divides each tile into one or more slices.

[0301] The division section 5011 may also divide the position information (Geometry) and attribute information (Attribute) into slices using the same method.

[0302] Furthermore, location information and attribute information may each consist of one type or two or more types. Also, in the case of point cloud data that does not contain attribute information, attribute information is not required.

[0303] Figure 43 is a block diagram of the division section 5011. The division section 5011 includes a tile division section 5031 (Tile Divider), a geometry slice division section 5032 (Geometry Slice Divider), and an attribute slice division section 5033 (Attribute Slice Divider).

[0304] The tile division unit 5031 generates multiple tile position information by dividing position information (Geometry) into tiles. The tile division unit 5031 also generates multiple tile attribute information by dividing attribute information into tiles. Furthermore, the tile division unit 5031 outputs tile additional information (TileMetaData) which includes information related to tile division and information generated during tile division.

[0305] The position information slice division unit 5032 generates multiple divided position information (multiple slice position information) by dividing multiple tile position information into slices. The position information slice division unit 5032 also outputs position slice additional information (Geometry Slice MetaData) which includes information related to the slice division of the position information and information generated in the slice division of the position information.

[0306] The attribute information slice division unit 5033 generates multiple divided attribute information (multiple slice attribute information) by dividing multiple tile attribute information into slices. The attribute information slice division unit 5033 also outputs attribute slice additional information (Attribute Slice MetaData) which includes information related to the slice division of attribute information and information generated in the slice division of attribute information.

[0307] Next, we will describe examples of tile shapes. The entire three-dimensional map is divided into multiple tiles. The data from multiple tiles is selectively sent to the three-dimensional data decoding device. Alternatively, the data from multiple tiles is sent to the three-dimensional data decoding device in order of importance. Depending on the situation, the tile shape may be selected from multiple shapes.

[0308] Figure 44 shows an example of a map viewed from above using point cloud data obtained by LiDAR. The example shown in Figure 44 is point cloud data of a highway, including overpasses (flyovers).

[0309] Figure 45 shows an example of dividing the point cloud data shown in Figure 44 into square tiles. Such division into squares can be easily performed on the map server. In addition, the tile height is set low for normal roads. In the case of overpasses, the tile height is set higher than for normal roads so that the tiles cover the overpass.

[0310] Figure 46 shows an example of dividing the point cloud data shown in Figure 44 into circular tiles. In this case, adjacent tiles may overlap in a plan view. The three-dimensional data encoding device transmits point cloud data of a cylindrical area (a circle in a top view) around the vehicle to the vehicle when the vehicle requires point cloud data of the surrounding area.

[0311] Also, similar to the example in Figure 45, the tile height is set lower for normal roads. In overpasses, the tile height is set higher than for normal roads so that the tiles cover the overpass.

[0312] The 3D data encoding device may change the height of the tiles according to, for example, the shape or height of a road or building. Alternatively, the 3D data encoding device may change the height of the tiles according to location information or area information. Furthermore, the 3D data encoding device may change the height of each tile individually. Or, the 3D data encoding device may change the height of each tile in sections containing multiple tiles. In other words, the 3D data encoding device may make the height of multiple tiles within a section the same. Also, tiles of different heights may overlap in a top view.

[0313] Figure 47 shows examples of tile divisions using tiles of various shapes, sizes, or heights. The tile shapes can be any shape, any size, or any combination thereof.

[0314] For example, in addition to the examples of dividing with non-overlapping square tiles and dividing with overlapping circular tiles as described above, the three-dimensional data encoding device may also divide with overlapping square tiles. Furthermore, the shape of the tiles does not have to be square or circular; polygons with three or more vertices may be used, or shapes without vertices may be used.

[0315] Furthermore, there may be two or more types of tile shapes, and tiles of different shapes may overlap. Also, there may be one or more types of tile shapes, and within the same shape being divided, shapes of different sizes may be combined, and these may overlap.

[0316] For example, in areas without objects such as roads, larger tiles are used than in areas where objects exist. Furthermore, the 3D data encoding device may adaptively change the shape or size of the tiles depending on the objects.

[0317] Furthermore, for example, a three-dimensional data encoding device may set the tiles in the direction of travel to be larger because it is highly likely that it will need to read tiles far ahead of the vehicle, which is the direction of travel of the vehicle. Conversely, since it is unlikely that the vehicle will move to the side, the tiles to the side may be set to be smaller than the tiles in the direction of travel.

[0318] Figure 48 shows an example of tile data stored on the server. For example, point cloud data is pre-divided into tiles and encoded, and the resulting encoded data is stored on the server. The user retrieves the desired tile data from the server when needed. Alternatively, the server (three-dimensional data encoding device) may perform tile division and encoding to include the data desired by the user, according to the user's instructions.

[0319] For example, if the moving object (vehicle) is moving at a high speed, it is conceivable that a wider range of point cloud data will be required. Therefore, the server may determine the shape and size of the tiles and perform tiling based on the pre-estimated speed of the vehicle (e.g., the legal speed limit on the road, the speed of the vehicle that can be estimated from the width and shape of the road, or statistical speed). Alternatively, as shown in Figure 48, the server may pre-encode tiles of multiple shapes or sizes and store the obtained data. The moving object may acquire tile data of an appropriate shape and size according to the direction and speed of the moving object.

[0320] Figure 49 shows an example of a system related to tile division. As shown in Figure 49, the shape and area of ​​the tiles may be determined based on the position of an antenna (base station), which is a communication means for transmitting point cloud data, or the communication area supported by the antenna. Alternatively, if point cloud data is generated by a sensor such as a camera, the shape and area of ​​the tiles may be determined based on the position of the sensor or the target range (detection range) of the sensor.

[0321] One tile may be assigned to one antenna or sensor, or one tile may be assigned to multiple antennas or sensors. Multiple tiles may be assigned to one antenna or sensor. The antenna or sensor may be fixed or movable.

[0322] For example, encoded data divided into tiles may be managed by a server connected to an antenna or sensor for the area assigned to the tile. The server may manage the encoded data for its own area and the tile information for adjacent areas. Multiple encoded data for multiple tiles may be managed in a centralized management server (cloud) that manages multiple servers corresponding to each tile. Alternatively, there may be no server corresponding to each tile, and the antenna or sensor may be directly connected to the centralized management server.

[0323] The coverage area of ​​the antenna or sensor may vary depending on the radio wave power, equipment differences, and installation conditions, and the shape and size of the tiles may also be changed accordingly. Based on the coverage area of ​​the antenna or sensor, slices or PCC frames may be assigned instead of tiles.

[0324] Next, we will explain a technique for dividing tiles into slices. By assigning similar objects to the same slice, encoding efficiency can be improved.

[0325] For example, a three-dimensional data encoding device may use the features of the point cloud data to recognize objects (roads, buildings, trees, etc.) and perform slicing by clustering the point cloud for each object.

[0326] Alternatively, the three-dimensional data encoding device may perform slicing by grouping objects with the same attributes and assigning slices to each group. Here, attributes refer to information about movement, for example, and objects are grouped by classifying them into dynamic information such as pedestrians and cars, semi-dynamic information such as accidents and traffic jams, semi-static information such as traffic regulations and road construction, and static information such as road surfaces and structures.

[0327] Note that data may be duplicated across multiple slices. For example, when slicing into multiple object groups, any object may belong to one object group or to two or more object groups.

[0328] Figure 50 shows an example of this slicing method. For example, in the example shown in Figure 50, the tile is a rectangular prism. However, the tile may be cylindrical or of other shapes.

[0329] The point cloud contained in a tile is grouped into object groups, such as roads, buildings, and trees. Then, the tile is sliced ​​so that each object group is contained in a single slice. Each slice is then encoded individually.

[0330] Next, the method for encoding the divided data will be described. The three-dimensional data encoding device (first encoding unit 5010) encodes each of the divided data. When encoding attribute information, the three-dimensional data encoding device generates dependency information as additional information, indicating which configuration information (location information, additional information, or other attribute information) was used as the basis for encoding. In other words, the dependency information indicates, for example, the configuration information of the reference (dependent). In this case, the three-dimensional data encoding device generates the dependency information based on the configuration information corresponding to the division shape of the attribute information. Note that the three-dimensional data encoding device may generate dependency information based on configuration information corresponding to multiple division shapes.

[0331] Dependency relationship information is generated by a three-dimensional data encoding device, and the generated dependency relationship information may be sent to a three-dimensional data decoding device. Alternatively, the three-dimensional data decoding device may generate the dependency relationship information, and the three-dimensional data encoding device may not need to send the dependency relationship information. Also, the dependency relationship used by the three-dimensional data encoding device may be predetermined, and the three-dimensional data encoding device may not need to send the dependency relationship information.

[0332] FIG. 51 is a diagram showing an example of the dependency relationship of each data. The tip of the arrow in the figure indicates the dependency destination, and the origin of the arrow indicates the dependency source. The three-dimensional data decoding device decodes the data in the order from the dependency destination to the dependency source. Also, the data indicated by the solid line in the figure is the data actually sent, and the data indicated by the dotted line is the data not sent.

[0333] Also, in the same figure, G indicates position information, and A indicates attribute information. G t1 indicates the position information of tile number 1, G t2 indicates the position information of tile number 2. G t1s1 indicates the position information of tile number 1 and slice number 1, G t1s2 indicates the position information of tile number 1 and slice number 2, G t2s1 indicates the position information of tile number 2 and slice number 1, G t2s2 indicates the position information of tile number 2 and slice number 2. Similarly, A t1 indicates the attribute information of tile number 1, A t2 indicates the attribute information of tile number 2. A t1s1 indicates the attribute information of tile number 1 and slice number 1, A t1s2 indicates the attribute information of tile number 1 and slice number 2, A t2s1 indicates the attribute information of tile number 2 and slice number 1, A t2s2 indicates the attribute information of tile number 2 and slice number 2.

[0334] Mtile indicates tile additional information, MGslice indicates position slice additional information, and MAslice indicates attribute slice additional information. D t1s1 is the attribute information At1s1 This shows the dependency information of D t2s1 Attribute information A t2s1 This shows the dependency information.

[0335] Depending on the application, different tile or slice partitioning structures may be used.

[0336] Furthermore, the three-dimensional data encoding device may rearrange the data in the order of decoding so that the three-dimensional data decoding device does not need to rearrange the data. Alternatively, the data may be rearranged in the three-dimensional data decoding device, or both the three-dimensional data encoding device and the three-dimensional data decoding device may rearrange the data.

[0337] Figure 52 shows an example of the data decoding order. In the example in Figure 52, decoding is performed sequentially from left to right. The 3D data decoding device decodes dependent data first among dependent data. For example, the 3D data encoding device pre-arranges and sends the data in this order. Any order is acceptable as long as the dependent data comes first. The 3D data encoding device may also send additional information and dependency information before the data.

[0338] Furthermore, the three-dimensional data decoding device may selectively decode tiles based on requests from the application and information obtained from the NAL unit header. Figure 53 shows an example of encoded tile data. For example, the decoding order of tiles is arbitrary; that is, there may be no dependencies between tiles.

[0339] Next, the configuration of the combiner 5025 included in the first decoding unit 5020 will be described. Figure 54 is a block diagram showing the configuration of the combiner 5025. The combiner 5025 includes a geoometry slice combiner 5041, an attribute slice combiner 5042, and a tile combiner.

[0340] The position information slice joining unit 5041 generates multiple tile position information by joining multiple divided position information using position slice additional information. The attribute information slice joining unit 5042 generates multiple tile attribute information by joining multiple divided attribute information using attribute slice additional information.

[0341] The tile joining unit 5043 generates position information by combining multiple tile position information using tile additional information. Furthermore, the tile joining unit 5043 generates attribute information by combining multiple tile attribute information using tile additional information.

[0342] The number of slices or tiles to be divided must be one or more. In other words, the slices or tiles do not need to be divided at all.

[0343] Next, the structure of the sliced ​​or tiled encoded data and the method of storing the encoded data in the NAL unit (multiplexing method) will be explained. Figure 55 shows the structure of the encoded data and the method of storing the encoded data in the NAL unit.

[0344] The encoded data (splitting position information and splitting attribute information) is stored in the NAL unit's payload.

[0345] Encoded data includes a header and a payload. The header includes identification information to identify the data contained in the payload. This identification information includes, for example, the type of slice or tile division (slice_type, tile_type), index information to identify the slice or tile (slice_idx, tile_idx), location information of the data (slice or tile), or the address of the data (address). Index information to identify a slice is also written as SliceIndex. Index information to identify a tile is also written as TileIndex. The type of division can be, for example, a method based on the object shape as described above, a method based on map information or location information, or a method based on the amount of data or processing amount.

[0346] Furthermore, the header of the encoded data includes identification information indicating dependencies. In other words, if there are dependencies between data, the header includes identification information for referencing the dependent data from the dependent data source. For example, the header of the dependent data includes identification information to identify that data. The header of the dependent data includes identification information indicating the dependent data. Note that if the identification information for identifying the data, additional information related to slicing or tiling, and identification information indicating dependencies can be identified or derived from other information, this information may be omitted.

[0347] Next, the flow of the point cloud data encoding and decoding processes according to this embodiment will be described. Figure 56 is a flowchart of the point cloud data encoding process according to this embodiment.

[0348] First, the three-dimensional data encoding device determines the division method to be used (S5011). This division method includes whether or not to perform tiling division or slicing division. The division method may also include the number of divisions if tiling or slicing division is performed, and the type of division. The type of division refers to methods based on object shape, methods based on map information or location information, or methods based on data volume or processing volume, as described above. The division method may also be predetermined.

[0349] If tile division is performed (Yes in S5012), the three-dimensional data encoding device generates multiple tile position information and multiple tile attribute information by dividing the position information and attribute information together (S5013). The three-dimensional data encoding device also generates tile addition information related to tile division. The three-dimensional data encoding device may divide the position information and attribute information independently.

[0350] If slicing is performed (Yes in S5014), the three-dimensional data encoding device generates multiple slicing position information and multiple slicing attribute information by independently slicing multiple tile position information and multiple tile attribute information (or position information and attribute information) (S5015). The three-dimensional data encoding device also generates position slice addition information and attribute slice addition information related to the slicing. The three-dimensional data encoding device may also slice the tile position information and tile attribute information together.

[0351] Next, the three-dimensional data encoding device generates multiple encoded location information and multiple encoded attribute information by encoding each of the multiple division location information and multiple division attribute information (S5016). The three-dimensional data encoding device also generates dependency information.

[0352] Next, the three-dimensional data encoding device generates encoded data (encoded stream) by NAL unitizing (multiplexing) multiple encoded position information, multiple encoded attribute information, and additional information (S5017). The three-dimensional data encoding device then transmits the generated encoded data.

[0353] Figure 57 is a flowchart of the point cloud data decoding process according to this embodiment. First, the three-dimensional data decoding device determines the division method by analyzing the additional information related to the division method (tile additional information, position slice additional information, and attribute slice additional information) contained in the encoded data (encoded stream) (S5021). This division method includes whether or not to perform tile division and whether or not to perform slice division. The division method may also include the number of divisions and the type of division when tile division or slice division is performed.

[0354] Next, the three-dimensional data decoding device generates partitioning location information and partitioning attribute information by decoding multiple encoded position information and multiple encoded attribute information contained in the encoded data using dependency information contained in the encoded data (S5022).

[0355] If the additional information indicates that slice division has been performed (Yes in S5023), the three-dimensional data decoding device generates multiple tile position information and multiple tile attribute information by combining multiple division position information and multiple division attribute information in their respective methods, based on the position slice additional information and attribute slice additional information (S5024). The three-dimensional data decoding device may combine the multiple division position information and multiple division attribute information in the same method.

[0356] If the additional information indicates that tile division has been performed (Yes in S5025), the three-dimensional data decoding device generates position information and attribute information by combining multiple tile position information and multiple tile attribute information (multiple division position information and multiple division attribute information) in the same way based on the tile additional information (S5026). The three-dimensional data decoding device may combine the multiple tile position information and multiple tile attribute information in different ways.

[0357] Next, we will explain the tile appending information. The three-dimensional data encoding device generates tile appending information, which is metadata about the tile division method, and transmits the generated tile appending information to the three-dimensional data decoding device.

[0358] Figure 58 shows an example of the syntax for tile metadata (TileMetaData). As shown in Figure 58, for example, tile metadata includes division method information (type_of_divide), shape information (topview_shape), overlap flag (tile_overlap_flag), overlap information (type_of_overlap), height information (tile_height), number of tiles (tile_number), and tile position information (global_position, relative_position).

[0359] The division method information (type_of_divide) indicates how the tiles are divided. For example, the division method information indicates whether the tiles are divided based on map information, i.e., based on a top view (top_view), or something else (other).

[0360] Shape information (topview_shape) is included in the tile information when, for example, the tile division method is based on a top view. Shape information indicates the shape of the tile when viewed from above. For example, this shape includes squares and circles. This shape may also include polygons other than ellipses, rectangles, or quadrilaterals, or other shapes. Furthermore, shape information is not limited to the shape of the tile when viewed from above, but may also indicate the three-dimensional shape of the tile (for example, cubes and cylinders).

[0361] The tile_overlap_flag indicates whether tiles overlap or not. For example, the tile overlap flag is included in the tile information when the tile division method is based on a top view. In this case, the tile overlap flag indicates whether tiles overlap in a top view. The tile overlap flag may also indicate whether tiles overlap in three-dimensional space.

[0362] The overlap information (type_of_overlap) is included in the tile information when tiles overlap, for example. The overlap information indicates how the tiles overlap, such as the size of the overlapping area.

[0363] The height information (tile_height) indicates the height of the tile. The height information may also include information indicating the shape of the tile. For example, if the tile's shape when viewed from above is rectangular, this information may indicate the lengths of the sides (vertical and horizontal lengths) of that rectangle. Alternatively, if the tile's shape when viewed from above is circular, this information may indicate the diameter or radius of that circle.

[0364] Furthermore, the height information may indicate the height of each tile, or it may indicate a common height for multiple tiles. Alternatively, multiple height types for roads and overpasses may be predefined, and the height information may indicate the height of each height type and the height type of each tile. Or, the height of each height type may be predefined, and the height information may indicate the height type of each tile. In other words, the height of each height type does not necessarily have to be indicated by the height information.

[0365] The tile_number indicates the number of tiles. Note that tile information may also include information indicating the spacing between tiles.

[0366] Tile position information (global_position, relative_position) is information used to identify the location of each tile. For example, tile position information indicates the absolute or relative coordinates of each tile.

[0367] Furthermore, some or all of the above information may be provided for each tile, or for multiple tiles (for example, for each frame or for multiple frames).

[0368] The three-dimensional data encoding device may include the tile addition information in the SEI (Supplemental Enhancement Information) and send it. Alternatively, the three-dimensional data encoding device may store the tile addition information in an existing parameter set (PPS, GPS, or APS, etc.) and send it.

[0369] For example, if the tile information changes from frame to frame, the tile information may be stored in a parameter set for each frame (such as GPS or APS). If the tile information does not change within a sequence, the tile information may be stored in a parameter set for each sequence (location SPS or attribute SPS). Furthermore, if the same tile division information is used for both location information and attribute information, the tile information may be stored in the parameter set of the PCC stream (stream PS).

[0370] Furthermore, tile information may be stored in any of the parameter sets described above, or in multiple parameter sets. Additionally, tile information may be stored in the header of the encoded data. Furthermore, tile information may be stored in the header of the NAL unit.

[0371] Furthermore, all or part of the tile addition information may be stored in one of the headers of the division location information and the division attribute information, but not in the other. For example, if the same tile addition information is used for both location information and attribute information, the tile addition information may be included in one of the headers of the location information or attribute information. For example, if attribute information depends on location information, the location information is processed first. Therefore, the header of the location information may contain this tile addition information, while the header of the attribute information may not. In this case, the three-dimensional data decoding device will determine, for example, that the attribute information of the dependency belongs to the same tile as the tile of the location information to which it depends.

[0372] The 3D data decoding device reconstructs the tiled point cloud data based on the tile information. If there is duplicate point cloud data, the 3D data decoding device identifies the multiple duplicate point cloud data, selects one, or merges the multiple point cloud data.

[0373] Furthermore, the three-dimensional data decoding device may perform decoding using tile-added information. For example, if multiple tiles overlap, the three-dimensional data decoding device may decode each tile, perform processing using the decoded data (e.g., smoothing or filtering), and generate point cloud data. This may enable highly accurate decoding.

[0374] Figure 59 shows an example of a system configuration including a three-dimensional data encoding device and a three-dimensional data decoding device. The tile division unit 5051 divides point cloud data, including position information and attribute information, into a first tile and a second tile. The tile division unit 5051 also sends tile addition information related to tile division to the decoding unit 5053 and the tile joining unit 5054.

[0375] The encoding unit 5052 generates encoded data by encoding the first tile and the second tile.

[0376] The decoding unit 5053 reconstructs the first and second tiles by decoding the encoded data generated by the encoding unit 5052. The tile joining unit 5054 reconstructs the point cloud data (position information and attribute information) by joining the first and second tiles using the tile addition information.

[0377] Next, we will explain slice appending information. The three-dimensional data encoding device generates slice appending information, which is metadata about the slice division method, and transmits the generated slice appending information to the three-dimensional data decoding device.

[0378] Figure 60 shows an example of the syntax for slice metadata (SliceMetaData). As shown in Figure 60, for example, slice metadata includes division method information (type_of_divide), overlap flag (slice_overlap_flag), overlap information (type_of_overlap), number of slices (slice_number), slice position information (global_position, relative_position), and slice size information (slice_bounding_box_size).

[0379] The division method information (type_of_divide) indicates how the slice is divided. For example, the division method information indicates whether the slice is divided based on object information as shown in Figure 50 (object). Note that the slice supplement information may also include information indicating how the object is divided. For example, this information indicates whether one object is divided into multiple slices or assigned to one slice. This information may also indicate the number of divisions if one object is divided into multiple slices.

[0380] The overlap flag (slice_overlap_flag) indicates whether or not the slices overlap. The overlap information (type_of_overlap) is included in the slice append information, for example, if the slices overlap. The overlap information indicates how the slices overlap, for example, the size of the overlapping area.

[0381] The slice_number indicates the number of slices.

[0382] Slice position information (global_position, relative_position) and slice size information (slice_bounding_box_size) are information about the region of the slice. Slice position information is information used to identify the position of each slice. For example, slice position information indicates the absolute or relative coordinates of each slice. Slice size information (slice_bounding_box_size) indicates the size of each slice. For example, slice size information indicates the size of the bounding box of each slice.

[0383] The three-dimensional data encoding device may include slice addition information in the SEI and send it out. Alternatively, the three-dimensional data encoding device may store the slice addition information in an existing parameter set (PPS, GPS, or APS, etc.) and send it out.

[0384] For example, if slice addition information changes from frame to frame, the slice addition information may be stored in a parameter set for each frame (such as GPS or APS). If the slice addition information does not change within a sequence, the slice addition information may be stored in a parameter set for each sequence (location SPS or attribute SPS). Furthermore, if the same slice division information is used for both location information and attribute information, the slice addition information may be stored in the parameter set of the PCC stream (stream PS).

[0385] Furthermore, slice addition information may be stored in any of the parameter sets mentioned above, or in multiple parameter sets. Also, slice addition information may be stored in the header of the encoded data. Additionally, slice addition information may be stored in the header of the NAL unit.

[0386] Furthermore, all or part of the slice addition information may be stored in one of the headers of the division location information and the division attribute information, but not in the other. For example, if the same slice addition information is used for both location information and attribute information, the slice addition information may be included in one of the headers of the location information or attribute information. For example, if attribute information depends on location information, the location information is processed first. Therefore, the header of the location information may contain this slice addition information, while the header of the attribute information may not. In this case, the three-dimensional data decoding device will determine, for example, that the attribute information that depends on the location information belongs to the same slice as the slice of the location information it depends on.

[0387] The 3D data decoding device reconstructs the sliced ​​point cloud data based on the slice addition information. If there is duplicate point cloud data, the 3D data decoding device identifies the multiple duplicate point cloud data, selects one, or merges the multiple point cloud data.

[0388] Furthermore, the three-dimensional data decoding device may perform decoding using slice-added information. For example, if multiple slices overlap, the three-dimensional data decoding device may decode each slice, perform processing (e.g., smoothing or filtering) using the decoded data, and generate point cloud data. This may enable highly accurate decoding.

[0389] Figure 61 is a flowchart of the three-dimensional data encoding process, including the generation of tile-added information, using the three-dimensional data encoding device according to this embodiment.

[0390] First, the three-dimensional data encoding device determines the method for dividing the tiles (S5031). Specifically, the three-dimensional data encoding device determines whether to use a top-view-based division method or another method. The three-dimensional data encoding device also determines the shape of the tiles when using the top-view-based division method. Furthermore, the three-dimensional data encoding device determines whether or not a tile overlaps with other tiles.

[0391] If the tile division method determined in step S5031 is a division method based on a top view (Yes in S5032), the three-dimensional data encoding device indicates in the tile addition information that the tile division method is a division method based on a top view (top_view) (S5033).

[0392] On the other hand, if the tile division method determined in step S5031 is other than the division method based on the top view (No in S5032), the three-dimensional data encoding device indicates in the tile addition information that the tile division method is other than the division method based on the top view (top_view) (S5034).

[0393] Furthermore, if the shape of the tile viewed from above, as determined in step S5031, is a square (square in S5035), the three-dimensional data encoding device records that the shape of the tile viewed from above is a square in the tile supplement information (S5036). On the other hand, if the shape of the tile viewed from above, as determined in step S5031, is a circle (circle in S5035), the three-dimensional data encoding device records that the shape of the tile viewed from above is a circle in the tile supplement information (S5037).

[0394] Next, the three-dimensional data encoding device determines whether a tile overlaps with another tile (S5038). If a tile overlaps with another tile (Yes in S5038), the three-dimensional data encoding device records that the tile overlaps in the tile information (S5039). On the other hand, if a tile does not overlap with another tile (No in S5038), the three-dimensional data encoding device records that the tile does not overlap in the tile information (S5040).

[0395] Next, the three-dimensional data encoding device divides the tiles based on the tile division method determined in step S5031, encodes each tile, and sends out the generated encoded data and tile-related information (S5041).

[0396] Figure 62 is a flowchart of the three-dimensional data decoding process using tile-added information by the three-dimensional data decoding device according to this embodiment.

[0397] First, the three-dimensional data decoding device analyzes the tile addition information contained in the bitstream (S5051).

[0398] If the tile information indicates that a tile does not overlap with other tiles (No in S5052), the 3D data decoding device generates point cloud data for each tile by decoding each tile (S5053). Next, the 3D data decoding device reconstructs point cloud data from the point cloud data of each tile based on the tile division method and tile shape indicated in the tile information (S5054).

[0399] On the other hand, if the tile addition information indicates that a tile overlaps with other tiles (Yes in S5052), the 3D data decoding device generates point cloud data for each tile by decoding each tile. The 3D data decoding device also identifies the overlapping portion of the tiles based on the tile addition information (S5055). The 3D data decoding device may use multiple pieces of overlapping information to perform the decoding process for the overlapping portion. Next, the 3D data decoding device reconstructs point cloud data from the point cloud data of each tile based on the tile division method, tile shape, and overlapping information indicated in the tile addition information (S5056).

[0400] The following describes variations related to slicing. The three-dimensional data encoding device may transmit additional information indicating the type of object (road, building, tree, etc.) or attributes (dynamic information, static information, etc.). Alternatively, encoding parameters may be predetermined according to the object, and the three-dimensional data encoding device may notify the three-dimensional data decoding device of the encoding parameters by sending the type of object or attributes.

[0401] The following methods may be used for the encoding order and transmission order of slice data. For example, the 3D data encoding device may encode slice data in order from data that is easy to recognize or cluster. Alternatively, the 3D data encoding device may encode slice data in order from slice data that has been clustered first. The 3D data encoding device may also transmit the encoded slice data in order. Alternatively, the 3D data encoding device may transmit slice data in order of the decoding priority in the application. For example, if the decoding priority of dynamic information is high, the 3D data encoding device may transmit slice data in order from slices grouped by dynamic information.

[0402] Furthermore, if the order of encoded data differs from the order of decoding priority, the three-dimensional data encoding device may rearrange the encoded data before sending it out. Also, when storing encoded data, the three-dimensional data encoding device may rearrange the encoded data before storing it.

[0403] The application (3D data decoding device) requests the server (3D data encoding device) to send slices containing the desired data. The server sends the slice data required by the application, and does not need to send unnecessary slice data.

[0404] The application requests the server to send tiles containing the desired data. The server sends the tile data that the application needs, and does not need to send any tile data that is not needed.

[0405] As described above, the three-dimensional data encoding device according to this embodiment performs the processing shown in Figure 63. First, the three-dimensional data encoding device generates multiple encoded data by encoding multiple subspaces (e.g., tiles) obtained by dividing the target space containing multiple three-dimensional points (S5061). The three-dimensional data encoding device generates a bitstream that includes the multiple encoded data and first information (e.g., topview_shape) indicating the shape of the multiple subspaces (S5062).

[0406] According to this, a three-dimensional data encoding device can improve encoding efficiency because it can select any shape from multiple types of subspace shapes.

[0407] For example, the shape is the two-dimensional or three-dimensional shape of the plurality of subspaces. For example, the shape is the shape of the plurality of subspaces viewed from above. In other words, the first information indicates the shape of the subspaces viewed from a specific direction (e.g., from above). To put it another way, the first information indicates the shape of the subspaces viewed from above. For example, the shape is a rectangle or a circle.

[0408] For example, the bitstream includes second information (e.g., tile_overlap_flag) indicating whether the multiple sub-intervals overlap or not.

[0409] According to this, a three-dimensional data encoding device can duplicate subspaces, thus enabling the generation of subspaces without complicating their shapes.

[0410] For example, the bitstream includes third information (e.g., type_of_divide) indicating whether the method of dividing the plurality of sub-sections is a top-down view method.

[0411] For example, the bitstream includes a fourth piece of information (e.g., tile_height) indicating at least one of the height, width, depth, and radius of the plurality of subsections.

[0412] For example, the bitstream includes a fifth piece of information (e.g., global_position or relative_position) indicating the position of each of the plurality of sub-sections.

[0413] For example, the bitstream includes a sixth piece of information (e.g., tile_number) indicating the number of sub-intervals.

[0414] For example, the bitstream includes a seventh piece of information indicating the interval between the plurality of sub-intervals.

[0415] For example, a three-dimensional data encoding device comprises a processor and memory, and the processor uses the memory to perform the above processing.

[0416] Furthermore, the three-dimensional data decoding device according to this embodiment performs the processing shown in Figure 64. First, the three-dimensional data decoding device reconstructs the multiple subspaces by decoding multiple encoded data generated by encoding multiple subspaces (e.g., tiles) which are obtained by dividing the target space containing multiple three-dimensional points included in the bitstream (S5071). The three-dimensional data decoding device reconstructs the target space by combining the multiple subspaces using first information (e.g., topview_shape) which is included in the bitstream and indicates the shape of the multiple subspaces (S5072). For example, the three-dimensional data decoding device can grasp the position and range of each subspace within the target space by recognizing the shape of the multiple subspaces using the first information. The three-dimensional data decoding device can combine the multiple subspaces based on the grasped position and range of the multiple subspaces. As a result, the three-dimensional data decoding device can correctly combine the multiple subspaces.

[0417] For example, the shape is the shape in two dimensions or three dimensions of the plurality of subspaces. For example, the shape is a rectangle or a circle.

[0418] For example, the bitstream includes second information (e.g., tile_overlap_flag) indicating whether the multiple sub-intervals overlap. In reconstructing the target space, the three-dimensional data decoder further uses the second information to combine the multiple sub-spaces. For example, the three-dimensional data decoder uses the second information to determine whether the sub-spaces overlap. If the sub-spaces overlap, the three-dimensional data decoder identifies the overlapping region and performs a predetermined action on the identified overlapping region.

[0419] For example, the bitstream includes third information (e.g., type_of_divide) indicating whether the method of dividing the plurality of sub-sections is a top-view method. If the third information indicates that the method of dividing the plurality of sub-sections is a top-view method, the three-dimensional data decoder combines the plurality of sub-spaces using the first information.

[0420] For example, the bitstream includes fourth information (e.g., tile_height) indicating at least one of the height, width, depth, and radius of the plurality of sub-sections. In reconstructing the target space, the three-dimensional data decoder further uses the fourth information to combine the plurality of sub-spaces. For example, by using the fourth information to recognize the height of the plurality of sub-spaces, the three-dimensional data decoder can determine the position and extent of each sub-space within the target space. Based on the determined position and extent of the plurality of sub-spaces, the three-dimensional data decoder can combine the plurality of sub-spaces.

[0421] For example, the bitstream includes fifth information (e.g., global_position or relative_position) indicating the position of each of the multiple sub-sections. In reconstructing the target space, the three-dimensional data decoder further uses the fifth information to combine the multiple sub-spaces. For example, by using the fifth information to recognize the positions of the multiple sub-spaces, the three-dimensional data decoder can grasp the position of each sub-space within the target space. Based on the grasped positions of the multiple sub-spaces, the three-dimensional data decoder can combine the multiple sub-spaces.

[0422] For example, the bitstream includes a sixth piece of information (e.g., tile_number) indicating the number of sub-intervals. The three-dimensional data decoder further uses the sixth piece of information to combine the sub-spaces in the reconstruction of the target space.

[0423] For example, the bitstream includes seventh information indicating the intervals between the multiple sub-intervals. In reconstructing the target space, the three-dimensional data decoder further uses the seventh information to combine the multiple sub-spaces. For example, by using the seventh information to recognize the intervals between the multiple sub-spaces, the three-dimensional data decoder can grasp the position and range of each sub-space within the target space. Based on the grasped positions and ranges of the multiple sub-spaces, the three-dimensional data decoder can combine the multiple sub-spaces.

[0424] For example, a three-dimensional data decoding device comprises a processor and memory, and the processor uses the memory to perform the above processing.

[0425] (Embodiment 7) This embodiment describes the processing of division units that do not contain points (e.g., tiles or slices). First, the method for dividing point cloud data will be described.

[0426] In video encoding standards such as HEVC, data exists for every pixel in a two-dimensional image. Therefore, even if the two-dimensional space is divided into multiple data regions, data exists in all data regions. On the other hand, in encoding three-dimensional point cloud data, the points themselves, which are elements of the point cloud data, are the data, and it is possible that data does not exist in some regions.

[0427] There are various methods for spatially dividing point cloud data, but these methods can be classified based on whether the divided data unit (e.g., tile or slice) always contains one or more point data points.

[0428] A division method in which each of the multiple division units contains at least one point data is called the first division method. One example of the first division method is dividing point cloud data while considering the encoding processing time or the size of the encoded data. In this case, the number of points in each division unit is approximately equal.

[0429] Figure 65 shows examples of division methods. For example, as a first division method, as shown in Figure 65(a), a method may be used to divide points belonging to the same space into two identical spaces. Alternatively, as shown in Figure 65(b), a space may be divided into multiple subspaces (division units) such that each division unit contains a point.

[0430] Because these methods involve point-based divisions, every division unit always contains at least one point.

[0431] A division method in which one or more division units may contain no point data is called a second division method. For example, as a second division method, a method of dividing the space equally can be used, as shown in Figure 65(c). In this case, points are not necessarily present in the division units. In other words, there may be cases where no points exist in the division units.

[0432] When a three-dimensional data encoding device divides point cloud data, it may indicate in the division supplement information (metadata) related to the division (e.g., tile supplement information or slice supplement information) whether (1) a division method was used in which all of the multiple division units contain one or more point data, (2) a division method was used in which one or more of the multiple division units do not contain point data, or (3) a division method was used in which one or more of the multiple division units may contain one or more division units that do not contain point data, and transmit the division supplement information.

[0433] Furthermore, the three-dimensional data encoding device may indicate the above information as the type of division method. Alternatively, the three-dimensional data encoding device may perform division using a predetermined division method and not transmit additional division information. In this case, the three-dimensional data encoding device will clearly indicate in advance whether the division method is the first division method or the second division method.

[0434] The following describes the second partitioning method and an example of generating and transmitting encoded data. While tiling is used as an example of a three-dimensional space partitioning method, it is not limited to tiling, and the following techniques can be applied to partitioning methods using different units than tiles. For example, tiling may be replaced with slicing.

[0435] Figure 66 shows an example of dividing point cloud data into six tiles. Figure 66 shows an example where the smallest unit is a point, and demonstrates how to divide both geometry and attribute information together. The same applies when geometry and attribute information are divided using separate division methods or numbers, when there is no attribute information, and when there are multiple attribute information.

[0436] In the example shown in Figure 66, after tile division, there are tiles that contain dots (#1, #2, #4, #6) and tiles that do not contain dots (#3, #5). Tiles that do not contain dots are called null tiles.

[0437] Furthermore, any method of division is acceptable, not limited to dividing into six tiles. For example, the division unit may be a cube, a rectangular prism, a cylinder, or any other shape that is not cubic. Multiple division units may be the same shape, or they may include different shapes. In addition, a predetermined method of division may be used, or different methods may be used for each predetermined unit (e.g., PCC frame).

[0438] In this partitioning method, when point cloud data is divided into tiles, if there is no data within a tile, a bitstream is generated that includes information indicating that the tile is a null tile.

[0439] The following describes the method for sending null tiles and the method for signaling null tiles. The three-dimensional data encoding device may generate and send the following information as additional information (metadata) related to data division. Figure 67 shows an example of the syntax of tile additional information (TileMetaData). The tile additional information includes division method information (type_of_divide), division method null information (type_of_divide_null), number of tile divisions (number_of_tiles), and tile null flag (tile_null_flag).

[0440] The division method information (type_of_divide) is information about the division method or division type. For example, the division method information indicates one or more division methods or division types. For example, division methods include top view division and equal division. Note that if there is only one definition for the division method, the division method information does not need to be included in the tile addition information.

[0441] The null information for the division method (type_of_divide_null) indicates whether the division method used is the first division method or the second division method described below. Here, the first division method is a division method in which all of the multiple division units always contain at least one point data. The second division method is a division method in which at least one of the multiple division units does not contain point data, or a division method in which there is a possibility that at least one of the multiple division units does not contain point data.

[0442] Furthermore, the tile information may include at least one of the following as division information for the entire tile: (1) information indicating the number of divisions of the tile (number of tile divisions (number_of_tiles)), or information for identifying the number of divisions of the tile; (2) information indicating the number of null tiles, or information for identifying the number of null tiles; and (3) information indicating the number of tiles other than null tiles, or information for identifying the number of tiles other than null tiles. In addition, the tile information may include information indicating the shape of the tile, or information indicating whether or not tiles overlap, as division information for the entire tile.

[0443] Furthermore, the tile information sequentially indicates the division information for each tile. For example, the order of the tiles is predetermined for each division method and is known to the three-dimensional data encoding device and the three-dimensional data decoding device. If the order of the tiles is not predetermined, the three-dimensional data encoding device may send information indicating the order to the three-dimensional data decoding device.

[0444] The tile division information includes a tile null flag (tile_null_flag) which indicates whether or not data (points) exist within the tile. Note that if there is no data within a tile, the tile null flag may also be included in the tile division information.

[0445] Furthermore, if a tile is not a null tile, the tile information includes segmentation information for each tile (position information (e.g., coordinates of the origin (origin_x, origin_y, origin_z)) and tile height information, etc.). However, if a tile is a null tile, the tile information does not include segmentation information for each tile.

[0446] For example, if the tile division information stores slice division information for each tile, the 3D data encoding device does not need to store slice division information for null tiles in the additional information.

[0447] In this example, the number of tiles (number_of_tiles) refers to the number of tiles, including null tiles. Figure 68 shows an example of tile index information (idx). In the example shown in Figure 68, index information is also assigned to null tiles.

[0448] Next, we will explain the data structure and transmission method of encoded data including null tiles. Figures 69 to 71 show the data structure when location information and attribute information are divided into six tiles, and no data exists in the third and fifth tiles.

[0449] Figure 69 shows an example of the dependencies between data. The tip of the arrow in the figure indicates the dependent, and the base of the arrow indicates the dependent. Also, in the same figure, G tn (n is 1-6) indicates the location information of tile number n, A tn This indicates the attribute information of tile number n. tile This indicates additional information about the tile.

[0450] Figure 70 shows an example of the configuration of the output data, which is encoded data sent from a three-dimensional data encoding device. Figure 71 shows the configuration of the encoded data and the method of storing the encoded data in the NAL unit.

[0451] As shown in Figure 71, the headers of the location information (divided location information) and attribute information (divided attribute information) data each contain the tile index information (tile_idx).

[0452] Furthermore, as shown in Structure 1 of Figure 70, the three-dimensional data encoding device does not need to transmit location information or attribute information that constitutes a null tile. Alternatively, as shown in Structure 2 of Figure 70, the three-dimensional data encoding device may transmit information indicating that the tile is a null tile as data for the null tile. For example, the three-dimensional data encoding device may indicate that the type of the data is a null tile in the tile_type stored in the header of the NAL unit or in the header within the NAL unit payload (nal_unit_payload), and transmit the header. Note that the following explanation will assume Structure 1.

[0453] In Structure 1, if a null tile exists, the value of the tile index information (tile_idx) included in the header of the location information data or attribute information data in the transmitted data will be gappy and not continuous.

[0454] Furthermore, the three-dimensional data encoding device sends out data so that the referenced data is decoded before the referenced data when there are dependencies between data. Note that attribute information tiles have dependencies on location information tiles. Attribute information and location information with dependencies are assigned the same tile index number.

[0455] The tile addition information related to tile division may be stored in both the location information parameter set (GPS) and the attribute information parameter set (APS), or in either one. If tile addition information is stored in either GPS or APS, the other GPS or APS may store reference information indicating the referenced GPS or APS. Furthermore, if the tile division method differs between location information and attribute information, different tile addition information will be stored in GPS and APS, respectively. Also, if the tile division method is the same for a sequence (multiple PCC frames), the tile addition information may be stored in GPS, APS, or SPS (sequence parameter set).

[0456] For example, if tile information is stored in both GPS and APS, the GPS will store tile information for location information, and the APS will store tile information for attribute information. Also, if tile information is stored in common information such as SPS, tile information used in common for both location information and attribute information may be stored, or tile information for location information and tile information for attribute information may be stored separately.

[0457] The following explains the combination of tile partitioning and slice partitioning. First, we will explain the data structure and data transmission when tile partitioning is performed after slice partitioning.

[0458] Figure 72 shows an example of the data dependencies when tiling is performed after slicing. In the figure, the tip of the arrow indicates the dependent data, and the base of the arrow indicates the dependent data. Also, in the figure, data shown with solid lines is data that is actually sent, and data shown with dotted lines is data that is not sent.

[0459] In the same figure, G represents location information and A represents attribute information. s1 This indicates the position information of slice number 1, G s2 This indicates the location information for slice number 2. G s1t1 This indicates the location information of slice number 1 and tile number 1, G s2t2 This indicates the location information for slice number 2 and tile number 2. Similarly, A s1 This shows the attribute information of slice number 1, A s2 This shows the attribute information for slice number 2. s1t1 This indicates the attribute information of slice number 1 and tile number 1, A s2t1 This indicates the attribute information for slice number 2 and tile number 1.

[0460] Mslice indicates slice information, MGtile indicates position tile information, and MAtile indicates attribute tile information. s1t1 Attribute information A s1t1 This shows the dependency information of D s2t1 Attribute information As2t1 This shows the dependency information.

[0461] The three-dimensional data encoding device does not need to generate and transmit positional information and attribute information related to null tiles.

[0462] Furthermore, even if the number of tile divisions is the same for all slices, the number of tiles generated and sent between slices may differ. For example, if the number of tile divisions for location information and attribute information differs, a null tile may exist in one of the location information or attribute information while it does not exist in the other. In the example shown in Figure 72, the location information of slice 1 (G s1 ) is G s1t1 and G s1t2 It is divided into two tiles, and of these, G s1t2 This is a null tile. On the other hand, the attribute information of slice 1 (A s1 ) is not divided and becomes one A s1t1 Null tiles exist, but null tiles do not.

[0463] Furthermore, the three-dimensional data encoding device generates and transmits attribute information dependency information if data exists in at least one attribute information tile, regardless of whether a null tile is included in the slice of position information. For example, when the three-dimensional data encoding device stores slice division information for each tile in the slice division information included in the slice addition information related to slice division, it stores information in this information indicating whether or not the tile is a null tile.

[0464] Figure 73 shows an example of the data decoding order. In the example in Figure 73, decoding is performed sequentially from left to right. The 3D data decoding device decodes dependent data first among dependent data. For example, the 3D data encoding device pre-arranges and sends the data in this order. Any order is acceptable as long as the dependent data comes first. The 3D data encoding device may also send additional information and dependency information before the data.

[0465] Next, we will explain the data structure and data transmission when performing slice division after tile division.

[0466] Figure 74 shows an example of the data dependencies when slicing is performed after tiling. In the figure, the tip of the arrow indicates the dependent data, and the base of the arrow indicates the dependent data. Also, in the figure, data shown with solid lines is data that is actually sent, and data shown with dotted lines is data that is not sent.

[0467] In the same figure, G represents location information and A represents attribute information. t1 This indicates the location information for tile number 1. G t1s1 This indicates the location information of tile number 1 and slice number 1, G t1s2 This indicates the location information for tile number 1 and slice number 2. Similarly, A t1 This shows the attribute information of tile number 1, A t1s1 This indicates the attribute information for tile number 1 and slice number 1.

[0468] Mtile indicates tile information, MGslice indicates position slice information, and MAslice indicates attribute slice information. t1s1 Attribute information A t1s1 This shows the dependency information of D t2s1 Attribute information A t2s1 This shows the dependency information.

[0469] The three-dimensional data encoding device does not slice null tiles. Furthermore, it does not need to generate or transmit positional information, attribute information, or dependency information related to the null tiles.

[0470] Figure 75 shows an example of the data decoding order. In the example in Figure 75, decoding is performed sequentially from left to right. The 3D data decoding device decodes dependent data first among dependent data. For example, the 3D data encoding device pre-arranges and sends the data in this order. Any order is acceptable as long as the dependent data comes first. The 3D data encoding device may also send additional information and dependency information before the data.

[0471] Next, we will explain the process of dividing and joining point cloud data. While we will explain tile and slice division as examples, similar methods can be applied to other spatial divisions.

[0472] Figure 76 is a flowchart of a three-dimensional data encoding process, including data partitioning by a three-dimensional data encoding device. First, the three-dimensional data encoding device determines the partitioning method to be used (S5101). Specifically, the three-dimensional data encoding device decides whether to use the first partitioning method or the second partitioning method. For example, the three-dimensional data encoding device may determine the partitioning method based on a specification from the user or an external device (e.g., a three-dimensional data decoding device), or it may determine the partitioning method according to the input point cloud data. The partitioning method to be used may also be predetermined.

[0473] Here, the first division method is a division method in which all of the multiple division units (tiles or slices) always contain at least one point data. The second division method is a division method in which there is at least one division unit that does not contain point data among the multiple division units, or a division method in which there is a possibility that there is at least one division unit that does not contain point data among the multiple division units.

[0474] If the determined division method is the first division method (first division method in S5102), the three-dimensional data encoding device indicates that the division method used is the first division method in the division addition information (e.g., tile addition information or slice addition information), which is metadata related to data division (S5103). Then, the three-dimensional data encoding device encodes all division units (S5104).

[0475] On the other hand, if the determined division method is the second division method (second division method in S5102), the three-dimensional data encoding device indicates that the division method used in the division supplement information is the second division method (S5105). Then, the three-dimensional data encoding device encodes the division units from among the multiple division units, excluding division units that do not contain point data (e.g., null tiles) (S5106).

[0476] Figure 77 is a flowchart of the three-dimensional data decoding process, including data merging by the three-dimensional data decoding device. First, the three-dimensional data decoding device refers to the partitioning information contained in the bitstream and determines whether the partitioning method used is the first partitioning method or the second partitioning method (S5111).

[0477] If the division method used is the first division method (first division method in S5112), the three-dimensional data decoder receives encoded data for all division units and decodes the received encoded data to generate decoded data for all division units (S5113). Next, the three-dimensional data decoder reconstructs the three-dimensional point cloud using the decoded data for all division units (S5114). For example, the three-dimensional data decoder reconstructs the three-dimensional point cloud by combining multiple division units.

[0478] On the other hand, if the division method used is the second division method (second division method in S5112), the three-dimensional data decoding device receives encoded data of division units containing point data and encoded data of division units not containing point data, and generates decoded data by decoding the received encoded data of division units (S5115). Note that if no division units without point data are transmitted, the three-dimensional data decoding device does not need to receive and decode division units without point cloud data. Next, the three-dimensional data decoding device reconstructs the three-dimensional point cloud using the decoded data of division units containing point data (S5116). For example, the three-dimensional data decoding device reconstructs the three-dimensional point cloud by combining multiple division units.

[0479] The following describes other methods for dividing point cloud data. When dividing space equally, as shown in Figure 65(c), there may be cases where no points exist in the divided space. In this case, the three-dimensional data encoding device combines the space without points with other spaces that do contain points. This allows the three-dimensional data encoding device to form multiple division units such that all division units contain at least one point.

[0480] Figure 78 is a flowchart of the data partitioning in this case. First, the three-dimensional data encoding device partitions the data in a specific way (S5121). For example, the specific way is the second partitioning method described above.

[0481] Next, the three-dimensional data encoding device determines whether or not the target division unit, which is the division unit to be processed, contains a point (S5122). If the target division unit contains a point (Yes in S5122), the three-dimensional data encoding device encodes the target division unit (S5123). On the other hand, if the target division unit does not contain a point (No in S5122), the three-dimensional data encoding device combines the target division unit with other division units that contain a point, and encodes the combined division unit (S5124). In other words, the three-dimensional data encoding device encodes the target division unit together with other division units that contain a point.

[0482] Here, we have described an example in which judgment and merging are performed for each division unit, but the processing method is not limited to this. For example, the three-dimensional data encoding device may determine whether each of the multiple division units contains a point, merge them so that there are no division units that do not contain a point, and then encode each of the merged multiple division units.

[0483] Next, we will explain how to send data that includes null tiles. The three-dimensional data encoding device does not send data for a target tile if that tile is a null tile. Figure 79 is a flowchart of the data transmission process.

[0484] First, the three-dimensional data encoding device determines the tile division method and divides the point cloud data into tiles using the determined division method (S5131).

[0485] Next, the three-dimensional data encoding device determines whether the target tile is a null tile (S5132). In other words, the three-dimensional data encoding device determines whether or not there is no data in the target tile.

[0486] If the target tile is a null tile (Yes in S5132), the 3D data encoding device indicates in the tile addition information that the target tile is a null tile, and does not indicate the information of the target tile (such as the tile's position and size) (S5133). In addition, the 3D data encoding device does not send out the target tile (S5134).

[0487] On the other hand, if the target tile is not a null tile (No in S5132), the 3D data encoding device indicates in the tile addition information that the target tile is not a null tile and displays information for each tile (S5135). The 3D data encoding device also sends out the target tile (S5136).

[0488] In this way, by not including information about null tiles in the tile appending information, the amount of information in the tile appending information can be reduced.

[0489] The following describes how to decode encoded data that includes null tiles. First, we will explain how to handle cases where there is no packet loss.

[0490] Figure 80 shows an example of transmitted data, which is encoded data sent from a three-dimensional data encoding device, and received data, which is input to a three-dimensional data decoding device. Note that this assumes a system environment without packet loss, and the received data is the same as the transmitted data.

[0491] In a system environment without packet loss, the three-dimensional data decoding device receives all of the transmitted data. Figure 81 is a flowchart of the processing performed by the three-dimensional data decoding device.

[0492] First, the three-dimensional data decoding device refers to the tile addition information (S5141) and determines whether each tile is a null tile or not (S5142).

[0493] If the tile information indicates that the target tile is not a null tile (No in S5142), the 3D data decoding device determines that the target tile is not a null tile and decodes the target tile (S5143). Next, the 3D data decoding device obtains tile information (tile position information (origin coordinates, etc.) and size, etc.) from the tile information and reconstructs the 3D data by combining multiple tiles using the obtained information (S5144).

[0494] On the other hand, if the tile information indicates that the target tile is not a null tile (Yes in S5142), the 3D data decoding device determines that the target tile is a null tile and does not decode the target tile (S5145).

[0495] Furthermore, the three-dimensional data decoding device may determine that missing data points are null tiles by sequentially analyzing the index information shown in the header of the encoded data. Alternatively, the three-dimensional data decoding device may combine a determination method using tile addition information with a determination method using index information.

[0496] Next, we will explain how to handle cases where packet loss occurs. Figure 82 shows an example of transmitted data sent from a three-dimensional data encoding device and received data input to a three-dimensional data decoding device. Here, we assume a system environment where packet loss occurs.

[0497] In a system environment with packet loss, the 3D data decoder may not be able to receive all of the transmitted data. In this example, G t2 and A t2 The packets have been lost.

[0498] Figure 83 is a flowchart of the processing of the three-dimensional data decoding device in this case. First, the three-dimensional data decoding device analyzes the continuity of the index information shown in the header of the encoded data (S5151) and determines whether or not the index number of the target tile exists (S5152).

[0499] If an index number exists for the target tile (Yes in S5152), the 3D data decoding device determines that the target tile is not a null tile and performs the decoding process for the target tile (S5153). Next, the 3D data decoding device obtains tile information (tile position information (origin coordinates, etc.) and size, etc.) from the tile addition information and reconstructs the 3D data by combining multiple tiles using the obtained information (S5154).

[0500] On the other hand, if index information for the target tile does not exist (No in S5152), the three-dimensional data decoding device determines whether the target tile is a null tile by referring to the tile addition information (S5155).

[0501] If the target tile is not a null tile (No in S5156), the 3D data decoding device determines that the target tile is lost (packet loss) and performs error decoding (S5157). Error decoding is, for example, a process that attempts to decode the original data as if the data had been there. In this case, the 3D data decoding device may regenerate the 3D data and perform reconstruction of the 3D data (S5154).

[0502] On the other hand, if the target tile is a null tile (Yes in S5156), the 3D data decoding device does not perform decoding or reconstruction of the 3D data, treating the target tile as a null tile (S5158).

[0503] Next, we will explain the encoding method when null tiles are not explicitly defined. The three-dimensional data encoding device may generate encoded data and additional information using the following method.

[0504] The 3D data encoding device does not include information about null tiles in the tile appending information. The 3D data encoding device adds the index numbers of tiles excluding null tiles to the data header. The 3D data encoding device does not transmit null tiles.

[0505] In this case, the number of tile divisions (number_of_tiles) indicates the number of divisions excluding null tiles. The 3D data encoding device may also store information indicating the number of null tiles separately in the bitstream. Furthermore, the 3D data encoding device may include information about null tiles in the additional information, or include some information about null tiles.

[0506] Figure 84 is a flowchart of the three-dimensional data encoding process performed by the three-dimensional data encoding device in this case. First, the three-dimensional data encoding device determines the tile division method and divides the point cloud data into tiles using the determined division method (S5161).

[0507] Next, the three-dimensional data encoding device determines whether the target tile is a null tile (S5162). In other words, the three-dimensional data encoding device determines whether or not there is no data in the target tile.

[0508] If the target tile is not a null tile (No in S5162), the 3D data encoding device adds index information for the tiles excluding null tiles to the data header (S5163). Then, the 3D data encoding device sends out the target tile (S5164).

[0509] On the other hand, if the target tile is a null tile (Yes in S5162), the three-dimensional data encoding device does not add index information for the target tile to the data header, nor does it send out the target tile.

[0510] Figure 85 shows an example of index information (idx) added to the data header. As shown in Figure 85, index information is not added to null tiles, and sequential numbers are added to tiles other than null tiles.

[0511] Figure 86 shows an example of the dependencies between data. The tip of the arrow in the figure indicates the dependent, and the base of the arrow indicates the dependent. Also, in the same figure, G tn (n is 1-4) indicates the location information of tile number n, A tn This indicates the attribute information of tile number n. tile This indicates additional information about the tile.

[0512] Figure 87 shows an example of the configuration of the transmitted data, which is encoded data sent out from a three-dimensional data encoding device.

[0513] The following describes the decoding method when null tiles are not explicitly specified. Figure 88 shows an example of transmitted data sent from a three-dimensional data encoding device and received data input to a three-dimensional data decoding device. This assumes a system environment with packet loss.

[0514] Figure 89 is a flowchart of the processing of the three-dimensional data decoding device in this case. First, the three-dimensional data decoding device analyzes the tile index information shown in the header of the encoded data and determines whether or not the index number of the target tile exists. The three-dimensional data decoding device also obtains the number of tile divisions from the tile addition information (S5171).

[0515] If an index number for the target tile exists (Yes in S5172), the 3D data decoding device performs the decoding process for the target tile (S5173). Next, the 3D data decoding device obtains tile information (tile position information (origin coordinates, etc.) and size, etc.) from the tile addition information and reconstructs the 3D data by combining multiple tiles using the obtained information (S5175).

[0516] On the other hand, if the index number of the target tile does not exist (No in S5172), the three-dimensional data decoding device determines that the target tile is a packet loss and performs error decoding (S5174). In addition, the three-dimensional data decoding device determines that any space not present in the data is a null tile and reconstructs the three-dimensional data.

[0517] Furthermore, by explicitly indicating null tiles, the three-dimensional data encoding device can appropriately determine that there are no points within a tile, rather than due to measurement errors, data loss due to data processing, or packet loss.

[0518] Furthermore, the three-dimensional data encoding device may use both a method that explicitly indicates null packets and a method that does not explicitly indicate null packets. In this case, the three-dimensional data encoding device may indicate in the tile addition information whether or not to explicitly indicate null packets. Alternatively, depending on the type of partitioning method, the device may decide in advance whether or not to explicitly indicate null packets, and the three-dimensional data encoding device may indicate whether or not to explicitly indicate null packets by indicating the type of partitioning method.

[0519] Furthermore, while Figure 67 and other figures show an example where the tile information includes information relating to all tiles, the tile information may also include information relating to some of the tiles among a group of tiles, or it may include information relating to the null tiles of some of the tiles among a group of tiles.

[0520] Furthermore, while we have described an example in which information related to segmented data (tiles), such as whether or not segmented data (tiles) exist, is stored in the tile addition information, some or all of this information may be stored in the parameter set or as data. If this information is stored as data, for example, a nal_unit_type may be defined to indicate whether or not segmented data exists, and this information may be stored in the NAL unit. Alternatively, this information may be stored in both the addition information and the data.

[0521] As described above, the three-dimensional data encoding device according to this embodiment performs the processing shown in Figure 90. First, the three-dimensional data encoding device generates multiple encoded data by encoding multiple subspaces (e.g., tiles or slices) obtained by dividing the target space containing multiple three-dimensional points (S5181). The three-dimensional data encoding device generates a bitstream containing the multiple encoded data and first information (e.g., tile_null_flag) corresponding to each of the multiple subspaces (S5182). Each of the multiple first information indicates whether or not second information showing the structure of the corresponding subspace is included in the bitstream.

[0522] According to this, for example, the second piece of information can be omitted for subspaces that do not contain points, thus reducing the amount of data in the bitstream.

[0523] For example, the second piece of information includes information indicating the coordinates of the origin of the corresponding subspace. For example, the second piece of information includes information indicating at least one of the height, width, and depth of the corresponding subspace.

[0524] According to this, the three-dimensional data encoding device can reduce the amount of data in the bitstream.

[0525] Furthermore, as shown in Figure 78, the three-dimensional data encoding device may divide the target space containing multiple three-dimensional points into multiple subspaces (e.g., tiles or slices), combine the multiple subspaces according to the number of three-dimensional points contained in each subspace, and encode the combined subspaces. For example, the three-dimensional data encoding device may combine multiple subspaces such that the number of three-dimensional points contained in each of the combined subspaces is equal to or greater than a predetermined number. For example, the three-dimensional data encoding device may combine a subspace that does not contain three-dimensional points with a subspace that contains three-dimensional points.

[0526] According to this, the three-dimensional data encoding device can suppress the generation of subspaces with a small number of points or no points at all, thereby improving encoding efficiency.

[0527] For example, a three-dimensional data encoding device comprises a processor and memory, and the processor uses the memory to perform the above processing.

[0528] Furthermore, the three-dimensional data decoding device according to this embodiment performs the processing shown in Figure 91. First, the three-dimensional data decoding device obtains a plurality of first pieces of information (e.g., tile_null_flag) from the bitstream, each corresponding to a plurality of subspaces (e.g., tiles or slices) obtained by dividing the target space containing a plurality of three-dimensional points, and each piece of first piece of information indicates whether or not the bitstream contains second piece of information indicating the structure of the corresponding subspace (S5191). Using the plurality of first pieces of information, the three-dimensional data decoding device (i) reconstructs the plurality of subspaces by decoding the plurality of encoded data generated by encoding the plurality of subspaces contained in the bitstream, and (ii) reconstructs the target space by combining the plurality of subspaces (S5192). For example, the three-dimensional data decoding device uses the first piece of information to determine whether or not the second piece of information is contained in the bitstream, and if the bitstream contains the second piece of information, it uses the second piece of information to combine the plurality of decoded subspaces.

[0529] According to this, for example, the second piece of information can be omitted for subspaces that do not contain points, thus reducing the amount of data in the bitstream.

[0530] For example, the second piece of information includes information indicating the coordinates of the origin of the corresponding subspace. For example, the second piece of information includes information indicating at least one of the height, width, and depth of the corresponding subspace.

[0531] According to this, a three-dimensional data decoding device can reduce the amount of data in the bitstream.

[0532] Furthermore, the three-dimensional data decoding device may divide a target space containing multiple three-dimensional points into multiple subspaces (e.g., tiles or slices), combine the multiple subspaces according to the number of three-dimensional points contained in each subspace, and receive encoded data generated by encoding the combined subspaces, and decode the received encoded data. For example, the encoded data may be generated by combining multiple subspaces such that the number of three-dimensional points contained in each of the combined subspaces is greater than or equal to a predetermined number. For example, the three-dimensional data may be generated by combining a subspace that does not contain three-dimensional points with a subspace that contains three-dimensional points.

[0533] According to this, the three-dimensional data device can decode encoded data with improved encoding efficiency by suppressing the generation of subspaces with a small number of points or no points at all.

[0534] For example, a three-dimensional data decoding device comprises a processor and memory, and the processor uses the memory to perform the above processing.

[0535] (Embodiment 8) Figure 92 is a block diagram showing an example of the configuration of a three-dimensional data encoding device according to this embodiment. Figure 93 is a diagram illustrating the schematic of the encoding method using the three-dimensional data encoding device according to this embodiment.

[0536] The 3D data encoding device 6800 generates segmented data, in which point cloud data is divided into multiple parts such as tiles or slices, and encodes each of these segmented data. Segmented data is also called sub-point cloud data. Point cloud data is data that indicates multiple three-dimensional positions in three-dimensional space. Multiple segmented data consists of multiple sub-point cloud data obtained by dividing the three-dimensional space in which the point cloud data is located into multiple sub-spaces. The number of divisions, i.e., the number of segmented data, may be 1, indicating no division, or it may be 2 or more.

[0537] Figure 92 illustrates the configuration of a three-dimensional data encoding device 6800 that divides data into two parts. Figure 93 shows an example of dividing point cloud data into four parts. In Figure 93, the space to be divided is explained using a two-dimensional space as an example, but it may also be a one-dimensional or three-dimensional space.

[0538] The three-dimensional data encoding device 6800 comprises a division method determination unit 6801, a division unit 6802, quantization units 6803a and 6803b, shift amount calculation units 6804a and 6804b, common position shift units 6805a and 6805b, individual position shift units 6806a and 6806b, and encoding units 6807a and 6807b.

[0539] The division method determination unit 6801 determines the division method for the point cloud data. The division method determination unit 6801 outputs division method information indicating the division method to the division unit 6802 and the shift amount calculation units 6804a and 6804b. Specific examples of division methods will be described later. The three-dimensional data encoding device 6800 does not necessarily have a division method determination unit 6801. In this case, the three-dimensional data encoding device 6800 may divide the point cloud data into multiple divided data using a predetermined division method.

[0540] The division unit 6802 divides the point cloud data into multiple divided data according to the division method determined by the division method determination unit 6801. The multiple divided data divided by the division unit 6802 are processed individually. For this reason, the three-dimensional data encoding device 6800 includes a processing unit that performs subsequent processing for each divided data. Specifically, the three-dimensional data encoding device 6800 includes a quantization unit 6803a, a shift amount calculation unit 6804a, a common position shift unit 6805a, an individual position shift unit 6806a, and an encoding unit 6807a for processing the first divided data. The three-dimensional data encoding device 6800 also includes a quantization unit 6803b, a shift amount calculation unit 6804b, a common position shift unit 6805b, an individual position shift unit 6806b, and an encoding unit 6807b for processing the second divided data. This allows the three-dimensional data encoding device 6800 to perform processing for each of the multiple divided data sets in parallel. While Figure 92 shows an example of a processing unit that processes two divided data sets in parallel, the three-dimensional data encoding device 6800 may have processing units that process three or more divided data sets in parallel. Furthermore, the three-dimensional data encoding device may be configured to process each of the multiple divided data sets with a separate processing unit.

[0541] Each of the quantization units 6803a and 6803b performs scaling (dividing the position information by an arbitrary value) and quantization on the corresponding segmented data. If multiple points overlap, each of the quantization units 6803a and 6803b may delete at least one of the overlapping points, or may not process that at least one point.

[0542] Each of the shift amount calculation units 6804a and 6804b calculates at least one of a common position shift amount and an individual position shift amount for shifting, that is, moving, the position of the corresponding divided data according to the division method determined by the division method determination unit 6801. Depending on the division method, the shift amount calculation units 6804a and 6804b may calculate only the common position shift amount, only the individual position shift amount, or both the common position shift amount and the individual position shift amount.

[0543] The common position shift amount is the amount of shift (movement) that moves the positions of multiple divided data points in common. In other words, the common position shift amount is the same among multiple divided data points. The common position shift amount includes the direction in which the positions of the multiple divided data points are moved, and the distance by which they are moved. The common position shift amount is an example of the first movement amount.

[0544] Individual position shift amounts are the amount of shift (movement) that moves the position of each of the multiple divided data individually. Each individual position shift amount is defined in a one-to-one correspondence with each of the multiple divided data, and in most cases, it differs from one another among the multiple divided data. The individual position shift amount includes the direction and distance of the movement of the corresponding divided data. Individual position shift amounts are an example of a second-order movement amount.

[0545] Each of the common position shift units 6805a and 6805b performs a position shift on the corresponding divided data using the common position shift amount calculated by the shift amount calculation units 6804a and 6804b. As a result, the multiple divided data 6811 to 6814, which are derived from the point cloud data 6810 shown in Figure 93(a), move in the direction and distance indicated by the common position shift amount, as shown in Figure 93(b).

[0546] Each of the individual position shift units 6806a and 6806b performs a position shift on the corresponding segmented data using the individual position shift amount calculated by the shift amount calculation units 6804a and 6804b. As a result, the multiple segmented data 6811 to 6814 shown in Figure 93(c) move in the direction and distance indicated by their respective individual shift amounts.

[0547] Encoding units 6807a and 6807b each encode the corresponding segmented data from among the multiple segmented data moved by the individual position shift units 6806a and 6806b.

[0548] The processing order of the division unit 6802, quantization units 6803a, 6803b, shift amount calculation units 6804a, 6804b, common position shift units 6805a, 6805b, and individual position shift units 6806a, 6806b may be changed. For example, the shift amount calculation units 6804a, 6804b and the common position shift units 6805a, 6805b may be processed before the division unit 6802. In this case, the shift amount calculation units 6804a, 6804b may be merged into a single processing unit, or the common position shift units 6805a, 6805b may be merged into a single processing unit. In this case, the shift amount calculation units 6804a and 6804b only need to calculate at least the common position shift amount among the common position shift amount and the individual position shift amount before the common position shift units 6805a and 6805b, and the individual position shift amount only needs to be calculated before the individual position shift units 6806a and 6806b perform their processing. In other words, the processing unit that calculates the common position shift amount may be configured to calculate the common position shift amount before the common position shift unit, separately from the processing unit that calculates the individual position shift amount. Also, two or more of the above-mentioned processing units may be merged.

[0549] Next, we will explain an example of calculating the common position shift amount and the individual position shift amount using Figure 94. Figure 94 is a diagram illustrating the first example of position shifting. The first example is one in which the point cloud data 6810 is shifted by the common position shift amount, and then multiple divided data 6811 to 6814 are shifted by their respective individual position shift amounts.

[0550] As shown in Figure 94(a), the three-dimensional data encoding device generates a bounding box 6820 that includes all the segmented data 6811-6814 of the point cloud data 6810, and calculates the point of minimum value within the generated bounding box 6820. The three-dimensional data encoding device then calculates the direction and distance of the vector represented by the difference between the calculated minimum value point and the origin as the common position shift amount. Since the origin is 0, the difference is represented by the coordinates of the minimum value point. The origin may be a predetermined reference point that is not 0. The bounding box 6820 may be the smallest rectangular area that encloses all the segmented data 6811-6814. The point of minimum value within the bounding box 6820 is the point closest to the origin within the area of ​​the bounding box 6820. The bounding box 6820 is also called the common bounding box. The bounding box is also called the encoded bounding box.

[0551] As shown in Figure 94(b), the three-dimensional data encoding device moves the multiple divided data 6811 to 6814 by the calculated common position shift amount. Alternatively, the three-dimensional data encoding device may also move the point cloud data 6810 before division by the common position shift amount.

[0552] Next, as shown in Figure 94(b), the three-dimensional data encoding device generates bounding boxes 6821 to 6824 for each of the multiple divided data 6811 to 6814 after a common position shift, and calculates the minimum value point of each generated bounding box 6821 to 6824. Then, for each of the multiple divided data 6811 to 6814, the three-dimensional data encoding device calculates the distance between the minimum value point of the bounding box corresponding to that divided data and the origin as the individual position shift amount of that divided data. Each bounding box 6821 to 6824 may be the smallest rectangular area surrounding the corresponding divided data 6811 to 6814. The minimum value point of the bounding boxes 6821 to 6824 is the point closest to the origin in each area of ​​the bounding boxes 6821 to 6824. Each bounding box, from 6821 to 6824, can also be called an individual bounding box.

[0553] As shown in Figure 94(c), the three-dimensional data encoding device moves each of the multiple divided data 6811 to 6814 by the calculated corresponding individual position shift amount.

[0554] The three-dimensional data encoding device generates a bitstream by encoding each of the multiple segmented data 6811-6814, which have been moved by individual positional shift amounts, using the corresponding bounding boxes 6821-6824. At this time, the three-dimensional data encoding device stores second bounding box information, which indicates the position and size of the minimum point of each bounding box 6821-6824, in the metadata included in the bitstream. Hereafter, the bounding box will also be referred to as the encoded bounding box (encoded BB).

[0555] The common position shift amount and the first bounding box information, which indicates the position and size of the minimum point of bounding box 6820, are stored in the SPS of the bitstream data structure shown in Figure 94(e). The individual position shift amounts are stored in the position information header of the corresponding segmented data. The second bounding box information of bounding boxes 6821 to 6824 used to encode each segmented data 6811 to 6814 is also stored in the position information header of the corresponding segmented data.

[0556] Here, if we denote the common position shift amount as Shift_A and the individual position shift amount as Shift_B(i) (where i is the index of the divided data), then the shift amount Shift(i) of the divided data (i) can be calculated using the following formula.

[0557] Shift(i) = Shift_A + Shift_B(i)

[0558] In other words, as shown in Figure 94(d), the total shift amount for each divided data can be calculated by adding the common position shift amount and the corresponding individual position shift amount.

[0559] The 3D data encoding device performs a positional shift on the point cloud data of the i-th segmented data by subtracting Shift(i) before encoding the point cloud data.

[0560] The 3D data decoding device obtains Shift_A and Shift_B(i) from the SPS and split data headers, calculates Shift(i), and then adds Shift(i) to the decoded split data(i) to return the split data to its original position. This allows for the correct restoration of multiple split data sets.

[0561] Next, a second example of position shifting that performs common position shifting and does not perform individual position shifting will be explained using Figure 95. In this second example, since the individual position shift amount is not transmitted, the amount of information in the bitstream can be reduced.

[0562] Figure 95 is a diagram illustrating a second example of position shifting. In this second example, point cloud data 6810 is shifted using a common position shift amount, while each segmented data 6811-6814 is not shifted using individual position shift amounts.

[0563] As shown in Figure 95(a), the three-dimensional data encoding device generates a bounding box 6820 that includes all of the segmented data 6811-6814 of the point cloud data 6810, and uses the generated bounding box 6820 to calculate the common position shift amount. The method for calculating the common position shift amount is the same as the method explained using Figure 94.

[0564] As shown in Figure 95(b), the three-dimensional data encoding device moves the multiple segmented data 6811 to 6814 by the calculated common position shift amount and encodes them using the bounding box 6820 that includes all the segmented data 6811 to 6814 after the common position shift.

[0565] Thus, in the second example, after dividing the point cloud data 6810, the individual position shift amount or bounding box information for each of the multiple divided data 6811 to 6814 is not calculated. Note that in the second example, the three-dimensional data encoding device shifts the positions of the multiple divided data 6811 to 6814 by a common position shift amount and encodes them using a common bounding box, but this is not limited to this. For example, the three-dimensional data encoding device may shift the positions of the multiple divided data 6811 to 6814 by a common position shift amount and encode them using the individual bounding boxes for each divided data 6811 to 6814. Alternatively, the three-dimensional data encoding device may shift each divided data 6811 to 6814 by an individual position shift amount and encode them using a common bounding box.

[0566] As shown in Figure 95(c), the total shift amount for each divided data is the common position shift amount.

[0567] Furthermore, the common position shift amount, and the first bounding box information indicating the position and size of the minimum point of the bounding box 6820 which includes all the divided data, are stored in the SPS of the bitstream data structure shown in Figure 95(d).

[0568] On the other hand, the header of the corresponding segmented data does not store the individual position shift amount or the second bounding box information used to encode each segmented data.

[0569] Furthermore, a flag (identification information) indicating that the data was encoded using a bounding box containing a common position shift amount and all segmented data 6811-6814, and a flag (identification information) indicating that the individual position shift amounts and the size information of the bounding boxes used to encode the segmented data are not stored in the position information header for each segmented data, are stored in SPS or GPS.

[0570] The three-dimensional data decoding device determines whether the data is encoded with common information or individual information based on the flags stored in the SPS or GPS, and calculates the location information and bounding box size to be used for decoding.

[0571] Hereafter, BB refers to the bounding box. Common information refers to the common position shift amount and first BB information common to multiple divided data. The common position shift amount may be indicated by the minimum value point of the common bounding box. Individual information refers to the individual position shift amount for each divided data and the second BB information of the bounding box for each divided data used for encoding. The individual position shift amount may be indicated by the minimum value point of the bounding box for each divided data. Divided region information is information that indicates the division boundary in space when dividing the data, and may include the minimum value point of the BB and BB information indicating the size of the BB.

[0572] Figure 96 is a flowchart showing an example of an encoding method when switching between the first and second examples. Figure 97 is a flowchart showing an example of a decoding method when switching between the first and second examples.

[0573] As shown in Figure 96, the three-dimensional data encoding device determines the common position shift amount of the point cloud data 6810 and the size of the common broadband surrounding the point cloud data 6810 (S6801).

[0574] The three-dimensional data encoding device determines whether or not to individually shift each segmented data 6811 to 6814 using an individual position shift amount (S6802). The three-dimensional data encoding device may also make this decision based on the amount of header information reduction or the result of calculating the encoding efficiency.

[0575] The three-dimensional data encoding device decides to send a common position shift amount and a common BB size if each segmented data 6811 to 6814 is not shifted individually (No in S6802) (S6803), and decides not to send individual position shift amounts and common BB sizes (S6804). As a result, the three-dimensional data encoding device generates a bitstream that includes the common position shift amount and the common BB size, but does not include individual position shift amounts and individual BB sizes. Identification information indicating that the data is not shifted individually may be stored in the bitstream.

[0576] On the other hand, the three-dimensional data encoding device decides to send a common position shift amount and a common BB size if each of the divided data 6811 to 6814 is shifted individually (Yes in S6802) (S6805), and also decides to send individual position shift amounts and individual BB sizes (S6806). As a result, the three-dimensional data encoding device generates a bitstream that includes the common position shift amount and the common BB size, and also includes the individual position shift amount and individual BB sizes. Identification information indicating that the data is shifted individually may be stored in the bitstream. In addition, the three-dimensional data encoding device may calculate the individual position shift amount and individual BB size for each of the divided data 6811 to 6814 in step S6806.

[0577] As shown in Figure 97, the three-dimensional data decoding device acquires identification information indicating whether or not individual shifts have been performed by acquiring the bitstream (S6811).

[0578] The three-dimensional data decoding device uses identification information to determine whether or not individual shifts were performed during encoding (S6812).

[0579] If the three-dimensional data decoding device determines that individual shifts have not been performed for each segmented data 6811 to 6814 (No in S6812), it obtains the common position shift amount and the common BB size from the bitstream (S6813).

[0580] On the other hand, if the three-dimensional data decoding device determines that individual shifts have been performed for each segmented data 6811 to 6814 (Yes in S6812), it obtains the common position shift amount and the common BB size from the bitstream (S6814), and also obtains the individual position shift amount and the individual BB size (S6815).

[0581] Next, we will explain a third example using Figure 98, in which the position shift is performed using a position shift amount determined using a divided region of the space in which the point cloud data exists. In this third example, the amount of information required for the position shift amount can be further reduced.

[0582] Figure 98 is a diagram illustrating a third example of position shift. In this third example, the total shift amount for each segment data 6811-6814 is represented by three shift amounts.

[0583] The three-dimensional data encoding device calculates the common position shift amount using the bounding box 6820, as shown in Figure 98(a). The method for calculating the common position shift amount is the same as the method explained using Figure 94. At this time, the three-dimensional data encoding device determines multiple division regions 6831 to 6834 (Figure 98(b)) for dividing the point cloud data 6810 into multiple parts, and divides the point cloud data 6810 into multiple divided data 6811 to 6814 according to the determined multiple division regions 6831 to 6834. Each of the multiple division regions 6831 to 6834 corresponds to a region of multiple divided data 6811 to 6814. The multiple division regions 6831 to 6834 are also called division region bounding boxes.

[0584] As shown in Figure 98(b), the three-dimensional data encoding device calculates the position shift amount of the divided regions by the direction and distance of the vector, which is represented by the difference between the minimum value point of the bounding box 6820 containing all the divided data and the minimum value points of each divided region 6831 to 6834.

[0585] As shown in Figure 98(c), the three-dimensional data encoding device generates bounding boxes 6821 to 6824 for each of the multiple divided data 6811 to 6814, with a size that includes the divided data 6811 to 6814, and calculates the minimum value point of each of the generated bounding boxes 6821 to 6824. Then, for each of the multiple divided data 6811 to 6814, the three-dimensional data encoding device calculates the direction and distance of the vector represented by the difference between the minimum value point of the bounding box corresponding to the divided data and the minimum value point of the corresponding divided region as the individual position shift amount of the divided data.

[0586] The three-dimensional data encoding device stores in the bitstream a common position shift amount, individual position shift amounts for each divided data, and bounding box information indicating the size of the bounding box 6820.

[0587] The common position shift amount and the first bounding box information, which indicates the position and size of the minimum point of the bounding box 6820, are stored in the SPS of the bitstream data structure shown in Figure 98(e). The individual position shift amounts are stored in the position information header of the corresponding segmented data. Segmented region information, including the position shift amount for each segmented region, is stored, for example, in the parameter set where the segmentation metadata is stored. The second bounding box information for the bounding boxes 6821 to 6824 used to encode each segmented data 6811 to 6814 is stored in the position information header of the corresponding segmented data. Here, each bounding box 6821 to 6824 used to encode each segmented data 6811 to 6814 is included in each segmented region 6831 to 6834.

[0588] Here, if we denote the common position shift amount as Shift_A, the individual position shift amount as Shift_B(i), and the position shift amount of the divided region as Shift_C(i) (where i is the index of the divided data), then the shift amount Shift(i) of the divided data (i) can be calculated using the following formula.

[0589] Shift(i) = Shift_A + Shift_B(i) + Shift_C(i) In other words, as shown in Figure 98(d), the total shift amount for each divided data can be calculated by adding up three shift amounts: the common position shift amount, the position shift amount for the divided region, and the individual position shift amount.

[0590] The 3D data encoding device performs a positional shift on the point cloud data of the i-th segmented data by subtracting Shift(i) before encoding the point cloud data.

[0591] The 3D data decoding device obtains Shift_A, Shift_B(i), and Shift_C(i) from the SPS and partitioned data headers, calculates Shift(i), and then adds Shift(i) to the decoded partitioned data(i) to return the partitioned data to its original position. This allows for the correct decoding of multiple partitioned data sets.

[0592] This method has the effect of reducing the amount of information required for each divided data by showing the individual position shift amount as the difference from the divided region when transmitting divided region information.

[0593] Furthermore, depending on whether or not the segmented region information is transmitted, the individual position shift amount may be switched between being the difference between the minimum value point of the common bounding box and the minimum value point of the individual bounding box, or the difference between the minimum value point of each segmented region and the minimum value point of the individual bounding box. In the latter case, the individual position shift amount is expressed as the sum of the position shift amount of the segmented region and the calculated difference.

[0594] Figure 99 is a flowchart showing an example of an encoding method when switching between the first and third examples when performing individual position shifts. Figure 100 is a flowchart showing an example of a decoding method when switching between the first and third examples when performing individual position shifts.

[0595] As shown in Figure 99, the three-dimensional data encoding device determines how to divide the point cloud data 6810 (S6821). Specifically, the three-dimensional data encoding device decides whether to shift the point cloud data using the first example or the third example.

[0596] The three-dimensional data encoding device determines, based on the determined division method, whether or not it is a third example method using a divided region (S6822).

[0597] If the three-dimensional data encoding device determines that the method is that of the first example (No in S6822), it sets the individual position shift amount to the difference between the point of minimum value of the common BB and the point of minimum value of the individual BB (S6823).

[0598] The three-dimensional data encoding device generates a bitstream containing common information and individual information (S6824). The bitstream may also contain identification information indicating that it is the method of the first example.

[0599] If the three-dimensional data encoding device determines that the method is that of the third example (Yes in S6822), it sets the individual position shift amount to the difference between the minimum value point of the divided region BB and the minimum value point of the individual BB (S6825).

[0600] The three-dimensional data encoding device sends out a bitstream containing common information, individual information, and partitioned region information (S6826). The bitstream may also contain identification information indicating that it is the method of the third example.

[0601] As shown in Figure 100, the three-dimensional data decoding device acquires a bitstream and determines whether or not the bitstream contains segmented region information (S6831). This allows the three-dimensional data decoding device to determine whether the acquired bitstream contains point cloud data encoded in the first example or in the third example. Specifically, if the bitstream contains segmented region information, it determines that the bitstream contains point cloud data encoded in the third example; otherwise, it determines that the bitstream contains point cloud data encoded in the first example. Alternatively, the three-dimensional data decoding device may determine whether the bitstream was encoded in the first example or the third example by acquiring identification information contained in the bitstream.

[0602] The three-dimensional data decoder obtains common and individual information from the bitstream if the bitstream does not contain partitioning region information (No in S6831), i.e., if it is encoded in the first example (S6832).

[0603] The three-dimensional data decoding device calculates the position shift amount for each segmented data, i.e., the position shift amount Shift(i) in the first example, the common BB, and the individual BB, based on the acquired common and individual information, and uses this information to decode the point cloud data (S6833).

[0604] The three-dimensional data decoder obtains common information, individual information, and partitioned region information from the bitstream if the bitstream contains partitioned region information (Yes in S6831), that is, if it is encoded in the third example (S6834).

[0605] The three-dimensional data decoding device calculates the position shift amount for each divided data, i.e., the position shift amount Shift(i) in the third example, common BB, individual BB, and divided region, based on the acquired common information, individual information, and divided region information, and uses this information to decode the point cloud data (S6835).

[0606] Next, we will explain a fourth example using Figure 101, in which the divided region of the space where the point cloud data exists is shifted by a position shift amount determined as the encoding bounding box. In the fourth example, since the divided region and the individual bounding boxes coincide, the amount of information in the bounding boxes can be reduced compared to the third example.

[0607] Figure 101 is a diagram illustrating a fourth example of position shift. In the fourth example, the total shift amount for each segment data 6811-6814 is represented by a two-stage shift amount: a common position shift amount and a position shift amount for the segmented region.

[0608] The three-dimensional data encoding device calculates the common position shift amount using the bounding box 6820, as shown in Figure 101(a). The method for calculating the common position shift amount is the same as the method explained using Figure 94. At this time, the three-dimensional data encoding device divides the point cloud data 6810 into multiple divided data 6811 to 6814. The method for dividing the point cloud data 6810 is the same as the method explained using Figure 98(a).

[0609] The three-dimensional data encoding device calculates the positional shift amount of the divided region, as shown in Figure 101(b). The method for calculating the positional shift amount of the divided region is the same as the method explained using Figure 98(b).

[0610] The three-dimensional data encoding device calculates the position shift amount of each divided region as an individual position shift amount for each divided data 6811 to 6814. For this reason, the three-dimensional data encoding device stores the common position shift amount, the individual position shift amounts (position shift amounts of the divided regions), and bounding box information indicating the size of the bounding box (divided region) in the bitstream.

[0611] The common position shift amount and the first bounding box information, which indicates the position and size of the minimum point of the bounding box 6820, are stored in the SPS of the bitstream data structure shown in Figure 101(d).

[0612] Furthermore, the individual position shift amount is stored in at least one of the location information header and the division metadata of the corresponding division data. If the individual position shift amount is stored in either the location information header or the division metadata, identification information indicating that the individual position shift amount is stored in the location information header of the division data, or identification information indicating that it is stored in the division metadata, may be stored in the GPS or SPS. Alternatively, the individual position shift amount may be stored in the division data header, and identification information (flag) indicating whether the position shift amount and bounding box information stored in the division data header match the division region may be stored in the GPS or SPS. This allows the 3D data decoder to determine from the above flag that the division region information is stored in the division data header, and to use the position shift amount and bounding box information stored in the division data header as division region information. Alternatively, the above flag may be stored in the division metadata, and if the above flag indicates 1, that is, if the division region information is stored in the division data header, the 3D data decoder may obtain the division region information by referring to the division data header.

[0613] Here, if we denote the common position shift amount as Shift_A, the individual position shift amount as Shift_B(i), and the position shift amount of the divided region as Shift_C(i) (where i is the index of the divided data), then the shift amount Shift(i) of the divided data (i) can be calculated using the following formula.

[0614] Shift_B(i) = Shift_C(i)

[0615] Shift(i) = Shift_A + Shift_B(i)

[0616] In other words, as shown in Figure 101(c), the total shift amount for each divided data can be calculated by adding the common position shift amount and the position shift amount of the divided region (i.e., the individual position shift amount).

[0617] The 3D data encoding device performs a positional shift on the point cloud data of the i-th segmented data by subtracting Shift(i) before encoding the point cloud data.

[0618] The 3D data decoding device obtains Shift_A, Shift_B(i), or Shift_C(i) from the SPS and the header of the divided data, calculates Shift(i), and then adds Shift(i) to the decoded divided data(i) to return the divided data to its original position. This allows for the correct decoding of multiple divided data sets.

[0619] This method uses the position shift amount of the divided region as the individual position shift amount, which eliminates the need to send new divided region information even when it is necessary to send divided region information, thus reducing the amount of information required.

[0620] Furthermore, depending on whether or not the segmented region information is sent, it is possible to switch between using the bounding box information of the individual bounding boxes of each segmented data or using the segmented region information for the individual position shift amount.

[0621] Figure 102 is a flowchart showing an example of an encoding method when switching between the third and fourth examples when storing divided region information.

[0622] The three-dimensional data encoding device determines how to divide the point cloud data 6810, as shown in Figure 102 (S6841). Specifically, the three-dimensional data encoding device decides whether to shift the point cloud data using the third example or the fourth example. The three-dimensional data encoding device may make this decision based on the amount of header information reduction or the result of determining the encoding efficiency.

[0623] The three-dimensional data encoding device determines whether or not to perform a position shift using individual bounding boxes based on the determined division method (S6842). In other words, the three-dimensional data encoding device determines whether to use the method of the third example or the method of the fourth example.

[0624] If the three-dimensional data encoding device determines that the method is that of the third example (Yes in S6842), it sets the individual position shift amount to the difference between the minimum value point of the divided region BB and the minimum value point of the individual BB (S6843).

[0625] The three-dimensional data encoding device generates a bitstream containing common information, individual information, and partitioned region information (S6844). The bitstream may also contain identification information indicating that it is the method of the third example.

[0626] If the three-dimensional data encoding device determines that the method is that of the fourth example (No in S6842), it sets the individual position shift amount to the difference between the point of minimum value of the common BB and the point of minimum value of the divided region BB (S6845).

[0627] The three-dimensional data encoding device sends out a bitstream containing common information and partitioned region information (S6846). The bitstream may also contain identification information indicating that it is the method of the fourth example.

[0628] Next, we will explain a fifth example using Figure 103, in which the position shift is performed using a position shift amount determined using a divided region of the space in which the point cloud data exists. In this fifth example, the amount of information required for the shift amount can be reduced.

[0629] Figure 103 is a diagram illustrating a fifth example of position shifting. The fifth example differs from the third example in that the divided region in the third example is an equally divided region.

[0630] The three-dimensional data encoding device calculates a common position shift amount using the bounding box 6820, as shown in Figure 103(a). The method for calculating the common position shift amount is the same as the method explained using Figure 94. At this time, the three-dimensional data encoding device determines multiple division regions 6841 to 6844 (Figure 103(b)) for dividing the point cloud data 6810 into multiple parts, for example, based on the bounding box 6820 according to predetermined rules. For example, if the bounding box 6820 is divided into N equal parts and N=4, the three-dimensional data encoding device can generate four division regions 6841 to 6844 of equal size. Here, for example, if it is decided that the region number is defined in Morton order as an identifier to identify each division region 6841 to 6844, the position shift amount of the division region can be calculated from the number of divisions and the region number because the size of each division region is the same. For this reason, instead of sending the position shift amount of the division region, the three-dimensional data encoding device may send the number of divisions and identification information to identify the region. As mentioned above, the identification information is the Morton order corresponding to each of the multiple partitioned regions 6841 to 6844.

[0631] The example in Figure 103(c) shows an example where the encoding region is a separate bounding box for each of the multiple divided data. In this case, the position shift amount for each divided data is shown as the sum of the common position shift amount, the position shift amount for the corresponding divided region, and the individual position shift amount from the reference position (the position of the smallest point) of the corresponding divided region. In this case, the corresponding divided region does not have to include the entire region of the encoding region (i.e., the individual BB).

[0632] The example in Figure 103(d) shows how the encoded region of each divided data is matched with each divided region. In this case, the position shift amount of each divided data is shown as the sum of the common position shift amount and the position shift amount of the corresponding divided region. In this case, the corresponding divided region includes the encoded region because it matches the encoded region (i.e., the individual BB).

[0633] Thus, the three-dimensional data encoding device may or may not match the encoding region of each divided data with the division region when dividing the point cloud data. When the encoding region of each divided data and the division region are to be matched, either method in Figure 103(c) or Figure 103(d) can be used. When the encoding region of each divided data and the division region are not to be matched, method in Figure 103(c) can be used.

[0634] When using the methods shown in Figure 103(c) and Figure 103(d), the common position shift amount and bounding box information indicating the position and size of the minimum point of the bounding box 6820 are stored in the SPS of the bitstream data structure. In addition, information regarding the division method and the number of divisions, which are part of the division region information, are stored in the SPS or GPS as information common to all division data, and the number (identification information) of each division region in a predetermined order (Morton order) is stored in the header of the position information of the division data as information for each division region.

[0635] Furthermore, the individual position shift amount is stored in the position information header of the corresponding segmented data. In the case of Figure 103(c), the individual position shift amount is also stored in the position information header of the segmented data.

[0636] Furthermore, when the divided regions and divided data are matched (i.e., when the divided regions and the individual bounding boxes of the divided data are matched), the number of the divided data (tile ID) included in the position information header of the divided data may be treated as a number in a predetermined order for each divided region. In this way, by indicating the amount of individual position shift with identification information for each data (a number in a predetermined order), it is possible to reduce the amount of information in the header.

[0637] Furthermore, when the three-dimensional data encoding device uses the division method of the fifth example, it may indicate the position shift amount of the divided regions using region information that includes the number of divisions and identification information of the divided regions, by using a division method and a predetermined order that equally divides the common bounding box. When using a different division method, it may switch to indicating the position shift amount in a way that does not use the above-mentioned region information.

[0638] Here, if we denote the common position shift amount as Shift_A, the individual position shift amount as Shift_B(i), and the position shift amount of the divided region as Shift_D(i) (where i is the index of the divided data), then the shift amount Shift(i) of the divided data (i) can be calculated using the following formula.

[0639] Shift(i) = Shift_A + Shift_B(i) + Shift_D(i)

[0640] In other words, the total shift amount for each divided data can be calculated by adding up three shift amounts: the common position shift amount, the position shift amount for the divided region, and the individual position shift amount.

[0641] The 3D data encoding device performs a positional shift on the point cloud data of the i-th segmented data by subtracting Shift(i) before encoding the point cloud data.

[0642] The three-dimensional data decoding device obtains Shift_A and Shift_B(i) from the SPS and the header of the divided data. Furthermore, it obtains information about the division method, the number of divisions, and a predetermined number for each divided region as division region information. It derives the position shift amount Shift_D(i) in a predetermined way, calculates Shift(i), and then adds Shift(i) to the decoded divided data(i) to return the divided data to its original position. This allows for the correct decoding of multiple divided data.

[0643] This section describes the rules and specific examples for reducing header size when partitioning point cloud data using an octree. Figure 104 is a diagram illustrating the encoding method when partitioning a three-dimensional space using an octree.

[0644] First, the three-dimensional data encoding device may divide the point cloud data in three-dimensional space into an octree after offsetting (shifting, moving) it by a common position shift amount. The three-dimensional data encoding device divides the bounding box 6850 of the point cloud data into eight division regions using an octree, and sets the number of divisions by the Depth of the octree. For example, the number of divisions for a given Depth is given by N = 2^(Depth * 3), where the number of divisions is 8 when Depth = 1, and 64 when Depth = 2. The order of the division regions is in Morton order. The positional information of the division regions can be calculated from the Morton order by applying the fifth example to three-dimensional space. Note that the Morton order has the characteristic that it can be calculated from the positional information of the division regions using a predetermined method.

[0645] The bounding box information for point cloud data is stored in an SPS or GPS, which contains metadata common to multiple segmented data sets. The bounding box information includes the minimum point (initial position) and size of the bounding box.

[0646] The three-dimensional data encoding device stores identification information indicating that the partitioning method uses an octree, and, if it is an octree partition, depth information indicating the depth of the octree, in the SPS or GPS of the bitstream data structure shown in Figure 104(e). The header of each partitioned data contains a Morton order number as the partitioned data number. Furthermore, if the partitioning method is an octree, the position shift amount of the partitioned data and the bounding box information for encoding are assumed to be derived from the Morton order and are not stored in the bitstream.

[0647] In other words, as shown in Figure 104(f), the three-dimensional data encoding device calculates identification information indicating that the data was divided by an octree, the depth of the octree, and the Morton order of each of the multiple partitioned regions divided by the octree using a predetermined method. The three-dimensional data decoding device obtains the identification information indicating that the data was divided by an octree, the depth of the octree, and the Morton order, as shown in Figure 104(g), and restores the position information of each of the multiple partitioned regions divided by the octree using a predetermined method.

[0648] Note that if there is no point cloud data within a divided region, the information for that divided region is not required. For example, if there is no point cloud data in the region with divided data number 2, that divided data number will not be sent, and the sequence of divided data numbers will skip 2 and become 1, 3, 4.

[0649] The division metadata may store the division data number, and may also store information on all division regions, including those without point cloud data. In this case, the division region information may indicate whether or not point cloud data exists in that division region.

[0650] Figure 105 shows an example of GPS syntax.

[0651] octree_partition_flag is a flag that indicates whether or not the point cloud data is partitioned using an octree partition.

[0652] `depth` indicates the depth of the octree division when the point cloud data is divided into octrees.

[0653] gheader_BBmin_present_flag is a flag that indicates whether the bounding box location information field for the encoded point cloud data is present in the location information header.

[0654] gheader_BBsize_present_flag is a flag that indicates whether the bounding box size information field for the encoded point cloud data is present in the location information header.

[0655] If octree_partition_flag=1, then gheader_BBmin_present_flag and gheader_BBsize_present_flag should be set to 0.

[0656] Figure 106 shows an example of the syntax for a location information header.

[0657] The `partition_id` indicates the identification information of the partitioned data. In the case of an octave tree partition, the `partition_id` indicates a unique position in Morton order.

[0658] BBmin indicates the amount of shift when encoding data or segmented data.

[0659] BBsize indicates the size of the bounding box when encoding data or segmented data.

[0660] Although this explanation uses an octave tree partitioning as an example, this method can be applied to other partitioning methods by defining a predetermined partitioning method, the order of the partitioned data, and the method for calculating positional information. For example, when viewing point cloud data from above and partitioning it at equal intervals in the xy plane, by defining and transmitting the partitioning method, the number or size of the partitions, and the order of the partitions, both the 3D data encoding device and the 3D data decoding device can calculate positional information and positional shift amounts based on this information, and the amount of data can be reduced by not transmitting positional information. Alternatively, information on the plane to be partitioned or information on partitioning with a quadave tree may be transmitted.

[0661] Figure 107 is a flowchart showing an example of an encoding method that switches processing depending on whether or not an octree division is performed. Figure 108 is a flowchart showing an example of a decoding method that switches processing depending on whether or not an octree division is performed.

[0662] The three-dimensional data encoding device determines the method for dividing the point cloud data, as shown in Figure 107 (S6851).

[0663] The three-dimensional data encoding device determines whether or not to perform an octave tree partition based on the determined partitioning method (S6852).

[0664] If the three-dimensional data encoding device determines that an octree division is not necessary (No in S6852), it determines the individual position shift amount to the point of the minimum value of the individual BB for each divided data, shifts the position of the divided data, and encodes the divided data using the individual BB (S6853).

[0665] The three-dimensional data encoding device stores common BB information in common metadata (S6854).

[0666] The three-dimensional data encoding device stores the individual position shift amount for each divided data in the header of the divided data (S6855).

[0667] If the three-dimensional data encoding device determines that an octree division is necessary (Yes in S6852), it determines the individual position shift amount to the point of the minimum value of the divided region of the octree division, shifts the position of the divided data, and encodes the divided data using the divided region of the octree (S6856).

[0668] The three-dimensional data encoding device stores common BB information, identification information indicating that it has been divided into an octave tree, and depth information in common metadata (S6857).

[0669] The three-dimensional data encoding device stores order information in the header of the divided data indicating the Morton order for identifying the individual position shift amount for each divided data (S6858).

[0670] The three-dimensional data decoding device obtains information indicating the method of dividing the point cloud data from common metadata (S6861).

[0671] The three-dimensional data decoding device determines whether the partitioning method is an octree partition based on the acquired information indicating the partitioning method (S6862). Specifically, the three-dimensional data decoding device determines whether the partitioning method is an octree based on identification information indicating whether or not an octree partition was performed.

[0672] If the partitioning method is not an octave partition (No in S6862), the 3D data decoding device acquires information on the common broadband, individual position shift amounts, and individual broadbands, and decodes the point cloud data (S6863).

[0673] The three-dimensional data decoding device, when the partitioning method is an octave tree partition (Yes in S6862), acquires common broadband information, depth information, and individual sequence information, calculates individual position shift amounts and encoded broadband information, and decodes the point cloud data (S6864).

[0674] The methods described in the multiple examples in this embodiment are expected to potentially reduce the amount of code regardless of which method is used. It is also possible to switch to any of the multiple example methods using a predetermined method.

[0675] For example, a three-dimensional data encoding device may calculate the code amount and decide to perform the first of the above multiple methods under predetermined conditions according to the calculated code amount. Alternatively, it may determine if the quantization coefficient is greater than a predetermined value in the case of irreversible compression, if the number of data divisions is greater than a predetermined number, or if the number of point clouds is less than a predetermined number, and switch from the first method to the second of the above multiple methods if the overhead may change.

[0676] Furthermore, while the header of each divided data set stores the difference from the common information as the individual information corresponding to that divided data, it is not limited to this, and the difference from the individual information of the previous divided data set may also be stored as the individual information of that divided data.

[0677] Figure 109 shows an example of a bitstream data structure when the divided data is classified into A data, which is randomly accessible, and B data, which is not. Figure 110 shows an example where the divided data from Figure 109 is used as frames.

[0678] In this case, that is, when there are one or more random access units, the data headers of data A and data headers of data B may each store different information. For example, the segmented data of data A may store individual difference information from common information stored in the GPS, and the segmented data of data B may store difference information with data A in the random access unit. Furthermore, if there are multiple segmented data of data B in a random access unit, each of the multiple segmented data of data B included in the same random access unit may store difference information with data A, or multiple segmented data of data B may store difference information with the previous data A or B data for that data.

[0679] The above explanation focused on segmented data, but the same principles apply to frames.

[0680] In this embodiment, the division region or boundary in data partitioning was mainly explained using an example (Figure 111) in which the BB of all partitioned data is targeted and that region is partitioned. However, even if the division region is in a different case, a similar code reduction effect can be expected by using the method of this embodiment.

[0681] As shown in Figure 111, when dividing a broadband (BB) containing all the divided data, the three-dimensional data encoding device may shift the position of the BB based on its minimum value, and then use the shifted point cloud data as the target for division. Furthermore, if the point cloud data has been scaled or quantized, the three-dimensional data encoding device may use the scaled or quantized point cloud data as the target for division, or it may use the point cloud data before scaling or quantization as the target for division.

[0682] As shown in Figure 112, the three-dimensional data encoding device may set the division region in the coordinate system of the input point cloud data. In this case, the three-dimensional data encoding device does not shift the point of the minimum value of BB in the point cloud data.

[0683] As shown in Figure 113, the three-dimensional data encoding device may define the division region in a higher-level coordinate system of the point cloud data. For example, the three-dimensional data encoding device may define the division region based on GPS coordinates such as those of map data.

[0684] In this case, the three-dimensional data encoding device may store and transmit the relative position information of the point cloud coordinate system with respect to a higher-level coordinate system in a bitstream. For example, when a sensor such as an in-vehicle LiDAR senses point cloud data while moving, the three-dimensional data encoding device may transmit the sensor's position information (GPS coordinates, acceleration, velocity, distance traveled) as relative position information. Also, in the case of point cloud data with a time-series frame structure, the three-dimensional data encoding device may store and transmit the time-series sensor position information for each frame in a bitstream. The higher-level coordinate system may be an absolute coordinate system or a relative coordinate system based on absolute coordinates. The higher-level coordinate system may also be a coordinate system even higher than the higher-level coordinate system.

[0685] As shown in Figure 114, the three-dimensional data encoding device may determine the division regions based on the object or data attributes of the point cloud data. For example, the three-dimensional data encoding device may determine the division regions based on the image recognition results. In this case, the division regions may include overlapping regions. The three-dimensional data encoding device may cluster each point in the point cloud based on predetermined attributes, divide the data based on the number of points, or divide the regions so that each region has approximately equal numbers of points.

[0686] A three-dimensional data encoding device may signal the number of points to be encoded, and a three-dimensional data decoding device may decode the point cloud data using the number of points.

[0687] The data to be encoded is set for each frame or for each segmented data obtained by dividing a frame, and the number of points to be encoded is basically stored in the respective data header. The number of points to be signaled may be determined by storing a reference value in common metadata, and the difference information from the reference value may be stored in the data header.

[0688] The three-dimensional data encoding device stores the number of reference points A (i.e., the reference value) in a GPS or SPS, as shown in Figure 115, and the number of points B(i) is the difference from the number of reference points A. The index of the divided data (i) may be stored in each data header. The number of points in divided data (i) is obtained by adding number B(i) to number A. Therefore, the three-dimensional data decoder calculates A+B(i) and decodes the calculated value as the number of points in divided data (i).

[0689] Alternatively, as shown in Figure 116, the three-dimensional data encoding device may store the difference information from the number of points in the previous segmented data in each data header. In this case, the number of points in segmented data (i) is obtained by adding the number B(1) to the number B(i) of the first segmented data to number A.

[0690] Furthermore, in the case of time-series point cloud data having multiple frame structures, as shown in Figure 117, for example, the three-dimensional data encoding device may store a reference value for the number of points to be encoded in a common SPS, and store the relative values ​​of each frame with respect to the reference value in the GPS or data header. Alternatively, the three-dimensional data encoding device may store the difference (relative value) with respect to the reference value stored in the SPS in the GPS, and further store the difference (relative value) from the sum of the reference value in the SPS and the difference value in the GPS in the data header. This is expected to reduce the overhead coding amount.

[0691] Furthermore, one possible method for dividing the point cloud data is to divide it so that the number of points is approximately equal. In that case, the same method as in Figure 115 can be used.

[0692] For example, as shown in Figure 118, in a method of dividing data so that the difference between points included in the divided data is 1 or less, Δ may be represented by 1 bit of data, and if Δ=0, the difference information does not need to be shown.

[0693] For example, as shown in Figure 119, the data may basically consist of data where the reference value recorded in the GPS is used as the number of divisions, and data where the remainder is used. If Δ=0, the difference information does not need to be shown.

[0694] Using any of the methods shown in Figures 115 to 119 above can be expected to reduce the amount of overhead information.

[0695] As described above, the three-dimensional data encoding device according to this embodiment performs the processing shown in Figure 120. The three-dimensional data encoding device is a device that encodes point cloud data indicating multiple three-dimensional positions in three-dimensional space. The three-dimensional data encoding device moves the point cloud data by a first displacement amount (S6871). Next, the three-dimensional data encoding device divides the point cloud data into multiple sub-point cloud data by dividing the three-dimensional space into multiple sub-spaces (S6872). For each of the multiple sub-point cloud data included in the point cloud data after it has been moved by the first displacement amount, the three-dimensional data encoding device moves the sub-point cloud data by a second displacement amount based on the position in the sub-space containing the sub-point cloud data (S6873). The three-dimensional data encoding device generates a bitstream by encoding the multiple sub-point cloud data after the displacement (S6874). The bitstream includes first displacement information for calculating the first displacement amount and multiple second displacement information for calculating multiple second displacement amounts obtained by moving the multiple sub-point cloud data. According to this method, since the divided sub-point cloud data is moved before encoding, the amount of positional information in each sub-point cloud data can be reduced, thereby improving encoding efficiency.

[0696] For example, multiple subspaces have equal size to each other. Each of the multiple second movement information entries includes the number of multiple subspaces and first identification information for identifying the corresponding subspace. Therefore, the amount of information in the second movement information can be reduced, and coding efficiency can be improved.

[0697] For example, the first identifying information is the Morton order corresponding to each of the multiple subspaces.

[0698] For example, each of the multiple subspaces is a space obtained by dividing a single three-dimensional space using an octree. The bitstream includes a second identification information indicating that the multiple subspaces are spaces divided using an octree, and depth information indicating the depth of the octree. Therefore, since the point cloud data in three-dimensional space is divided using an octree, the amount of information in the positional information of each subpoint cloud data can be reduced, and encoding efficiency can be improved.

[0699] For example, the splitting is performed after the point cloud data has been moved by the first displacement amount.

[0700] For example, a three-dimensional data encoding device comprises a processor and memory, and the processor uses the memory to perform the above processing.

[0701] Furthermore, the three-dimensional data decoding device according to this embodiment performs the processing shown in Figure 121. The three-dimensional data decoding device decodes from the bitstream multiple sub-point cloud data obtained by dividing the three-dimensional space into multiple sub-spaces, each of which is a sub-point cloud data representing multiple three-dimensional positions, and each sub-point cloud data is moved by a first movement amount and a corresponding second movement amount, along with first movement information for calculating the first movement amount and multiple second movement information for calculating the multiple second movement amounts obtained by moving the multiple sub-point cloud data (S6881). The three-dimensional data decoding device restores the point cloud data by moving each of the multiple sub-point cloud data by a movement amount obtained by adding the first movement amount and the corresponding second movement amount (S6882). With this, point cloud data can be correctly decoded using a bitstream with improved encoding efficiency.

[0702] For example, multiple subspaces have equal size to each other. Each of the multiple second movement information entries includes the number of multiple subspaces and first identification information for identifying the corresponding subspace. Therefore, the amount of information in the second movement information can be reduced, and coding efficiency can be improved.

[0703] For example, the first identifying information is the Morton order corresponding to each of the multiple subspaces.

[0704] For example, each of the multiple subspaces is a space obtained by dividing a single three-dimensional space using an octree. The bitstream includes a second identification information indicating that the multiple subspaces are spaces divided using an octree, and depth information indicating the depth of the octree. Therefore, since the point cloud data in three-dimensional space is divided using an octree, the amount of information in the positional information of each subpoint cloud data can be reduced, and encoding efficiency can be improved.

[0705] For example, a three-dimensional data decoding device comprises a processor and memory, and the processor uses the memory to perform the above processing.

[0706] The three-dimensional data encoding device and three-dimensional data decoding device, etc., according to embodiments of the present disclosure have been described above, but the present disclosure is not limited to these embodiments.

[0707] Furthermore, each processing unit included in the three-dimensional data encoding device and the three-dimensional data decoding device, etc., according to the above embodiment is typically implemented as an integrated circuit (LSI). These may be individually integrated into a single chip, or some or all of them may be integrated into a single chip.

[0708] Furthermore, integrated circuit implementation is not limited to LSIs; it may also be achieved using dedicated circuits or general-purpose processors. Field-Programmable Gate Arrays (FPGAs), which can be programmed after LSI manufacturing, or reconfigurable processors, which allow for the reconfiguration of the connections and settings of circuit cells within the LSI, may also be used.

[0709] Furthermore, in each of the above embodiments, each component may be implemented by being composed of dedicated hardware or by executing a software program suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.

[0710] Furthermore, this disclosure may be implemented as a three-dimensional data encoding method or a three-dimensional data decoding method, etc., performed by a three-dimensional data encoding device and a three-dimensional data decoding device, etc.

[0711] Furthermore, the division of functional blocks in the block diagram is just one example; multiple functional blocks can be implemented as a single functional block, a single functional block can be divided into multiple parts, or some functions can be moved to other functional blocks. In addition, the functions of multiple functional blocks with similar functions can be processed in parallel or time-sharing by a single piece of hardware or software.

[0712] Furthermore, the order in which each step in the flowchart is performed is illustrative for the purpose of specifically illustrating this disclosure, and may be in a different order. Also, some of the above steps may be performed simultaneously (in parallel) with other steps.

[0713] Although a three-dimensional data encoding device and a three-dimensional data decoding device, etc., relating to one or more embodiments have been described above based on embodiments, this disclosure is not limited to these embodiments. Without departing from the spirit of this disclosure, various modifications that a person skilled in the art can conceive of may be applied to these embodiments, and forms constructed by combining components from different embodiments may also be included within the scope of one or more embodiments. [Industrial applicability]

[0714] This disclosure is applicable to three-dimensional data encoding devices and three-dimensional data decoding devices. [Explanation of Symbols]

[0715] 2400 3D data encoding device 2401 Quantization section 2402, 2411 8-ary tree generation part 2403 Merge Decision Unit 2404 Entropy coding unit 2410 Three-dimensional data decoding device 2412 Merge Information Decoding Unit 2413 Entropy Decoder 2414 Inverse quantization section 4601 Three-Dimensional Data Encoding System 4602 Three-dimensional data decoding system 4603 Sensor terminal 4604 External connection section 4611 Point Cloud Data Generation System 4612 Presentation section 4613 Encoding section 4614 Multiplexer 4615 Input / output section 4616 Control Unit 4617 Sensor Information Acquisition Unit 4618 Point Cloud Data Generation Unit 4621 Sensor Information Acquisition Unit 4622 Input / output section 4623 Demultiplexer 4624 Decoding section 4625 Presentation section 4626 User Interface 4627 Control Unit 4630 First encoding unit 4631 Location information encoder 4632 Attribute information encoder 4633 Additional Information Encoding Unit 4634 Multiplexer 4640 First Decoding Unit 4641 Demultiplexer 4642 Location Information Decoding Unit 4643 Attribute Information Decoding Unit 4644 Additional Information Decoding Unit 4650 Second encoding unit 4651 Additional Information Generation Unit 4652 Position image generation unit 4653 Attribute Image Generation Unit 4654 Video Encoding Unit 4655 Additional Information Encoding Unit 4656 Multiplexer 4660 Second decoding unit 4661 Demultiplexer 4662 Video Decoding Unit 4663 Additional Information Decoding Unit 4664 Location information generator 4665 Attribute information generation section 4670 Encoding section 4680 Decoding Unit 4710 First Multiplexing Unit 4711 File Conversion Unit 4720 First demultiplexing unit 4721 File Reverse Conversion Unit 4730 Second Multiplexing Unit 4731 File Conversion Unit 4740 Second demultiplexing unit 4741 File Reverse Conversion Unit 4750 Third Multiplexing Unit 4751 File Conversion Unit 4760 Third demultiplexing unit 4761 File Reverse Conversion Unit 4801 Encoding section 4802 Multiplexer 5010 First encoding unit 5011 Split section 5012 Location information encoder 5013 Attribute information encoder 5014 Additional Information Encoding Unit 5015 Multiplexer 5020 First Decoding Unit 5021 Demultiplexer 5022 Location Information Decoding Unit 5023 Attribute Information Decoding Unit 5024 Additional Information Decoding Unit 5025 Joint 5031 Tile division section 5032 Position information slice division section 5033 Attribute Information Slice Division Section 5041 Position information slice joining section 5042 Attribute Information Slice Joining Section 5043 Tile joint 5051 Tile division section 5052 Encoding section 5053 Decoding Unit 5054 Tile joint 6800 3D data encoding device 6801 Division method determination section 6802 Split section 6803a, 6803b quantization section 6804a, 6804b Shift amount calculation unit 6805a, 6805b Common position shift section 6806a, 6806b Individual position shift section 6807a, 6807b encoding section

Claims

1. A three-dimensional data encoding method for encoding three-dimensional data, which is performed by a three-dimensional data encoding device, Each of these generates multiple sub-three-dimensional data sets that include positional information of a portion of the aforementioned three-dimensional data, The common displacement amount for the aforementioned multiple sub-three-dimensional data is calculated, Multiple individual displacement amounts corresponding to each of the aforementioned multiple sub-three-dimensional data are calculated, Each of the aforementioned sub-three-dimensional data is moved using the common movement amount and the corresponding individual movement amount. Each of the multiple sub-three-dimensional data moved using the common movement amount and the corresponding individual movement amount is encoded. The aforementioned common displacement is a movement of the same distance in the plurality of sub-three-dimensional data, and is a movement in the direction from the first point to the second point in three-dimensional space. The individual movement amounts are different distances in the plurality of sub-three-dimensional data, and are movements in the direction from the second point to the third point in the three-dimensional space. Three-dimensional data encoding method.

2. Common movement information relating to the common movement amount and individual movement information relating to each of the multiple individual movement amounts are generated. A bitstream is generated that includes the common movement information and a plurality of individual movement information. The three-dimensional data encoding method according to claim 1.

3. Each of the aforementioned sub-three-dimensional data corresponds to one of the multiple sub-spaces obtained by dividing the three-dimensional space corresponding to the aforementioned three-dimensional data. The three-dimensional data encoding method according to claim 1 or 2.

4. The plurality of subspaces have equal size to each other. The three-dimensional data encoding method according to claim 3.

5. The aforementioned subspaces have different sizes from each other. The three-dimensional data encoding method according to claim 3.

6. The plurality of subspaces have the same shape as each other. A three-dimensional data encoding method according to any one of claims 3 to 5.

7. The plurality of subspaces have different shapes from each other. A three-dimensional data encoding method according to any one of claims 3 to 5.

8. The aforementioned subspaces can overlap with each other. A three-dimensional data encoding method according to any one of claims 3 to 7.

9. A three-dimensional data decoding method performed by a three-dimensional data decoding device, By acquiring common movement information, multiple individual movement information, and multiple sub-3D data, each containing positional information of a portion of the 3D data, The aforementioned multiple sub-three-dimensional data are decoded to obtain multiple decoded positional information. Using the aforementioned common movement information, the common movement amount is calculated, Using each of the aforementioned multiple individual movement information, the individual movement amount corresponding to each of the aforementioned multiple sub-three-dimensional data is calculated. Each of the aforementioned sub-three-dimensional data is moved using the common movement amount and the corresponding individual movement amount. Three-dimensional data decoding method.

10. The common movement amount and the corresponding individual movement amount are used to move each of the multiple sub-three-dimensional data, and the positional information of the three-dimensional data is restored. The method for decoding three-dimensional data according to claim 9.

11. Each of the aforementioned sub-three-dimensional data corresponds to one of the multiple sub-spaces obtained by dividing the three-dimensional space corresponding to the aforementioned three-dimensional data. The method for decoding three-dimensional data according to claim 9 or 10.

12. The plurality of subspaces have equal size to each other. The method for decoding three-dimensional data according to claim 11.

13. The aforementioned subspaces have different sizes from each other. The method for decoding three-dimensional data according to claim 11.

14. The plurality of subspaces have the same shape as each other. A method for decoding three-dimensional data according to any one of claims 11 to 13.

15. The plurality of subspaces have different shapes from each other. A method for decoding three-dimensional data according to any one of claims 11 to 13.

16. The aforementioned subspaces can overlap with each other. A method for decoding three-dimensional data according to any one of claims 11 to 15.

17. A three-dimensional data encoding device for encoding three-dimensional data, Processor and Equipped with memory, The processor uses the memory to: Each of these generates multiple sub-three-dimensional data sets that include positional information of a portion of the aforementioned three-dimensional data, The common displacement amount for the aforementioned multiple sub-three-dimensional data is calculated, Multiple individual displacement amounts corresponding to each of the aforementioned multiple sub-three-dimensional data are calculated, Each of the aforementioned sub-three-dimensional data is moved using the common movement amount and the corresponding individual movement amount. Each of the multiple sub-three-dimensional data moved using the common movement amount and the corresponding individual movement amount is encoded. The aforementioned common displacement is a movement of the same distance in the plurality of sub-three-dimensional data, and is a movement in the direction from the first point to the second point in three-dimensional space. The individual movement amounts are different distances in the plurality of sub-three-dimensional data, and are movements in the direction from the second point to the third point in the three-dimensional space. Three-dimensional data encoding device.

18. Processor and Equipped with memory, The processor uses the memory to: By acquiring common movement information, multiple individual movement information, and multiple sub-3D data, each containing positional information of a portion of the 3D data, The aforementioned multiple sub-three-dimensional data are decoded to obtain multiple decoded positional information. Using the aforementioned common movement information, the common movement amount is calculated, Using each of the aforementioned multiple individual movement information, the individual movement amount corresponding to each of the aforementioned multiple sub-three-dimensional data is calculated. Each of the aforementioned sub-three-dimensional data is moved using the common movement amount and the corresponding individual movement amount. Three-dimensional data decoding device.

19. A program for causing a computer to execute the three-dimensional data encoding method according to any one of claims 1 to 8.

20. A program for causing a computer to execute the three-dimensional data decoding method according to any one of claims 9 to 16.