Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding apparatus, and three-dimensional data decoding apparatus
By combining and encoding point cloud datasets with geometric transformations and movement information, the method addresses inefficiencies in existing three-dimensional data coding, improving encoding efficiency and supporting multiplexing and storage in standardized formats.
Patent Information
- Application Number
- JP2023206742
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-11-13
- Filing Date
- 2023-12-07
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2039-11-12
AI Technical Summary
Existing three-dimensional data coding methods lack efficiency in compressing and transmitting large amounts of point cloud data, particularly in systems that combine multiple codecs like PCC, and there is a lack of defined formats for multiplexing and storage.
A method that combines and encodes multiple point cloud datasets using geometric transformations and movement information, generating a bitstream with position and attribute information, and supports multiplexing and storage in ISOBMFF formats.
Improves encoding efficiency by allowing simultaneous encoding of multiple point cloud datasets and supports multiplexing and storage in standardized formats, enhancing data transmission and decoding processes.
Smart Images

Figure 0007711155000001 
Figure 0007711155000002 
Figure 0007711155000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding apparatus, and a three-dimensional data decoding apparatus.
Background Art
[0002] In the future, the spread of devices or services that utilize three-dimensional data is expected in a wide range of fields such as computer vision, map information, monitoring, infrastructure inspection, or video distribution for autonomous operation of automobiles or robots. Three-dimensional data is acquired by various methods such as a distance sensor such as a range finder, a stereo camera, or a combination of a plurality of monocular cameras.
[0003] As one of the methods for expressing three-dimensional data, there is a method called point cloud that represents the shape of a three-dimensional structure by a point group in a three-dimensional space. In a point cloud, the position and color of the point group are stored. Although the point cloud is expected to become the mainstream as a method for expressing three-dimensional data, the amount of data of the point group is very large. Therefore, in the accumulation or transmission of three-dimensional data, as in the case of two-dimensional moving images (for example, MPEG-4 AVC or HEVC standardized by MPEG), compression of the amount of data by encoding is essential.
[0004] Also, regarding the compression of the point cloud, it is partially supported by a public library (Point Cloud Library) that performs point cloud-related processing.
[0005] Also, a technique is known for searching and displaying facilities located around a vehicle using three-dimensional map data (see, for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0006]
Patent Document 1
[0007] In the coding process for three-dimensional data, it is desirable to be able to improve the coding efficiency.
[0008] An object of the present disclosure is to provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device that can improve encoding efficiency. [Means for solving the problem]
[0009] A three-dimensional data encoding method according to an embodiment of the present disclosure includes: Three-dimensional Data and second Three-dimensional Of the data, the second Three-dimensional The data is converted into the first Three-dimensional data and the second data after the conversion process Three-dimensional By combining this data, Three-dimensional 3. Generate data Three-dimensional a bitstream is generated by encoding the data, the bitstream being Three-dimensional Included in the data Multiple position information or multiple attribute information Each of the first Three-dimensional Data and the second Three-dimensional First information indicating whether the data belongs to the Movement and second information indicating the content of the
[0010] A three-dimensional data decoding method according to an embodiment of the present disclosure includes: Three-dimensional Data and second Three-dimensional Of the data, the second Three-dimensional The data is converted into the first Three-dimensional data and the second data after the conversion process Three-dimensional The third one was generated by combining the Three-dimensional The third bitstream is generated by encoding the data. Three-dimensional and decoding the data from the bitstream to obtain the thirdThree-dimensional Contained in the data Multiple position information or multiple attribute information each of which is the first Three-dimensional data and the second Three-dimensional Obtain first information indicating to which of the data and the second data each belongs, and second information indicating the content of the Movement Using the first information and the second information, restore the first data and the second data from the decoded third data. Three-dimensional data from the Three-dimensional data and the second Three-dimensional Restore data.
Advantages of the Invention
[0011] The present disclosure can provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device that can improve encoding efficiency.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35
Figure 36
Figure 37
Figure 38
Figure 39
Figure 40
Figure 41
Figure 42
Figure 43
Figure 44
Figure 45
Figure 46
Figure 47
Figure 48
Figure 49
Figure 50
Figure 51
Figure 52
Figure 53
Figure 54
Figure 55
Figure 56
Figure 57
Figure 58
Figure 59
Figure 60
Figure 61
Figure 62
Figure 63
Figure 64
Figure 65
Figure 66
Figure 67
Figure 68
Figure 69
Figure 70
Figure 71
Figure 72
Figure 73
Figure 74
Figure 75
Figure 76
Figure 77
Figure 78
Figure 79
Figure 80
Figure 81
Figure 82
Figure 83
Figure 84
Figure 85
Figure 86
Figure 87
Figure 88
Figure 89
Figure 90
Figure 91
Figure 92
Figure 93
Figure 94
Figure 95
Figure 96
Figure 97
Figure 98
Figure 99
Figure 100
Figure 101
Figure 102
Figure 103
Figure 104
Figure 105
Figure 106
Figure 107
Figure 108
Figure 109
Figure 110
Figure 111
Figure 112
Figure 113
Figure 114
Figure 115
Figure 116
Figure 117
Embodiments for Carrying Out the Invention
[0013] In the three-dimensional data encoding method according to one aspect of the present disclosure, among the first point cloud data and the second point cloud data at the same time, a conversion process including movement is performed on the second point cloud data, and the first point cloud data and the second point cloud data after the conversion process are combined to generate third point cloud data, and the third point cloud data is encoded to generate a bit stream. The bit stream includes first information indicating to which of the first point cloud data and the second point cloud data each of a plurality of three-dimensional points included in the third point cloud data belongs, and second information indicating the content of the movement.
[0014] According to this, the encoding efficiency can be improved by encoding a plurality of point cloud data at the same time collectively.
[0015] For example, the conversion process may include a geometric transformation in addition to the movement, and the bitstream may include third information indicating the content of the geometric transformation.
[0016] According to this, the encoding efficiency can be improved by performing a geometric transformation.
[0017] For example, the geometric transformation may include at least one of a shift, a rotation, and a reversal.
[0018] For example, the content of the geometric transformation may be determined based on the number of overlapping points that are three-dimensional points with the same position information in the third point cloud data.
[0019] For example, the second information may include information indicating the amount of movement of the movement.
[0020] For example, the second information may include information indicating whether or not to move the origin of the second point cloud data to the origin of the first point cloud data.
[0021] For example, the first point cloud data and the second point cloud data may be generated by spatially dividing the fourth point cloud data.
[0022] For example, the bitstream includes the position information of each of the plurality of three-dimensional points included in the third point cloud data and one or more pieces of attribute information, and one of the one or more pieces of attribute information may include the first information.
[0023] A 3D data decoding method according to an aspect of the present disclosure performs a conversion process including movement on the second point cloud data among the first point cloud data and the second point cloud data at the same time, and combines the first point cloud data and the second point cloud data after the conversion process. Decode the third point cloud data generated by encoding the generated third point cloud data from the bitstream, and from the bitstream, for each of a plurality of 3D points included in the third point cloud data, which of the first point cloud data and the second point cloud data it belongs to Obtain first information indicating and second information indicating the content of the movement, and use the first information and the second information to restore the first point cloud data and the second point cloud data from the decoded third point cloud data.
[0024] According to this, the encoding efficiency can be improved by encoding a plurality of point cloud data at the same time together.
[0025] For example, the conversion process includes a geometric transformation in addition to the movement, obtains third information indicating the content of the geometric transformation from the bitstream, and in the restoration of the first point cloud data and the second point cloud data, the first information, the second information, and the third information may be used to restore the first point cloud data and the second point cloud data from the decoded third point cloud data.
[0026] According to this, the encoding efficiency can be improved by performing a geometric transformation.
[0027] For example, the geometric transformation may include at least one of a shift, a rotation, and an inversion.
[0028] For example, the second information may include information indicating the amount of movement of the movement.
[0029] For example, the second information may include information indicating whether to move the origin of the second point cloud data to the origin of the first point cloud data.
[0030] For example, the fourth point cloud data may be generated by spatially combining the restored first point cloud data and the second point cloud data.
[0031] For example, the bit stream includes position information of each of the plurality of three-dimensional points included in the third point cloud data and one or more pieces of attribute information, and one of the one or more pieces of attribute information may include the first information.
[0032] Further, a three-dimensional data encoding apparatus according to an aspect of the present disclosure includes a processor and a memory. The processor uses the memory to perform a conversion process including movement on the second point cloud data among the first point cloud data and the second point cloud data at the same time, and combines the first point cloud data and the second point cloud data after the conversion process to generate third point cloud data, and encodes the third point cloud data to generate a bit stream. The bit stream includes first information indicating to which of the first point cloud data and the second point cloud data each of the plurality of three-dimensional points included in the third point cloud data belongs, and second information indicating the content of the movement.
[0033] According to this, the encoding efficiency can be improved by encoding a plurality of point cloud data at the same time together.
[0034] Further, a three-dimensional data decoding apparatus according to an aspect of the present disclosure includes a processor and a memory. The processor uses the memory to perform a conversion process including movement on the second point cloud data among the first point cloud data and the second point cloud data at the same time, and decodes the third point cloud data from the bit stream generated by encoding the third point cloud data generated by combining the first point cloud data and the second point cloud data after the conversion process. From the bit stream, first information indicating to which of the first point cloud data and the second point cloud data each of the plurality of three-dimensional points included in the third point cloud data belongs, and second information indicating the content of the movement are acquired, and the first point cloud data and the second point cloud data are restored from the decoded third point cloud data using the first information and the second information.
[0035] According to this, the encoding efficiency can be improved by collectively encoding a plurality of point cloud data at the same time.
[0036] Note that these general or specific aspects may be implemented in a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.
[0037] Hereinafter, embodiments will be specifically described with reference to the drawings. Note that all of the embodiments described below show specific examples of the present disclosure. The numerical values, shapes, materials, components, arrangement positions and connection forms of the components, steps, order of steps, etc. shown in the following embodiments are merely examples and are not intended to limit the present disclosure. In addition, among the components in the following embodiments, the components not described in the independent claims are described as optional components.
[0038] (Embodiment 1) When using the encoded data of the point cloud in an actual device or service, it is desirable to transmit and receive necessary information according to the application in order to suppress the network bandwidth. However, until now, such a function has not existed in the encoding structure of three-dimensional data, and there has been no encoding method therefor.
[0039] In the present embodiment, a three-dimensional data encoding method and a three-dimensional data encoding device for providing a function of transmitting and receiving necessary information according to the application in the encoded data of a three-dimensional point cloud, a three-dimensional data decoding method and a three-dimensional data decoding device for decoding the encoded data, a three-dimensional data multiplexing method for multiplexing the encoded data, and a three-dimensional data transmission method for transmitting the encoded data will be described.
[0040] In particular, currently, a first encoding method and a second encoding method are being considered as encoding methods (encoding formats) for point cloud data, but the configuration of the encoded data and the method of storing the encoded data in the system format have not been defined, and there is a problem that MUX processing (multiplexing), or transmission or storage, cannot be performed in the encoding unit as it is.
[0041] Also, there has been no method to support a format in which two codecs, the first encoding method and the second encoding method, are mixed like PCC (Point Cloud Compression).
[0042] In the present embodiment, the configuration of PCC encoded data in which two codecs, the first encoding method and the second encoding method, are mixed, and the method of storing the encoded data in the system format will be described.
[0043] First, the configuration of a three-dimensional data (point cloud data) encoding / decoding system according to the present embodiment will be described. FIG. 1 is a diagram showing a configuration example of a three-dimensional data encoding / decoding system according to the present embodiment. As shown in FIG. 1, the three-dimensional data encoding / decoding system includes a three-dimensional data encoding system 4601, a three-dimensional data decoding system 4602, a sensor terminal 4603, and an external connection unit 4604.
[0044] The three-dimensional data encoding system 4601 generates encoded data or multiplexed data by encoding point cloud data, which is three-dimensional data. Note that the three-dimensional data encoding system 4601 may be a three-dimensional data encoding device realized by a single device, or may be a system realized by a plurality of devices. Also, the three-dimensional data encoding device may include a part of a plurality of processing units included in the three-dimensional data encoding system 4601.
[0045] The three-dimensional data encoding system 4601 includes a point cloud data generation system 4611, a presentation unit 4612, an encoding unit 4613, a multiplexing unit 4614, an input / output unit 4615, and a control unit 4616. The point cloud data generation system 4611 includes a sensor information acquisition unit 4617 and a point cloud data generation unit 4618.
[0046] The sensor information acquisition unit 4617 acquires sensor information from the sensor terminal 4603 and outputs the sensor information to the point cloud data generation unit 4618. The point cloud data generation unit 4618 generates point cloud data from the sensor information and outputs the point cloud data to the encoding unit 4613.
[0047] The presentation unit 4612 presents the sensor information or the point cloud data to the user. For example, the presentation unit 4612 displays information or an image based on the sensor information or the point cloud data.
[0048] The encoding unit 4613 encodes (compresses) the point cloud data and outputs the obtained encoded data, the control information obtained in the encoding process, and other additional information to the multiplexing unit 4614. The additional information includes, for example, the sensor information.
[0049] The multiplexing unit 4614 generates multiplexed data by multiplexing the encoded data, the control information, and the additional information input from the encoding unit 4613. The format of the multiplexed data is, for example, a file format for storage or a packet format for transmission.
[0050] The input / output unit 4615 (for example, a communication unit or an interface) outputs the multiplexed data to the outside. Alternatively, the multiplexed data is stored in a storage unit such as an internal memory. The control unit 4616 (or an application execution unit) controls each processing unit. That is, the control unit 4616 performs control such as encoding and multiplexing.
[0051] Note that the sensor information may be input to the encoding unit 4613 or the multiplexing unit 4614. Also, the input / output unit 4615 may output the point cloud data or the encoded data to the outside as it is.
[0052] The transmission signal (multiplexed data) output from the three-dimensional data encoding system 4601 is input to the three-dimensional data decoding system 4602 via the external connection unit 4604.
[0053] The three-dimensional data decoding system 4602 generates point cloud data, which is three-dimensional data, by decoding the encoded data or multiplexed data. Note that the three-dimensional data decoding system 4602 may be a three-dimensional data decoding device realized by a single device, or may be a system realized by a plurality of devices. Further, the three-dimensional data decoding device may include a part of a plurality of processing units included in the three-dimensional data decoding system 4602.
[0054] The three-dimensional data decoding system 4602 includes a sensor information acquisition unit 4621, an input / output unit 4622, a demultiplexing unit 4623, a decoding unit 4624, a presentation unit 4625, a user interface 4626, and a control unit 4627.
[0055] The sensor information acquisition unit 4621 acquires sensor information from the sensor terminal 4603.
[0056] The input / output unit 4622 acquires the transmission signal, decodes the multiplexed data (file format or packet) from the transmission signal, and outputs the multiplexed data to the demultiplexing unit 4623.
[0057] The demultiplexing unit 4623 acquires the encoded data, control information, and additional information from the multiplexed data, and outputs the encoded data, control information, and additional information to the decoding unit 4624.
[0058] The decoding unit 4624 reconstructs the point cloud data by decoding the encoded data.
[0059] The presentation unit 4625 presents the point cloud data to the user. For example, the presentation unit 4625 displays information or an image based on the point cloud data. The user interface 4626 acquires an instruction based on the user's operation. The control unit 4627 (or the app execution unit) controls each processing unit. That is, the control unit 4627 performs controls such as demultiplexing, decoding, and presentation.
[0060] Note that the input / output unit 4622 may directly acquire the point cloud data or the encoded data from the outside. Further, the presentation unit 4625 may acquire additional information such as sensor information and present information based on the additional information. Further, the presentation unit 4625 may perform presentation based on the user's instruction acquired by the user interface 4626.
[0061] The sensor terminal 4603 generates sensor information which is information obtained by a sensor. The sensor terminal 4603 is a terminal equipped with a sensor or a camera, and examples thereof include a moving body such as an automobile, a flying object such as an airplane, a mobile terminal, or a camera.
[0062] The sensor information that can be acquired by the sensor terminal 4603 is, for example, (1) the distance between the sensor terminal 4603 and an object or the reflectivity of the object obtained from a LIDAR, a millimeter-wave radar, or an infrared sensor, (2) the distance between a camera and an object or the reflectivity of the object obtained from a plurality of monocular camera images or stereo camera images. Further, the sensor information may include the attitude, orientation, gyro (angular velocity), position (GPS information or altitude), speed, or acceleration of the sensor. Further, the sensor information may include temperature, atmospheric pressure, humidity, or magnetism.
[0063] The external connection unit 4604 is realized by an integrated circuit (LSI or IC), an external storage unit, communication with a cloud server via the Internet, or broadcasting.
[0064] Next, the point cloud data will be described. FIG. 2 is a diagram showing the configuration of the point cloud data. FIG. 3 is a diagram showing a configuration example of a data file in which the information of the point cloud data is described.
[0065] Point cloud data includes data of a plurality of points. The data of each point includes position information (three-dimensional coordinates) and attribute information for the position information. A collection of such points is called a point cloud. For example, a point cloud represents the three-dimensional shape of an object.
[0066] Position information such as three-dimensional coordinates (Position) may also be referred to as geometry. Further, the data of each point may include attribute information (attribute) of a plurality of attribute types. The attribute types are, for example, color or reflectance.
[0067] One piece of attribute information may be associated with one piece of position information, or attribute information having a plurality of different attribute types may be associated with one piece of position information. Also, a plurality of pieces of attribute information of the same attribute type may be associated with one piece of position information.
[0068] The configuration example of the data file shown in FIG. 3 is an example in the case where the position information and the attribute information correspond one-to-one, and shows the position information and the attribute information of N points constituting the point cloud data.
[0069] The position information is, for example, information on three axes of x, y, and z. The attribute information is, for example, RGB color information. A typical data file is a ply file or the like.
[0070] Next, the types of point cloud data will be described. FIG. 4 is a diagram showing the types of point cloud data. As shown in FIG. 4, the point cloud data includes a static object and a dynamic object.
[0071] The static object is three-dimensional point cloud data at an arbitrary time (a certain time). The dynamic object is three-dimensional point cloud data that changes over time. Hereinafter, the three-dimensional point cloud data at a certain time is called a PCC frame or a frame.
[0072] The object may be a point cloud with a limited area like normal video data, or a large-scale point cloud with an unlimited area like map information.
[0073] Also, there may be point cloud data with various densities, including sparse point cloud data and dense point cloud data.
[0074] Hereinafter, the details of each processing unit will be described. Sensor information is obtained in various ways, such as a distance sensor like LIDAR or a range finder, a stereo camera, or a combination of multiple monocular cameras. The point cloud data generation unit 4618 generates point cloud data based on the sensor information obtained by the sensor information acquisition unit 4617. The point cloud data generation unit 4618 generates position information as the point cloud data and adds attribute information for the position information to the position information.
[0075] The point cloud data generation unit 4618 may process the point cloud data when generating the position information or adding the attribute information. For example, the point cloud data generation unit 4618 may reduce the data volume by deleting overlapping point clouds. Also, the point cloud data generation unit 4618 may convert the position information (such as position shift, rotation, or normalization), or render the attribute information.
[0076] In FIG. 1, the point cloud data generation system 4611 is included in the three-dimensional data encoding system 4601, but it may be provided independently outside the three-dimensional data encoding system 4601.
[0077] The encoding unit 4613 generates encoded data by encoding the point cloud data based on a predefined encoding method. There are roughly two types of encoding methods as follows. The first is an encoding method using position information, which will be described hereinafter as the first encoding method. The second is an encoding method using a video codec, which will be described hereinafter as the second encoding method.
[0078] The decoding unit 4624 decodes the point cloud data by decoding the encoded data based on a predefined encoding method.
[0079] The multiplexing unit 4614 generates multiplexed data by multiplexing the encoded data using an existing multiplexing method. The generated multiplexed data is transmitted or stored. In addition to the PCC encoded data, the multiplexing unit 4614 multiplexes other media such as video, audio, subtitles, applications, files, or reference time information. Further, the multiplexing unit 4614 may also multiplex sensor information or attribute information related to the point cloud data.
[0080] Examples of the multiplexing method or file format include ISOBMFF, MPEG-DASH which is an ISOBMFF-based transmission method, MMT, MPEG-2 TS Systems, RMP, and the like.
[0081] The demultiplexing unit 4623 extracts the PCC encoded data, other media, time information, etc. from the multiplexed data.
[0082] The input / output unit 4615 transmits the multiplexed data using a method suitable for the medium for transmission such as broadcasting or communication, or the medium for storage. The input / output unit 4615 may communicate with other devices via the Internet, or may communicate with a storage unit such as a cloud server.
[0083] Examples of the communication protocol include http, ftp, TCP, or UDP. A PULL-type communication method may be used, or a PUSH-type communication method may be used.
[0084] Either wired transmission or wireless transmission may be used. Examples of wired transmission include Ethernet (registered trademark), USB, RS-232C, HDMI (registered trademark), or coaxial cable. Examples of wireless transmission include wireless LAN, Wi-Fi (registered trademark), Bluetooth (registered trademark), or millimeter wave.
[0085] In addition, as the broadcasting method, for example, DVB-T2, DVB-S2, DVB-C2, ATSC 3.0, or ISDB-S3, etc. are used.
[0086] FIG. 5 is a diagram showing the configuration of a first encoding unit 4630 which is an example of an encoding unit 4613 that performs encoding according to a first encoding method. FIG. 6 is a block diagram of the first encoding unit 4630. The first encoding unit 4630 generates encoded data (encoded stream) by encoding point cloud data according to the first encoding method. The first encoding unit 4630 includes a position information encoding unit 4631, an attribute information encoding unit 4632, an additional information encoding unit 4633, and a multiplexing unit 4634.
[0087] The first encoding unit 4630 is characterized in that encoding is performed while being aware of the three-dimensional structure. In addition, the first encoding unit 4630 is characterized in that the attribute information encoding unit 4632 performs encoding using the information obtained from the position information encoding unit 4631. The first encoding method is also called GPCC (Geometry based PCC).
[0088] The point cloud data is PCC point cloud data such as a PLY file, or PCC point cloud data generated from sensor information, and includes position information (Position), attribute information (Attribute), and other additional information (MetaData). The position information is input to the position information encoding unit 4631, the attribute information is input to the attribute information encoding unit 4632, and the additional information is input to the additional information encoding unit 4633.
[0089] The position information encoding unit 4631 generates encoded position information (Compressed Geometry), which is encoded data, by encoding the position information. For example, the position information encoding unit 4631 encodes the position information using an N-ary tree structure such as an octree. Specifically, in the octree, the target space is divided into eight nodes (sub-spaces), and 8-bit information (occupancy code) indicating whether or not each node contains a point cloud is generated. Further, the node containing the point cloud is further divided into eight nodes, and 8-bit information indicating whether or not each of the eight nodes contains a point cloud is generated. This process is repeated until the number of point clouds included in a predetermined hierarchy or node is equal to or less than a threshold value.
[0090] The attribute information encoding unit 4632 generates encoded attribute information (Compressed Attribute), which is encoded data, by encoding using the configuration information generated by the position information encoding unit 4631. For example, the attribute information encoding unit 4632 determines a reference point (reference node) to be referred to in the encoding of the target point (target node) to be processed based on the octree structure generated by the position information encoding unit 4631. For example, the attribute information encoding unit 4632 refers to a node in which the parent node in the octree is the same as the target node among the surrounding nodes or adjacent nodes. Note that the method for determining the reference relationship is not limited to this.
[0091] Further, the encoding process of the attribute information may include at least one of quantization processing, prediction processing, and arithmetic encoding processing. In this case, reference means using the reference node to calculate the predicted value of the attribute information, or using the state of the reference node (for example, occupancy information indicating whether or not the reference node contains a point cloud) to determine the encoding parameter. For example, the encoding parameter is a quantization parameter in quantization processing or a context in arithmetic encoding.
[0092] The additional information encoding unit 4633 generates encoded additional information (Compressed MetaData), which is encoded data, by encoding compressible data among the additional information.
[0093] The multiplexing unit 4634 generates a compressed stream, which is encoded data, by multiplexing the encoded position information, encoded attribute information, encoded additional information, and other additional information. The generated compressed stream is output to a processing unit in a system layer (not shown).
[0094] Next, a first decoding unit 4640, which is an example of a decoding unit 4624 that decodes using the first encoding method, will be described. FIG. 7 is a diagram showing the configuration of the first decoding unit 4640. FIG. 8 is a block diagram of the first decoding unit 4640. The first decoding unit 4640 generates point cloud data by decoding encoded data (compressed stream) encoded using the first encoding method with the first encoding method. This first decoding unit 4640 includes a demultiplexing unit 4641, a position information decoding unit 4642, an attribute information decoding unit 4643, and an additional information decoding unit 4644.
[0095] A compressed stream, which is encoded data, is input to the first decoding unit 4640 from a processing unit in a system layer (not shown).
[0096] The demultiplexing unit 4641 separates the encoded position information (compressed geometry), encoded attribute information (compressed attribute), encoded additional information (compressed metadata), and other additional information from the encoded data.
[0097] The position information decoding unit 4642 generates position information by decoding the encoded position information. For example, the position information decoding unit 4642 restores the position information of the point cloud represented by three-dimensional coordinates from the encoded position information represented by an N-ary tree structure such as an octree.
[0098] The attribute information decoding unit 4643 decodes the encoded attribute information based on the configuration information generated by the position information decoding unit 4642. For example, the attribute information decoding unit 4643 determines a reference point (reference node) to be referred to in decoding the target point (target node) to be processed based on the octree structure obtained by the position information decoding unit 4642. For example, the attribute information decoding unit 4643 refers to a node among the peripheral nodes or adjacent nodes whose parent node in the octree is the same as the target node. Note that the method for determining the reference relationship is not limited to this.
[0099] Also, the decoding process of the attribute information may include at least one of inverse quantization processing, prediction processing, and arithmetic decoding processing. In this case, reference means using the reference node to calculate the predicted value of the attribute information, or using the state of the reference node (for example, occupancy information indicating whether the reference node includes a point cloud) to determine the decoding parameter. For example, the decoding parameter is a quantization parameter in the inverse quantization process or a context in arithmetic decoding.
[0100] The additional information decoding unit 4644 generates additional information by decoding the encoded additional information. Also, the first decoding unit 4640 uses the additional information required for the decoding processes of the position information and the attribute information during decoding and outputs the additional information required for the application to the outside.
[0101] Next, a second encoding unit 4650, which is an example of an encoding unit 4613 that performs encoding using the second encoding method, will be described. FIG. 9 is a diagram showing the configuration of the second encoding unit 4650. FIG. 10 is a block diagram of the second encoding unit 4650.
[0102] The second encoding unit 4650 generates encoded data (encoded stream) by encoding the point cloud data using the second encoding method. This second encoding unit 4650 includes an additional information generation unit 4651, a position image generation unit 4652, an attribute image generation unit 4653, a video encoding unit 4654, an additional information encoding unit 4655, and a multiplexing unit 4656.
[0103] The second encoding unit 4650 generates a position image and an attribute image by projecting a three-dimensional structure onto a two-dimensional image, and encodes the generated position image and attribute image using an existing video encoding method. The second encoding method is also called VPCC (Video based PCC).
[0104] The point cloud data is PCC point cloud data such as a PLY file, or PCC point cloud data generated from sensor information, and includes position information (Position), attribute information (Attribute), and other additional information (MetaData).
[0105] The additional information generation unit 4651 generates map information of a plurality of two-dimensional images by projecting a three-dimensional structure onto a two-dimensional image.
[0106] The position image generation unit 4652 generates a position image (Geometry Image) based on the position information and the map information generated by the additional information generation unit 4651. This position image is, for example, a depth image in which distance (Depth) is indicated as a pixel value. Note that this depth image may be an image of a plurality of point clouds viewed from one viewpoint (an image obtained by projecting a plurality of point clouds onto one two-dimensional plane), or a plurality of images of a plurality of point clouds viewed from a plurality of viewpoints, or a single image obtained by integrating these plurality of images.
[0107] The attribute image generation unit 4653 generates an attribute image based on the attribute information and the map information generated by the additional information generation unit 4651. This attribute image is, for example, an image in which attribute information (e.g., color (RGB)) is indicated as a pixel value. Note that this image may be an image of a plurality of point clouds viewed from one viewpoint (an image obtained by projecting a plurality of point clouds onto one two-dimensional plane), or a plurality of images of a plurality of point clouds viewed from a plurality of viewpoints, or a single image obtained by integrating these plurality of images.
[0108] The video encoding unit 4654 generates an encoded position image (Compressed Geometry Image) and an encoded attribute image (Compressed Attribute Image), which are encoded data, by encoding the position image and the attribute image using a video encoding method. Note that any known encoding method may be used as the video encoding method. For example, the video encoding method may be AVC, HEVC, or the like.
[0109] The additional information encoding unit 4655 generates encoded additional information (Compressed MetaData) by encoding the additional information included in the point cloud data, the map information, and the like.
[0110] The multiplexing unit 4656 generates an encoded stream (Compressed Stream), which is encoded data, by multiplexing the encoded position image, the encoded attribute image, the encoded additional information, and other additional information. The generated encoded stream is output to a processing unit of a system layer (not shown).
[0111] Next, a second decoding unit 4660, which is an example of a decoding unit 4624 that decodes the second encoding method, will be described. FIG. 11 is a diagram showing the configuration of the second decoding unit 4660. FIG. 12 is a block diagram of the second decoding unit 4660. The second decoding unit 4660 generates point cloud data by decoding encoded data (encoded stream) encoded by the second encoding method using the second encoding method. The second decoding unit 4660 includes a demultiplexing unit 4661, a video decoding unit 4662, an additional information decoding unit 4663, a position information generation unit 4664, and an attribute information generation unit 4665.
[0112] An encoded stream (Compressed Stream), which is encoded data, is input from a processing unit of a system layer (not shown) to the second decoding unit 4660.
[0113] The inverse multiplexing unit 4661 separates from the encoded data an encoded position image (Compressed Geometry Image), an encoded attribute image (Compressed Attribute Image), encoded additional information (Compressed MetaData), and other additional information.
[0114] The video decoding unit 4662 generates a position image and an attribute image by decoding the encoded position image and the encoded attribute image using a video encoding method. Note that any known encoding method may be used as the video encoding method. For example, the video encoding method is AVC or HEVC or the like.
[0115] The additional information decoding unit 4663 generates additional information including map information and the like by decoding the encoded additional information.
[0116] The position information generation unit 4664 generates position information using the position image and the map information. The attribute information generation unit 4665 generates attribute information using the attribute image and the map information.
[0117] The second decoding unit 4660 uses the additional information necessary for decoding during decoding and outputs the additional information necessary for the application to the outside.
[0118] Hereinafter, problems in the PCC encoding method will be described. FIG. 13 is a diagram showing a protocol stack related to PCC encoded data. FIG. 13 shows an example in which data of other media such as video (for example, HEVC) or audio is multiplexed with the PCC encoded data and transmitted or stored.
[0119] The multiplexing method and the file format have functions for multiplexing various encoded data and transmitting or storing it. In order to transmit or store the encoded data, the encoded data must be converted into the format of the multiplexing method. For example, in HEVC, a technique of storing encoded data in a data structure called a NAL unit and storing the NAL unit in ISOBMFF is defined.
[0120] On the one hand, currently, as encoding methods for point cloud data, a first encoding method (Codec1) and a second encoding method (Codec2) are being considered. However, the configuration of the encoded data and the method of storing the encoded data into the system format are not defined, and there is a problem that MUX processing (multiplexing), transmission, and storage in the encoding unit cannot be performed as it is.
[0121] Note that hereinafter, unless otherwise specified for a particular encoding method, it shall indicate either the first encoding method or the second encoding method.
[0122] (Embodiment 2) In this embodiment, a method of storing NAL units into an ISOBMFF file will be described.
[0123] ISOBMFF (ISO based media file format) is a file format standard defined in ISO / IEC 14496-12. ISOBMFF defines a format that can multiplex and store various media such as video, audio, and text, and is a media-independent standard.
[0124] The basic structure (file) of ISOBMFF will be described. The basic unit in ISOBMFF is a box. A box is composed of type, length, and data, and a set of boxes of various types combined is a file.
[0125] FIG. 14 is a diagram showing the basic structure (file) of ISOBMFF. The ISOBMFF file mainly includes boxes such as ftyp that indicates the file brand in 4CC (4-character code), moov that stores metadata such as control information, and mdat that stores data.
[0126] The storage method for each media in an ISOBMFF file is specified separately. For example, the storage methods for AVC video and HEVC video are specified in ISO / IEC 14496-15. Here, in order to store or transmit PCC encoded data, it is conceivable to extend and use the functions of ISOBMFF, but there is still no regulation on storing PCC encoded data in an ISOBMFF file. Therefore, in this embodiment, a method for storing PCC encoded data in an ISOBMFF file will be described.
[0127] FIG. 15 is a diagram showing a protocol stack when storing NAL units common to the PCC codec in an ISOBMFF file. Here, NAL units common to the PCC codec are stored in the ISOBMFF file. Although the NAL units are common to the PCC codec, since a plurality of PCC codecs are stored in the NAL units, it is desirable to specify storage methods (Carriage of Codec1, Carriage of Codec2) according to each codec.
[0128] (Embodiment 3) In this embodiment, the types of encoded data (Geometry (position information), Attribute (attribute information), Metadata (additional information)) generated by the above-described first encoding unit 4630 or second encoding unit 4650, the generation method of the additional information (metadata), and the multiplexing process in the multiplexing unit will be described. Note that the additional information (metadata) may also be referred to as a parameter set or control information.
[0129] In this embodiment, the dynamic object (three-dimensional point cloud data that changes over time) described with reference to FIG. 4 will be described as an example, but the same method may be used for a static object (three-dimensional point cloud data at an arbitrary time).
[0130] FIG. 16 is a diagram showing the configurations of an encoding unit 4801 and a multiplexing unit 4802 included in the three-dimensional data encoding apparatus according to the present embodiment. The encoding unit 4801 corresponds to, for example, the above-described first encoding unit 4630 or second encoding unit 4650. The multiplexing unit 4802 corresponds to the above-described multiplexing unit 4634 or 4656.
[0131] The encoding unit 4801 encodes point cloud data of a plurality of PCC (Point Cloud Compression) frames, and generates encoded data (Multiple Compressed Data) of a plurality of pieces of position information, attribute information, and additional information.
[0132] The multiplexing unit 4802 converts data into a data configuration considering data access in a decoding apparatus by NAL unitizing data of a plurality of data types (position information, attribute information, and additional information).
[0133] FIG. 17 is a diagram showing a configuration example of the encoded data generated by the encoding unit 4801. The arrows in the figure indicate the dependency relationships related to the decoding of the encoded data, and the source of the arrow depends on the data at the destination of the arrow. That is, the decoding apparatus decodes the data at the destination of the arrow, and uses the decoded data to decode the data at the source of the arrow. In other words, to depend means that the data at the destination is referenced (used) in the processing (encoding or decoding, etc.) of the data at the source.
[0134] First, the generation process of the encoded data of the position information will be described. The encoding unit 4801 generates encoded position data (Compressed Geometry Data) for each frame by encoding the position information of each frame. The encoded position data is represented by G(i). Here, i indicates the frame number, the time of the frame, or the like.
[0135] Further, the encoding unit 4801 generates a position parameter set (GPS(i)) corresponding to each frame. The position parameter set includes parameters that can be used for decoding the encoded position data. Also, the encoded position data for each frame depends on the corresponding position parameter set.
[0136] Also, the encoded position data consisting of a plurality of frames is defined as a geometry sequence. The encoding unit 4801 generates a geometry sequence parameter set (Geometry Sequence PS, also denoted as position SPS) that stores parameters commonly used for the decoding process of a plurality of frames within the geometry sequence. The geometry sequence depends on the position SPS.
[0137] Next, the generation process of the encoded data of the attribute information will be described. The encoding unit 4801 generates encoded attribute data (Compressed Attribute Data) for each frame by encoding the attribute information of each frame. Also, the encoded attribute data is represented by A(i). In FIG. 17, an example where there are attribute X and attribute Y is shown, and the encoded attribute data of attribute X is represented by AX(i), and the encoded attribute data of attribute Y is represented by AY(i).
[0138] Further, the encoding unit 4801 generates an attribute parameter set (APS(i)) corresponding to each frame. Also, the attribute parameter set of attribute X is represented by AXPS(i), and the attribute parameter set of attribute Y is represented by AYPS(i). The attribute parameter set includes parameters that can be used for decoding the encoded attribute information. The encoded attribute data depends on the corresponding attribute parameter set.
[0139] Also, the encoded attribute data consisting of a plurality of frames is defined as an Attribute Sequence. The encoding unit 4801 generates an Attribute Sequence Parameter Set (also denoted as Attribute SPS) that stores parameters commonly used for the decoding process for a plurality of frames within the Attribute Sequence. The Attribute Sequence depends on the Attribute SPS.
[0140] Also, in the first encoding method, the encoded attribute data depends on the encoding position data.
[0141] Also, FIG. 17 shows an example when there are two types of attribute information (Attribute X and Attribute Y). When there are two types of attribute information, for example, each data and metadata are generated by two encoding units. Also, for example, an Attribute Sequence is defined for each type of attribute information, and an Attribute SPS is generated for each type of attribute information.
[0142] Note that FIG. 17 shows an example where there is one type of position information and two types of attribute information, but it is not limited to this. The attribute information may be one type or three or more types. In this case as well, the encoded data can be generated in the same way. Also, in the case of point cloud data without attribute information, the attribute information may not be present. In that case, the encoding unit 4801 may not generate a parameter set related to the attribute information.
[0143] Next, the generation process of additional information (metadata) will be described. The encoding unit 4801 generates a PCC Stream Parameter Set (also denoted as Stream PS) which is a parameter set for the entire PCC stream. The encoding unit 4801 stores in the Stream PS parameters that can be commonly used for the decoding process for one or more position sequences and one or more attribute sequences. For example, the Stream PS includes identification information indicating the codec of the point cloud data, information indicating the algorithm used for encoding, etc. The position sequence and the attribute sequence depend on the Stream PS.
[0144] Next, the access unit and GOF will be described. In this embodiment, the concepts of a new access unit (Access Unit: AU) and GOF (Group of Frame) are introduced.
[0145] The access unit is a basic unit for accessing data during decoding, and is composed of one or more data and one or more metadata. For example, the access unit is composed of position information at the same time and one or more attribute information. The GOF is a random access unit and is composed of one or more access units.
[0146] The encoding unit 4801 generates an access unit header (AU Header) as identification information indicating the start of the access unit. The encoding unit 4801 stores parameters related to the access unit in the access unit header. For example, the access unit header includes the composition or information of the encoded data included in the access unit. In addition, the access unit header includes parameters commonly used for the data included in the access unit, such as parameters related to the decoding of the encoded data.
[0147] Note that the encoding unit 4801 may generate an access unit delimiter that does not include parameters related to the access unit instead of the access unit header. This access unit delimiter is used as identification information indicating the start of the access unit. The decoding device identifies the start of the access unit by detecting the access unit header or the access unit delimiter.
[0148] Next, the generation of the identification information at the beginning of the GOF will be described. The encoding unit 4801 generates a GOF header as the identification information indicating the beginning of the GOF. The encoding unit 4801 stores the parameters related to the GOF in the GOF header. For example, the GOF header includes the configuration or information of the encoded data included in the GOF. Also, the GOF header includes parameters commonly used for the data included in the GOF, such as parameters related to the decoding of the encoded data.
[0149] Note that the encoding unit 4801 may generate a GOF delimiter that does not include the parameters related to the GOF instead of the GOF header. This GOF delimiter is used as the identification information indicating the beginning of the GOF. The decoding device identifies the beginning of the GOF by detecting the GOF header or the GOF delimiter.
[0150] In the PCC encoded data, for example, an access unit is defined in units of PCC frames. The decoding device accesses the PCC frame based on the identification information at the beginning of the access unit.
[0151] Also, for example, the GOF is defined as one random access unit. The decoding device accesses the random access unit based on the identification information at the beginning of the GOF. For example, if the PCC frames are independent of each other and can be decoded alone, the PCC frames may be defined as random access units.
[0152] Note that two or more PCC frames may be assigned to one access unit, or a plurality of random access units may be assigned to one GOF.
[0153] Also, the encoding unit 4801 may define and generate parameter sets or metadata other than the above. For example, the encoding unit 4801 may generate SEI (Supplemental Enhancement Information) that stores parameters (optional parameters) that may not necessarily be used during decoding.
[0154] Next, the configuration of the encoded data and the method of storing the encoded data in the NAL unit will be described.
[0155] For example, a data format is defined for each type of encoded data. FIG. 18 is a diagram showing an example of encoded data and an NAL unit.
[0156] For example, as shown in FIG. 18, the encoded data includes a header and a payload. Note that the encoded data may include length information indicating the length (data amount) of the encoded data, the header, or the payload. Also, the encoded data may not include a header.
[0157] The header includes, for example, identification information for specifying the data. This identification information indicates, for example, the data type or the frame number.
[0158] The header includes, for example, identification information indicating a reference relationship. This identification information is stored in the header, for example, when there is a dependency relationship between data, and is information for referring from a reference source to a reference destination. For example, the header of the reference destination includes identification information for specifying the data. The header of the reference source includes identification information indicating the reference destination.
[0159] Note that when the reference destination or the reference source can be identified or derived from other information, the identification information for specifying the data or the identification information indicating the reference relationship may be omitted.
[0160] The multiplexing unit 4802 stores the encoded data in the payload of the NAL unit. The NAL unit header includes pcc_nal_unit_type, which is identification information for the encoded data. FIG. 19 is a diagram showing an example of the semantics of pcc_nal_unit_type.
[0161] As shown in FIG. 19, when pcc_codec_type is Codec1 (the first encoding method), the values 0 to 10 of pcc_nal_unit_type are assigned to the encoding position data (Geometry), encoding attribute X data (AttributeX), encoding attribute Y data (AttributeY), position PS (Geom.PS), attribute XPS (AttrX.PS), attribute YPS (AttrX.PS), position SPS (Geometry Sequence PS), attribute XSPS (AttributeX Sequence PS), attribute YSPS (AttributeY Sequence PS), AU header (AU Header), and GOF header (GOF Header) in Codec1. Also, values 11 and above are assigned for future use in Codec1.
[0162] When pcc_codec_type is Codec2 (the second encoding method), the values 0 to 2 of pcc_nal_unit_type are assigned to the codec's DataA, MetaDataA, and MetaDataB. Also, values 3 and above are assigned for future use in Codec2.
[0163] Next, the data transmission order will be described. Hereinafter, the constraints on the transmission order of NAL units will be described.
[0164] The multiplexing unit 4802 transmits NAL units in groups of GOF or AU. The multiplexing unit 4802 places the GOF header at the beginning of the GOF and the AU header at the beginning of the AU.
[0165] Even if data is lost due to packet loss or the like, the multiplexing unit 4802 may arrange the sequence parameter set (SPS) for each AU so that the decoder can decode from the next AU.
[0166] If there is a dependency related to decoding in the symbolized data, the decoding device decodes the reference data first and then decodes the referenced data. In the decoding device, in order to be able to decode in the order received without rearranging the data, the multiplexing unit 4802 sends out the reference data first.
[0167] FIG. 20 is a diagram showing an example of the transmission order of NAL units. FIG. 20 shows three examples: position information priority, parameter priority, and data integration.
[0168] The transmission order with position information priority is an example of sending together each of the information regarding position information and the information regarding attribute information. In this transmission order, the transmission of the information regarding position information is completed earlier than the transmission of the information regarding attribute information.
[0169] For example, by using this transmission order, a decoding device that does not decode the attribute information may be able to provide a time for not processing by ignoring the decoding of the attribute information. Also, for example, in the case of a decoding device that wants to decode the position information earlier, it may be possible to decode the position information earlier by obtaining the encoded data of the position information earlier.
[0170] Note that in FIG. 20, the attribute XSPS and the attribute YSPS are integrated and described as the attribute SPS, but the attribute XSPS and the attribute YSPS may be arranged individually.
[0171] In the transmission order with parameter set priority, the parameter set is sent first and the data is sent later.
[0172] As described above, according to the constraints of the NAL unit transmission order, the multiplexing unit 4802 may send out the NAL units in any order. For example, order identification information is defined, and the multiplexing unit 4802 may have a function of sending out the NAL units in a plurality of pattern orders. For example, the order identification information of the NAL units is stored in the stream PS.
[0173] The three-dimensional data decoding device may perform decoding based on the order identification information. The desired transmission order may be instructed from the three-dimensional data decoding device to the three-dimensional data encoding device, and the three-dimensional data encoding device (multiplexing unit 4802) may control the transmission order according to the instructed transmission order.
[0174] Note that the multiplexing unit 4802 may generate encoded data in which a plurality of functions are merged as long as it is within the range that follows the constraints of the transmission order, such as the transmission order of data integration. For example, as shown in FIG. 20, the GOF header and the AU header may be integrated, or AXPS and AYPS may be integrated. In this case, an identifier indicating that the data has a plurality of functions is defined for pcc_nal_unit_type.
[0175] Hereinafter, a modification example of the present embodiment will be described. PS has levels such as frame-level PS, sequence-level PS, and PCC sequence-level PS. If the PCC sequence level is the upper level and the frame level is the lower level, the following method may be used for the parameter storage method.
[0176] The value of the default PS is indicated by the higher-level PS. Also, when the value of the lower-level PS is different from the value of the higher-level PS, the value of the PS is indicated by the lower-level PS. Alternatively, the value of the PS is not described at the higher level, and the value of the PS is described at the lower-level PS. Alternatively, information indicating whether the value of the PS is indicated by the lower-level PS, the higher-level PS, or both is indicated in either one or both of the lower-level PS and the higher-level PS. Alternatively, the lower-level PS may be merged into the higher-level PS. Alternatively, when the lower-level PS and the higher-level PS overlap, the multiplexing unit 4802 may omit either one of the transmissions.
[0177] Note that the encoding unit 4801 or the multiplexing unit 4802 may divide data into slices or tiles, etc., and send out the divided data. The divided data includes information for identifying the divided data, and parameters used for decoding the divided data are included in the parameter set. In this case, an identifier indicating that the pcc_nal_unit_type is data storing data or parameters related to tiles or slices is defined.
[0178] (Embodiment 4) In HEVC encoding, there are data splitting tools such as slices or tiles to enable parallel processing in the decoding device, but there are none in PCC (Point Cloud Compression) encoding yet.
[0179] In PCC, various data splitting methods can be considered depending on parallel processing, compression efficiency, and compression algorithms. Here, the definitions of slices and tiles, data structures, and transmission / reception methods will be described.
[0180] FIG. 21 is a block diagram showing the configuration of a first encoding unit 4910 included in the three-dimensional data encoding device according to the present embodiment. The first encoding unit 4910 generates encoded data (encoded stream) by encoding point cloud data using a first encoding method (GPCC (Geometry based PCC)). This first encoding unit 4910 includes a splitting unit 4911, a plurality of position information encoding units 4912, a plurality of attribute information encoding units 4913, an additional information encoding unit 4914, and a multiplexing unit 4915.
[0181] The splitting unit 4911 generates a plurality of split data by splitting the point cloud data. Specifically, the splitting unit 4911 generates a plurality of split data by splitting the space of the point cloud data into a plurality of sub-spaces. Here, the sub-space is either one of the tile and the slice, or a combination of the tile and the slice. More specifically, the point cloud data includes position information, attribute information, and additional information. The splitting unit 4911 splits the position information into a plurality of split position information and splits the attribute information into a plurality of split attribute information. Also, the splitting unit 4911 generates additional information regarding the splitting.
[0182] The plurality of position information encoding units 4912 generate a plurality of encoded position information by encoding the plurality of split position information. For example, the plurality of position information encoding units 4912 perform parallel processing on the plurality of split position information.
[0183] The plurality of attribute information encoding units 4913 generate a plurality of encoded attribute information by encoding the plurality of split attribute information. For example, the plurality of attribute information encoding units 4913 perform parallel processing on the plurality of split attribute information.
[0184] The additional information encoding unit 4914 generates encoded additional information by encoding the additional information included in the point cloud data and the additional information regarding data splitting generated during splitting by the splitting unit 4911.
[0185] The multiplexing unit 4915 generates encoded data (encoded stream) by multiplexing the plurality of encoded position information, the plurality of encoded attribute information, and the encoded additional information, and transmits the generated encoded data. Also, the encoded additional information is used during decoding.
[0186] Note that in Fig. 21, an example is shown where the number of the position information encoding units 4912 and the number of the attribute information encoding units 4913 are each two. However, the number of the position information encoding units 4912 and the number of the attribute information encoding units 4913 may each be one, or may be three or more. Further, the plurality of divided data may be processed in parallel within the same chip like a plurality of cores in the CPU, may be processed in parallel by the cores of a plurality of chips, or may be processed in parallel by the plurality of cores of a plurality of chips.
[0187] Fig. 22 is a block diagram showing the configuration of the first decoding unit 4920. The first decoding unit 4920 restores the point cloud data by decoding the encoded data (encoded stream) generated by encoding the point cloud data by the first encoding method (GPCC). The first decoding unit 4920 includes a demultiplexing unit 4921, a plurality of position information decoding units 4922, a plurality of attribute information decoding units 4923, an additional information decoding unit 4924, and a combining unit 4925.
[0188] The demultiplexing unit 4921 generates a plurality of encoded position information, a plurality of encoded attribute information, and encoded additional information by demultiplexing the encoded data (encoded stream).
[0189] The plurality of position information decoding units 4922 generate a plurality of divided position information by decoding the plurality of encoded position information. For example, the plurality of position information decoding units 4922 process the plurality of encoded position information in parallel.
[0190] The plurality of attribute information decoding units 4923 generate a plurality of divided attribute information by decoding the plurality of encoded attribute information. For example, the plurality of attribute information decoding units 4923 process the plurality of encoded attribute information in parallel.
[0191] The plurality of additional information decoding units 4924 generate additional information by decoding the encoded additional information.
[0192] The combining unit 4925 generates position information by combining a plurality of divided position information using additional information. The combining unit 4925 generates attribute information by combining a plurality of divided attribute information using additional information.
[0193] In addition, in FIG. 22, examples where the number of the position information decoding unit 4922 and the attribute information decoding unit 4923 is two are shown, but the number of the position information decoding unit 4922 and the attribute information decoding unit 4923 may be one or three or more respectively. Further, the plurality of divided data may be processed in parallel within the same chip like a plurality of cores in the CPU, may be processed in parallel by the cores of a plurality of chips, or may be processed in parallel by the plurality of cores of a plurality of chips.
[0194] Next, the configuration of the dividing unit 4911 will be described. FIG. 23 is a block diagram of the dividing unit 4911. The dividing unit 4911 includes a slice dividing unit 4931 (Slice Divider), a position information tile dividing unit 4932 (Geometry Tile Divider), and an attribute information tile dividing unit 4933 (Attribute Tile Divider).
[0195] The slice dividing unit 4931 generates a plurality of slice position information by dividing the position information (Position(Geometry)) into slices. In addition, the slice dividing unit 4931 generates a plurality of slice attribute information by dividing the attribute information (Attribute) into slices. Further, the slice dividing unit 4931 outputs slice additional information (Slice MetaData) including information related to the slice division and information generated in the slice division.
[0196] The position information tile dividing unit 4932 generates a plurality of divided position information (a plurality of tile position information) by dividing the plurality of slice position information into tiles. In addition, the position information tile dividing unit 4932 outputs position tile additional information (Geometry Tile MetaData) including information related to the tile division of the position information and information generated in the tile division of the position information.
[0197] The attribute information tile splitting unit 4933 generates a plurality of split attribute information (a plurality of tile attribute information) by splitting a plurality of slice attribute information into tiles. Further, the attribute information tile splitting unit 4933 outputs attribute tile additional information (Attribute Tile MetaData) including information related to the tile splitting of the attribute information and information generated in the tile splitting of the attribute information.
[0198] Note that the number of slices or tiles to be split is 1 or more. That is, it is not necessary to perform the splitting of slices or tiles.
[0199] Also, here, an example in which tile splitting is performed after slice splitting is shown, but slice splitting may be performed after tile splitting. Further, in addition to slices and tiles, a new splitting type may be defined, and splitting may be performed with three or more splitting types.
[0200] Hereinafter, a method for splitting point cloud data will be described. FIG. 24 is a diagram showing an example of slice and tile splitting.
[0201] First, a method for slice splitting will be described. The splitting unit 4911 splits the three-dimensional point cloud data into arbitrary point clouds in units of slices. The splitting unit 4911 does not split the position information and the attribute information that constitute the points in the slice splitting, but splits the position information and the attribute information together. That is, the splitting unit 4911 performs slice splitting so that the position information and the attribute information at an arbitrary point belong to the same slice. Note that according to these, the number of splits and the splitting method may be any method. Further, the minimum unit of splitting is a point. For example, the number of splits of the position information and the attribute information is the same. For example, the three-dimensional points corresponding to the position information after slice splitting and the three-dimensional points corresponding to the attribute information are included in the same slice.
[0202] Further, the splitting unit 4911 generates slice additional information, which is additional information related to the number of splits and the splitting method during slice splitting. The slice additional information is the same for the position information and the attribute information. For example, the slice additional information includes information indicating the reference coordinate position, size, or side length of the bounding box after splitting. The slice additional information also includes information indicating the number of splits, the split type, and the like.
[0203] Next, the method of tile splitting will be described. The splitting unit 4911 splits the slice-split data into slice position information (G slice) and slice attribute information (A slice), and splits the slice position information and the slice attribute information into tile units respectively.
[0204] Note that FIG. 24 shows an example of splitting in an octree structure, but the number of splits and the splitting method may be any method.
[0205] Also, the splitting unit 4911 may split the position information and the attribute information by different splitting methods, or by the same splitting method. Also, the splitting unit 4911 may split a plurality of slices into tiles by different splitting methods, or by the same splitting method.
[0206] Further, the splitting unit 4911 generates tile additional information related to the number of splits and the splitting method during tile splitting. The tile additional information (position tile additional information and attribute tile additional information) is independent of the position information and the attribute information. For example, the tile additional information includes information indicating the reference coordinate position, size, or side length of the bounding box after splitting. The tile additional information also includes information indicating the number of splits, the split type, and the like.
[0207] Next, an example of a method for splitting point cloud data into slices or tiles will be described. The splitting unit 4911 may use a predetermined method as the method of slice or tile splitting, or may adaptively switch the method used according to the point cloud data.
[0208] At the time of slice division, the division unit 4911 divides the three-dimensional space in a batch with respect to the position information and the attribute information. For example, the division unit 4911 determines the shape of the object and divides the three-dimensional space into slices according to the shape of the object. For example, the division unit 4911 extracts an object such as a tree or a building and performs division in units of objects. For example, the division unit 4911 performs slice division so that the whole of one or a plurality of objects is included in one slice. Or, the division unit 4911 divides one object into a plurality of slices.
[0209] In this case, the encoding device may change the encoding method for each slice, for example. For example, the encoding device may use a high-quality compression method for a specific object or a specific part of the object. In this case, the encoding device may store information indicating the encoding method for each slice in additional information (metadata).
[0210] Also, the division unit 4911 may perform slice division so that each slice corresponds to a predetermined coordinate space based on the map information or the position information.
[0211] At the time of tile division, the division unit 4911 divides the position information and the attribute information independently. For example, the division unit 4911 divides the slice into tiles according to the data amount or the processing amount. For example, the division unit 4911 determines whether the data amount of the slice (for example, the number of three-dimensional points included in the slice) is more than a predetermined threshold value. The division unit 4911 divides the slice into tiles when the data amount of the slice is more than the threshold value. The division unit 4911 does not divide the slice into tiles when the data amount of the slice is less than the threshold value.
[0212] For example, the division unit 4911 divides the slice into tiles so that the processing amount or the processing time in the decoding device is within a certain range (equal to or less than a predetermined value). Thereby, the processing amount per tile in the decoding device becomes constant, and distributed processing in the decoding device becomes easy.
[0213] In addition, when the processing amounts differ between the position information and the attribute information, for example, when the processing amount of the position information is larger than that of the attribute information, the dividing unit 4911 increases the number of divisions of the position information compared to the number of divisions of the attribute information.
[0214] Also, for example, when the content allows the decoding device to quickly decode and display the position information and then slowly decode and display the attribute information later, the dividing unit 4911 may increase the number of divisions of the position information compared to the number of divisions of the attribute information. Thereby, the decoding device can increase the number of parallel data of the position information, so that the processing of the position information can be speeded up compared to the processing of the attribute information.
[0215] Note that the decoding device does not necessarily have to perform parallel processing on the sliced or tiled data, and may determine whether to perform parallel processing on them according to the number or capabilities of the decoding processing units.
[0216] By dividing in the above-described manner, adaptive encoding according to the content or object can be realized. Also, parallel processing in the decoding process can be realized. Thereby, the flexibility of the point cloud encoding system or the point cloud decoding system is improved.
[0217] FIG. 25 is a diagram showing an example of a pattern of slice and tile divisions. DU in the figure is a data unit (DataUnit), indicating the data of a tile or a slice. Each DU includes a slice index (SliceIndex) and a tile index (TileIndex). The numerical value at the upper right of the DU in the figure indicates the slice index, and the numerical value at the lower left of the DU indicates the tile index.
[0218] In pattern 1, in slice division, the number of divisions and the division method are the same for the G slice and the A slice. In tile division, the number of divisions and the division method for the G slice are different from those for the A slice. Also, the same number of divisions and division method are used among multiple G slices. The same number of divisions and division method are used among multiple A slices.
[0219] In Pattern 2, in slice division, the number of divisions and the division method are the same for the G slice and the A slice. In tile division, the number of divisions and the division method for the G slice are different from those for the A slice. Also, the number of divisions and the division method are different among multiple G slices. The number of divisions and the division method are different among multiple A slices.
[0220] Next, the encoding method of the divided data will be described. The three-dimensional data encoding device (the first encoding unit 4910) encodes the divided data respectively. When encoding the attribute information, the three-dimensional data encoding device generates dependency relationship information indicating based on which configuration information (position information, additional information, or other attribute information) the encoding is performed, as additional information. That is, the dependency relationship information indicates, for example, the configuration information of the reference destination (dependency destination). In this case, the three-dimensional data encoding device generates the dependency relationship information based on the configuration information corresponding to the divided shape of the attribute information. Note that the three-dimensional data encoding device may generate the dependency relationship information based on the configuration information corresponding to a plurality of divided shapes.
[0221] The dependency relationship information is generated by the three-dimensional data encoding device, and the generated dependency relationship information may be sent to the three-dimensional data decoding device. Alternatively, the three-dimensional data decoding device generates the dependency relationship information, and the three-dimensional data encoding device does not have to send the dependency relationship information. Also, the dependency relationship used by the three-dimensional data encoding device may be determined in advance, and the three-dimensional data encoding device does not have to send the dependency relationship information.
[0222] FIG. 26 is a diagram showing an example of the dependency relationship of each data. The tip of the arrow in the figure indicates the dependency destination, and the origin of the arrow indicates the dependency source. The three-dimensional data decoding device decodes the data in the order from the dependency destination to the dependency source. Also, the data indicated by the solid line in the figure is the actually sent data, and the data indicated by the dotted line is the data not sent.
[0223] Also, in the same figure, G indicates position information, and A indicates attribute information. G s1 indicates the position information of slice number 1, G s2indicates the position information of slice number 2. G s1t1 indicates the position information of slice number 1 and tile number 1, G s1t2 indicates the position information of slice number 1 and tile number 2, G s2t1 indicates the position information of slice number 2 and tile number 1, G s2t2 indicates the position information of slice number 2 and tile number 2. Similarly, A s1 indicates the attribute information of slice number 1, A s2 indicates the attribute information of slice number 2. A s1t1 indicates the attribute information of slice number 1 and tile number 1, A s1t2 indicates the attribute information of slice number 1 and tile number 2, A s2t1 indicates the attribute information of slice number 2 and tile number 1, A s2t2 indicates the attribute information of slice number 2 and tile number 2.
[0224] Mslice indicates slice additional information, MGtile indicates position tile additional information, and MAtile indicates attribute tile additional information. D s1t1 is the attribute information A s1t1 indicates the dependency information of, D s2t1 is the attribute information A s2t1 indicates the dependency information of.
[0225] Also, the three-dimensional data encoding device may rearrange the data in the decoding order so that the three-dimensional data decoding device does not need to rearrange the data. Note that the three-dimensional data decoding device may rearrange the data, or both the three-dimensional data encoding device and the three-dimensional data decoding device may rearrange the data.
[0226] FIG. 27 is a diagram showing an example of the decoding order of data. In the example of FIG. 27, decoding is performed in order from the left data. The three-dimensional data decoding apparatus decodes the data that is the destination of the dependency first among the data in a dependency relationship. For example, the three-dimensional data encoding apparatus rearranges the data in advance and sends it out so as to be in this order. Any order may be used as long as the data that is the destination of the dependency comes first. Further, the three-dimensional data encoding apparatus may send out the additional information and the dependency relationship information before the data.
[0227] FIG. 28 is a flowchart showing the processing flow by the three-dimensional data encoding apparatus. First, the three-dimensional data encoding apparatus encodes the data of a plurality of slices or tiles as described above (S4901). Next, the three-dimensional data encoding apparatus rearranges the data so that the data that is the destination of the dependency comes first as shown in FIG. 27 (S4902). Next, the three-dimensional data encoding apparatus multiplexes (NAL unitizes) the rearranged data (S4903).
[0228] Next, the configuration of the combining unit 4925 included in the first decoding unit 4920 will be described. FIG. 29 is a block diagram showing the configuration of the combining unit 4925. The combining unit 4925 includes a position information tile combining unit 4941 (Geometry Tile Combiner), an attribute information tile combining unit 4942 (Attribute Tile Combiner), and a slice combining unit (Slice Combiner).
[0229] The position information tile combining unit 4941 generates a plurality of slice position information by combining a plurality of divided position information using the position tile additional information. The attribute information tile combining unit 4942 generates a plurality of slice attribute information by combining a plurality of divided attribute information using the attribute tile additional information.
[0230] The slice combining unit 4943 generates position information by combining a plurality of slice position information using the slice additional information. Further, the slice combining unit 4943 generates attribute information by combining a plurality of slice attribute information using the slice additional information.
[0231] Note that the number of slices or tiles to be divided is 1 or more. That is, the slices or tiles may not be divided.
[0232] Also, here, an example in which tile division is performed after slice division is shown, but slice division may be performed after tile division. Also, in addition to slices and tiles, a new division type may be defined, and division may be performed using three or more division types.
[0233] Next, the configuration of the sliced or tiled encoded data and the method of storing the encoded data in the NAL unit (multiplexing method) will be described. FIG. 30 is a diagram showing the configuration of the encoded data and the method of storing the encoded data in the NAL unit.
[0234] The encoded data (division position information and division attribute information) is stored in the payload of the NAL unit.
[0235] The encoded data includes a header and a payload. The header includes identification information for identifying the data included in the payload. This identification information includes, for example, the type of slice division or tile division (slice_type, tile_type), index information for identifying a slice or tile (slice_idx, tile_idx), position information of the data (slice or tile), or the address of the data. The index information for identifying a slice is also referred to as a slice index (SliceIndex). The index information for identifying a tile is also referred to as a tile index (TileIndex). Also, the type of division is, for example, a method based on the object shape as described above, a method based on map information or position information, or a method based on the data amount or processing amount.
[0236] Note that all or part of the above information may be stored in one of the headers of the segmentation position information and the segmentation attribute information, and may not be stored in the other. For example, when the same segmentation method is used for the position information and the attribute information, the segmentation types (slice_type, tile_type) and index information (slice_idx, tile_idx) are the same for the position information and the attribute information. Therefore, these information may be included in one of the headers of the position information and the attribute information. For example, when the attribute information depends on the position information, the position information is processed first. Therefore, these information may be included in the header of the position information and may not be included in the header of the attribute information. In this case, the three-dimensional data decoding device determines, for example, that the dependent attribute information belongs to the same slice or tile as the slice or tile of the dependent position information.
[0237] In addition, additional information related to slice segmentation or tile segmentation (slice additional information, position tile additional information, or attribute tile additional information), dependency relationship information indicating the dependency relationship, etc. may be stored in and transmitted by an existing parameter set (GPS, APS, position SPS, or attribute SPS, etc.). When the segmentation method changes for each frame, information indicating the segmentation method may be stored in a parameter set (GPS or APS, etc.) for each frame. When the segmentation method does not change within a sequence, information indicating the segmentation method may be stored in a parameter set (position SPS or attribute SPS) for each sequence. Furthermore, when the same segmentation method is used for the position information and the attribute information, information indicating the segmentation method may be stored in a parameter set (stream PS) of the PCC stream.
[0238] In addition, the above information may be stored in any of the above parameter sets, or may be stored in multiple parameter sets. Also, a parameter set for tile segmentation or slice segmentation may be defined, and the above information may be stored in the parameter set. Also, these information may be stored in the header of the encoded data.
[0239] Also, the header of the encoded data includes identification information indicating dependencies. That is, when there are dependencies between data, the header includes identification information for referring from the source of the dependency to the destination. For example, the header of the destination data includes identification information for specifying the data. The header of the source data includes identification information indicating the destination. Note that if the identification information for specifying data, additional information related to slice division or tile division, and the identification information indicating dependencies can be identified or derived from other information, these pieces of information may be omitted.
[0240] Next, the flow of the encoding process and decoding process of the point cloud data according to this embodiment will be described. FIG. 31 is a flowchart of the encoding process of the point cloud data according to this embodiment.
[0241] First, the three-dimensional data encoding device determines the division method to be used (S4911). This division method includes whether to perform slice division or tile division. Also, the division method may include the number of divisions in the case of performing slice division or tile division, and the type of division, etc. The type of division is a method based on the object shape as described above, a method based on map information or position information, or a method based on the data amount or processing amount, etc. Note that the division method may be predetermined.
[0242] When slice division is performed (Yes in S4912), the three-dimensional data encoding device generates a plurality of slice position information and a plurality of slice attribute information by collectively dividing the position information and the attribute information (S4913). Also, the three-dimensional data encoding device generates slice additional information related to slice division. Note that the three-dimensional data encoding device may divide the position information and the attribute information independently.
[0243] When tile division is performed (Yes in S4914), the three-dimensional data encoding device generates a plurality of divided position information and a plurality of divided attribute information by independently dividing a plurality of slice position information and a plurality of slice attribute information (or position information and attribute information) (S4915). Further, the three-dimensional data encoding device generates position tile addition information and attribute tile addition information related to tile division. Note that the three-dimensional data encoding device may divide the slice position information and the slice attribute information together.
[0244] Next, the three-dimensional data encoding device generates a plurality of encoded position information and a plurality of encoded attribute information by encoding each of the plurality of divided position information and the plurality of divided attribute information (S4916). Further, the three-dimensional data encoding device generates dependency information.
[0245] Next, the three-dimensional data encoding device generates encoded data (encoded stream) by NAL unitizing (multiplexing) the plurality of encoded position information, the plurality of encoded attribute information, and the addition information (S4917). Further, the three-dimensional data encoding device transmits the generated encoded data.
[0246] FIG. 32 is a flowchart of the decoding process of the point cloud data according to the present embodiment. First, the three-dimensional data decoding device determines the division method by analyzing the addition information (slice addition information, position tile addition information, and attribute tile addition information) related to the division method included in the encoded data (encoded stream) (S4921). This division method includes whether to perform slice division or not, and whether to perform tile division or not. Further, the division method may include the number of divisions when performing slice division or tile division, and the type of division, etc.
[0247] Next, the three-dimensional data decoding device generates divided position information and divided attribute information by decoding the plurality of encoded position information and the plurality of encoded attribute information included in the encoded data using the dependency information included in the encoded data (S4922).
[0248] When it is shown that tile splitting is performed based on additional information (Yes in S4923), the three-dimensional data decoding device generates a plurality of slice position information and a plurality of slice attribute information by combining a plurality of split position information and a plurality of split attribute information in respective ways based on the position tile additional information and the attribute tile additional information (S4924). Note that the three-dimensional data decoding device may combine the plurality of split position information and the plurality of split attribute information in the same way.
[0249] When it is shown that slice splitting is performed based on additional information (Yes in S4925), the three-dimensional data decoding device generates position information and attribute information by combining a plurality of slice position information and a plurality of slice attribute information (a plurality of split position information and a plurality of split attribute information) in the same way based on the slice additional information (S4926). Note that the three-dimensional data decoding device may combine the plurality of slice position information and the plurality of slice attribute information in different ways.
[0250] As described above, the three-dimensional data encoding device according to the present embodiment performs the processing shown in FIG. 33. First, the three-dimensional data encoding device divides the target space including a plurality of three-dimensional points into a plurality of divided data (for example, tiles) each including one or more three-dimensional points and included in a plurality of sub-spaces (for example, slices) into which the target space is divided. Here, the divided data is a data aggregate including one or more data included in a sub-space and including one or more three-dimensional points. Further, the divided data is also a space and may include a space that does not include three-dimensional points. Further, a plurality of divided data may be included in one sub-space, or one divided data may be included in one sub-space. Note that a plurality of sub-spaces may be set in the target space, or one sub-space may be set in the target space.
[0251] Next, the three-dimensional data encoding device generates a plurality of encoded data corresponding to each of the plurality of divided data by encoding each of the plurality of divided data (S4931). The three-dimensional data encoding device generates a bitstream including the plurality of encoded data and a plurality of control information (for example, the header shown in FIG. 30) for each of the plurality of encoded data (S4932). Each of the plurality of control information stores a first identifier (for example, slice_idx) indicating a subspace corresponding to the encoded data corresponding to the control information and a second identifier (for example, tile_idx) indicating the divided data corresponding to the encoded data corresponding to the control information.
[0252] According to this, a three-dimensional data decoding device that decodes the bitstream generated by the three-dimensional data encoding device can easily restore the target space by combining the data of the plurality of divided data using the first identifier and the second identifier. Therefore, the processing amount in the three-dimensional data decoding device can be reduced.
[0253] For example, in the encoding by the three-dimensional data encoding device, the position information and the attribute information of the three-dimensional points included in each of the plurality of divided data are encoded. Each of the plurality of encoded data includes encoded data of the position information and encoded data of the attribute information. Each of the plurality of control information includes control information of the encoded data of the position information and control information of the encoded data of the attribute information. The first identifier and the second identifier are stored in the control information of the encoded data of the position information.
[0254] For example, in the bitstream, each of the plurality of control information is arranged before the encoded data corresponding to the control information.
[0255] In addition, the three-dimensional data encoding device sets a target space including a plurality of three-dimensional points into one or more sub-spaces, includes one or more divided data including one or more three-dimensional points in the sub-space, generates a plurality of encoded data corresponding to each of the plurality of divided data by encoding each of the divided data, generates a bit stream including the plurality of encoded data and a plurality of control information for each of the plurality of encoded data, and each of the plurality of control information may store a first identifier indicating the sub-space corresponding to the encoded data corresponding to the control information and a second identifier indicating the divided data corresponding to the encoded data corresponding to the control information.
[0256] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor performs the above processing using the memory.
[0257] In addition, the three-dimensional data decoding device according to the present embodiment performs the process shown in FIG. 34. First, the three-dimensional data decoding device is included in a plurality of sub-spaces (for example, slices) obtained by dividing a target space including a plurality of three-dimensional points, and each of a plurality of divided data (for example, tiles) each including one or more three-dimensional points is encoded. From a bit stream including a plurality of encoded data generated thereby and a plurality of control information (for example, the header shown in FIG. 30) for each of the plurality of encoded data, a first identifier (for example, slice_idx) indicating a sub-space corresponding to the encoded data corresponding to the control information stored in the plurality of control information, and a second identifier (for example, tile_idx) indicating divided data corresponding to the encoded data corresponding to the control information are acquired (S4941). Next, the three-dimensional data decoding device restores a plurality of divided data by decoding the plurality of encoded data (S4942). Next, the three-dimensional data decoding device restores the target space by combining the plurality of divided data using the first identifier and the second identifier (S4943). For example, the three-dimensional data encoding device restores a plurality of sub-spaces by combining the plurality of divided data using the second identifier, and restores the target space (a plurality of three-dimensional points) by combining the plurality of sub-spaces using the first identifier. Note that the three-dimensional data decoding device may acquire the encoded data of a desired sub-space or divided data from the bit stream using at least one of the first identifier and the second identifier, and selectively decode or preferentially decode the acquired encoded data.
[0258] According to this, the three-dimensional data decoding device can easily restore the target space by combining the data of the plurality of divided data using the first identifier and the second identifier. Therefore, the processing amount in the three-dimensional data decoding device can be reduced.
[0259] For example, each of a plurality of encoded data is generated by encoding the position information and the attribute information of three-dimensional points included in corresponding divided data, and includes the encoded data of the position information and the encoded data of the attribute information. Each of a plurality of control information includes the control information of the encoded data of the position information and the control information of the encoded data of the attribute information. The first identifier and the second identifier are stored in the control information of the encoded data of the position information.
[0260] For example, in a bit stream, the control information is arranged before the corresponding encoded data.
[0261] For example, a three-dimensional data decoding device includes a processor and a memory, and the processor performs the above processing using the memory.
[0262] (Embodiment 5) In the position information encoding using adjacency dependence, the higher the density of the point cloud, the more likely the encoding efficiency is improved. In this embodiment, the three-dimensional data encoding device combines the point cloud data of consecutive frames, and encodes the point cloud data of consecutive frames together. At this time, the three-dimensional data encoding device generates encoded data with information added for identifying the frame to which each leaf node included in the combined point cloud data belongs.
[0263] Here, the point cloud data of consecutive frames is likely to be similar. Therefore, in consecutive frames, the upper levels of the occupancy codes are likely to be the same. That is, by encoding consecutive frames together, the upper levels of the occupancy codes can be shared.
[0264] Also, the distinction of which frame the point cloud belongs to is made at the leaf node by encoding the index of the frame.
[0265] FIG. 35 is a diagram showing an image of generating a tree structure and an occupancy code from point cloud data of N PCC (Point Cloud Compression) frames. In the figure, the points in the arrows indicate the points belonging to each PCC frame. First, each point belonging to each PCC frame is assigned a frame index for identifying the frame.
[0266] Next, the points belonging to the N frames are converted into a tree structure, and an occupancy code is generated. Specifically, for each point, it is determined which leaf node in the tree structure the point belongs to. In the figure, the tree structure shows a set of nodes. Starting from the upper nodes in order, it is determined which node the point belongs to. The determination result for each node is encoded as an occupancy code. The occupancy code is common to the N frames.
[0267] Nodes may contain points from different frames with different frame indices assigned. Note that when the resolution of the octree is low, points from the same frame with the same frame index may also be mixed.
[0268] The bottom - layer nodes (leaf nodes) may contain points belonging to multiple frames mixed (duplicated).
[0269] In the tree structure and the occupancy code, the tree structure and the occupancy code at the upper levels may be common components in all frames, and the tree structure and the occupancy code at the lower levels may be individual components for each frame, or may be a mixture of common components and individual components.
[0270] For example, at the bottom - layer nodes such as leaf nodes, 0 or more points with frame indices are generated, and information indicating the number of points and information on the frame indices for each point are generated. These information can be said to be individual information for each frame.
[0271] FIG. 36 is a diagram showing an example of frame combination. As shown in FIG. 36(a), by generating a tree structure by grouping a plurality of frames together, the point density of the frames included in the same node increases. Also, by sharing the tree structure, the data amount of the occupancy code can be reduced. Thereby, there is a possibility of improving the coding rate.
[0272] Also, as shown in FIG. 36(b), since the arithmetic coding effect is enhanced by the individual components of the occupancy code in the tree structure becoming denser, there is a possibility of improving the coding rate.
[0273] Hereinafter, the combination of a plurality of temporally different PCC frames will be described as an example, but it is also applicable when there are not a plurality of frames, that is, when frames are not combined (N = 1). Also, the plurality of point cloud data to be combined is not limited to a plurality of frames, that is, point cloud data at different times of the same object. That is, the following method is also applicable to the combination of a plurality of spatially or spatio-temporally different point cloud data. Also, the following method is also applicable to the combination of point cloud data or point cloud files with different contents.
[0274] FIG. 37 is a diagram showing an example of the combination of a plurality of temporally different PCC frames. FIG. 37 shows an example of acquiring point cloud data with a sensor such as LiDAR while an automobile is moving. The dotted line indicates the acquisition range of the sensor for each frame, that is, the area of the point cloud data. When the acquisition range of the sensor is large, the range of the point cloud data also becomes large.
[0275] The method of combining and encoding point cloud data is effective for the following types of point cloud data. For example, in the example shown in FIG. 37, the automobile is moving, and the frame is identified by a 360° scan around the automobile. That is, frame 2, which is the next frame, corresponds to another 360° scan after the vehicle has moved in the X direction.
[0276] In this case, since there are overlapping regions between Frame 1 and Frame 2, there is a possibility that they contain the same point cloud data. Therefore, it may be possible to improve the encoding efficiency by combining and encoding Frame 1 and Frame 2. Note that it is also conceivable to combine more frames. However, increasing the number of frames to be combined increases the number of bits required for encoding the frame indexes added to the leaf nodes.
[0277] Also, point cloud data may be acquired by sensors at different positions. Thereby, each point cloud data acquired from each position may be used as a frame. That is, the plurality of frames may be point cloud data acquired by a single sensor, or may be point cloud data acquired by a plurality of sensors. Also, among the plurality of frames, some or all of the objects may be the same or different.
[0278] Next, the flow of the three-dimensional data encoding process according to the present embodiment will be described. FIG. 38 is a flowchart of the three-dimensional data encoding process. The three-dimensional data encoding device reads the point cloud data of all N frames based on the number of frames to be combined, which is the number of combined frames N.
[0279] First, the three-dimensional data encoding device determines the number of combined frames N (S5401). For example, this number of combined frames N is specified by the user.
[0280] Next, the three-dimensional data encoding device acquires point cloud data (S5402). Next, the three-dimensional data encoding device records the frame indexes of the acquired point cloud data (S5403).
[0281] If the N frames have not been processed (No in S5404), the three-dimensional data encoding device designates the next point cloud data (S5405) and performs the processes after step S5402 on the designated point cloud data.
[0282] On the other hand, when the N frames have been processed (Yes in S5404), the three-dimensional data encoding device combines the N frames and encodes the combined frame (S5406).
[0283] FIG. 39 is a flowchart of the encoding process (S5406). First, the three-dimensional data encoding device generates common information common to the N frames (S5411). For example, the common information includes an occupancy code and information indicating the number of combined frames N.
[0284] Next, the three-dimensional data encoding device generates individual information, which is individual information for each frame (S5412). For example, the individual information includes the number of points included in the leaf node and the frame index of the points included in the leaf node.
[0285] Next, the three-dimensional data encoding device combines the common information and the individual information, and generates encoded data by encoding the combined information. Next, the three-dimensional data encoding device generates additional information (metadata) related to frame combination, and encodes the generated additional information (S5414).
[0286] Next, the flow of the three-dimensional data decoding process according to the present embodiment will be described. FIG. 40 is a flowchart of the three-dimensional data decoding process.
[0287] First, the three-dimensional data decoding device acquires the number of combined frames N from the bit stream (S5421). Next, the three-dimensional data encoding device acquires the encoded data from the bit stream (S5422). Next, the three-dimensional data decoding device decodes the encoded data to acquire the point cloud data and the frame index (S5423). Finally, the three-dimensional data decoding device divides the decoded point cloud data using the frame index (S5424).
[0288] Figure 41 is a flowchart of the decoding and splitting processes (S5423 and S5424). First, the three-dimensional data decoding device decodes (acquires) common information and individual information from the encoded data (bitstream) (S5431).
[0289] Next, the three-dimensional data decoding device determines whether to decode a single frame or a plurality of frames (S5432). For example, whether to decode a single frame or a plurality of frames may be specified from the outside. Here, the plurality of frames may be all the frames of the combined frames or some of the frames. For example, the three-dimensional data decoding device may determine to decode a specific frame required by the application and determine not to decode the frames that are not required. Or, when real-time decoding is required, the three-dimensional data decoding device may determine to decode a single frame among the plurality of combined frames.
[0290] When decoding a single frame (Yes in S5432), the three-dimensional data decoding device extracts the individual information corresponding to the specified single frame index from the decoded individual information, and restores the point cloud data of the frame corresponding to the specified frame index by decoding the extracted individual information (S5433).
[0291] On the other hand, when decoding a plurality of frames (No in S5432), the three-dimensional data decoding device extracts the individual information corresponding to the frame indices of the specified plurality of frames (or all frames), and restores the point cloud data of the specified plurality of frames by decoding the extracted individual information (S5434). Next, the three-dimensional data decoding device divides the decoded point cloud data (individual information) based on the frame index (S5435). That is, the three-dimensional data decoding device divides the decoded point cloud data into a plurality of frames.
[0292] Note that the three-dimensional data decoding device may decode the data of all the combined frames at once and divide the decoded data into each frame, or may decode any partial frames among all the combined frames at once and divide the decoded data into each frame. Further, the three-dimensional data decoding device may decode a predetermined unit frame composed of a plurality of frames alone.
[0293] Hereinafter, the configuration of the three-dimensional data encoding device according to the present embodiment will be described. FIG. 42 is a block diagram showing the configuration of an encoding unit 5410 included in the three-dimensional data encoding device according to the present embodiment. The encoding unit 5410 generates encoded data (encoded stream) by encoding point cloud data (point cloud). This encoding unit 5410 includes a division unit 5411, a plurality of position information encoding units 5412, a plurality of attribute information encoding units 5413, an additional information encoding unit 5414, and a multiplexing unit 5415.
[0294] The division unit 5411 generates a plurality of divided data of a plurality of frames by dividing the point cloud data of the plurality of frames. Specifically, the division unit 5411 generates a plurality of divided data by dividing the space of the point cloud data of each frame into a plurality of subspaces. Here, the subspace is either one of a tile and a slice, or a combination of a tile and a slice. More specifically, the point cloud data includes position information, attribute information (such as color or reflectance), and additional information. Further, a frame number is input to the division unit 5411. The division unit 5411 divides the position information of each frame into a plurality of divided position information, and divides the attribute information of each frame into a plurality of divided attribute information. Further, the division unit 5411 generates additional information regarding the division.
[0295] For example, the division unit 5411 first divides the point cloud into tiles. Next, the division unit 5411 further divides the obtained tiles into slices.
[0296] A plurality of position information encoding units 5412 generate a plurality of encoded position information by encoding a plurality of divided position information. For example, the position information encoding unit 5412 encodes the divided position information using an N-ary tree structure such as an octree. Specifically, in an octree, the target space is divided into eight nodes (sub-spaces), and 8-bit information (occupancy code) indicating whether or not a point cloud is included in each node is generated. Further, a node including a point cloud is further divided into eight nodes, and 8-bit information indicating whether or not a point cloud is included in each of the eight nodes is generated. This process is repeated until it becomes equal to or less than a threshold value of the number of point clouds included in a predetermined layer or node. For example, the plurality of position information encoding units 5412 perform parallel processing on the plurality of divided position information.
[0297] The attribute information encoding unit 4632 generates encoded attribute information, which is encoded data, by encoding using the configuration information generated by the position information encoding unit 4631. For example, the attribute information encoding unit 4632 determines a reference point (reference node) to be referred to in the encoding of a target point (target node) to be processed based on the octree structure generated by the position information encoding unit 4631. For example, the attribute information encoding unit 4632 refers to a node in which the parent node in the octree is the same as the target node among the surrounding nodes or adjacent nodes. Note that the method for determining the reference relationship is not limited to this.
[0298] Further, the encoding process of the position information or the attribute information may include at least one of quantization processing, prediction processing, and arithmetic encoding processing. In this case, reference means using the reference node to calculate the predicted value of the attribute information, or using the state of the reference node (for example, occupancy information indicating whether or not a point cloud is included in the reference node) to determine the encoding parameter. For example, the encoding parameter is a quantization parameter in quantization processing or a context in arithmetic encoding.
[0299] A plurality of attribute information encoding units 5413 generate a plurality of encoded attribute information by encoding a plurality of divided attribute information. For example, the plurality of attribute information encoding units 5413 perform parallel processing on the plurality of divided attribute information.
[0300] The additional information encoding unit 5414 generates encoded additional information by encoding the additional information included in the point cloud data and the additional information regarding data division generated at the time of division by the division unit 5411.
[0301] The multiplexing unit 5415 generates encoded data (encoded stream) by multiplexing the plurality of encoded position information, the plurality of encoded attribute information, and the encoded additional information of a plurality of frames, and transmits the generated encoded data. Also, the encoded additional information is used at the time of decoding.
[0302] FIG. 43 is a block diagram of the division unit 5411. The division unit 5411 includes a tile division unit 5421 and a slice division unit 5422.
[0303] The tile division unit 5421 generates a plurality of tile position information by dividing each of the position information (Position(Geometry)) of a plurality of frames into tiles. Also, the tile division unit 5421 generates a plurality of tile attribute information by dividing each of the attribute information (Attribute) of a plurality of frames into tiles. Also, the tile division unit 5421 outputs tile additional information (Tile MetaData) including information related to tile division and information generated in tile division.
[0304] The slice division unit 5422 generates a plurality of divided position information (a plurality of slice position information) by dividing the plurality of tile position information into slices. Also, the slice division unit 5422 generates a plurality of divided attribute information (a plurality of slice attribute information) by dividing the plurality of tile attribute information into slices. Also, the slice division unit 5422 outputs slice additional information (Slice MetaData) including information related to slice division and information generated in slice division.
[0305] Also, the division unit 5411 uses a frame number (frame index) in order to indicate the origin coordinates, the attribute information, etc. in the division process.
[0306] FIG. 44 is a block diagram of the position information encoding unit 5412. The position information encoding unit 5412 includes a frame index generation unit 5431 and an entropy encoding unit 5432.
[0307] The frame index generation unit 5431 determines the value of the frame index based on the frame number, and adds the determined frame index to the position information. The entropy encoding unit 5432 generates encoded position information by entropy encoding the divided position information to which the frame index is added.
[0308] FIG. 45 is a block diagram of the attribute information encoding unit 5413. The attribute information encoding unit 5413 includes a frame index generation unit 5441 and an entropy encoding unit 5442.
[0309] The frame index generation unit 5441 determines the value of the frame index based on the frame number, and adds the determined frame index to the attribute information. The entropy encoding unit 5442 generates encoded attribute information by entropy encoding the divided attribute information to which the frame index is added.
[0310] Next, the flow of the encoding process and the decoding process of the point cloud data according to the present embodiment will be described. FIG. 46 is a flowchart of the encoding process of the point cloud data according to the present embodiment.
[0311] First, the three-dimensional data encoding device determines the division method to be used (S5441). This division method includes whether to perform slice division or not, and whether to perform tile division or not. Further, the division method may include the number of divisions in the case of performing slice division or tile division, and the type of division and the like.
[0312] When tile splitting is performed (Yes in S5442), the three-dimensional data encoding device generates a plurality of tile position information and a plurality of tile attribute information by splitting the position information and the attribute information (S5443). Further, the three-dimensional data encoding device generates tile addition information related to tile splitting.
[0313] When slice splitting is performed (Yes in S5444), the three-dimensional data encoding device generates a plurality of split position information and a plurality of split attribute information by splitting the plurality of tile position information and the plurality of tile attribute information (or the position information and the attribute information) (S5445). Further, the three-dimensional data encoding device generates slice addition information related to slice splitting.
[0314] Next, the three-dimensional data encoding device generates a plurality of encoded position information and a plurality of encoded attribute information by encoding each of the plurality of split position information and the plurality of split attribute information with a frame index (S5446). Further, the three-dimensional data encoding device generates dependency information.
[0315] Next, the three-dimensional data encoding device generates encoded data (encoded stream) by NAL unitizing (multiplexing) the plurality of encoded position information, the plurality of encoded attribute information, and the addition information (S5447). Further, the three-dimensional data encoding device transmits the generated encoded data.
[0316] FIG. 47 is a flowchart of the encoding process (S5446). First, the three-dimensional data encoding device encodes the split position information (S5451). Next, the three-dimensional data encoding device encodes the frame index for the split position information (S5452).
[0317] When there is segmentation attribute information (Yes in S5453), the 3D data encoding device encodes the segmentation attribute information (S5454) and encodes the frame index for the segmentation attribute information (S5455). On the other hand, when there is no segmentation attribute information (No in S5453), the 3D data encoding device does not perform the encoding of the segmentation attribute information and the encoding of the frame index for the segmentation attribute information. Note that the frame index may be stored in either or both of the segmentation position information and the segmentation attribute information.
[0318] Note that the 3D data encoding device may encode the attribute information using the frame index or may encode it without using the frame index. That is, the 3D data encoding device may use the frame index to identify the frame to which each point belongs and encode each frame, or may encode the points belonging to all frames without identifying the frames.
[0319] Hereinafter, the configuration of the 3D data decoding device according to the present embodiment will be described. FIG. 48 is a block diagram showing the configuration of the decoding unit 5450. The decoding unit 5450 restores the point cloud data by decoding the encoded data (encoded stream) generated by encoding the point cloud data. The decoding unit 5450 includes a demultiplexing unit 5451, a plurality of position information decoding units 5452, a plurality of attribute information decoding units 5453, an additional information decoding unit 5454, and a combining unit 5455.
[0320] The demultiplexing unit 5451 generates a plurality of encoded position information, a plurality of encoded attribute information, and encoded additional information by demultiplexing the encoded data (encoded stream).
[0321] The plurality of position information decoding units 5452 generate a plurality of segmentation position information by decoding the plurality of encoded position information. For example, the plurality of position information decoding units 5452 perform parallel processing on the plurality of encoded position information.
[0322] A plurality of attribute information decoding units 5453 generate a plurality of divided attribute information by decoding a plurality of encoded attribute information. For example, the plurality of attribute information decoding units 5453 perform parallel processing on the plurality of encoded attribute information.
[0323] A plurality of additional information decoding units 5454 generate additional information by decoding the encoded additional information.
[0324] The combining unit 5455 generates position information by combining a plurality of divided position information using the additional information. The combining unit 5455 generates attribute information by combining a plurality of divided attribute information using the additional information. Also, the combining unit 5455 divides the position information and the attribute information into position information of a plurality of frames and attribute information of a plurality of frames using the frame index.
[0325] FIG. 49 is a block diagram of the position information decoding unit 5452. The position information decoding unit 5452 includes an entropy decoding unit 5461 and a frame index acquisition unit 5462. The entropy decoding unit 5461 generates divided position information by performing entropy decoding on the encoded position information. The frame index acquisition unit 5462 acquires a frame index from the divided position information.
[0326] FIG. 50 is a block diagram of the attribute information decoding unit 5453. The attribute information decoding unit 5453 includes an entropy decoding unit 5471 and a frame index acquisition unit 5472. The entropy decoding unit 5471 generates divided attribute information by performing entropy decoding on the encoded attribute information. The frame index acquisition unit 5472 acquires a frame index from the divided attribute information.
[0327] FIG. 51 is a diagram showing the configuration of the combining unit 5455. The combining unit 5455 generates position information by combining a plurality of divided position information. The combining unit 5455 generates attribute information by combining a plurality of divided attribute information. Also, the combining unit 5455 divides the position information and the attribute information into position information of a plurality of frames and attribute information of a plurality of frames using the frame index.
[0328] FIG. 52 is a flowchart of the decoding process of the point cloud data according to the present embodiment. First, the three-dimensional data decoding device determines the splitting method by analyzing the additional information (slice additional information and tile additional information) related to the splitting method included in the encoded data (encoded stream) (S5461). This splitting method includes whether to perform slice splitting or tile splitting. Further, the splitting method may include the number of splits and the type of split when performing slice splitting or tile splitting.
[0329] Next, the three-dimensional data decoding device generates split position information and split attribute information by decoding a plurality of encoded position information and a plurality of encoded attribute information included in the encoded data using the dependency information included in the encoded data (S5462).
[0330] When it is shown by the additional information that slice splitting is performed (Yes in S5463), the three-dimensional data decoding device generates a plurality of tile position information by combining the plurality of split position information and generates a plurality of tile attribute information by combining the plurality of split attribute information based on the slice additional information (S5464). Here, the plurality of split position information, the plurality of split attribute information, the plurality of tile position information, and the plurality of tile attribute information include frame indexes.
[0331] When it is shown by the additional information that tile splitting is performed (Yes in S5465), the three-dimensional data decoding device generates position information by combining the plurality of tile position information (the plurality of split position information) and generates attribute information by combining the plurality of tile attribute information (the plurality of split attribute information) based on the tile additional information (S5466). Here, the plurality of tile position information, the plurality of tile attribute information, the position information, and the attribute information include frame indexes.
[0332] FIG. 53 is a flowchart of the decoding process (S5464 or S5466). First, the three-dimensional data decoding device decodes the division position information (slice position information) (S5471). Next, the three-dimensional data decoding device decodes the frame index for the division position information (S5472).
[0333] When the division attribute information exists (Yes in S5473), the three-dimensional data decoding device decodes the division attribute information (S5474) and decodes the frame index for the division attribute information (S5475). On the other hand, when the division attribute information does not exist (No in S5473), the three-dimensional data decoding device does not decode the division attribute information and does not decode the frame index for the division attribute information.
[0334] Note that the three-dimensional data decoding device may decode the attribute information using the frame index or may decode it without using the frame index.
[0335] Hereinafter, the encoding unit in frame combination will be described. FIG. 54 is a diagram showing an example of a frame combination pattern. The example in this figure is, for example, an example in the case where PCC frames are in time series and data is generated and encoded in real time.
[0336] FIG. 54(a) shows the case of fixedly combining 4 frames. The three-dimensional data encoding device generates encoded data after waiting for the generation of data for 4 frames.
[0337] FIG. 54(b) shows the case where the number of frames changes adaptively. For example, the three-dimensional data encoding device changes the number of combined frames in order to adjust the amount of code of the encoded data in rate control.
[0338] Note that the three-dimensional data encoding device may not combine frames if there is a possibility that there is no effect due to frame combination. Also, the three-dimensional data encoding device may switch between the case of combining frames and the case of not combining frames.
[0339] FIG. 54(c) is an example in which a part of a plurality of frames to be combined overlaps with a part of a plurality of frames to be combined next. This example is useful when real-time performance or low latency is required, such as sequentially transmitting the encoded data.
[0340] FIG. 55 is a diagram showing a configuration example of a PCC frame. The three-dimensional data encoding device may be configured such that the frames to be combined include at least data units that can be decoded independently. For example, as shown in FIG. 55(a), when all the PCC frames are intra-encoded and the PCC frames can be decoded independently, any of the above patterns can be applied.
[0341] Also, as shown in FIG. 55(b), when a random access unit such as a GOF (Group of Frames) is set, such as when inter prediction is applied, the three-dimensional data encoding device may combine the data with the GOF unit as the minimum unit.
[0342] Note that the three-dimensional data encoding device may encode the common information and the individual information together, or may encode them separately. Also, the three-dimensional data encoding device may use a common data structure for the common information and the individual information, or may use different data structures.
[0343] Also, after generating the occupancy code for each frame, the three-dimensional data encoding device compares the occupancy codes of a plurality of frames, and for example, determines whether there is a large common part between the occupancy codes of the plurality of frames based on a predetermined criterion. If there is a large common part, common information may be generated. Alternatively, the three-dimensional data encoding device may determine whether to combine frames, which frames to combine, or the number of frames to combine based on whether there is a large common part.
[0344] Next, the configuration of the encoding position information will be described. FIG. 56 is a diagram showing the configuration of the encoding position information. The encoding position information includes a header and a payload.
[0345] FIG. 57 is a diagram showing an example of the syntax of the header (Geometry_header) of the encoded position information. The header of the encoded position information includes a GPS index (gps_idx), offset information (offset), other information (other_geometry_information), a frame combination flag (combine_frame_flag), and the number of combined frames (number_of_combine_frame).
[0346] The GPS index indicates the identifier (ID) of the parameter set (GPS) corresponding to the encoded position information. GPS is a parameter set of the encoded position information for one frame or a plurality of frames. In addition, when there is a parameter set for each frame, the identifiers of a plurality of parameter sets may be shown in the header.
[0347] The offset information indicates the offset position for obtaining combined data. The other information indicates other information related to the position information (for example, the difference value of the quantization parameter (QPdelta), etc.). The frame combination flag is a flag indicating whether the encoded data is frame-combined. The number of combined frames indicates the number of combined frames.
[0348] In addition, some or all of the above information may be described in the SPS or GPS. Note that the SPS is a parameter set in units of a sequence (a plurality of frames), and is a parameter set commonly used for the encoded position information and the encoded attribute information.
[0349] FIG. 58 is a diagram showing an example of the syntax of the payload (Geometry_data) of the encoded position information. The payload of the encoded position information includes common information and leaf node information.
[0350] The common information is data combined from one or more frames and includes an occupancy code (occupancy_Code), etc.
[0351] Leaf node information (combine_information) is the information of each leaf node. As a loop of the number of frames, leaf node information may be shown for each frame.
[0352] As a method for indicating the frame index of points included in a leaf node, either Method 1 or Method 2 can be used. FIG. 59 is a diagram showing an example of leaf node information in the case of Method 1. The leaf node information shown in FIG. 59 includes the three-dimensional number of points (NumberOfPoints) indicating the number of points included in the node and the frame index (FrameIndex) for each point.
[0353] FIG. 60 is a diagram showing an example of leaf node information in the case of Method 2. In the example shown in FIG. 60, the leaf node information includes bitmap information (bitmapIsFramePointsFlag) indicating the frame indices of a plurality of points by a bitmap. FIG. 61 is a diagram showing an example of the bitmap information. In this example, it is shown by the bitmap that the leaf node includes three-dimensional points with frame indices 1, 3, and 5.
[0354] Note that when the quantization resolution is low, there may be duplicate points in the same frame. In this case, the three-dimensional number of points (NumberOfPoints) may be shared, and the number of three-dimensional points in each frame and the total number of three-dimensional points in multiple frames may be shown.
[0355] Also, when irreversible compression is used, the three-dimensional data encoding device may delete duplicate points and reduce the amount of information. The three-dimensional data encoding device may delete duplicate points before frame combination or after frame combination.
[0356] Next, the configuration of the encoding attribute information will be described. FIG. 62 is a diagram showing the configuration of the encoding attribute information. The encoding attribute information includes a header and a payload.
[0357] FIG. 63 is a diagram showing an example of the syntax of the header (Attribute_header) of the encoded attribute information. The header of the encoded attribute information includes an APS index (aps_idx), offset information (offset), other information (other_attribute_information), a frame combination flag (combine_frame_flag), and the number of combined frames (number_of_combine_frame).
[0358] The APS index indicates the identifier (ID) of the parameter set (APS) corresponding to the encoded attribute information. The APS is a parameter set of the encoded attribute information for one frame or a plurality of frames. Note that when there is a parameter set for each frame, the identifiers of a plurality of parameter sets may be shown in the header.
[0359] The offset information indicates the offset position for obtaining the combined data. The other information indicates other information regarding the attribute information (for example, the differential value of the quantization parameter (QPdelta), etc.). The frame combination flag is a flag indicating whether the encoded data is frame-combined. The number of combined frames indicates the number of combined frames.
[0360] Note that some or all of the above information may be described in the SPS or APS.
[0361] FIG. 64 is a diagram showing an example of the syntax of the payload (Attribute_data) of the encoded attribute information. The payload of the encoded attribute information includes leaf node information (combine_information). For example, the configuration of this leaf node information is the same as the leaf node information included in the payload of the encoded position information. That is, the leaf node information (frame index) may be included in the attribute information.
[0362] In addition, the leaf node information (frame index) may be stored in one of the encoding position information and the encoding attribute information, and may not be stored in the other. In this case, the leaf node information (frame index) stored in one of the encoding position information and the encoding attribute information is referred to when decoding the other information. Further, information indicating the reference destination may be included in the encoding position information or the encoding attribute information.
[0363] Next, an example of the transmission order and decoding order of the encoded data will be described. FIG. 65 is a diagram showing the configuration of the encoded data. The encoded data includes a header and a payload.
[0364] FIGS. 66 to 68 are diagrams showing the transmission order of data and the reference relationship of data. In the figure, G(1) etc. indicate the encoding position information, GPS(1) etc. indicate the parameter set of the encoding position information, and SPS indicates the parameter set of the sequence (multiple frames). Also, the number in () indicates the value of the frame index. Note that the three-dimensional data encoding device may transmit the data in the decoding order.
[0365] FIG. 66 is a diagram showing an example of the transmission order when frames are not combined. FIG. 67 is a diagram showing an example when frames are combined and metadata (parameter set) is added for each PCC frame. FIG. 68 is a diagram showing an example when frames are combined and metadata (parameter set) is added in units of combination.
[0366] In the header of the frame-combined data, an identifier of the metadata of the reference destination is stored in order to obtain the metadata of the frame. As shown in FIG. 68, the metadata for each of a plurality of frames may be grouped together. Parameters common to a plurality of frame-combined frames may be grouped into one. Parameters not common to the frames indicate values for each frame.
[0367] Information for each frame (parameters not common to frames) is, for example, a time stamp indicating the generation time, encoding time, or decoding time of frame data. Further, the information for each frame may include information on the sensor that acquired the frame data (such as the speed of the sensor, acceleration, position information, the orientation of the sensor, and other sensor information).
[0368] FIG. 69 is a diagram showing an example of decoding some frames in the example shown in FIG. 67. As shown in FIG. 69, in the frame combined data, if there is no dependency between frames, the three-dimensional data decoding apparatus can decode each data independently.
[0369] When the point cloud data has attribute information, the three-dimensional data encoding apparatus may combine the attribute information. The attribute information is encoded and decoded with reference to the position information. The position information to be referred to may be the position information before frame combination or the position information after frame combination. The number of combined frames of the position information and the number of combined frames of the attribute information may be the same or independent (different).
[0370] FIGS. 70 to 73 are diagrams showing the transmission order of data and the reference relationship of data. FIGS. 70 and 71 show an example of combining both the position information and the attribute information in 4 frames. In FIG. 70, metadata (parameter set) is added for each PCC frame. In FIG. 71, metadata (parameter set) is added in units of combination. In the figure, A(1) etc. indicate encoded attribute information, and APS(1) etc. indicate parameter sets of the encoded attribute information. Also, the numbers in parentheses indicate the values of the frame indices.
[0371] FIG. 72 shows an example of combining the position information in 4 frames and not combining the attribute information. As shown in FIG. 72, the position information may be combined in frames without combining the attribute information.
[0372] FIG. 73 shows an example combining frame combination and tile division. When performing tile division as shown in FIG. 73, the header of each tile position information includes information such as GPS index (gps_idx) and the number of combined frames (number_of_combine_frame). Further, the header of each tile position information includes a tile index (tile_idx) for identifying a tile.
[0373] As described above, the three-dimensional data encoding device according to the present embodiment performs the processing shown in FIG. 74. First, the three-dimensional data encoding device generates third point cloud data by combining the first point cloud data and the second point cloud data (S5481). Next, the three-dimensional data encoding device generates encoded data by encoding the third point cloud data (S5482). Further, the encoded data includes identification information (e.g., frame index) indicating to which of the first point cloud data and the second point cloud data each of the plurality of three-dimensional points included in the third point cloud data belongs.
[0374] According to this, the three-dimensional data encoding device can improve the encoding efficiency by encoding a plurality of point cloud data together.
[0375] For example, the first point cloud data and the second point cloud data are point cloud data (e.g., PCC frames) at different times. For example, the first point cloud data and the second point cloud data are point cloud data (e.g., PCC frames) at different times of the same object.
[0376] The encoded data includes the position information and attribute information of each of the plurality of three-dimensional points included in the third point cloud data, and the identification information is included in the attribute information.
[0377] For example, the encoded data includes position information (e.g., occupancy code) representing the position of each of the plurality of three-dimensional points included in the third point cloud data using an N-ary tree (where N is an integer of 2 or more).
[0378] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor performs the above processing using the memory.
[0379] Also, the three-dimensional data decoding device according to the present embodiment performs the processing shown in FIG. 75. First, the three-dimensional data decoding device decodes the encoded data to obtain the third point cloud data generated by combining the first point cloud data and the second point cloud data, and identification information indicating to which of the first point cloud data and the second point cloud data each of the plurality of three-dimensional points included in the third point cloud data belongs (S5491). Next, the three-dimensional data decoding device separates the first point cloud data and the second point cloud data from the third point cloud data using the identification information (S5492).
[0380] According to this, the three-dimensional data decoding device can decode the encoded data with improved encoding efficiency by encoding a plurality of point cloud data together.
[0381] For example, the first point cloud data and the second point cloud data are point cloud data (for example, PCC frames) at different times. For example, the first point cloud data and the second point cloud data are point cloud data (for example, PCC frames) at different times of the same object.
[0382] The encoded data includes the position information and attribute information of each of the plurality of three-dimensional points included in the third point cloud data, and the identification information is included in the attribute information.
[0383] For example, the encoded data includes position information (for example, occupancy code) representing the position of each of the plurality of three-dimensional points included in the third point cloud data using an N-ary tree (where N is an integer of 2 or more).
[0384] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor performs the above processing using the memory.
[0385] (Embodiment 6) In this embodiment, spatial joining for encoding by joining point cloud data in different spaces will be described. FIGS. 76 and 77 are diagrams showing examples of joining a plurality of point cloud data that are spatially different. Hereinafter, an example in the case where the point cloud data is static point cloud data will be described.
[0386] In the example shown in FIG. 76, the three-dimensional data encoding device divides one point cloud data into four point clouds, and joins (merges) the four divided spaces into one space.
[0387] In the example shown in FIG. 77, the three-dimensional data encoding device divides point cloud data such as a large-scale map into tiles based on position information, and further divides each tile into slices according to the object to which the three-dimensional points belong. The three-dimensional data encoding device further joins (merges) a part of the data divided into tiles and slices. Here, a tile is, for example, a sub-space obtained by dividing the original space based on position information. A slice is, for example, a point cloud obtained by classifying three-dimensional points according to the type of object to which the three-dimensional points belong. Note that the slice may be a sub-space obtained by dividing the original space or a tile.
[0388] In the example shown in FIG. 77, among the data divided into slices, the three-dimensional data encoding device joins (merges) slices having the same object for buildings and trees. Note that the three-dimensional data encoding device may divide one or more point cloud data and merge two or more point cloud data after division. Further, the three-dimensional data encoding device may join a part of the divided data and not join the rest.
[0389] Next, the configuration of the three-dimensional data encoding device according to this embodiment will be described. FIG. 78 is a block diagram showing the configuration of an encoding unit 6100 included in the three-dimensional data encoding device.
[0390] The symbolization unit 6100 generates encoded data (encoded stream) by encoding point cloud data (point cloud). This symbolization unit 6100 includes a splitting and combining unit 6101, a plurality of position information encoding units 6102, a plurality of attribute information encoding units 6103, an additional information encoding unit 6104, and a multiplexing unit 6105.
[0391] The splitting and combining unit 6101 generates a plurality of split data by splitting the point cloud data, and generates combined data by combining the generated plurality of split data. Specifically, the splitting and combining unit 6101 generates a plurality of split data by splitting the space of the point cloud data into a plurality of sub-spaces. Here, the sub-space is either one of a tile and a slice, or a combination of a tile and a slice. More specifically, the point cloud data includes position information, attribute information (such as color or reflectance), and additional information. The splitting and combining unit 6101 splits the position information into a plurality of split position information, and generates combined position information by combining the plurality of split position information. Also, the splitting and combining unit 6101 splits the attribute information into a plurality of split attribute information, and generates combined attribute information by combining the plurality of split attribute information. Further, the splitting and combining unit 6101 generates split additional information regarding splitting and combined additional information regarding combining.
[0392] The plurality of position information encoding units 6102 generate encoded position information by encoding the combined position information. For example, the position information encoding unit 6102 encodes the split position information using an N-ary tree structure such as an octree. Specifically, in an octree, the target space is split into eight nodes (sub-spaces), and 8-bit information (occupancy code) indicating whether or not each node contains a point cloud is generated. Further, a node containing a point cloud is split into eight nodes, and 8-bit information indicating whether or not each of the eight nodes contains a point cloud is generated. This process is repeated until the number of point clouds included in a predetermined layer or node reaches a threshold value or less.
[0393] The attribute information encoding unit 6103 generates encoded attribute information by encoding combined attribute information using the configuration information generated by the position information encoding unit 6102. For example, the attribute information encoding unit 6103 determines a reference point (reference node) to be referred to in the encoding of a target point (target node) to be processed based on the octree structure generated by the position information encoding unit 6102. For example, the attribute information encoding unit 6103 refers to a node among peripheral nodes or adjacent nodes whose parent node in the octree is the same as the target node. Note that the method for determining the reference relationship is not limited to this.
[0394] Also, the encoding process of the position information or the attribute information may include at least one of quantization processing, prediction processing, and arithmetic encoding processing. In this case, reference means using the reference node to calculate the predicted value of the attribute information, or using the state of the reference node (for example, occupancy information indicating whether a point cloud is included in the reference node) to determine the encoding parameter. For example, the encoding parameter is a quantization parameter in quantization processing or a context in arithmetic encoding.
[0395] The additional information encoding unit 6104 generates encoded additional information by encoding the divided additional information and the combined additional information.
[0396] The multiplexing unit 6105 generates encoded data (encoded stream) by multiplexing the encoded position information, the encoded attribute information, and the encoded additional information, and sends out the generated encoded data. Also, the encoded additional information is used at the time of decoding.
[0397] FIG. 79 is a block diagram of the division and combination unit 6101. The division and combination unit 6101 includes a division unit 6111 and a combination unit 6112.
[0398] The splitting unit 6111 generates a plurality of split position information by splitting each of the position information (Position(Geometry)) into tiles. Also, the splitting unit 6111 generates a plurality of split attribute information by splitting each of the attribute information (Attribute) into tiles. Further, the splitting unit 6111 outputs split additional information including the information related to the splitting and the information generated in the splitting. Note that slices may be used instead of tiles, or a combination of tiles and slices may be used.
[0399] The combining unit 6112 generates combined position information by combining the split position information. Also, the combining unit 6112 adds a tile ID to each point included in the combined position information. The tile ID indicates the tile to which the corresponding point belongs. Further, the combining unit 6112 generates combined attribute information by combining the split attribute information.
[0400] FIG. 80 is a block diagram of the position information encoding unit 6102 and the attribute information encoding unit 6103. The position information encoding unit 6102 includes a tile ID encoding unit 6121 and an entropy encoding unit 6122.
[0401] The tile ID encoding unit 6121 generates a spatial index which is information indicating to which tile the overlapping points in the leaf node belong, and encodes the generated spatial index. Here, an example where the spatial index is a tile ID is shown. The entropy encoding unit 6122 generates encoded position information by entropy encoding the combined position information.
[0402] The attribute information encoding unit 6103 includes a tile ID encoding unit 6123 and an entropy encoding unit 6124.
[0403] The tile ID encoding unit 6123 generates a spatial index and encodes the generated spatial index. Here, an example where the spatial index is a tile ID is shown. The entropy encoding unit 6124 generates encoded attribute information by entropy encoding the split attribute information to which a frame index is added.
[0404] Note that the spatial index may be encoded as position information or as attribute information.
[0405] Thus, the three-dimensional data encoding device divides the point cloud data into tiles or slices, etc., for each of the position information and the attribute information, and then combines the divided data. The three-dimensional data encoding device arithmetic-encodes each of the combined data. Further, the three-dimensional data encoding device arithmetic-encodes the additional information (metadata) (for example, additional data and attribute information, etc.) generated during data division and data combination.
[0406] Next, the operations of spatial division and spatial combination will be described. FIG. 81 is a flowchart of three-dimensional data encoding processing by a three-dimensional data encoding device. First, the three-dimensional data encoding device determines a division method and a combination method (S6101). For example, the division method includes whether to perform slice division or not, and whether to perform tile division or not. Further, the division method may include the number of divisions and the type of division, etc., when performing slice division or tile division. Also, the combination method includes whether to perform combination or not. Further, when performing combination, the number of combinations and the type of combination, etc., may be included. Also, the division method and the combination method may be instructed from the outside, or may be determined based on the information of three-dimensional points (for example, the number or density, etc., of three dimensions).
[0407] Next, the three-dimensional data encoding device reads the point cloud data (S6102). Next, the three-dimensional data encoding device determines which tile each point belongs to (S6103). That is, tile division (or / and slice division) is performed. For example, the three-dimensional data encoding device divides the point cloud data into tiles or slices based on the information of the division region. FIG. 82 is a diagram schematically showing the division process and the combination process. The information of the division region is, for example, the spatial regions indicated by A, B, D, C, E, F in FIG. 82. For example, the information of the division region when dividing the point cloud may be information indicating the boundary of the space to be divided, or may be information indicating the reference position of the space and information indicating the range of the bounding box with respect to the reference position.
[0408] Also, the above information may be known, or the above information may be specified or derived using other information that can specify or derive the information. For example, when dividing point cloud data into tiles based on the GPS coordinates of a map, the reference position and size of each tile may be determined in advance. Or, a table showing the reference position and size associated with a tile identifier (tile ID) may be provided separately, and the three-dimensional data encoding device may determine the reference position and size of the tile, etc. based on the tile ID. Hereinafter, the description will be made on the premise of using the tile ID.
[0409] Next, the three-dimensional data encoding device shifts the position information for each tile (S6104). Also, the three-dimensional data encoding device assigns a tile ID to the position information (S6105). Thereby, data combination is performed. That is, the three-dimensional data encoding device combines the divided tiles or slices. Specifically, for each tile, the position information of the point cloud is shifted to the three-dimensional reference position of the spatial combination. In the example shown in FIG. 82, the three-dimensional data encoding device shifts the points belonging to tile E so that the area of tile E overlaps the area of tile A. For example, the three-dimensional data encoding device shifts the points belonging to tile E so that the difference between the reference position of tile E and the reference position of tile A is 0. Similarly, the three-dimensional data encoding device shifts the points belonging to F, B, and C to the position of tile A, respectively. Note that FIG. 82 shows an example using two-dimensional shift, but three-dimensional shift may also be used.
[0410] Also, the shifted data (combined data) includes a pair of the shifted position information and the tile ID. There may be a plurality of points having the same position information and different tile IDs in the combined data.
[0411] Also, for example, the shift amount is determined for each tile ID. The three-dimensional data encoding device can determine the shift amount based on the tile ID.
[0412] Next, the three-dimensional data encoding device encodes the position information (S6106). At this time, the three-dimensional data encoding device encodes the combined position information and the tile ID together.
[0413] Also, when there is attribute information such as color or reflectance for the position information, the three-dimensional data encoding device converts (S6107) and encodes (S6108) the attribute information. Note that the three-dimensional data encoding device may use the combined position information or the position information before combination for the conversion and encoding of the attribute information. Further, the three-dimensional data encoding device may store information indicating which of the combined and pre-combined position information is used in additional information (metadata).
[0414] Also, in the conversion of the attribute information, for example, when the position of the three-dimensional point changes due to quantization or the like after the encoding of the position information, the three-dimensional data encoding device re-assigns the attribute information of the original three-dimensional point to the changed three-dimensional point.
[0415] FIG. 83 is a flowchart of the encoding process (S6106) of the position information. The three-dimensional data encoding device encodes the combined position information and the tile ID.
[0416] First, the three-dimensional data encoding device generates common information common to a plurality of tiles (S6111). This common information includes tile division information, tile combination information, occupancy codes, etc. Specifically, the three-dimensional data encoding device converts the position information of the points belonging to a plurality of tiles into a tree structure and generates an occupancy code. More specifically, the three-dimensional data encoding device determines for each point to which leaf node in the tree structure the point belongs. Also, the three-dimensional data encoding device sequentially determines from the upper nodes to which node the point to be processed belongs, and encodes the determination result for each node as an occupancy code. The occupancy code is common to a plurality of tiles.
[0417] Next, the three-dimensional data encoding device generates individual information for each tile (S6112). This individual information includes information indicating the number of duplicate points included in the node, and a spatial index (tile ID), etc.
[0418] Specifically, in the node, there may be a mixture of points belonging to different tiles with different tile IDs assigned. Also, in the lowest-level nodes (leaf nodes), there may be a mixture (duplication) of points belonging to multiple tiles. The three-dimensional data encoding device generates information (spatial index) indicating to which tile the duplicate points in the leaf node belong, and encodes this information.
[0419] Here, information related to the entire tile, such as information on the division region in tile division, the number of divisions, and the method of combining divided data, is called common information. In contrast, information for each tile, such as duplicate points in the leaf node, is called individual information.
[0420] In this way, the three-dimensional data encoding device processes a plurality of tiles with a position shift in a batch, generates a common occupancy code, and generates a spatial index.
[0421] Next, the three-dimensional data encoding device combines the common information and the individual information to generate encoded data (S6113). For example, the three-dimensional data encoding device arithmetically encodes the common information and the individual information. Next, the three-dimensional data encoding device generates additional information (metadata) related to spatial combination, and encodes this additional information (S6114).
[0422] Note that the three-dimensional data encoding device may encode the spatial index together with the occupancy code, or encode it as attribute information, or encode it for both.
[0423] In addition, when the resolution of the octree is low, points within the same tile with the same tile ID may be included in one leaf node. In that case, multiple points of the same tile with the same tile ID may overlap, and multiple points of different tiles with different tile IDs may also overlap.
[0424] Note that the three-dimensional data encoding device may add an individual tile ID to a null tile, which is a tile that does not contain point cloud data, or may not add an individual tile ID to the null tile and add the tile IDs of individual tiles to tiles other than the null tile.
[0425] Hereinafter, the preprocessing of tile combination will be described. FIG. 84 is a diagram showing an example of this preprocessing. That is, the method of tile combination is not limited to the method of performing a position shift so that the above-mentioned multiple tiles overlap. For example, the three-dimensional data encoding device may perform preprocessing to convert the position of the tile data as shown in FIG. 84 and then combine them. In the following, mainly, the case where the above-mentioned shift (for example, the shift that aligns the origins of the tiles) is performed after the preprocessing will be described as an example, but the above-mentioned shift may be included in the preprocessing.
[0426] The conversion methods include, for example, circular shift (B in the figure), shift (C in the figure), rotation (E in the figure), inversion (F in the figure), etc. Note that the conversion method may include other conversion methods. For example, the conversion method may include other geometric conversions. For example, the conversion method may include enlargement or reduction. Note that a circular shift is a conversion that shifts the position of the point cloud data and moves the data that has overflowed from the original space in the reverse direction. In addition, the three-dimensional data encoding device may determine the shift amount at the time of combination for each tile. In addition, the three-dimensional data encoding device may use any one of the above-mentioned multiple conversion methods or may use a plurality of them.
[0427] For example, the three-dimensional data encoding device arithmetic-encodes and transmits, as tile-by-tile information, information indicating the type and content of these conversion methods (such as the amount of rotation or shift amount, etc.) and the amount of position shift during combination. In the three-dimensional data decoding device, if this information is known or derivable, the three-dimensional data encoding device may not transmit this information.
[0428] Also, the three-dimensional data encoding device may perform conversion and position shift of the original data so that there are many overlapping points in the point cloud data after combination. Alternatively, the three-dimensional data encoding device may perform conversion and position shift of the original data so that there are few overlapping points. Also, the three-dimensional data encoding device may perform conversion based on sensor information.
[0429] Next, the configuration of the additional information (metadata) will be described. FIG. 85 is a diagram showing an example of the syntax of GPS, which is a parameter set of position information. GPS includes at least one of gps_idx indicating the frame number and sps_idx indicating the sequence number.
[0430] Also, GPS includes GPS information (gps_information), tile information (tile_information), tile combine flag (tile_combine_flag), and tile combine information (tile_combine_information).
[0431] GPS information (gps_information) indicates the coordinates and size of the bounding box in the PCC frame, as well as quantization parameters, etc. Tile information (tile_information) indicates information related to tile division. The tile combine flag (tile_combine_flag) is a flag indicating whether to combine data. Tile combine information (tile_combine_information) indicates information related to data combination.
[0432] FIG. 86 is a diagram showing a syntax example of tile combine information. The tile combine information includes combine type information (type_of_combine), an origin shift flag (shift_equal_origin_flag), and shift amounts (shift_x, shift_y, shift_z).
[0433] The combine type information (type_of_combine) indicates the type of data combination. For example, type_of_combine = A indicates combining split tiles without deforming them.
[0434] The origin shift flag (shift_equal_origin_flag) is a flag indicating whether to perform an origin shift that shifts the origin of the tile to the origin of the target space. In the case of an origin shift, the shift amounts can be omitted because they are the same values as the origin included in the tile information (tile_information).
[0435] When the shift amounts for combination are not known, such as when no origin shift is performed, the tile combine information includes the shift amounts (shift_x, shift_y, shift_z) for each tile. Also, the shift amounts are not described when the tile is a null tile.
[0436] In addition, the tile combine information includes information() indicating information on other combination methods. For example, this information is the information on the preprocessing when preprocessing is performed before combination. For example, this information includes the type of preprocessing used (circular shift, shift, rotation, inversion, etc., enlargement, reduction, etc.) and the content thereof (shift amount, rotation amount, etc., enlargement ratio, reduction ratio, etc.).
[0437] FIG. 87 is a diagram showing an example of the syntax of tile information. The tile information includes divide type information (type_of_divide), number of tiles information (number_of_tiles), tile null flag (tile_null_flag), origin information (origin_x, origin_y, origin_z), and node size (node_size).
[0438] The divide type information (type_of_divide) indicates the type of tile division. The number of tiles information (number_of_tiles) indicates the number of divisions (the number of tiles). The tile null flag (tile_null_flag) is information about a null tile, indicating whether the tile is a null tile. The origin information (origin_x, origin_y, origin_z) is information indicating the position (coordinates) of the tile, for example, indicating the coordinates of the origin of the tile. The node size (node_size) indicates the size of the tile.
[0439] Also, the tile information may include information indicating quantization parameters and the like. Also, the information on the coordinates and size of the tile may be omitted if it is known in the three-dimensional data decoding device or can be derived by the three-dimensional data decoding device.
[0440] Note that the tile information (tile_information) and the tile combine information (tile_combine_information) may be merged. Also, the divide type information (type_of_devide) and the combine type information (type_of_combine) may be merged.
[0441] Next, the configuration of the encoded position information will be described. FIG. 88 is a diagram showing the configuration of the encoded position information. The encoded position information includes a header and a payload.
[0442] FIG. 89 is a diagram showing an example of the syntax of the header (Geometry_header) of the encoded position information. The header of the encoded position information includes a GPS index (gps_idx), offset information (offset), position information (geometry_information), a combine flag (combine_flag), the number of combined tiles (number_of_combine_tile), and a combine index (combine_index).
[0443] The GPS index (gps_idx) indicates the ID of the parameter set (GPS) corresponding to the encoded attribute information. Note that when there is a parameter set for each tile, a plurality of IDs may be indicated in the header. The offset information (offset) indicates the offset position (address) for acquiring data. The position information (geometry_information) indicates information such as the bounding box of the PCC frame (for example, the coordinates and size of the tile).
[0444] The combine flag (combine_flag) is a flag indicating whether the encoded data is spatially combined. The number of combined tiles (number_of_combine_tile) indicates the number of tiles to be combined. Note that this information may be described in the SPS or GPS. The combine index (combine_index) indicates the index of the combined data when the data is combined. When the number of combined data in the frame is 1, the combine index may be omitted.
[0445] FIG. 90 is a diagram showing an example of the syntax of the payload (Geometry_data) of the encoded position information. The payload of the encoded position information includes an occupancy code (occupancy_Code(depth, i)) and combine information (combine_information).
[0446] The occupancy code (occupancy_Code(depth, i)) is data in which one or more frames are combined and is common information.
[0447] The combined information is the information of the leaf node. The information of the leaf node may be described as a loop of the number of tiles. That is, the information of the leaf node may be shown for each tile.
[0448] Also, as a method for indicating the spatial index of a plurality of points included in the leaf node, for example, there are a method (Method 1) of respectively indicating the number of duplicate points included in the node and the value of the spatial index for each point, and a method (Method 2) using a bitmap.
[0449] FIG. 91 is a diagram showing an example of the syntax of the combined information in Method 1. The combined information includes the number of duplicate points (NumberOfPoints) indicating the number of duplicate points and the tile ID (TileID(i)).
[0450] Note that when the quantization resolution is low, there may be duplicate points in the same frame. In this case, the number of duplicate points (NumberOfPoints) may be shared, and the number of duplicate points included in the same frame and the total number of duplicate points included in a plurality of frames may be described.
[0451] FIG. 92 is a diagram showing an example of the syntax of the combined information in Method 2. The combined information includes bitmap information (bitmap). FIG. 93 is a diagram showing an example of the bitmap information. In this example, it is shown by the bitmap that the leaf node includes three-dimensional points of tile numbers 1, 3, and 5.
[0452] Note that when the three-dimensional data encoding device performs irreversible compression, it may delete duplicate points to reduce the amount of information. Also, the three-dimensional data encoding device may delete duplicate points before data combination or after data combination.
[0453] Next, other transmission methods of the spatial index will be described. The three-dimensional data encoding device may convert the spatial index into RANK information and transmit it. Rank conversion is a method of converting a bitmap into a rank (Ranking) and the number of bits (RANKbit), and transmitting the converted rank and the number of bits.
[0454] FIG. 94 is a block diagram of the tile ID encoding unit 6121. This tile ID encoding unit 6121 includes a bitmap generation unit 6131, a bitmap inversion unit 6132, a lookup table reference unit 6133, and a bit number acquisition unit 6134.
[0455] The bitmap generation unit 6131 generates a bitmap based on the combined tile number, the number of overlapping points, and the tile ID for each three-dimensional point. For example, the bitmap generation unit 6131 generates a bitmap indicating to which of the plurality of three-dimensional data the three-dimensional points included in the leaf node (unit space) in the combined three-dimensional data obtained by combining a plurality of three-dimensional data belong. The bitmap is digital data represented by 0 and 1. The bitmap, the conversion code, and the number of three-dimensional points included in the unit space are in one-to-one correspondence.
[0456] The bitmap inversion unit 6132 inverts the bitmap generated by the bitmap generation unit 6131 based on the number of overlapping points.
[0457] Specifically, for example, the bitmap inversion unit 6132 determines whether the number of overlapping points exceeds half of the combined tile number, and if it exceeds half, it inverts each bit of the bitmap (in other words, exchanges 0 and 1). Also, for example, if the number of overlapping points does not exceed half of the combined tile number, the bitmap inversion unit 6132 does not invert the bitmap.
[0458] In this way, the bitmap inversion unit 6132 determines whether the number of overlapping points exceeds a predetermined threshold value, and if it exceeds the predetermined threshold value, it inverts each bit of the bitmap. Also, for example, if the number of overlapping points does not exceed the predetermined threshold value, the bitmap inversion unit 6132 does not invert the bitmap. For example, the predetermined threshold value is half of the number of combined tiles.
[0459] The lookup table reference unit 6133 converts the bitmap (inverted bitmap) after the bitmap inversion unit 6132 inverts it, or the bitmap whose number of overlapping points does not exceed half of the number of combined tiles, into a rank (also referred to as ranking) using a predetermined lookup table. That is, the lookup table reference unit 6133, for example, uses the lookup table to convert the bitmap generated by the bitmap generation unit 6131 or the inverted bitmap generated by the bitmap inversion unit 6132 into a rank (in other words, generates a rank). The lookup table reference unit 6133 has, for example, a memory that stores the lookup table.
[0460] The rank is a numerical value that classifies the bitmap according to the number of 1s included in the bitmap (that is, for each number indicated by N), and indicates the tile ID or order within the classified group. For example, N is also the number of overlapping points.
[0461] The bit number acquisition unit 6134 acquires the number of bits of the rank from the number of combined tiles and the number of overlapping points. Note that the required number of bits of the rank determined from the maximum number of ranks (the number of bits required to represent the rank in binary) is a number uniquely determined from the number of combined tiles and the number of overlapping points. The three-dimensional data encoding device generates information (rank information) indicating the number of bits of the rank by the bit number acquisition unit 6134, and calculates and encodes it at a stage subsequent to the bit number acquisition unit 6134 (not shown).
[0462] Note that when the value of the rank is 0, the bit number acquisition unit 6134 may not send the rank information. Also, when the rank is not 0, the bit number acquisition unit 6134 may encode the value of rank - 1. Alternatively, when the rank is not 0, the bit number acquisition unit 6134 may output the value of rank - 1 as the rank information.
[0463] FIG. 95 is a block diagram of the tile ID acquisition unit 6162. This tile ID acquisition unit 6162 includes a bit number acquisition unit 6141, a rank acquisition unit 6142, a look-up table reference unit 6143, and a bitmap inversion unit 6144.
[0464] The tile ID acquisition unit 6140 decodes the number of combined tiles from the additional information (metadata) included in the encoded data acquired from the three-dimensional data encoding device, and further extracts the number of duplicate points for each leaf node from the encoded data.
[0465] The bit number acquisition unit 6141 acquires the number of bits of the rank from the number of combined tiles and the number of duplicate points.
[0466] Note that the maximum number of ranks and the number of bits of the rank (required number of bits) are numbers uniquely determined from the number of combined tiles and the number of duplicate points. Also, the process by which the three-dimensional data decoding device acquires the rank is the same as the process in the above-described three-dimensional data encoding device.
[0467] The rank acquisition unit 6142 acquires the rank for the number of bits acquired above from the decoded encoded data.
[0468] The look-up table reference unit 6143 acquires a bitmap using a predetermined look-up table from the number of duplicate points of the leaf node and the rank acquired by the rank acquisition unit 6142. The look-up table reference unit 6143 includes, for example, a memory that stores the look-up table.
[0469] Note that the look-up table stored in the look-up table reference unit 6143 is a table corresponding to the look-up table used in the three-dimensional data encoding device. That is, it is a table for the look-up table reference unit 6143 to output an inverted bitmap in which the number of 1s does not exceed half based on the rank. For example, the look-up table used by the look-up table reference unit 6143 included in the three-dimensional data decoding device is the same as the look-up table included in the look-up table reference unit 6133 included in the three-dimensional data encoding device.
[0470] Next, the bitmap inversion unit 6144 inverts the bitmap output by the look-up table reference unit 6143 based on the number of overlapping points.
[0471] Specifically, the bitmap inversion unit 6144 determines whether the number of overlapping points exceeds half of the number of combined tiles. If it exceeds half, it inverts each bit of the bitmap. On the other hand, if it does not exceed half, the bitmap inversion unit 6144 does not invert the bitmap. That is, an inverted bitmap in which the number of 1s does not exceed half is converted into a bitmap in which the number of 1s exceeds half.
[0472] Thereby, the three-dimensional data decoding device can obtain the tile ID of the overlapping points from the bitmap.
[0473] Also, for example, the tile ID acquisition unit 6140 outputs three-dimensional data in which three-dimensional points are associated with tile IDs from the combined data (for example, combined three-dimensional data) included in the acquired encoded data and the bitmap.
[0474] Also, here, an example in which the method of inverting the bitmap when the number of overlapping points exceeds half of the number of combined tiles is used is shown, but the method of not inverting the bitmap may also be used.
[0475] Next, the configuration of the encoded attribute information will be described. FIG. 96 is a diagram showing the configuration of the encoded attribute information. The encoded attribute information includes a header and a payload.
[0476] FIG. 97 is a diagram showing a syntax example of the header (Attribute_header) of the encoded attribute information. The header (Attribute_header) of the encoded attribute information includes an APS index (aps_idx), offset information (offset), other information (other_attribute_information), the number of combined tiles (number_of_combine_tile), and reference information (refer_different_tile).
[0477] The APS index (aps_idx) indicates the ID of the parameter set (APS) corresponding to the encoded attribute information. Note that when there is a parameter set for each tile, a plurality of IDs may be indicated in the header. The offset information (offset) indicates the offset position (address) for obtaining the combined data. The other information (other_attribute_information) indicates other attribute data, for example, the difference value of the quantization parameter (QPdelta), etc. The number of combined tiles (number_of_combine_tile) indicates the number of tiles to be combined. Note that these pieces of information may be described in the SPS or APS.
[0478] The reference information (refer_different_tile) is a flag indicating whether to encode or decode the attribute information of the target point, which is a three-dimensional point to be encoded or decoded, using the attribute information of the three-dimensional points in the same tile, or the attribute information of the three-dimensional points belonging to the same tile and other than the same tile. For example, the following value assignments are conceivable. refer_different_tile = 0 indicates that the attribute information of the target point is encoded or decoded using the attribute information of the three-dimensional points within the same tile as the target point. refer_different_tile = 1 indicates that the attribute information of the target point is encoded or decoded using the attribute information of the three-dimensional points within the same tile and other than the same tile as the target point.
[0479] FIG. 98 is a diagram showing an example of the syntax of the payload (Attribute_data) of the encoded attribute information. The payload of the encoded attribute information includes combine_information. For example, this combine_information is the same as the combine_information shown in FIG. 90.
[0480] Next, an example of the transmission order and decoding order of the encoded data will be described. FIG. 99 is a diagram showing the configuration of the encoded data. The encoded data includes a header and a payload.
[0481] FIGS. 100 to 102 are diagrams showing the transmission order of the data and the reference relationship of the data. In the figure, G(1) etc. indicate the encoded position information, GPS(1) etc. indicate the parameter set of the encoded position information, and SPS indicates the parameter set of the sequence (multiple frames). Also, the numbers in () indicate the values of the tile IDs. Note that the three-dimensional data encoding device may transmit the data in the decoding order.
[0482] FIG. 100 is a diagram showing an example of the transmission order when no spatial combination is performed. FIG. 101 is a diagram showing an example when spatial combination is performed and metadata (parameter set) is added for each space (tile). FIG. 102 is a diagram showing an example when spatial combination is performed and metadata (parameter set) is added in the unit of combination.
[0483] In the header of the spatially combined data, an identifier of the reference metadata for obtaining the metadata of the tile is stored. As shown in FIG. 102, the metadata of a plurality of tiles may be grouped together. The common parameters for the plurality of tiles combined by tiles may be grouped into one. The parameters not common to the plurality of tiles may indicate values for each tile.
[0484] The information for each tile (parameters not common to multiple tiles) includes, for example, the position information of the tile, the relative position with respect to other tiles, the encoding time, the shape and size of the tile, information indicating whether the tile overlaps, the quantization parameter for each tile, and the information of preprocessing, etc. The information of preprocessing indicates the type and content of preprocessing, and includes, for example, information indicating that rotation has been performed and information indicating the rotation amount when rotation is performed after tile division.
[0485] FIG. 103 is a diagram showing an example of decoding some frames in the example shown in FIG. 101. As shown in FIG. 103, in the spatially combined data, if there is no dependency between tiles, the three-dimensional data decoding device can extract and decode some data. That is, the three-dimensional data decoding device can decode each data independently.
[0486] Note that when there is a dependency between tiles, the three-dimensional data encoding device sends the reference destination data first. The three-dimensional data decoding device decodes the reference destination data first. Note that the three-dimensional data encoding device does not consider the sending order, and the three-dimensional data decoding device may rearrange the data.
[0487] When the point cloud data has attribute information, the three-dimensional data encoding device may spatially combine the attribute information. The attribute information is encoded or decoded with reference to the position information. The position information to be referred to may be the position information before spatial combination or the position information after spatial combination.
[0488] FIGS. 104 to 106 are diagrams showing the sending order of data and the reference relationship of data. In the figure, A(1) etc. indicate encoded attribute information, APS(1) etc. indicate parameter sets of the encoded attribute information, and SPS indicates a parameter set of a sequence (multiple frames). Also, the numbers in () indicate the values of tile IDs.
[0489] FIG. 104 is a diagram showing the data transmission order and the data reference relationship when refer_different_tile = 1. In this example, A(1-4) may refer to each other. Also, A(1-4) is encoded or decoded using the information of G(1-4).
[0490] Further, the three-dimensional data decoding device divides G(1-4) and A(1-4) into tiles 1-4 using the tile IDs 1-4 decoded together with G(1-4).
[0491] FIG. 105 is a diagram showing the data transmission order and the data reference relationship when refer_different_tile = 1. In this example, A(1-4) do not refer to each other. For example, A(1) refers to A(1) but does not refer to A(2-4). Also, A(1-4) is encoded or decoded using the information of G(1-4).
[0492] The three-dimensional data decoding device divides G(1-4) and A(1-4) into tiles 1-4 using the tile IDs 1-4 decoded together with G(1-4).
[0493] FIG. 106 is a diagram showing another example of the data transmission order and the data reference relationship when refer_different_tile = 0. In this example, APS is added to the header for each of A(1-4). Also, the number of combined tiles for the position information and the number of combined tiles for the attribute information may be the same, or may be independent (different).
[0494] For example, as shown in FIG. 106, the position information may be tile-combined and the attribute information may not be tile-combined.
[0495] The configuration of the three-dimensional data decoding apparatus according to the present embodiment will be described below. FIG. 107 is a block diagram showing the configuration of a decoding unit 6150 included in the three-dimensional data decoding apparatus according to the present embodiment. The decoding unit 6150 restores point cloud data by decoding encoded data (encoded stream) generated by encoding the point cloud data. The decoding unit 6150 includes a de-multiplexing unit 6151, a position information decoding unit 6152, an attribute information decoding unit 6153, an additional information decoding unit 6154, and a splitting / merging unit 6155.
[0496] The de-multiplexing unit 6151 generates encoded position information, encoded attribute information, and encoded additional information by de-multiplexing the encoded data (encoded stream).
[0497] The position information decoding unit 6152 generates combined position information with a tile ID added by decoding the encoded position information. The attribute information decoding unit 6153 generates combined attribute information by decoding the encoded attribute information. The additional information decoding unit 6154 generates combined additional information and split additional information by decoding the encoded additional information.
[0498] The splitting / merging unit 6155 generates position information by splitting and merging the combined position information using the combined additional information and the split additional information. The splitting / merging unit 6155 generates attribute information by splitting and merging the combined attribute information using the combined additional information and the split additional information.
[0499] FIG. 108 is a block diagram of the position information decoding unit 6152 and the attribute information decoding unit 6153. The position information decoding unit 6152 includes an entropy decoding unit 6161 and a tile ID acquisition unit 6162. The entropy decoding unit 6161 generates combined position information by entropy decoding the encoded position information. The tile ID acquisition unit 6162 acquires the tile ID from the combined position information when the tile ID is added to the position information.
[0500] The attribute information decoding unit 6153 includes an entropy decoding unit 6163 and a tile ID acquisition unit 6164. The entropy decoding unit 6163 generates combined attribute information by entropy decoding the encoded attribute information. The tile ID acquisition unit 6164 acquires the tile ID from the combined attribute information when the tile ID is added to the attribute information.
[0501] FIG. 109 is a diagram showing the configuration of the split-join unit 6155. The split-join unit 6155 includes a split unit 6171 and a join unit 6172. The split unit 6171 generates a plurality of split position information (tile position information) by splitting the join position information using the tile ID and the join additional information. The split unit 6171 generates a plurality of split attribute information (tile attribute information) by splitting the combined attribute information using the tile ID and the join additional information.
[0502] The join unit 6172 generates position information by joining the plurality of split position information using the split additional information. The join unit 6172 generates attribute information by joining the plurality of split attribute information using the split additional information.
[0503] Next, the three-dimensional data decoding process will be described. FIG. 110 is a flowchart of the three-dimensional data decoding process according to the present embodiment.
[0504] First, the three-dimensional data decoding device acquires the split additional information, which is metadata related to the splitting method, and the join additional information, which is metadata related to the joining method, by decoding them from the bit stream (S6121). Next, the three-dimensional data decoding device acquires the decoded data (decoded position information) and the tile ID by decoding the position information from the bit stream (S6122).
[0505] Next, the three-dimensional data decoding device determines whether refer_different_tile included in the APS is 1 (S6123). If refer_different_tile is 1 (Yes in S6123), the three-dimensional data decoding device decodes the attribute information of the target point based on the decoded position information and the attribute information of the three-dimensional points included in the same tile as well as tiles other than the tile containing the target point, which is the three-dimensional point to be processed (S6124). On the other hand, if refer_different_tile is not 1 (0) (No in S6123), the three-dimensional data decoding device decodes the attribute information of the target point based on the decoded position information and the attribute information of the three-dimensional points within the same tile as the tile containing the target point (S6125). That is, the three-dimensional data decoding device does not refer to the attribute information of the three-dimensional points included in tiles different from the tile containing the target point.
[0506] Through the above decoding process, the three-dimensional data decoding device decodes the tile ID, position information, and attribute information.
[0507] Also, when the tile ID is encoded as the position information, the three-dimensional data decoding device acquires the tile ID during the decoding of the position information. Also, when there is attribute information, the three-dimensional data decoding device decodes the attribute information of the target point based on the decoded attribute information. On the other hand, when there is no attribute information, the three-dimensional data decoding device does not decode the attribute information. Note that the three-dimensional data decoding device does not have to decode the attribute information even when there is attribute information if decoding is not required.
[0508] Also, when refer_different_tile = 1 included in the APS, the three-dimensional data decoding device decodes the attribute information based on the tile-combined attribute information. That is, the three-dimensional data decoding device decodes using the attribute information regardless of the tile ID.
[0509] On the one hand, when refer_different_tile = 0 included in APS, the three-dimensional data decoding device decodes the attribute information based on the non-tile-combined attribute information. That is, the three-dimensional data decoding device filters the three-dimensional points using the tile ID, and decodes the target point using the attribute information of the three-dimensional points with the same tile ID as the target point. That is, the three-dimensional data decoding device does not use the attribute information of the three-dimensional points with a tile ID different from the target point.
[0510] Next, the three-dimensional data decoding device separates the combined data. Specifically, the three-dimensional data decoding device divides the position information for each tile based on the tile ID (S6126).
[0511] Next, the three-dimensional data decoding device combines the tiles or slices. Specifically, the three-dimensional data decoding device inversely shifts the position information for each tile (S6127), and combines the processed position information (S6128).
[0512] Hereinafter, a method of transmitting the spatial index as attribute information will be described. When multiple point cloud data are spatially combined, the tile ID (spatial index) added to each three-dimensional point may be encoded using the predictive coding method described in the present disclosure as new attribute information instead of as the position information of each three-dimensional point. Specific examples of this method are shown below.
[0513] FIG. 111 and FIG. 112 are diagrams showing the data transmission order and the data reference relationship in this case. Further, FIG. 111 shows an example when refer_different_tile = 1, and FIG. 112 shows an example when refer_different_tile = 0.
[0514] For example, when the point cloud data has position information and first attribute information (e.g., color), when applying spatial combination, as shown in FIG. 111, the three-dimensional data encoding device encodes the spatial index (e.g., tile ID) as the second attribute information (A2 in FIG. 111), and stores in the SPS identification information indicating that the second attribute information includes the spatial index.
[0515] In addition, when transmitting the spatial index as attribute information, it may be determined that a reversible encoding method must be used. For example, the reversible encoding method is an encoding method that does not perform quantization. Also, when the spatial index is attribute information, constraints may be imposed on the quantization parameters so that reversible encoding is achieved. For example, when the attribute information is color or reflectance, etc., the parameters for reversible encoding and the parameters for irreversible encoding are included in the applicable range. When the attribute information is a frame index or a spatial index, the parameters for reversible encoding are included in the applicable range, and the parameters for irreversible encoding may not be included in the applicable range.
[0516] In addition, the three-dimensional data encoding device may transmit information for each space (tile) as second attribute information in addition to the spatial index. For example, the information for each tile may indicate at least one of the tile position, the relative position between the target tile and other tiles, the encoding time, the tile shape, the tile size, information indicating whether the tiles overlap, the quantization parameters for each tile, and the preprocessing information. The preprocessing information indicates the type and content of the preprocessing. For example, when rotation is performed after tile division, it includes information indicating that rotation has been performed and information indicating the amount of rotation. Also, the three-dimensional data encoding device may store the information of a plurality of tiles together.
[0517] In the example of refer_different_tile = 1 shown in FIG. 111, the additional information (metadata) of the second attribute information indicates that the type of the attribute information is a spatial index or a tile ID. Note that the additional information may indicate the type of the spatial index. For example, this additional information is included in the SPS.
[0518] Also, A(1-4) may refer to each other. For example, the first attribute information indicates color information. Also, A(1-4) is encoded or decoded using the information of G(1-4). The three-dimensional data decoding device divides G(1-4) and A(1-4) into tiles 1-4 using the tile IDs 1-4 decoded together with G(1-4).
[0519] The second attribute information indicates the tile IDs related to tiles 1 to 4. If there is no point belonging to the tile ID, that is, if it is a null tile, the tile ID may not be indicated.
[0520] In the example of refer_different_tile = 0 shown in FIG. 112, the three-dimensional data decoding device determines the attribute information to be referred to based on the spatial index. Therefore, the three-dimensional data encoding device sends out the spatial index before the attribute information such as color or reflectance. In FIG. 112, APS2 and A2 are sent out before APS1 and A1. The three-dimensional data decoding device decodes APS2 and A2 before APS1 and A1.
[0521] Note that the transmission order may be changed according to the value of refer_different_tile. Also, the three-dimensional data decoding device may rearrange the data if the data is not in order.
[0522] For example, A(1 - 4) do not refer to each other. For example, A(1) refers to A(1). Also, A(1 - 4) are encoded or decoded using the information of G(1 - 4).
[0523] Hereinafter, an example in the case where the point cloud data has color attribute information and spatial index attribute information will be described. FIG. 113 is a diagram showing the data configuration in this case.
[0524] When performing spatial combination, the three-dimensional data encoding device generates attribute information and sends out the attribute information. Note that when the number of spatial combinations is variable, if the number of spatial combinations is 1, the point cloud data may not have a spatial index. The three-dimensional data decoding device may determine that there is no attribute information corresponding to the spatial index when the number of spatial combinations is 1.
[0525] The three-dimensional data encoding device may set the IDs of the parameter sets of GPS, APS1, and APS2 to the same value. Also, the three-dimensional data encoding device may set the value of this ID for each frame and set it to consecutive values in a plurality of frames. When there is no attribute information including a spatial index, the ID of the parameter set may not be consecutive, and the ID of the frame without attribute information including a spatial index may be skipped.
[0526] When the combination number is 2 or more, there is attribute information of the spatial index corresponding to the position information. When the combination number is 1, there is no attribute information of the spatial index corresponding to the position information.
[0527] Hereinafter, other examples will be described. For example, in the above, tile division was described as an example of spatial division, but the present method is not limited to this, and the above method may be applied to other methods of dividing space or methods of clustering point groups.
[0528] Also, in the above, an example of encoding the tile ID as the spatial index was mainly described, but instead of the tile ID, information indicating the shift amount of the space at the time of combination may be used.
[0529] Also, in the above, an example of storing the spatial index as bitmap or rank information in the position information or the attribute information was shown. Any one of these multiple methods may be used, or a method combining multiple methods may be used. For example, the spatial index may be included in both the position information and the attribute information. Also, the formats (for example, bitmap or rank information) of the spatial indexes included in the position information and the attribute information may be the same or different.
[0530] Also, at least one of the method of spatial division, the number of divisions, the method of combination, and the combination number may change over time. For example, the three-dimensional data encoding device may change the number of divisions or the combination number for each space, or may change it adaptively. Also, the number of divisions and the combination number may or may not match.
[0531] Also, in the above, an example of combining split data and sending out the combined data was shown, but the following combinations may also be used. FIG. 114 is a diagram showing an example of a combination of splitting and combining. For example, as in frame 3 shown in FIG. 114, a part of the split data may be combined. In that case, the three-dimensional data encoding device sends out the split data and the combined data. At this time, the three-dimensional data encoding device generates an identifier for identifying the split data and the combined data, and sends out the identifier.
[0532] Also, as in frame 1 shown in FIG. 114, the three-dimensional data encoding device may split the data again after combining the split data. In this case, the three-dimensional data encoding device generates information regarding the splitting and sends out the information.
[0533] Also, as shown in FIG. 114, a plurality of combinations may be switched. In this case, there may be a frame in which splitting and combining are not used, such as frame 2.
[0534] FIG. 115 is a diagram showing an example of combining frame combining and spatial combining. As shown in the figure, frame combining and spatial combining may be combined.
[0535] As described above, the three-dimensional data encoding device according to the present embodiment performs the processes shown in FIG. 116. The three-dimensional data encoding device performs a conversion process including movement on the second point cloud data among the first point cloud data and the second point cloud data at the same time, and generates third point cloud data by combining the first point cloud data and the second point cloud data after the conversion process (S6131). Next, the three-dimensional data encoding device generates a bit stream by encoding the third point cloud data (S6132). The bit stream includes first information (for example, tile ID or spatial index) indicating to which of the first point cloud data and the second point cloud data each of a plurality of three-dimensional points included in the third point cloud data belongs, and second information (for example, origin shift flag (shift_equal_origin_flag), or shift amount (shift_x, shift_y, shift_z)) indicating the content of the movement. According to this, the encoding efficiency can be improved by collectively encoding a plurality of point cloud data at the same time.
[0536] For example, the conversion process includes a geometric transformation in addition to the movement, and the bit stream includes third information (for example, information() shown in FIG. 86, etc.) indicating the content of the geometric transformation. According to this, the encoding efficiency can be improved by performing the geometric transformation.
[0537] For example, the geometric transformation includes at least one of a shift, a rotation, and a reversal. For example, the three-dimensional data encoding device determines the content of the geometric transformation based on the number of overlapping points that are three-dimensional points with the same position information in the third point cloud data.
[0538] For example, the second information includes information indicating the amount of movement of the movement (for example, shift amount (shift_x, shift_y, shift_z)).
[0539] For example, the second information includes information (for example, including origin shift flag (shift_equal_origin_flag)) indicating whether to move the origin of the second point cloud data to the origin of the first point cloud data.
[0540] For example, the three-dimensional data encoding device generates the first point group data and the second point group data by spatially dividing the fourth point group data.
[0541] For example, the bit stream includes the position information of each of the plurality of three-dimensional points included in the third point group data and one or more pieces of attribute information, and one of the one or more pieces of attribute information includes the first information.
[0542] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor performs the above processing using the memory.
[0543] In addition, the three-dimensional data decoding device according to the present embodiment performs the processing shown in FIG. 117. The three-dimensional data decoding device performs a conversion process including movement on the second point group data among the first point group data and the second point group data at the same time, and combines the first point group data and the second point group data after the conversion process. The three-dimensional data decoding device decodes the third point group data from the bit stream generated by encoding the generated third point group data (S6141). Further, the three-dimensional data decoding device obtains, from the bit stream, first information (for example, a tile ID or a spatial index) indicating to which of the first point group data and the second point group data each of the plurality of three-dimensional points included in the third point group data belongs, and second information indicating the content of the movement (for example, a shift_equal_origin_flag or shift amounts (shift_x, shift_y, shift_z)) (S6142). Next, the three-dimensional data decoding device restores the first point group data and the second point group data from the decoded third point group data using the first information and the second information (S6143). For example, the three-dimensional data decoding device separates the first point group data and the second point group data after the conversion process from the third point group data using the first information. The three-dimensional data decoding device generates the second point group data by performing an inverse conversion on the second point group data after the conversion process using the second information. According to this, the encoding efficiency can be improved by encoding a plurality of point group data at the same time collectively.
[0544] For example, the conversion process includes geometric transformation in addition to translation. The three-dimensional data decoding device acquires third information (e.g., information() shown in FIG. 86, etc.) indicating the content of the geometric transformation from the bit stream. In restoring the first point cloud data and the second point cloud data, the three-dimensional data decoding device restores the first point cloud data and the second point cloud data from the decoded third point cloud data using the first information, the second information, and the third information. According to this, the encoding efficiency can be improved by performing geometric transformation.
[0545] For example, the geometric transformation includes at least one of shift, rotation, and inversion. For example, the second information includes information indicating the amount of translation (e.g., shift amounts (shift_x, shift_y, shift_z)). For example, the second information includes information indicating whether to move the origin of the second point cloud data to the origin of the first point cloud data (e.g., origin shift flag (shift_equal_origin_flag).
[0546] For example, the three-dimensional data decoding device generates fourth point cloud data by spatially combining the restored first point cloud data and second point cloud data.
[0547] For example, the bit stream includes position information of each of a plurality of three-dimensional points included in the third point cloud data and one or more attribute information, and one of the one or more attribute information includes the first information.
[0548] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor performs the above processing using the memory.
[0549] As described above, the three-dimensional data encoding device and the three-dimensional data decoding device according to the embodiments of the present disclosure have been described, but the present disclosure is not limited to these embodiments.
[0550] Further, each processing unit included in the three-dimensional data encoding device, the three-dimensional data decoding device, and the like according to the above embodiment is typically realized as an LSI which is an integrated circuit. These may be individually formed into one chip, or may be formed into one chip so as to include some or all of them.
[0551] Further, the integration into an integrated circuit is not limited to an LSI, and it may be realized by a dedicated circuit or a general-purpose processor. An FPGA (Field Programmable Gate Array) that can be programmed after manufacturing the LSI, or a reconfigurable processor that can reconfigure the connection and setting of circuit cells inside the LSI may be used.
[0552] Further, in each of the above embodiments, each component may be configured by dedicated hardware, or may be realized by executing a software program suitable for each component. Each component may be realized by a program execution unit such as a CPU or a processor reading and executing a software program recorded on a recording medium such as a hard disk or a semiconductor memory.
[0553] Further, the present disclosure may be realized as a three-dimensional data encoding method, a three-dimensional data decoding method, or the like executed by a three-dimensional data encoding device, a three-dimensional data decoding device, and the like.
[0554] Also, the division of functional blocks in the block diagram is an example, and a plurality of functional blocks may be realized as one functional block, one functional block may be divided into a plurality, or some functions may be transferred to other functional blocks. Further, the functions of a plurality of functional blocks having similar functions may be processed by a single piece of hardware or software in parallel or time-divisionally.
[0555] Also, the order in which each step in the flowchart is executed is for exemplification for specifically explaining the present disclosure, and may be an order other than the above. Further, some of the above steps may be executed simultaneously (in parallel) with other steps.
[0556] The three-dimensional data encoding device, the three-dimensional data decoding device, etc. according to one or more aspects have been described based on the embodiments. However, the present disclosure is not limited to these embodiments. Without departing from the spirit of the present disclosure, various modifications conceived by those skilled in the art applied to these embodiments, or forms constructed by combining components in different embodiments may also be included within the scope of one or more aspects.
Industrial Applicability
[0557] The present disclosure can be applied to a three-dimensional data encoding device and a three-dimensional data decoding device.
Explanation of Signs
[0558] 4601 Three-dimensional data encoding system 4602 Three-dimensional data decoding system 4603 Sensor terminal 4604 External connection part 4611 Point cloud data generation system 4612 Presentation part 4613 Encoding part 4614 Multiplexing part 4615 Input / output part 4616 Control part 4617 Sensor information acquisition part 4618 Point cloud data generation part 4621 Sensor information acquisition part 4622 Input / output part 4623 Demultiplexing part 4624 Decoding part 4625 Presentation part 4626 User interface 4627 Control part 4630 First encoding part 4631 Position information encoding part 4632 Attribute information encoding part 4633 Additional information encoding part 4634 Multiplexing part 4640 First decoding part 4641 Inverse Multiplexing Unit 4642 Location Information Decoding Unit 4643 Attribute Information Decoding Unit 4644 Additional Information Decoding Unit 4650 Second Encoding Unit 4651 Additional Information Generation Unit 4652 Location Image Generation Unit 4653 Attribute Image Generation Unit 4654 Video Encoding Unit 4655 Additional Information Encoding Unit 4656 Multiplexing Unit 4660 Second Decoding Unit 4661 Inverse Multiplexing Unit 4662 Video Decoding Unit 4663 Additional Information Decoding Unit 4664 Location Information Generation Unit 4665 Attribute Information Generation Unit 4801 Encoding Unit 4802 Multiplexing Unit 4910 First Encoding Unit 4911 Splitting Unit 4912 Location Information Encoding Unit 4913 Attribute Information Encoding Unit 4914 Additional Information Encoding Unit 4915 Multiplexing Unit 4920 First Decoding Unit 4921 Inverse Multiplexing Unit 4922 Location Information Decoding Unit 4923 Attribute Information Decoding Unit 4924 Additional Information Decoding Unit 4925 Combining Unit 4931 Slice Splitting Unit 4932 Location Information Tile Splitting Unit 4933 Attribute Information Tile Splitting Unit 4941 Location Information Tile Combining Unit 4942 Attribute Information Tile Combining Unit 4943 Slice Combining Unit 5410 Encoding Unit 5411 Splitting Unit 5412 Location Information Encoding Unit 5413 Attribute Information Encoding Unit 5414 Additional Information Encoding Unit 5415 Multiplexing Unit 5421 Tile Division Unit 5422 Slice Division Unit 5431, 5441 Frame Index Generation Unit 5432, 5442 Entropy Encoding Unit 5450 Decoding Unit 5451 Demultiplexing Unit 5452 Position Information Decoding Unit 5453 Attribute Information Decoding Unit 5454 Additional Information Decoding Unit 5455 Combining Unit 5461, 5471 Entropy Decoding Unit 5462, 5472 Frame Index Acquisition Unit 6100 Encoding Unit 6101 Splitting and Combining Unit 6102 Position Information Encoding Unit 6103 Attribute Information Encoding Unit 6104 Additional Information Encoding Unit 6105 Multiplexing Unit 6111 Splitting Unit 6112 Combining Unit 6121, 6123 Tile ID Encoding Unit 6122, 6124 Entropy Encoding Unit 6131 Bitmap Generation Unit 6132, 6144 Bitmap Inversion Unit 6133, 6143 Lookup Table Reference Unit 6134, 6141 Bit Number Acquisition Unit 6142 Rank Acquisition Unit 6150 Decoding Unit 6151 Demultiplexing Unit 6152 Position Information Decoding Unit 6153 Attribute Information Decoding Unit 6154 Additional Information Decoding Unit 6155 Splitting and Combining Unit 6161, 6163 Entropy Decoding Unit 6162, 6164 Tile ID Acquisition Unit 6171 Splitting Unit 6172 Junction
Claims
1. Among the first three-dimensional data and the second three-dimensional data at the same time, perform a conversion process including movement on the second three-dimensional data, and combine the first three-dimensional data and the second three-dimensional data after the conversion process to generate third three-dimensional data, generate a bitstream by encoding the third three-dimensional data, the bitstream includes first information indicating to which of the first three-dimensional data and the second three-dimensional data each of a plurality of position information or a plurality of attribute information included in the third three-dimensional data belongs, and second information indicating the content of the movement Three-dimensional data encoding method.
2. The conversion process includes geometric transformation in addition to the movement, the bitstream includes third information indicating the content of the geometric transformation The three-dimensional data encoding method according to Claim 1.
3. The geometric transformation includes at least one of shift, rotation, and inversion The three-dimensional data encoding method according to Claim 2.
4. Determine the content of the geometric transformation based on the number of overlapping points that are three-dimensional points with the same position information in the third three-dimensional data The three-dimensional data encoding method according to Claim 2 or 3.
5. The second information includes information indicating the amount of movement of the movement The three-dimensional data encoding method according to any one of Claims 1 to 4.
6. The second information includes information indicating whether to move the origin of the second three-dimensional data to the origin of the first three-dimensional data The three-dimensional data encoding method according to any one of Claims 1 to 5.
7. Generate the first three-dimensional data and the second three-dimensional data by spatially dividing the fourth three-dimensional data The three-dimensional data encoding method according to any one of Claims 1 to 6.
8. Among the first three-dimensional data and the second three-dimensional data at the same time, perform a conversion process including movement on the second three-dimensional data, and decode the third three-dimensional data from the bitstream generated by encoding the third three-dimensional data generated by combining the first three-dimensional data and the second three-dimensional data after the conversion process, obtain, from the bitstream, first information indicating to which of the first three-dimensional data and the second three-dimensional data each of a plurality of position information or a plurality of attribute information included in the third three-dimensional data belongs, and second information indicating the content of the movement Restoring the first three-dimensional data and the second three-dimensional data from the decoded third three-dimensional data by using the first information and the second information Three-dimensional data decoding method. **Claim 9**: The conversion process includes geometric transformation in addition to the movement, acquiring third information indicating the content of the geometric transformation from the bitstream, in the restoration of the first three-dimensional data and the second three-dimensional data, restoring the first three-dimensional data and the second three-dimensional data from the decoded third three-dimensional data by using the first information, the second information, and the third information The three-dimensional data decoding method according to claim 8. **Claim 10** The geometric transformation includes at least one of shift, rotation, and inversion The three-dimensional data decoding method according to claim 9. **Claim 11**: The second information includes information indicating the amount of movement of the movement The three-dimensional data decoding method according to any one of claims 8 to 10. **Claim 12**: The second information includes information indicating whether to move the origin of the second three-dimensional data to the origin of the first three-dimensional data The three-dimensional data decoding method according to any one of claims 8 to 11. **Claim 13** Generating fourth three-dimensional data by spatially combining the restored first three-dimensional data and the second three-dimensional data The three-dimensional data decoding method according to any one of claims 8 to 12. **Claim 14** A processor and, a memory, the processor uses the memory, performing a conversion process including movement on the second three-dimensional data among the first three-dimensional data and the second three-dimensional data at the same time, and generating third three-dimensional data by combining the first three-dimensional data and the second three-dimensional data after the conversion process, generating a bitstream by encoding the third three-dimensional data, the bitstream includes first information indicating to which of the first three-dimensional data and the second three-dimensional data each of a plurality of position information or a plurality of attribute information included in the third three-dimensional data belongs, and second information indicating the content of the movement Three-dimensional data encoding device. **Claim 15** A processor and, a memory, the processor uses the memory, Of the first three-dimensional data and the second three-dimensional data at the same time, perform a conversion process including movement on the second three-dimensional data, and decode the third three-dimensional data generated by encoding the third three-dimensional data generated by combining the first three-dimensional data and the second three-dimensional data after the conversion process. From the bitstream, acquire first information indicating to which of the first three-dimensional data and the second three-dimensional data each of a plurality of position information or a plurality of attribute information included in the third three-dimensional data belongs, and second information indicating the content of the movement. Restore the first three-dimensional data and the second three-dimensional data from the decoded third three-dimensional data using the first information and the second information. Three-dimensional data decoding device.
Citation Information
Patent Citations
Map display device
WO2014020663A1
Image processing device and image processing method
WO2017082078A1
Three-dimensional data generation method, three-dimensional data transmission method, three-dimensional data generation device, and three-dimensional data transmission device
WO2018016168A1