Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device

By summarizing and encoding multiple point group data, generating third point group data and encoding, the problem of low encoding efficiency of three-dimensional data in the prior art is solved, and more efficient data compression and transmission are achieved.

CN112368744BActive Publication Date: 2025-06-06PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN201980041047.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-10-02
Filing Date
2019-10-02
Publication Date
2025-06-06
Estimated Expiration
2039-10-02

AI Technical Summary

Technical Problem

In the prior art, the encoding efficiency of three-dimensional data is low, and it is difficult to effectively compress large-scale point cloud data.

Method used

By summarizing and encoding the multiple point group data, a third point group data containing multiple three-dimensional point position information and identification information is generated, and the data is encoded to generate encoded data.

Benefits of technology

It improves the coding efficiency of three-dimensional data, reduces the compression of data volume, and enhances the performance of data transmission and storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112368744B_ABST
    Figure CN112368744B_ABST
Patent Text Reader

Abstract

A three-dimensional data encoding method obtains third point group data, wherein the third point group data combines the first point group data and the second point group data and includes position information of each of a plurality of three-dimensional points included in the third point group data, and identification information indicating to which of the first point group data and the second point group data the plurality of three-dimensional points respectively belong (S5661). Encoding data (S5662) is generated by encoding the obtained third point group data. During the generation, for each of the plurality of three-dimensional points, the identification information of the three-dimensional point is encoded as attribute information of the three-dimensional point.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, and a three-dimensional data decoding device. Background Art

[0002] In the future, devices and services that make use of 3D data will become more common in large fields such as computer vision, map information, monitoring, infrastructure inspection, or image distribution, which are used for autonomous operation of cars or robots. 3D data is obtained by various methods such as distance sensors such as rangefinders, stereo cameras, or a combination of multiple single-lens cameras.

[0003] As a method of expressing three-dimensional data, there is a method called point cloud, which expresses the shape of a three-dimensional structure through a group of points in a three-dimensional space. The position and color of the point group are stored in the point cloud. Although point cloud is expected to become the mainstream method of expressing three-dimensional data, the amount of point group data is very large. Therefore, in the accumulation or transmission of three-dimensional data, it is necessary to compress the data volume through encoding, just like two-dimensional dynamic images (as an example, there are MPEG-4AVC or HEVC standardized by MPEG).

[0004] Furthermore, compression of point clouds is partially supported by a public library (PointCloud Library) that performs point cloud association processing.

[0005] Furthermore, there is a known technique for searching for facilities around a vehicle using three-dimensional map data and displaying the facilities (for example, refer to Patent Document 1).

[0006] Prior art literature

[0007] Patent Literature

[0008] Patent Document 1 International Publication No. 2014 / 020663 Summary of the invention

[0009] Problem that the invention aims to solve

[0010] It is desirable to improve encoding efficiency in the encoding process of three-dimensional data.

[0011] An object of the present disclosure is to provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device or a three-dimensional data decoding device that can improve encoding efficiency.

[0012] Means used to solve problems

[0013] A three-dimensional data encoding method according to one embodiment of the present invention obtains third point group data, wherein the third point group data combines the first point group data and the second point group data and includes position information of each of a plurality of three-dimensional points included in the third point group data, and identification information indicating to which of the first point group data and the second point group data the plurality of three-dimensional points respectively belong, and generates encoded data by encoding the obtained third point group data. In the generation, for each of the plurality of three-dimensional points, the identification information of the three-dimensional point is encoded as attribute information of the three-dimensional point.

[0014] A three-dimensional data decoding method according to one embodiment of the present invention obtains encoded data, and obtains position information and attribute information of each of a plurality of three-dimensional points contained in third point group data that is a combination of first point group data and second point group data by decoding the encoded data, wherein the attribute information includes identification information indicating to which of the first point group data and the second point group data the three-dimensional point corresponding to the attribute information belongs.

[0015] Effects of the Invention

[0016] The present disclosure can provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device that can improve encoding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is a diagram showing the structure of a three-dimensional data encoding and decoding system according to the first embodiment.

[0018] Figure 2 This is a diagram showing a structural example of point cloud data in Implementation Example 1.

[0019] Figure 3 This is a diagram showing a structural example of a data file that describes point group data information according to Implementation Example 1.

[0020] Figure 4 This is a diagram showing types of point cloud data in Implementation Example 1.

[0021] Figure 5 This is a diagram showing the structure of the first encoding unit according to implementation mode 1.

[0022] Figure 6 This is a block diagram of the first encoding unit according to implementation mode 1.

[0023] Figure 7 This is a diagram showing the structure of the first decoding unit according to Implementation Example 1.

[0024] Figure 8 This is a block diagram of the first decoding unit according to Implementation Example 1.

[0025] Fig. 9This is a diagram showing the structure of the second encoding unit according to implementation mode 1.

[0026] Fig.10 This is a block diagram of the second encoding unit according to implementation mode 1.

[0027] Fig.11 This is a diagram showing the structure of the second decoding unit according to implementation mode 1.

[0028] Fig.12 This is a block diagram of the second decoding unit according to implementation mode 1.

[0029] Fig.13 This is a diagram showing a protocol stack related to PCC coded data according to the first embodiment.

[0030] Fig.14 This is a diagram showing the basic structure of ISOBMFF according to implementation mode 2.

[0031] Fig.15 This is a diagram showing a protocol stack according to the second implementation mode.

[0032] Fig.16 This is a diagram showing the structure of the encoding unit and the multiplexing unit of Implementation Example 3.

[0033] Fig.17 This is a diagram showing a structural example of encoded data according to Implementation Example 3.

[0034] Fig.18 This is a diagram showing a structural example of coded data and a NAL unit according to the third embodiment.

[0035] Fig.19 This is a diagram showing a semantic example of pcc_nal_unit_type according to the third implementation mode.

[0036] Fig. 20 This is a diagram showing an example of the transmission order of NAL units in Implementation Example 3.

[0037] Fig.21 This is a block diagram of the first encoding unit according to implementation mode 4.

[0038] Fig. 22 This is a block diagram of the first decoding unit according to implementation mode 4.

[0039] Fig.23 This is a block diagram of a division unit according to the fourth embodiment.

[0040] Fig.24 This is a diagram showing an example of division of slices and tiles according to the fourth embodiment.

[0041] Fig.25 This is a diagram showing an example of a division pattern of slices and tiles according to the fourth embodiment.

[0042] Fig.26 This is a diagram showing an example of dependency relationship in implementation mode 4.

[0043] Fig. 27 This is a diagram showing an example of the decoding order of data according to Implementation Example 4.

[0044] Fig.28 This is a flowchart of the encoding process of implementation mode 4.

[0045] Fig.29 This is a block diagram of the combining part of implementation mode 4.

[0046] Fig.30 This is a diagram showing a structural example of coded data and a NAL unit according to the fourth embodiment.

[0047] Fig.31 This is a flowchart of the encoding process of implementation mode 4.

[0048] Fig.32 This is a flowchart of the decoding process of implementation mode 4.

[0049] Fig.33 This is a flowchart of the encoding process of implementation mode 4.

[0050] Fig.34 This is a flowchart of the decoding process of implementation mode 4.

[0051] Fig.35 This is a diagram showing an image of a tree structure and occupancy coding generated from point cloud data of a plurality of frames according to the fifth embodiment.

[0052] Fig.36 This is a diagram showing an example of frame combination in the fifth embodiment.

[0053] Fig.37 This is a diagram showing an example of combining a plurality of frames according to the fifth embodiment.

[0054] Fig.38 This is a flowchart of the three-dimensional data encoding process in the fifth embodiment.

[0055] Fig.39 This is a flowchart of the encoding process of implementation mode 5.

[0056] Fig.40 This is a flowchart of the three-dimensional data decoding process according to the fifth embodiment.

[0057] Fig.41 This is a flowchart of the decoding and segmentation processing of implementation mode 5.

[0058] Fig.42 This is a block diagram of the encoding unit of implementation mode 5.

[0059] Fig.43 This is a block diagram of a division unit in the fifth embodiment.

[0060] Fig.44 This is a block diagram of a position information encoding unit according to the fifth embodiment.

[0061] Fig.45 This is a block diagram of an attribute information encoding unit according to the fifth embodiment.

[0062] Fig.46 This is a flowchart of the encoding process of point group data in implementation mode 5.

[0063] Fig.47 This is a flowchart of the encoding process of implementation mode 5.

[0064] Fig.48 This is a block diagram of the decoding unit of implementation mode 5.

[0065] Fig.49 This is a block diagram of a position information decoding unit according to the fifth embodiment.

[0066] Fig.50 This is a block diagram of an attribute information decoding unit according to the fifth embodiment.

[0067] Fig.51 This is a block diagram of the combining part of implementation mode 5.

[0068] Fig.52 This is a flowchart of the decoding process of the point group data in the fifth embodiment.

[0069] Fig.53 This is a flowchart of the decoding process of implementation mode 5.

[0070] Fig.54 This is a diagram showing an example of a frame combination pattern according to the fifth embodiment.

[0071] Fig.55 This is a diagram showing a configuration example of a PCC frame according to the fifth embodiment.

[0072] Fig.56 This is a diagram showing the structure of the encoding position information of implementation mode 5.

[0073] Fig.57 This is a diagram showing a syntax example of a header of encoding position information according to the fifth embodiment.

[0074] Fig.58 This is a diagram showing a syntax example of a payload of encoding position information according to the fifth embodiment.

[0075] Fig.59 This is a diagram showing an example of leaf node information according to the fifth embodiment.

[0076] Fig.60This is a diagram showing an example of leaf node information according to the fifth embodiment.

[0077] Fig.61 This is a diagram showing an example of bit mapping information according to the fifth embodiment.

[0078] Fig.62 This is a diagram showing the structure of encoding attribute information in implementation mode 5.

[0079] Fig.63 This is a diagram showing a syntax example of a header of encoding attribute information according to the fifth embodiment.

[0080] Fig.64 This is a diagram showing a syntax example of a payload of coded attribute information according to the fifth embodiment.

[0081] Fig.65 This is a diagram showing the structure of encoded data in implementation mode 5.

[0082] Fig.66 This is a diagram showing the order of sending data and the reference relationship of data in implementation mode 5.

[0083] Fig.67 This is a diagram showing the order of sending data and the reference relationship of data in implementation mode 5.

[0084] Fig.68 This is a diagram showing the order of sending data and the reference relationship of data in implementation mode 5.

[0085] Fig.69 This is a diagram showing an example of decoding a part of frames according to the fifth embodiment.

[0086] Fig.70 This is a diagram showing the order of sending data and the reference relationship of data in implementation mode 5.

[0087] Fig.71 This is a diagram showing the order of sending data and the reference relationship of data in implementation mode 5.

[0088] Fig.72 This is a diagram showing the order of sending data and the reference relationship of data in implementation mode 5.

[0089] Fig.73 This is a diagram showing the order of sending data and the reference relationship of data in implementation mode 5.

[0090] Fig.74 This is a flowchart of the encoding process of implementation mode 5.

[0091] Fig.75 This is a flowchart of the decoding process of implementation mode 5.

[0092] Fig.76 This is a diagram showing an example of three-dimensional points according to the sixth embodiment.

[0093] Fig.77 This is a diagram showing an example of setting LoD in the sixth embodiment.

[0094] Fig.78 This is a diagram showing an example of a threshold value used in setting LoD in the sixth embodiment.

[0095] Fig.79 This is a diagram showing an example of attribute information used in the prediction value of the sixth embodiment.

[0096] Fig.80 This is a diagram showing an example of the Exponential Golomb code according to the sixth embodiment.

[0097] Fig.81 This is a diagram showing the processing of the Exponential Golomb code according to the sixth embodiment.

[0098] Fig.82 This is a diagram showing a syntax example of an attribute header according to the sixth embodiment.

[0099] Fig.83 This is a diagram showing a syntax example of attribute data according to the sixth embodiment.

[0100] Fig.84 This is a flowchart of the three-dimensional data encoding process in the sixth embodiment.

[0101] Fig.85 This is a flowchart of the attribute information encoding process in the sixth embodiment.

[0102] Fig.86 This is a diagram showing the processing of the Exponential Golomb code according to the sixth embodiment.

[0103] Fig.87 This is a diagram of an example of a reverse calculation table showing the relationship between the remaining codes and their values ​​in Implementation Example 6.

[0104] Fig.88 This is a flowchart of the three-dimensional data decoding process according to the sixth embodiment.

[0105] Fig.89 This is a flowchart of the attribute information decoding process in the sixth embodiment.

[0106] Fig.90 This is a block diagram of a three-dimensional data encoding device according to a sixth embodiment.

[0107] Fig.91 This is a block diagram of a three-dimensional data decoding device according to a sixth embodiment.

[0108] Fig.92 This is a diagram showing the structure of attribute information in the sixth embodiment.

[0109] Fig.93 This is a diagram for illustrating the encoded data of implementation mode 6.

[0110] Fig.94 This is a flowchart of the three-dimensional data encoding process in the sixth embodiment.

[0111] Fig.95 This is a flowchart of the three-dimensional data decoding process according to the sixth embodiment. DETAILED DESCRIPTION

[0112] A three-dimensional data encoding method according to one embodiment of the present invention obtains third point group data, wherein the third point group data combines the first point group data and the second point group data and includes position information of each of a plurality of three-dimensional points included in the third point group data, and identification information indicating to which of the first point group data and the second point group data the plurality of three-dimensional points respectively belong, and generates encoded data by encoding the obtained third point group data. In the generation, for each of the plurality of three-dimensional points, the identification information of the three-dimensional point is encoded as attribute information of the three-dimensional point.

[0113] Therefore, the three-dimensional data encoding method can improve the encoding efficiency by collectively encoding multiple point group data.

[0114] For example, in the generating, the attribute information of the first three-dimensional point may be encoded using the attribute information of the second three-dimensional points around the first three-dimensional point among the plurality of three-dimensional points.

[0115] For example, the attribute information of the first three-dimensional point may include first identification information indicating that the first three-dimensional point belongs to the first point group data, and the attribute information of the second three-dimensional point may include second identification information indicating that the second three-dimensional point belongs to the second point group data.

[0116] For example, in the generation, the attribute information of the second three-dimensional point may be used to calculate a predicted value of the attribute information of the first three-dimensional point, and the difference between the attribute information of the first three-dimensional point and the predicted value, i.e., the prediction residual, may be calculated to generate encoded data including the prediction residual.

[0117] For example, in the obtaining, the third point group data may be obtained by combining the first point group data with the second point group data to generate the third point group data.

[0118] For example, the encoded data may include the identification information in the same data format as other attribute information different from the identification information.

[0119] A three-dimensional data decoding method according to one embodiment of the present invention obtains encoded data, and obtains position information and attribute information of each of a plurality of three-dimensional points contained in third point group data that is a combination of first point group data and second point group data by decoding the encoded data, wherein the attribute information includes identification information indicating to which of the first point group data and the second point group data the three-dimensional point corresponding to the attribute information belongs.

[0120] Therefore, the three-dimensional data decoding method can decode the encoded data whose encoding efficiency is improved by collectively encoding a plurality of point group data.

[0121] For example, in the obtaining, the attribute information of the first three-dimensional point may be decoded using the attribute information of the second three-dimensional points around the first three-dimensional point among the plurality of three-dimensional points.

[0122] For example, the attribute information of the first three-dimensional point may include first identification information indicating that the first three-dimensional point belongs to the first point group data, and the attribute information of the second three-dimensional point may include second identification information indicating that the second three-dimensional point belongs to the second point group data.

[0123] For example, the encoded data may include a prediction residual, and in decoding the encoded data, the attribute information of the second three-dimensional point is used to calculate a predicted value of the attribute information of the first three-dimensional point, and the attribute information of the first three-dimensional point is calculated by adding the predicted value to the prediction residual.

[0124] For example, the third point group data may be divided into the first point group data and the second point group data by further using the identification information.

[0125] For example, the encoded data may include the identification information in the same data format as other attribute information different from the identification information.

[0126] In addition, a three-dimensional data encoding device of one embodiment of the present invention includes a processor and a memory, and the processor uses the memory to obtain third point group data, the third point group data combines the first point group data and the second point group data, and includes position information of each of a plurality of three-dimensional points included in the third point group data, and identification information indicating to which of the first point group data and the second point group data the plurality of three-dimensional points respectively belong, and generates encoded data by encoding the obtained third point group data, and in the generation, for each of the plurality of three-dimensional points, the identification information of the three-dimensional point is encoded as attribute information of the three-dimensional point.

[0127] Therefore, the three-dimensional data encoding device can improve the encoding efficiency by collectively encoding multiple point group data.

[0128] A three-dimensional data decoding device according to one embodiment of the present invention comprises a processor and a memory, wherein the processor uses the memory to obtain encoded data, and obtains position information and attribute information of each of a plurality of three-dimensional points contained in third point group data that is a combination of first point group data and second point group data by decoding the encoded data, wherein the attribute information includes identification information indicating to which of the first point group data and the second point group data the three-dimensional point corresponding to the attribute information belongs.

[0129] Thus, the three-dimensional data decoding device can decode coded data in which coding efficiency is improved by collectively coding a plurality of point group data.

[0130] In addition, these general or specific forms can be implemented by systems, methods, integrated circuits, computer programs, or computer-readable recording media such as CD-ROMs, and can be implemented by any combination of systems, methods, integrated circuits, computer programs, and recording media.

[0131] The following detailed description of the implementation mode is given with reference to the accompanying drawings. In addition, the implementation modes to be described below are all specific examples of the present disclosure. The numerical values, shapes, materials, constituent elements, configuration positions of constituent elements, connection forms, steps, order of steps, etc. shown in the following implementation modes are all examples, and the main purpose is not to limit the present disclosure. Furthermore, the constituent elements of the following implementation modes that are not recorded in the technical solution showing the highest concept are described as arbitrary constituent elements.

[0132] (Implementation Method 1)

[0133] When point cloud coded data is used in actual devices or services, it is desirable to transmit and receive required information according to the application in order to reduce network bandwidth. However, such a function does not exist in the existing 3D data coding structure, and therefore there is no corresponding coding method.

[0134] What will be described in this embodiment is a three-dimensional data encoding method and a three-dimensional data encoding device for providing the function of sending and receiving required information according to the purpose in the encoded data of a three-dimensional point cloud, a three-dimensional data decoding method and a three-dimensional data decoding device for decoding the encoded data, a three-dimensional data multiplexing method for multiplexing the encoded data, and a three-dimensional data transmission method for transmitting the encoded data.

[0135] In particular, currently, the first encoding method and the second encoding method have been studied as encoding methods (encoding methods) for point group data, but the structure of the encoded data and the method of storing the encoded data in the system format are not defined, and there is a problem that MUX processing (multiplexing) in the encoding unit, or transmission or accumulation cannot be directly performed.

[0136] Furthermore, there is no method that supports a format in which two codecs, namely the first encoding method and the second encoding method, exist in a mixed form, such as PCC (Point Cloud Compression).

[0137] In this embodiment, a structure of PCC coded data in which two codecs, namely, a first coding method and a second coding method, exist in a mixed state, and a method of storing the coded data in a system format will be described.

[0138] First, the configuration of the three-dimensional data (point cloud data) encoding and decoding system according to the present embodiment will be described. Figure 1 2 is a diagram showing a configuration example of a three-dimensional data encoding and decoding system according to the present embodiment. Figure 1 As shown, the three-dimensional data encoding and decoding system includes a three-dimensional data encoding system 4601, a three-dimensional data decoding system 4602, a sensor terminal 4603 and an external connection unit 4604.

[0139] The three-dimensional data encoding system 4601 generates encoded data or multiplexed data by encoding point group data as three-dimensional data. In addition, the three-dimensional data encoding system 4601 can be a three-dimensional data encoding device implemented by a single device, or a system implemented by multiple devices. In addition, the three-dimensional data encoding device can also be included in a part of the multiple processing units included in the three-dimensional data encoding system 4601.

[0140] The three-dimensional data encoding system 4601 includes a point cloud data generating system 4611, a presentation unit 4612, an encoding unit 4613, a multiplexing unit 4614, an input / output unit 4615, and a control unit 4616. The point cloud data generating system 4611 includes a sensor information obtaining unit 4617 and a point cloud data generating unit 4618.

[0141] The sensor information acquisition unit 4617 acquires sensor information from the sensor terminal 4603 and outputs the sensor information to the point cloud data generation unit 4618. The point cloud data generation unit 4618 generates point cloud data based on the sensor information and outputs the point cloud data to the encoding unit 4613.

[0142] The presentation unit 4612 presents the sensor information or point group data to the user. For example, the presentation unit 4612 displays information or an image based on the sensor information or point group data.

[0143] The encoding unit 4613 encodes (compresses) the point cloud data, and outputs the obtained encoded data, control information obtained in the encoding process, and other additional information to the multiplexing unit 4614. The additional information includes, for example, sensor information.

[0144] The multiplexing unit 4614 generates multiplexed data by multiplexing the coded data, control information, and additional information input from the coding unit 4613. The format of the multiplexed data is, for example, a file format for storage or a packet format for transmission.

[0145] The input / output unit 4615 (e.g., a communication unit or an interface) outputs the multiplexed data to the outside. Alternatively, the multiplexed data is stored in a storage unit such as an internal memory. The control unit 4616 (or an application execution unit) controls each processing unit. That is, the control unit 4616 performs control such as encoding and multiplexing.

[0146] Furthermore, the sensor information may be input to the encoding unit 4613 or the multiplexing unit 4614. Furthermore, the input / output unit 4615 may directly output the point cloud data or the encoded data to the outside.

[0147] The transmission signal (multiplexed data) output from the three-dimensional data encoding system 4601 is input to the three-dimensional data decoding system 4602 via the external connection unit 4604 .

[0148] The three-dimensional data decoding system 4602 generates point group data as three-dimensional data by decoding the coded data or the multiplexed data. In addition, the three-dimensional data decoding system 4602 may be a three-dimensional data decoding device implemented by a single device, or may be a system implemented by multiple devices. In addition, the three-dimensional data decoding device may also include a part of the multiple processing units included in the three-dimensional data decoding system 4602.

[0149] The three-dimensional data decoding system 4602 includes a sensor information acquisition unit 4621 , an input / output unit 4622 , an inverse multiplexing unit 4623 , a decoding unit 4624 , a prompting unit 4625 , a user interface 4626 , and a control unit 4627 .

[0150] The sensor information acquisition unit 4621 acquires sensor information from the sensor terminal 4603 .

[0151] The input / output unit 4622 obtains the transmission signal, decodes the multiplexed data (file format or packet) according to the transmission signal, and outputs the multiplexed data to the demultiplexing unit 4623 .

[0152] The inverse multiplexing unit 4623 obtains the encoded data, control information, and additional information from the multiplexed data, and outputs the encoded data, control information, and additional information to the decoding unit 4624 .

[0153] The decoding unit 4624 reconstructs the point cloud data by decoding the encoded data.

[0154] The prompt unit 4625 prompts the user with the point group data. For example, the prompt unit 4625 displays information or images based on the point group data. The user interface 4626 obtains instructions based on the user's operation. The control unit 4627 (or the application execution unit) controls each processing unit. That is, the control unit 4627 performs inverse multiplexing, decoding, prompting, etc.

[0155] In addition, the input / output unit 4622 may also directly obtain point group data or coded data from the outside. In addition, the prompt unit 4625 may also obtain additional information such as sensor information and prompt information based on the additional information. In addition, the prompt unit 4625 may also perform prompts based on the user's instructions obtained by the user interface 4626.

[0156] The sensor terminal 4603 generates sensor information, which is information obtained by the sensor. The sensor terminal 4603 is a terminal equipped with a sensor or a camera, for example, a moving object such as a car, a flying object such as an airplane, a mobile terminal or a camera.

[0157] The sensor information that can be obtained by the sensor terminal 4603 is, for example, (1) the distance between the sensor terminal 4603 and the object, or the reflectivity of the object, obtained by LIDAR, millimeter wave radar, or infrared sensor, (2) the distance between the camera and the object, or the reflectivity of the object, obtained from multiple monocular camera images or stereo camera images. In addition, the sensor information may also include the posture, orientation, angular motion (angular velocity), position (GPS information or altitude), speed or acceleration of the sensor. In addition, the sensor information may also include temperature, air pressure, humidity, or magnetism.

[0158] The external connection unit 4604 is implemented by an integrated circuit (LSI or IC), an external storage unit, communication with a cloud server via the Internet, or broadcasting.

[0159] Next, point group data will be described. Figure 2 It is a diagram showing the structure of point group data. Figure 3 It is a diagram showing a structural example of a data file that describes information on point group data.

[0160] Point group data includes data of multiple points. The data of each point includes position information (three-dimensional coordinates) and attribute information for the position information. A group of multiple points is called a point group. For example, a point group represents the three-dimensional shape of an object.

[0161] Sometimes, position information (Position) such as three-dimensional coordinates is also called geometry (Geometry). In addition, the data of each point can also include attribute information (attribute) of multiple attribute categories. Attribute categories are, for example, color or reflectivity.

[0162] One piece of attribute information may be associated with one piece of position information, or attribute information of a plurality of different attribute categories may be associated with one piece of position information. In addition, a plurality of attribute information of the same attribute category may be associated with one piece of position information.

[0163] Figure 3 The illustrated data file structure example is an example of a case where position information and attribute information correspond to each other on a one-to-one basis, and indicates position information and attribute information of N points constituting point cloud data.

[0164] The position information is, for example, information on three axes, x, y, and z. The attribute information is, for example, color information of RGB. Representative data files include ply files and the like.

[0165] Next, types of point cloud data will be described. Figure 4 is a graph showing the types of point group data. Figure 4 As shown, the point group data includes static objects and dynamic objects.

[0166] A static object is a 3D point group data at any time (a certain moment). A dynamic object is a 3D point group data that changes over time. Hereinafter, the 3D point group data at a certain moment is referred to as a PCC frame or frame.

[0167] The object may be a point group whose area is limited to a certain extent, such as normal image data, or a large-scale point group whose area is not limited, such as map information.

[0168] In addition, there may be point group data of various densities, such as sparse point group data and dense point group data.

[0169] The details of each processing unit are described below. The sensor information is obtained by various methods such as a distance sensor such as LIDAR or a rangefinder, a stereo camera, or a combination of multiple monocular cameras. The point group data generation unit 4618 generates point group data based on the sensor information obtained by the sensor information acquisition unit 4617. The point group data generation unit 4618 generates position information as point group data and adds attribute information for the position information to the position information.

[0170] The point group data generation unit 4618 may also process the point group data when generating the position information or additional attribute information. For example, the point group data generation unit 4618 may also reduce the amount of data by deleting point groups with repeated positions. In addition, the point group data generation unit 4618 may also transform the position information (position conversion, rotation or standardization, etc.) and may also render the attribute information.

[0171] In addition, Figure 1 In the figure, the point group data generating system 4611 is included in the three-dimensional data encoding system 4601, but can also be independently set outside the three-dimensional data encoding system 4601.

[0172] The coding unit 4613 encodes the point cloud data based on a predetermined coding method, thereby generating coded data. There are generally two types of coding methods. The first type is a coding method using position information, which is hereinafter referred to as the first coding method. The second type is a coding method using a video codec, which is hereinafter referred to as the second coding method.

[0173] The decoding unit 4624 decodes the encoded data based on a predetermined encoding method, thereby decoding the point cloud data.

[0174] The multiplexing unit 4614 generates multiplexed data by multiplexing the coded data using an existing multiplexing method. The generated multiplexed data is transmitted or stored. In addition to the PCC coded data, the multiplexing unit 4614 also multiplexes other media such as images, sounds, subtitles, applications, files, or reference time information. In addition, the multiplexing unit 4614 can also multiplex attribute information associated with sensor information or point group data.

[0175] As multiplexing methods or file formats, there are ISOBMFF, MPEG-DASH which is a transmission method based on ISOBMFF, MMT, MPEG-2TS Systems, RMP, etc.

[0176] The demultiplexing unit 4623 extracts PCC coded data, other media, time information, etc. from the multiplexed data.

[0177] The input / output unit 4615 transmits the multiplexed data using a method consistent with a transmission medium or a storage medium such as broadcasting or communication. The input / output unit 4615 can communicate with other devices via the Internet, or can communicate with a storage unit such as a cloud server.

[0178] As the communication protocol, http, ftp, TCP, UDP, etc. can be used. Either a PULL type communication method or a PUSH type communication method can be used.

[0179] Any of wired transmission and wireless transmission can be used. As wired transmission, Ethernet (registered trademark), USB, RS-232C, HDMI (registered trademark), or coaxial cable, etc. are used. As wireless transmission, wireless LAN, Wi-Fi (registered trademark), Bluetooth (registered trademark), or millimeter wave, etc. are used.

[0180] In addition, as a broadcasting method, for example, DVB-T2, DVB-S2, DVB-C2, ATSC3.0, or ISDB-S3 is used.

[0181] Figure 5 This diagram shows a structure of a first coding unit 4630 which is an example of a coding unit 4613 that performs coding using the first coding method. Figure 6 4 is a block diagram of the first coding unit 4630. The first coding unit 4630 generates coded data (coded stream) by coding the point cloud data using the first coding method. The first coding unit 4630 includes a position information coding unit 4631, an attribute information coding unit 4632, an additional information coding unit 4633, and a multiplexing unit 4634.

[0182] The first coding unit 4630 has a feature of performing coding in consideration of the three-dimensional structure. In addition, the first coding unit 4630 has a feature of performing coding by the attribute information coding unit 4632 using information obtained from the position information coding unit 4631. The first coding method is also called GPCC (Geometry based PCC).

[0183] The point group data is PCC point group data such as a PLY file, or PCC point group data generated based on sensor information, and includes position information (Position), attribute information (Attribute), and other additional information (MetaData). The position information is input to the position information encoding unit 4631, the attribute information is input to the attribute information encoding unit 4632, and the additional information is input to the additional information encoding unit 4633.

[0184] The position information encoding unit 4631 generates encoded position information (Compressed Geometry) as encoded data by encoding the position information. For example, the position information encoding unit 4631 uses an N-ary tree structure such as an octree to encode the position information. Specifically, in the octree, the object space is divided into 8 nodes (subspaces), and 8 bits of information (occupancy code) are generated to indicate whether each node contains a point group. In addition, the node containing the point group is further divided into 8 nodes, and 8 bits of information are generated to indicate whether each of the 8 nodes contains a point group. This process is repeated until it becomes below the threshold of the number of point groups contained in a predetermined layer or node.

[0185] The attribute information encoding unit 4632 generates the encoded attribute information (Compressed Attribute) as the encoded data by encoding using the structure information generated by the position information encoding unit 4631. For example, the attribute information encoding unit 4632 determines the reference point (reference node) to be referred to in the encoding of the object point (object node) of the processing object based on the octree structure generated by the position information encoding unit 4631. For example, the attribute information encoding unit 4632 refers to a node whose parent node in the octree is the same as the parent node of the object node among the surrounding nodes or adjacent nodes. In addition, the method of determining the reference relationship is not limited to this.

[0186] In addition, the encoding process of the attribute information may include at least one of a quantization process, a prediction process, and an arithmetic coding process. In this case, the reference means using the reference node in the calculation of the predicted value of the attribute information, or using the state of the reference node in the determination of the encoding parameter (for example, indicating whether the occupancy information of the point group is included in the reference node). For example, the encoding parameter is a quantization parameter in the quantization process, or a context in the arithmetic coding, etc.

[0187] The additional information encoding unit 4633 generates encoded additional information (Compressed MetaData) as encoded data by encoding compressible data in the additional information.

[0188] The multiplexing unit 4634 generates a coded stream (Compressed Stream) as coded data by multiplexing the coding position information, coding attribute information, coding additional information, and other additional information. The generated coded stream is output to a processing unit of the system layer (not shown).

[0189] Next, the first decoding unit 4640 which is an example of the decoding unit 4624 that performs decoding according to the first encoding method is described. Figure 7 This is a diagram showing the structure of the first decoding unit 4640. Figure 846 is a block diagram of a first decoding unit 4640. The first decoding unit 4640 generates point cloud data by decoding coded data (coded stream) coded by the first coding method by the first coding method. The first decoding unit 4640 includes an inverse multiplexing unit 4641, a position information decoding unit 4642, an attribute information decoding unit 4643, and an additional information decoding unit 4644.

[0190] A coded stream (Compressed Stream) as coded data is input to the first decoding unit 4640 from a processing unit of a system layer (not shown).

[0191] The inverse multiplexing unit 4641 separates the encoded position information (Compressed Geometry), the encoded attribute information (Compressed Attribute), the encoded additional information (Compressed MetaData), and other additional information from the encoded data.

[0192] The position information decoding unit 4642 generates position information by decoding the encoded position information. For example, the position information decoding unit 4642 restores the position information of the point group represented by three-dimensional coordinates based on the encoded position information represented by an N-ary tree structure such as an octree.

[0193] The attribute information decoding unit 4643 decodes the encoded attribute information based on the structure information generated by the position information decoding unit 4642. For example, the attribute information decoding unit 4643 determines a reference point (reference node) to be referred to in decoding the object point (object node) of the processing object based on the octree structure obtained by the position information decoding unit 4642. For example, the attribute information decoding unit 4643 refers to a node whose parent node in the octree is the same as the parent node of the object node among the surrounding nodes or adjacent nodes. In addition, the method of determining the reference relationship is not limited to this.

[0194] In addition, the decoding process of the attribute information may also include at least one of an inverse quantization process, a prediction process, and an arithmetic decoding process. In this case, the reference means that the reference node is used in the calculation of the predicted value of the attribute information, or the state of the reference node is used in the determination of the decoded parameter (for example, indicating whether the occupancy information of the point group is included in the reference node). For example, the decoded parameter is a quantization parameter in the inverse quantization process, or a context in the arithmetic decoding, etc.

[0195] The additional information decoding unit 4644 generates additional information by decoding the encoded additional information. In addition, the first decoding unit 4640 uses the additional information required for decoding processing of the position information and the attribute information during decoding, and outputs the additional information required by the application to the outside.

[0196] Next, the second encoding unit 4650 which is an example of the encoding unit 4613 that performs encoding using the second encoding method is described. Fig. 9 This is a diagram showing the structure of the second encoding unit 4650. Fig.10 This is a block diagram of the second encoding unit 4650.

[0197] The second coding unit 4650 generates coded data (coded stream) by coding the point cloud data using the second coding method. The second coding unit 4650 includes an additional information generating unit 4651 , a position image generating unit 4652 , an attribute image generating unit 4653 , a video coding unit 4654 , an additional information coding unit 4655 , and a multiplexing unit 4656 .

[0198] The second coding unit 4650 generates a position image and an attribute image by projecting a three-dimensional structure onto a two-dimensional image, and codes the generated position image and attribute image using an existing video coding method. The second coding method is also called VPCC (Video based PCC).

[0199] The point group data is PCC point group data such as a PLY file or PCC point group data generated based on sensor information, and includes position information (Position), attribute information (Attribute), and other additional information (MetaData).

[0200] The additional information generating unit 4651 generates mapping information of a plurality of two-dimensional images by projecting a three-dimensional structure onto a two-dimensional image.

[0201] The position image generation unit 4652 generates a position image (Geometry Image) based on the position information and the mapping information generated by the additional information generation unit 4651. The position image is, for example, a distance image representing the distance (Depth) as a pixel value. In addition, the distance image can be an image of a plurality of point groups observed from one viewpoint (an image of a plurality of point groups projected on a two-dimensional plane), or a plurality of images of a plurality of point groups observed from a plurality of viewpoints, or a single image formed by integrating these plurality of images.

[0202] The attribute image generation unit 4653 generates an attribute image based on the attribute information and the mapping information generated by the additional information generation unit 4651. The attribute image is, for example, an image that represents the attribute information (e.g., color (RGB)) as a pixel value. In addition, the image may be an image of a plurality of point groups observed from one viewpoint (an image in which a plurality of point groups are projected on a two-dimensional plane), or may be a plurality of images of a plurality of point groups observed from a plurality of viewpoints, or may be a single image formed by integrating these plurality of images.

[0203] The image encoding unit 4654 encodes the position image and the attribute image using an image encoding method to generate a compressed position image (Compressed Geometry Image) and a compressed attribute image (Compressed Attribute Image) as encoding data. In addition, as the image encoding method, any known encoding method can be used. For example, the image encoding method is AVC or HEVC.

[0204] The additional information encoding unit 4655 generates coded additional information (Compressed MetaData) by encoding the additional information and mapping information included in the point cloud data.

[0205] The multiplexing unit 4656 generates a coded stream (Compressed Stream) as coded data by multiplexing the coding position image, the coding attribute image, the coding additional information, and other additional information. The generated coded stream is output to a processing unit of the system layer (not shown).

[0206] Next, the second decoding unit 4660 which is an example of the decoding unit 4624 that performs decoding according to the second encoding method is described. Fig.11 This is a diagram showing the structure of the second decoding unit 4660. Fig.12 46 is a block diagram of a second decoding unit 4660. The second decoding unit 4660 generates point cloud data by decoding the coded data (coded stream) coded by the second coding method using the second coding method. The second decoding unit 4660 includes an inverse multiplexing unit 4661, a video decoding unit 4662, an additional information decoding unit 4663, a position information generating unit 4664, and an attribute information generating unit 4665.

[0207] A coded stream (Compressed Stream) as coded data is input to the second decoding unit 4660 from a processing unit of a system layer (not shown).

[0208] The inverse multiplexing unit 4661 separates the coded position image (Compressed Geometry Image), the coded attribute image (Compressed Attribute Image), the coded additional information (Compressed MetaData), and other additional information from the coded data.

[0209] The video decoding unit 4662 generates a position image and an attribute image by decoding the encoded position image and the encoded attribute image using a video encoding method. In addition, as the video encoding method, any known encoding method can be used. For example, the video encoding method is AVC or HEVC.

[0210] The additional information decoding unit 4663 generates additional information including mapping information and the like by decoding the encoded additional information.

[0211] The position information generating unit 4664 generates the position information using the position image and the mapping information. The attribute information generating unit 4665 generates the attribute information using the attribute image and the mapping information.

[0212] The second decoding unit 4660 uses the additional information required for decoding during decoding, and outputs the additional information required for the application to the outside.

[0213] The following describes the issues in the PCC coding method. Fig.13 It is a diagram showing a protocol stack related to PCC coded data. Fig.13 This shows an example of multiplexing, transmitting, or accumulating data of other media such as images (for example, HEVC) or audio into PCC coded data.

[0214] The multiplexing method and file format have functions for multiplexing, transmitting or accumulating various coded data. In order to transmit or accumulate coded data, the coded data must be converted into a format of the multiplexing method. For example, HEVC specifies a technology for storing coded data in a data structure called a NAL unit and storing the NAL unit in ISOBMFF.

[0215] On the other hand, currently, as coding methods for point group data, the first coding method (Codec1) and the second coding method (Codec2) are studied, but the structure of the coded data and the method of storing the coded data in the system format are not defined, and there is a problem that MUX processing (multiplexing), transmission and accumulation in the coding unit cannot be directly performed.

[0216] In the following, unless a specific encoding method is mentioned, either the first encoding method or the second encoding method is indicated.

[0217] (Implementation Method 2)

[0218] In this embodiment, a method of storing a NAL unit in an ISOBMFF file is described.

[0219] ISOBMFF (ISO based media file format) is a file format standard specified in ISO / IEC14496-12. ISOBMFF specifies a format that can reuse and store various media such as video, audio, and text, and is a standard that is independent of the media.

[0220] The basic structure (file) of ISOBMFF is explained. The basic unit in ISOBMFF is a box. A box consists of type, length, and data. A collection of boxes of various types is a file.

[0221] Fig.14 This is a diagram showing the basic structure (file) of ISOBMFF. The ISOBMFF file mainly includes boxes such as ftyp, which is represented by 4CC (4 character code) as the file brand, moov, which stores metadata such as control information, and mdat, which stores data.

[0222] In addition, the storage method for each media of the ISOBMFF file is specified. For example, the storage method of AVC video and HEVC video is specified by ISO / IEC14496-15. Here, in order to store or transmit PCC coded data, it is considered to expand the use of ISOBMFF functions, but there is no provision for storing PCC coded data in ISOBMFF files. Therefore, in this embodiment, a method for storing PCC coded data in an ISOBMFF file is described.

[0223] Fig.15 This is a diagram showing a protocol stack when a NAL unit common to a PCC codec is stored in an ISOBMFF file. Here, a NAL unit common to a PCC codec is stored in an ISOBMFF file. The NAL unit is common to the PCC codec, but since multiple PCC codecs are stored in the NAL unit, it is preferable to define a storage method (Carriage of Codec1, Carriage of Codec2) corresponding to each codec.

[0224] (Implementation method 3)

[0225] In this embodiment, the types of coded data (position information (Geometry), attribute information (Attribute), additional information (Metadata)) generated by the first coding unit 4630 or the second coding unit 4650, the generation method of the additional information (metadata), and the multiplexing process in the multiplexing unit are described. In addition, the additional information (metadata) is sometimes expressed as a parameter set or control information.

[0226] In this embodiment, Figure 4 The dynamic object (three-dimensional point group data that changes with time) described in the example is used for explanation, but the same method can also be used in the case of a static object (three-dimensional point group data at any time).

[0227] Fig.16 4 is a diagram showing the configuration of a coding unit 4801 and a multiplexing unit 4802 included in the three-dimensional data coding device of this embodiment. The coding unit 4801 corresponds to, for example, the first coding unit 4630 or the second coding unit 4650 described above. The multiplexing unit 4802 corresponds to the multiplexing unit 4634 or 4656 described above.

[0228] The encoding unit 4801 encodes the point cloud data of a plurality of PCC (Point Cloud Compression) frames to generate encoded data (Multiple Compressed Data) of a plurality of position information, attribute information, and additional information.

[0229] The multiplexing unit 4802 converts the data of a plurality of data types (position information, attribute information, and additional information) into a data structure that takes data access in the decoding device into consideration.

[0230] Fig.17 4801. The arrows in the figure represent the dependency relationship of decoding of the encoded data, and the direction of the arrow depends on the data of the direction of the arrow. That is, the decoding device decodes the data of the direction of the arrow, and uses the decoded data to decode the data of the direction of the arrow. In other words, the so-called dependency refers to the reference (use) of the data of the dependency target in the processing (encoding or decoding, etc.) of the data of the dependency source.

[0231] First, the generation process of the encoded data of the position information is described. The encoding unit 4801 generates the encoded position data (Compressed Geometry Data) of each frame by encoding the position information of each frame. In addition, G(i) represents the encoded position data. i represents the frame number or the time of the frame, etc.

[0232] In addition, the encoding unit 4801 generates a position parameter set (GPS(i)) corresponding to each frame. The position parameter set includes parameters that can be used in decoding the encoded position data. In addition, the encoded position data of each frame depends on the corresponding position parameter set.

[0233] In addition, the coded position data composed of multiple frames is defined as a position sequence (Geometry Sequence). The coding unit 4801 generates a position sequence parameter set (Geometry Sequence PS: also referred to as position SPS), which stores parameters commonly used in the decoding process for multiple frames in the position sequence. The position sequence depends on the position SPS.

[0234] Next, the generation process of the encoded attribute information data is described. The encoding unit 4801 generates the compressed attribute data (Compressed Attribute Data) of each frame by encoding the attribute information of each frame. In addition, A(i) represents the compressed attribute data. Fig.17 , an example in which attribute X and attribute Y exist is shown, and the encoded attribute data of attribute X is represented by AX(i), and the encoded attribute data of attribute Y is represented by AY(i).

[0235] In addition, the encoding unit 4801 generates an attribute parameter set (APS(i)) corresponding to each frame. In addition, AXPS(i) represents the attribute parameter set of attribute X, and AYPS(i) represents the attribute parameter set of attribute Y. The attribute parameter set includes parameters that can be used in decoding the encoded attribute information. The encoded attribute data depends on the corresponding attribute parameter set.

[0236] In addition, the coded attribute data consisting of multiple frames is defined as an attribute sequence (Attribute Sequence). The coding unit 4801 generates an attribute sequence parameter set (Attribute Sequence PS: also referred to as attribute SPS), which stores parameters commonly used in the decoding process for multiple frames in the attribute sequence. The attribute sequence depends on the attribute SPS.

[0237] Furthermore, in the first encoding method, the encoded attribute data depends on the encoded position data.

[0238] In addition, Fig.17 , an example of a case where there are two types of attribute information (attribute X and attribute Y) is shown. When there are two types of attribute information, for example, two encoding units generate respective data and metadata. In addition, for example, an attribute sequence is defined for each type of attribute information, and an attribute SPS is generated for each type of attribute information.

[0239] In addition, Fig.17 , an example is shown in which there is one type of position information and two types of attribute information, but the present invention is not limited thereto, and the attribute information may be one type or three or more types. In this case, the encoded data can also be generated by the same method. In addition, in the case of point cloud data without attribute information, the attribute information may not exist. In this case, the encoding unit 4801 may not generate a parameter set associated with the attribute information.

[0240] Next, the generation process of the additional information (metadata) is described. The coding unit 4801 generates a PCC stream PS (PCC Stream PS: also referred to as stream PS) which is a parameter set of the entire PCC stream. The coding unit 4801 stores parameters that can be used in common in the decoding process for one or more position sequences and one or more attribute sequences in the stream PS. For example, the stream PS includes identification information of the codec representing the point group data, information representing the algorithm used in the encoding, etc. The position sequence and the attribute sequence depend on the stream PS.

[0241] Next, the access unit and GOF are described. In this embodiment, new concepts of access unit (AccessUnit: AU) and GOF (Group of Frame) are introduced.

[0242] An access unit is a basic unit for accessing data during decoding, and is composed of one or more data and one or more metadata. For example, an access unit is composed of location information and one or more attribute information at the same time. A GOF is a random access unit, and is composed of one or more access units.

[0243] The coding unit 4801 generates an access unit header (AU Header) as identification information indicating the beginning of the access unit. The coding unit 4801 stores the parameters of the access unit in the access unit header. For example, the access unit header includes the structure or information of the coded data included in the access unit. In addition, the access unit header includes parameters commonly used by the data included in the access unit, such as parameters for decoding the coded data.

[0244] In addition, the encoder 4801 may generate an access unit delimiter that does not include access unit parameters instead of the access unit header. The access unit delimiter is used as identification information indicating the beginning of the access unit. The decoding device identifies the beginning of the access unit by detecting the access unit header or the access unit delimiter.

[0245] Next, the generation of identification information at the beginning of GOF is described. The coding unit 4801 generates a GOF header (GOFHeader) as identification information indicating the beginning of GOF. The coding unit 4801 stores the parameters of GOF in the GOF header. For example, the GOF header contains the structure or information of the encoded data contained in the GOF. In addition, the GOF header contains parameters commonly used in the data contained in the GOF, such as parameters for decoding the encoded data.

[0246] In addition, the encoder 4801 may generate a GOF delimiter that does not include GOF parameters instead of the GOF header. The GOF delimiter is used as identification information indicating the beginning of the GOF. The decoding device recognizes the beginning of the GOF by detecting the GOF header or the GOF delimiter.

[0247] In PCC coded data, for example, an access unit is defined as a PCC frame unit. The decoding device accesses the PCC frame based on identification information at the head of the access unit.

[0248] In addition, for example, GOF is defined as one random access unit. The decoding device accesses the random access unit based on the identification information at the beginning of GOF. For example, if PCC frames have no dependency on each other and can be decoded independently, the PCC frame can also be defined as a random access unit.

[0249] Furthermore, two or more PCC frames may be allocated to one access unit, and a plurality of random access units may be allocated to one GOF.

[0250] In addition, the encoder 4801 may define and generate parameter sets or metadata other than those described above. For example, the encoder 4801 may generate SEI (Supplemental Enhancement Information) storing parameters (optional parameters) that may not necessarily be used during decoding.

[0251] Next, the structure of the coded data and the method of storing the coded data in the NAL unit are described.

[0252] For example, a data format is defined for each type of coded data. Fig.18 This is a diagram showing examples of coded data and NAL units.

[0253] For example, Fig.18 As shown, the coded data includes a header and a payload. In addition, the coded data may also include length information indicating the length (data amount) of the coded data, the header or the payload. In addition, the coded data may not include a header.

[0254] The header includes, for example, identification information for specifying data. The identification information indicates, for example, the data type or the frame number.

[0255] The header includes, for example, identification information indicating a reference relationship. This identification information is stored in the header when there is a dependency relationship between data, and is information used to refer to the reference target based on the reference source. For example, the header of the reference target includes identification information for identifying the data. The header of the reference source includes identification information indicating the reference target.

[0256] Furthermore, when the reference destination or the reference source can be identified or derived from other information, identification information for specifying data or identification information indicating a reference relationship may be omitted.

[0257] The multiplexing unit 4802 stores the coded data in the payload of the NAL unit. The NAL unit header includes pcc_nal_unit_type which is identification information of the coded data. Fig.19 This is a diagram showing an example of the semantics of pcc_nal_unit_type.

[0258] like Fig.19 As shown, when pcc_codec_type is codec 1 (Codec 1: first coding method), pcc_nal_unit_type values ​​0 to 10 are allocated to the coded position data (Geometry), coded attribute X data (AttributeX), coded attribute Y data (AttributeY), position PS (Geom.PS), attribute XPS (AttrX.PS), attribute YPS (AttrY.PS), position SPS (Geometry Sequence PS), attribute XSPS (AttributeX Sequence PS), attribute YSPS (AttributeY Sequence PS), AU Header (AU Header), and GOF Header (GOF Header) in codec 1. In addition, values ​​11 and above are allocated to the spare of codec 1.

[0259] When pcc_codec_type is codec 2 (Codex2: second encoding method), pcc_nal_unit_type values ​​0 to 2 are assigned to codec data A (DataA), metadata A (MetaDataA), and metadata B (MetaDataB). In addition, values ​​3 and later are assigned to codec 2 spare.

[0260] Next, the order in which data is transmitted is described. Next, the constraints on the order in which NAL units are transmitted are described.

[0261] The multiplexing unit 4802 collects and transmits NAL units in units of GOF or AU. The multiplexing unit 4802 arranges a GOF header at the beginning of a GOF and arranges an AU header at the beginning of an AU.

[0262] Even when data is lost due to packet loss or the like, the multiplexing unit 4802 can configure a sequence parameter set (SPS) for each AU so that the decoding device can perform decoding from the next AU.

[0263] When there is a decoding dependency in the coded data, the decoding device decodes the reference source data after decoding the reference target data. In the decoding device, in order to decode the data in the order received without reordering the data, the multiplexing unit 4802 sends the reference target data first.

[0264] Fig. 20 This is a diagram showing an example of the order in which NAL units are transmitted. Fig. 20 It shows three examples: position information priority, parameter priority, and data integration.

[0265] The transmission order of location information priority is an example in which information related to location information and information related to attribute information are separately sent in a bundle. In the case of this transmission order, the transmission of information related to location information is completed faster than the transmission of information related to attribute information.

[0266] For example, a decoding device that does not decode attribute information by using the transmission order may set a time when no processing is performed by ignoring the decoding of attribute information. In addition, for example, in the case of a decoding device that wants to decode position information as quickly as possible, it is possible to decode the position information more quickly by obtaining the encoded data of the position information as quickly as possible.

[0267] In addition, Fig. 20 In the description, the attribute XSPS and the attribute YSPS are integrated and recorded as the attribute SPS, but the attribute XSPS and the attribute YSPS may be configured separately.

[0268] In the parameter set priority sending order, the parameter set is sent first, and then the data is sent.

[0269] As described above, the multiplexing unit 4802 may send the NAL units in any order according to the constraints of the NAL unit sending order. For example, the order identification information may be defined, and the multiplexing unit 4802 may have the function of sending the NAL units in a plurality of styles. For example, the order identification information of the NAL units may be stored in the stream PS.

[0270] The 3D data decoding device may also perform decoding based on the sequence identification information. The 3D data decoding device may also indicate a desired transmission sequence to the 3D data encoding device, and the 3D data encoding device (multiplexer 4802) may control the transmission sequence according to the indicated transmission sequence.

[0271] In addition, the multiplexing unit 4802 may generate coded data incorporating a plurality of functions, as long as it is within the range of constraints of the transmission order, such as the transmission order of data integration. Fig. 20As shown, the GOF header and the AU header may be integrated, or the AXPS and the AYPS may be integrated. In this case, in pcc_nal_unit_type, an identifier indicating that the data has multiple functions is defined.

[0272] The following describes a modified example of the present embodiment. PS has levels, such as frame-level PS, sequence-level PS, and PCC sequence-level PS. If the PCC sequence level is set as the upper level and the frame level is set as the lower level, the parameter storage method can also use the following method.

[0273] It is represented by a PS that is higher than the default PS value. In addition, when the value of the lower PS is different from the value of the upper PS, the value of the PS is represented by the lower PS. Alternatively, the value of the PS is not recorded in the upper position, and the value of the PS is recorded in the lower PS. Alternatively, the value of the PS is represented by the lower PS, or by the upper PS, or by the information represented by both parties in either or both of the lower PS and the upper PS. Alternatively, the lower PS may be merged into the upper PS. Alternatively, when the lower PS and the upper PS are repeated, the multiplexing unit 4802 may omit the transmission of either party.

[0274] In addition, the encoding unit 4801 or the multiplexing unit 4802 may divide the data into slices or tiles, etc., and transmit the divided data. The divided data includes information for identifying the divided data, and the parameter set includes parameters for decoding the divided data. In this case, in pcc_nal_unit_type, an identifier indicating that the data or parameters of the tile or slice are stored is defined.

[0275] (Implementation 4)

[0276] In HEVC encoding, there are tools for data segmentation such as slices and tiles to enable parallel processing in a decoding device, but it is not PCC (Point Cloud Compression) encoding.

[0277] In PCC, various data segmentation methods are considered by means of parallel processing, compression efficiency, and compression algorithms. Here, the definition, data structure, and transmission and reception methods of slices and tiles are explained.

[0278] Fig.211 is a block diagram showing the structure of the first coding unit 4910 included in the three-dimensional data coding device of this embodiment. The first coding unit 4910 generates coded data (coded stream) by coding the point cloud data using the first coding method (GPCC (Geometry based PCC)). The first coding unit 4910 includes a segmentation unit 4911, a plurality of position information coding units 4912, a plurality of attribute information coding units 4913, an additional information coding unit 4914, and a multiplexing unit 4915.

[0279] The segmentation unit 4911 generates a plurality of segmentation data by segmenting the point group data. Specifically, the segmentation unit 4911 generates a plurality of segmentation data by segmenting the space of the point group data into a plurality of subspaces. Here, the subspace refers to one of a tile and a slice or a combination of a tile and a slice. More specifically, the point group data includes position information, attribute information, and additional information. The segmentation unit 4911 segments the position information into a plurality of segmentation position information, and segments the attribute information into a plurality of segmentation attribute information. In addition, the segmentation unit 4911 generates additional information related to the segmentation.

[0280] The plurality of position information encoding units 4912 generate a plurality of coded position information by encoding a plurality of divided position information. For example, the plurality of position information encoding units 4912 processes a plurality of divided position information in parallel.

[0281] The multiple attribute information encoding unit 4913 generates multiple coded attribute information by encoding multiple pieces of split attribute information. For example, the multiple attribute information encoding unit 4913 processes multiple pieces of split attribute information in parallel.

[0282] The additional information encoding unit 4914 generates encoded additional information by encoding the additional information included in the point group data and the additional information related to data division generated when the data is divided in the dividing unit 4911.

[0283] The multiplexing unit 4915 generates coded data (coded stream) by multiplexing a plurality of coding position information, a plurality of coding attribute information, and coded additional information, and transmits the generated coded data. The coded additional information is used at the time of decoding.

[0284] In addition, Fig.21 , an example is shown in which the number of the position information coding unit 4912 and the number of the attribute information coding unit 4913 are 2, but the number of the position information coding unit 4912 and the number of the attribute information coding unit 4913 may be 1 or 3 or more. In addition, a plurality of divided data may be processed in parallel in the same chip like a plurality of cores in a CPU, may be processed in parallel by cores of a plurality of chips, or may be processed in parallel by a plurality of cores of a plurality of chips.

[0285] Fig. 22 4920 is a block diagram showing the structure of the first decoding unit 4920. The first decoding unit 4920 restores the point group data by decoding the coded data (coded stream) generated by encoding the point group data using the first coding method (GPCC). The first decoding unit 4920 includes a demultiplexing unit 4921, a plurality of position information decoding units 4922, a plurality of attribute information decoding units 4923, an additional information decoding unit 4924, and a combining unit 4925.

[0286] The demultiplexing unit 4921 generates a plurality of encoding position information, a plurality of encoding attribute information, and encoding additional information by demultiplexing the encoded data (encoded stream).

[0287] The plurality of position information decoding units 4922 generate a plurality of divided position information by decoding a plurality of coded position information. For example, the plurality of position information decoding units 4922 processes a plurality of coded position information in parallel.

[0288] The multiple attribute information decoding unit 4923 generates multiple pieces of segmented attribute information by decoding multiple pieces of coded attribute information. For example, the multiple attribute information decoding unit 4923 processes multiple pieces of coded attribute information in parallel.

[0289] The plurality of additional information decoding units 4924 generate additional information by decoding the encoded additional information.

[0290] The combining unit 4925 generates position information by combining a plurality of pieces of segment position information using the additional information. The combining unit 4925 generates attribute information by combining a plurality of pieces of segment attribute information using the additional information.

[0291] In addition, Fig. 22 , an example is shown in which the number of position information decoding units 4922 and attribute information decoding units 4923 is two, but the number of position information decoding units 4922 and attribute information decoding units 4923 may be one, or may be three or more. In addition, a plurality of divided data may be processed in parallel in the same chip like a plurality of cores in a CPU, may be processed in parallel by cores of a plurality of chips, or may be processed in parallel by a plurality of cores of a plurality of chips.

[0292] Next, the structure of the dividing portion 4911 is described. Fig.23 4 is a block diagram of the dividing unit 4911. The dividing unit 4911 includes a slice dividing unit 4931 (Slice Divider), a position information tile dividing unit 4932 (Geometry Tile Divider), and an attribute information tile dividing unit 4933 (Attribute Tile Divider).

[0293] The slice segmentation unit 4931 generates a plurality of slice position information by segmenting the position information (Position (Geometry)) into slices. In addition, the slice segmentation unit 4931 generates a plurality of slice attribute information by segmenting the attribute information (Attribute) into slices. In addition, the slice segmentation unit 4931 outputs slice additional information (Slice MetaData) including information on the slice segmentation and information generated in the slice segmentation.

[0294] The position information tile segmentation unit 4932 generates a plurality of segmentation position information (a plurality of tile position information) by segmenting a plurality of slice position information into tiles. In addition, the position information tile segmentation unit 4932 outputs position tile additional information (GeometryTile MetaData) including information of tile segmentation of position information and information generated in tile segmentation of position information.

[0295] The attribute information tile partitioning unit 4933 generates multiple partition attribute information (multiple tile attribute information) by partitioning multiple slice attribute information into tiles. In addition, the attribute information tile partitioning unit 4933 outputs attribute tile additional information (AttributeTile MetaData) including information of tile partitioning of attribute information and information generated in tile partitioning of attribute information.

[0296] In addition, the number of divided slices or tiles is greater than 1. In other words, the division of slices or tiles may not be performed.

[0297] In addition, here, an example of performing tile division after slice division is shown, but slice division may be performed after tile division. In addition, in addition to slices and tiles, new division categories may be defined, and division may be performed with three or more division categories.

[0298] The following describes a method for segmenting point cloud data. Fig.24 The diagram shows an example of slicing and tile division.

[0299] First, the method of slice segmentation is described. The segmentation unit 4911 segments the three-dimensional point group data into arbitrary point groups in slice units. In the slice segmentation, the segmentation unit 4911 does not segment the position information and attribute information of the constituent points, but segments the position information and the attribute information together. That is, the segmentation unit 4911 performs slice segmentation in a manner that the position information and the attribute information of any point belong to the same slice. In addition, according to these methods, any number of segmentation and segmentation method are possible. In addition, the minimum unit of segmentation is a point. For example, the number of segmentation of the position information and the attribute information is the same. For example, the three-dimensional point corresponding to the position information after the slice segmentation and the three-dimensional point corresponding to the attribute information are included in the same slice.

[0300] In addition, the segmentation unit 4911 generates additional information of the number of segments and the segmentation method, i.e., slice additional information, when segmenting the slice. The slice additional information is the same in the position information and the attribute information. For example, the slice additional information includes information indicating the reference coordinate position, size, or side length of the boundary box after segmentation. In addition, the slice additional information includes information indicating the number of segments and the segmentation type, etc.

[0301] Next, a method of tile division is described: The division unit 4911 divides the data after slice division into slice position information (G slice) and slice attribute information (A slice), and divides the slice position information and the slice attribute information into tile units.

[0302] In addition, Fig.24 In FIG. 1 , an example of partitioning using an octree structure is shown, but any number of partitions and method of partitioning may be used.

[0303] In addition, the segmentation unit 4911 may segment the position information and the attribute information using different segmentation methods or the same segmentation method. In addition, the segmentation unit 4911 may segment the plurality of slices into tiles using different segmentation methods or the same segmentation method.

[0304] In addition, the segmentation unit 4911 generates tile additional information of the number of segments and the segmentation method when the tile is segmented. The tile additional information (position tile additional information and attribute tile additional information) is independent of the position information and the attribute information. For example, the tile additional information includes information indicating the reference coordinate position, size or side length of the boundary box after segmentation. In addition, the tile additional information includes information indicating the number of segments and the segmentation type, etc.

[0305] Next, an example of a method of dividing the point cloud data into slices or tiles will be described. As a method of dividing into slices or tiles, the dividing unit 4911 may use a predetermined method or may adaptively switch the method to be used according to the point cloud data.

[0306] When performing slice segmentation, the segmentation unit 4911 segments the three-dimensional space together with the position information and the attribute information. For example, the segmentation unit 4911 determines the shape of the object and divides the three-dimensional space into slices according to the shape of the object. For example, the segmentation unit 4911 extracts objects such as trees or buildings and performs segmentation in object units. For example, the segmentation unit 4911 performs slice segmentation in a manner that the entirety of one or more objects is included in one slice. Alternatively, the segmentation unit 4911 segments one object into a plurality of slices.

[0307] In this case, the encoding device may, for example, change the encoding method for each slice. For example, the encoding device may use a high-quality compression method for a specific object or a specific part of an object. In this case, the encoding device may store information indicating the encoding method for each slice in the additional information (metadata).

[0308] Furthermore, the dividing unit 4911 may perform slice division based on map information or position information so that each slice corresponds to a predetermined coordinate space.

[0309] When segmenting the tiles, the segmentation unit 4911 segments the position information and the attribute information independently. For example, the segmentation unit 4911 segments the slice into tiles according to the data volume or the processing volume. For example, the segmentation unit 4911 determines whether the data volume of the slice (for example, the number of three-dimensional points included in the slice) is greater than a predetermined threshold. The segmentation unit 4911 segments the slice into tiles when the data volume of the slice is greater than the threshold. The segmentation unit 4911 does not segment the slice into tiles when the data volume of the slice is less than the threshold.

[0310] For example, the division unit 4911 divides the slice into tiles so that the processing amount or processing time in the decoding device is within a certain range (less than a predetermined value). Thus, the processing amount per tile in the decoding device becomes constant, and distributed processing in the decoding device becomes easy.

[0311] Furthermore, when the processing amounts of the position information and the attribute information are different, for example, when the processing amount of the position information is greater than that of the attribute information, the division unit 4911 divides the position information into a greater number of parts than the attribute information.

[0312] Furthermore, for example, when the position information is decoded and displayed quickly in the decoding device according to the content, and the attribute information can be decoded and displayed slowly afterwards, the division unit 4911 may divide the position information into more parts than the attribute information. Thus, the decoding device can increase the number of parallel positions of the position information, and thus can process the position information faster than the attribute information.

[0313] Furthermore, the decoding device does not necessarily need to process sliced ​​or tiled data in parallel, and may determine whether to process them in parallel based on the number or capabilities of decoding processing units.

[0314] By performing segmentation in the above-described manner, adaptive encoding corresponding to the content or object can be achieved. In addition, parallel processing in the decoding process can be achieved. As a result, the flexibility of the point group encoding system or the point group decoding system is improved.

[0315] Fig.25This is a diagram showing an example of the style of segmentation of slices and tiles. The DU in the figure is a data unit (DataUnit), which represents the data of a tile or slice. In addition, each DU contains a slice index (SliceIndex) and a tile index (TileIndex). The value on the upper right of the DU in the figure represents the slice index, and the value on the lower left of the DU represents the tile index.

[0316] In pattern 1, in slice partitioning, the number of partitions and the partitioning method are the same in G slices and A slices. In tile partitioning, the number of partitions and the partitioning method for G slices are different from those for A slices. In addition, the same number of partitions and the partitioning method are used among multiple G slices. The same number of partitions and the partitioning method are used among multiple A slices.

[0317] In pattern 2, in slice partitioning, the number of partitions and the partitioning method are the same for G slices and A slices. In tile partitioning, the number of partitions and the partitioning method for G slices are different from the number of partitions and the partitioning method for A slices. In addition, the number of partitions and the partitioning method are different between multiple G slices. The number of partitions and the partitioning method are different between multiple A slices.

[0318] Next, the encoding method of the segmented data is described. The three-dimensional data encoding device (first encoding unit 4910) encodes the segmented data separately. When encoding the attribute information, the three-dimensional data encoding device generates dependency information indicating which structural information (position information, additional information or other attribute information) is encoded as additional information. That is, the dependency information, for example, indicates structural information of a reference target (dependent target). In this case, the three-dimensional data encoding device generates the dependency information based on the structural information corresponding to the segmented shape of the attribute information. In addition, the three-dimensional data encoding device can also generate the dependency information based on the structural information corresponding to multiple segmented shapes.

[0319] Alternatively, the dependency information may be generated by the three-dimensional data encoding device, and the generated dependency information may be sent to the three-dimensional data decoding device. Alternatively, the three-dimensional data decoding device may generate the dependency information, and the three-dimensional data encoding device may not send the dependency information. In addition, the dependency used by the three-dimensional data encoding device may be predetermined, and the three-dimensional data encoding device may not send the dependency information.

[0320] Fig.26 The figure is an example of a dependency relationship between data. The direction of the arrow in the figure indicates the dependency target, and the direction of the arrow indicates the dependency source. The three-dimensional data decoding device decodes the data in the order from the dependency target to the dependency source. In addition, the data represented by the solid line in the figure is the data actually sent, and the data represented by the dotted line is the data not sent.

[0321] In this figure, G represents position information and A represents attribute information. s1 Indicates the location information of slice number 1, G s2 Indicates the position information of slice number 2. G s1t1 Indicates the location information of slice number 1 and tile number 1, G s1t2 Indicates the location information of slice number 1 and tile number 2, G s2t1 Indicates the location information of slice number 2 and tile number 1, G s2t2 Indicates the location information of slice number 2 and tile number 2. Similarly, A s1 Indicates the attribute information of slice number 1, A s2 Indicates the attribute information of slice number 2. s1t1 Indicates the attribute information of slice number 1 and tile number 1, A s1t2 Indicates the attribute information of slice number 1 and tile number 2, A s2t1 Indicates the attribute information of slice number 2 and tile number 1, A s2t2 Indicates the attribute information of slice number 2 and tile number 2.

[0322] Mslice indicates additional information about slices, MGtile indicates additional information about position tiles, and MAtile indicates additional information about attribute tiles. s1t1 Indicates attribute information A s1t1 Dependency information of D s2t1 Indicates attribute information A s2t1 Dependency information.

[0323] In addition, the 3D data encoding device may also reorder the data in the decoding order, so that the 3D data decoding device does not need to reorder the data. In addition, the data may be reordered in the 3D data decoding device, or the data may be reordered in both the 3D data encoding device and the 3D data decoding device.

[0324] Fig. 27 is a diagram showing an example of the order in which data is decoded. Fig. 27 In the example, the decoding is performed sequentially from the data on the left. The three-dimensional data decoding device decodes the data of the dependent target first among the data in a dependent relationship. For example, the three-dimensional data encoding device pre-arranges the data in such a way as to achieve the order and sends it. In addition, any order is acceptable as long as the order in which the data of the dependent target comes first. In addition, the three-dimensional data encoding device may also send the additional information and the dependency information before the data.

[0325] Fig.284 is a flowchart showing the process flow of the three-dimensional data encoding device. First, the three-dimensional data encoding device encodes the data of a plurality of slices or tiles as described above (S4901). Fig. 27 As shown, the 3D data encoding device reorders the data in such a way that the data dependent on the target is first (S4902). Next, the 3D data encoding device multiplexes (NAL unitizes) the reordered data (S4903).

[0326] Next, the structure of the combining unit 4925 included in the first decoding unit 4920 is described. Fig.29 4 is a block diagram showing the structure of the combining unit 4925. The combining unit 4925 includes a position information tile combining unit 4941 (Geometry Tile Combiner), an attribute information tile combining unit 4942 (Attribute Tile Combiner), and a slice combining unit (Slice Combiner).

[0327] The position information tile combining unit 4941 generates a plurality of slice position information by combining a plurality of partition position information using the position tile additional information. The attribute information tile combining unit 4942 generates a plurality of slice attribute information by combining a plurality of partition attribute information using the attribute tile additional information.

[0328] The slice combining unit 4943 generates position information by combining a plurality of slice position information using the slice additional information. In addition, the slice combining unit 4943 generates attribute information by combining a plurality of slice attribute information using the slice additional information.

[0329] In addition, the number of divided slices or tiles is greater than 1. In other words, the division of slices or tiles may not be performed.

[0330] In addition, here, an example of performing tile division after slice division is shown, but slice division may be performed after tile division. In addition, in addition to slices and tiles, new division categories may be defined, and division may be performed with three or more division categories.

[0331] Next, the structure of the coded data after slice division or tile division and the method of storing the coded data in the NAL unit (multiplexing method) are described. Fig.30 This diagram shows the structure of coded data and the method of storing coded data in NAL units.

[0332] The coded data (segmentation position information and segmentation attribute information) is stored in the payload of the NAL unit.

[0333] The encoded data includes a header and a payload. The header includes identification information for determining the data contained in the payload. The identification information includes, for example, the type of slice segmentation or tile segmentation (slice_type, tile_type), index information for determining the slice or tile (slice_idx, tile_idx), location information of the data (slice or tile), or the address (address) of the data. The index information for determining the slice is also recorded as a slice index (SliceIndex). The index information for determining the tile is also recorded as a tile index (TileIndex). In addition, the type of segmentation is, for example, a method based on the shape of the object as described above, a method based on map information or location information, or a method based on the amount of data or the amount of processing.

[0334] In addition, all or part of the above information may be stored in one of the headers of the segmented position information and the headers of the segmented attribute information, but not in the other. For example, when the same segmentation method is used in the position information and the attribute information, the segmentation category (slice_type, tile_type) and index information (slice_idx, tile_idx) in the position information and the attribute information are the same. Therefore, this information may also be included in the header of one of the position information and the attribute information. For example, when the attribute information depends on the position information, the position information is processed first. Therefore, this information may also be included in the header of the position information, but not in the header of the attribute information. In this case, the three-dimensional data decoding device, for example, determines that the attribute information of the dependent source belongs to the same slice or tile as the slice or tile of the position information of the dependent target.

[0335] In addition, additional information for slice segmentation or tile segmentation (slice additional information, position tile additional information or attribute tile additional information), and dependency information indicating dependencies, etc. may also be stored in an existing parameter set (GPS, APS, position SPS or attribute SPS, etc.) and sent. In the case where the segmentation method changes for each frame, information indicating the segmentation method may also be stored in the parameter set (GPS or APS, etc.) of each frame. In the case where the segmentation method does not change within a sequence, information indicating the segmentation method may also be stored in the parameter set (position SPS or attribute SPS) of each sequence. Furthermore, in the case where the same segmentation method is used in the position information and the attribute information, information indicating the segmentation method may also be stored in the parameter set (stream PS) of the PCC stream.

[0336] In addition, the above information can be stored in any of the above parameter sets, or in multiple parameter sets. In addition, a parameter set for tile segmentation or slice segmentation can also be defined, and the above information can be stored in the parameter set. In addition, this information can also be stored in the header of the encoded data.

[0337] In addition, the header of the encoded data contains identification information indicating the dependency relationship. That is, when there is a dependency relationship between the data, the header contains identification information for referencing the dependent target from the dependent source. For example, the header of the data of the dependent target contains identification information for determining the data. The header of the data of the dependent source contains identification information indicating the dependent target. In addition, in the case where the identification information for determining the data, the additional information of the slice segmentation or tile segmentation, and the identification information indicating the dependency relationship can be identified or derived from other information, this information can also be omitted.

[0338] Next, the flow of the encoding process and the decoding process of the point cloud data according to the present embodiment will be described. Fig.31 This is a flowchart of the encoding process of the point cloud data according to the present embodiment.

[0339] First, the three-dimensional data encoding device determines the segmentation method to be used (S4911). The segmentation method includes whether to perform slice segmentation and whether to perform tile segmentation. In addition, the segmentation method may also include the number of segmentations when performing slice segmentation or tile segmentation, and the type of segmentation. The type of segmentation is a method based on the shape of the object as described above, a method based on map information or location information, or a method based on the amount of data or the amount of processing. In addition, the segmentation method may also be predetermined.

[0340] When slice segmentation is performed (Yes in S4912), the three-dimensional data encoding device generates a plurality of slice position information and a plurality of slice attribute information by segmenting the position information and the attribute information together (S4913). In addition, the three-dimensional data encoding device generates slice additional information of the slice segmentation. In addition, the three-dimensional data encoding device may also segment the position information and the attribute information independently.

[0341] When tile segmentation is performed ("Yes" in S4914), the three-dimensional data encoding device generates a plurality of segmentation position information and a plurality of segmentation attribute information (S4915) by independently segmenting a plurality of slice position information and a plurality of slice attribute information (or position information and attribute information). In addition, the three-dimensional data encoding device generates position tile additional information and attribute tile additional information for tile segmentation. In addition, the three-dimensional data encoding device may also segment the slice position information and the slice attribute information together.

[0342] Next, the three-dimensional data encoding device generates a plurality of coded position information and a plurality of coded attribute information by encoding the plurality of segmentation position information and the plurality of segmentation attribute information respectively (S4916). In addition, the three-dimensional data encoding device generates dependency relationship information.

[0343] Next, the three-dimensional data encoding device generates encoded data (encoded stream) by NAL unitizing (multiplexing) the plurality of encoding position information, the plurality of encoding attribute information, and the additional information (S4917). In addition, the three-dimensional data encoding device transmits the generated encoded data.

[0344] Fig.32 4921 is a flowchart of the decoding process of the point cloud data of the present embodiment. First, the three-dimensional data decoding device determines the segmentation method by parsing the additional information (slice additional information, position tile additional information and attribute tile additional information) of the segmentation method contained in the coded data (coded stream). The segmentation method includes whether to perform slice segmentation and whether to perform tile segmentation. In addition, the segmentation method may also include the number of segmentations in the case of performing slice segmentation or tile segmentation, and the type of segmentation.

[0345] Next, the three-dimensional data decoding apparatus generates segmentation position information and segmentation attribute information by decoding a plurality of pieces of encoding position information and a plurality of pieces of encoding attribute information included in the encoded data using the dependency information included in the encoded data (S4922).

[0346] When tile segmentation is indicated by the additional information ("Yes" in S4923), the three-dimensional data decoding device combines the plurality of segmentation position information and the plurality of segmentation attribute information in respective methods based on the position tile additional information and the attribute tile additional information, thereby generating a plurality of slice position information and a plurality of slice attribute information (S4924). In addition, the three-dimensional data decoding device may also combine the plurality of segmentation position information and the plurality of segmentation attribute information in the same method.

[0347] When the additional information indicates that the slice segmentation is performed ("Yes" in S4925), the three-dimensional data decoding device combines the plurality of slice position information and the plurality of slice attribute information (the plurality of segmentation position information and the plurality of segmentation attribute information) in the same method based on the slice additional information, thereby generating the position information and the attribute information (S4926). Alternatively, the three-dimensional data decoding device may combine the plurality of slice position information and the plurality of slice attribute information in different methods.

[0348] As described above, the three-dimensional data encoding device of this embodiment performs Fig.33The processing shown. First, the three-dimensional data encoding device is divided into a plurality of segmentation data (such as tiles) each containing one or more three-dimensional points, and the plurality of segmentation data are contained in a plurality of subspaces (such as slices) after dividing the object space containing the plurality of three-dimensional points. Here, the segmentation data is contained in the subspace, which is one or more data sets containing one or more three-dimensional points. In addition, the segmentation data may also be a space, or may include a space that does not contain three-dimensional points. In addition, multiple segmentation data may be included in one subspace, or one segmentation data may be included in one subspace. In addition, multiple subspaces may be set in the object space, or one subspace may be set in the object space.

[0349] Next, the three-dimensional data encoding device generates a plurality of coded data corresponding to each of the plurality of segmented data by encoding each of the plurality of segmented data (S4931). The three-dimensional data encoding device generates a plurality of coded data and a plurality of control information for each of the plurality of coded data (e.g. Fig.30 In each of the plurality of control information, a first identifier (e.g., slice_idx) indicating a subspace corresponding to the coded data corresponding to the control information and a second identifier (e.g., tile_idx) indicating segmented data corresponding to the coded data corresponding to the control information are stored.

[0350] Thus, a 3D data decoding device that decodes a bit stream generated by a 3D data encoding device can easily restore the object space by combining data of a plurality of segmented data using the first identifier and the second identifier, thereby reducing the amount of processing in the 3D data decoding device.

[0351] For example, in the encoding, the three-dimensional data encoding device encodes the position information and attribute information of the three-dimensional point contained in each of the plurality of segmented data. Each of the plurality of encoded data contains encoded data of the position information and encoded data of the attribute information. Each of the plurality of control information contains control information of the encoded data of the position information and control information of the encoded data of the attribute information. The first identifier and the second identifier are stored in the control information of the encoded data of the position information.

[0352] For example, in a bit stream, a plurality of control information are respectively arranged before the coded data corresponding to the control information.

[0353] Alternatively, the three-dimensional data encoding device may set an object space including a plurality of three-dimensional points as one or more subspaces, wherein the subspace includes one or more segmented data including one or more three-dimensional points, encodes each of the segmented data to thereby generate a plurality of encoded data corresponding to each of the plurality of segmented data, generates a bit stream including the plurality of encoded data and a plurality of control information for each of the plurality of encoded data, and stores in each of the plurality of control information a first identifier representing the subspace corresponding to the encoded data corresponding to the control information, and a second identifier representing the segmented data corresponding to the encoded data corresponding to the control information.

[0354] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.

[0355] In addition, the three-dimensional data decoding device of this embodiment performs Fig.34 First, the three-dimensional data decoding device obtains a plurality of coded data and a plurality of control information (eg Fig.30 The 3D data decoding device obtains a bit stream of a header (shown in FIG. 4 ), obtains a first identifier (e.g., slice_idx) and a second identifier (e.g., tile_idx) stored in the plurality of control information (S4941), the plurality of coded data being generated by encoding each of a plurality of segmented data (e.g., tiles) each containing one or more three-dimensional points in a plurality of subspaces (e.g., slices) after the object space containing a plurality of three-dimensional points is divided, and the plurality of control information is for each of the plurality of coded data, the first identifier indicating the subspace corresponding to the coded data corresponding to the control information, and the second identifier indicating the segmented data corresponding to the coded data corresponding to the control information. Next, the 3D data decoding device restores the plurality of segmented data by decoding the plurality of coded data (S4942). Next, the 3D data decoding device restores the object space by combining the plurality of segmented data using the first identifier and the second identifier (S4943). For example, the 3D data coding device restores the plurality of subspaces by combining the plurality of segmented data using the second identifier, and restores the object space (a plurality of three-dimensional points) by combining the plurality of subspaces using the first identifier. Furthermore, the three-dimensional data decoding device may obtain the encoded data of the desired subspace or divided data from the bit stream using at least one of the first identifier and the second identifier, and selectively decode or preferentially decode the obtained encoded data.

[0356] Thus, the three-dimensional data decoding device can easily restore the target space by combining the data of a plurality of divided data using the first identifier and the second identifier, thereby reducing the amount of processing in the three-dimensional data decoding device.

[0357] For example, each of the plurality of coded data is generated by encoding the position information and the attribute information of the three-dimensional point included in the corresponding segmented data, and includes the coded data of the position information and the coded data of the attribute information. The plurality of control information includes the control information of the coded data of the position information and the control information of the coded data of the attribute information, respectively. The first identifier and the second identifier are stored in the control information of the coded data of the position information.

[0358] For example, in a bitstream, the control information is arranged before the corresponding coded data.

[0359] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.

[0360] (Implementation 5)

[0361] In the position information encoding using adjacent dependency, there is a possibility that the higher the density of the point group, the higher the encoding efficiency. In this embodiment, the three-dimensional data encoding device combines the point group data of consecutive frames to encode the point group data of consecutive frames. At this time, the three-dimensional data encoding device generates encoded data, and the encoded data is added with information for identifying the frames to which the leaf nodes contained in the combined point group data belong.

[0362] Here, the point cloud data of the continuous frames are likely to be similar. Therefore, in the continuous frames, the upper level of the occupancy rate coding is likely to be the same. That is, by collectively coding the continuous frames, the upper level of the occupancy rate coding can be shared.

[0363] Furthermore, the identification of which frame a point group belongs to is performed in the leaf node by encoding the index of the frame.

[0364] Fig.35 This is a diagram showing an image of a tree structure and occupancy code generated from point cloud data of N PCC (Point Cloud Compression) frames. In this figure, the points in the arrows indicate the points belonging to each PCC frame. First, a frame index for identifying the frame is assigned to the points belonging to each PCC frame.

[0365] Next, the points belonging to N frames are transformed into a tree structure to generate occupancy codes. Specifically, for each point, it is determined which leaf node in the tree structure the point belongs to. In this figure, the tree structure shows a set of nodes. From the upper node, it is determined which node the point belongs to. The determination result of each node is encoded as an occupancy code. The occupancy code is common in N frames.

[0366] In a node, points of different frames assigned with different frame indices may exist in a mixed state. In addition, when the resolution of the octree is small, points of the same frame assigned with the same frame indices may exist in a mixed state.

[0367] Sometimes, points belonging to a plurality of frames are mixed (repeated) in the nodes (leaf nodes) at the lowest level.

[0368] In the tree structure and occupancy coding, the upper-level tree structure and occupancy coding may be common components in all frames, and the lower-level tree structure and occupancy coding may be individual components of each frame, or common components and individual components may be mixed.

[0369] For example, points with frame indices of 0 or more are generated in the lowest level nodes such as leaf nodes, and information indicating the number of points and frame index information for each point are generated. These information can also be said to be individual information in the frame.

[0370] Fig.36 is a diagram showing an example of frame combination. Fig.36 As shown in (a), by aggregating multiple frames and generating a tree structure, the density of points of the frames contained in the same node increases. In addition, by sharing the tree structure, the amount of data for occupancy coding can be reduced. As a result, there is a possibility of increasing the coding rate.

[0371] In addition, if Fig.36 As shown in (b) of FIG. 1 , the individual components of the occupancy coding in the tree structure become denser, thereby improving the effect of arithmetic coding, so there is a possibility that the coding rate can be increased.

[0372] The following description is made by taking the combination of multiple PCC frames that are different in time as an example, but it can also be applied when it is not a multiple frame, that is, when the frame is not combined (N=1). In addition, the multiple point group data to be combined are not limited to multiple frames, that is, they are not limited to point group data of the same object at different times. That is, the following method can also be applied to the combination of multiple point group data that are different in space or time and space. In addition, the following method can also be applied to the combination of point group data or point group files with different contents.

[0373] Fig.37 This is a diagram showing an example of combining a plurality of PCC frames that are different in time. Fig.37 This shows an example of a car moving while acquiring point group data using a sensor such as LiDAR. The dotted line shows the acquisition range of the sensor for each frame, i.e., the area of ​​the point group data. When the acquisition range of the sensor is large, the range of the point group data also becomes large.

[0374] The method of combining point group data for encoding is effective for point group data such as the following. Fig.37 In the example shown, the car moves, and the frame is recognized by scanning 360° around the car. That is, frame 2, which is the next frame, corresponds to another 360° scan after the vehicle moves in the X direction.

[0375] In this case, since there are repeated areas in frame 1 and frame 2, there is a possibility that the same point group data is included. Therefore, by combining frame 1 and frame 2 for encoding, there is a possibility that the encoding efficiency can be improved. In addition, combining more frames is also considered. However, if the number of frames to be combined is increased, the number of bits required for encoding the frame index attached to the leaf node increases.

[0376] In addition, point group data may be obtained by sensors at different positions. Thus, each point group data obtained from each position may be used as a frame. That is, multiple frames may be point group data obtained by a single sensor or point group data obtained by multiple sensors. In addition, between multiple frames, some or all of the objects may be the same or different.

[0377] Next, the flow of the three-dimensional data encoding process according to the present embodiment will be described. Fig.38 3D data encoding processing flow chart The 3D data encoding device reads point cloud data of all N frames based on the number of combined frames N which is the number of frames to be combined.

[0378] First, the three-dimensional data encoding device determines the number of combined frames N (S5401). For example, the number of combined frames N is specified by the user.

[0379] Next, the three-dimensional data encoding device obtains point cloud data (S5402) and then records the frame index of the obtained point cloud data (S5403).

[0380] If N frames have not been processed (No in S5404), the three-dimensional data encoding device specifies the next point cloud data (S5405), and performs the processing after step S5402 on the specified point cloud data.

[0381] On the other hand, when N frames have been processed (Yes in S5404), the three-dimensional data encoding device combines the N frames and encodes the combined frames (S5406).

[0382] Fig.39 4 is a flowchart of the encoding process (S5406). First, the three-dimensional data encoding device generates common information common to N frames (S5411). For example, the common information includes information indicating the occupancy encoding and the number of combined frames N.

[0383] Next, the three-dimensional data encoding device generates individual information, that is, individual information for each frame (S5412). For example, the individual information includes the number of points included in the leaf node and the frame index of the point included in the leaf node.

[0384] Next, the 3D data encoding device combines the common information with the individual information and encodes the combined information to generate encoded data (S5413). Next, the 3D data encoding device generates frame-combined additional information (metadata) and encodes the generated additional information (S5414).

[0385] Next, the flow of three-dimensional data decoding processing according to this embodiment will be described. Fig.40 3D data decoding process.

[0386] First, the three-dimensional data decoding device obtains the combined frame number N from the bit stream (S5421). Then, the three-dimensional data encoding device obtains the encoded data from the bit stream (S5422). Then, the three-dimensional data decoding device obtains the point group data and the frame index by decoding the encoded data (S5423). Finally, the three-dimensional data decoding device uses the frame index to segment the decoded point group data (S5424).

[0387] Fig.41 1 is a flowchart of the decoding and segmentation processing (S5423 and S5424). First, the three-dimensional data decoding apparatus decodes (obtains) common information and individual information from the encoded data (bit stream) (S5431).

[0388] Next, the three-dimensional data decoding device determines whether to decode a single frame or a plurality of frames (S5432). For example, it can be specified externally whether to decode a single frame or a plurality of frames. Here, the plurality of frames can be all frames of the combined frames or a portion of the frames. For example, the three-dimensional data decoding device can also determine to decode specific frames that require an application and not to decode unnecessary frames. Alternatively, in the case of requiring real-time decoding, the three-dimensional data decoding device can also determine to decode a single frame of the combined plurality of frames.

[0389] When decoding a single frame ("Yes" in S5432), the three-dimensional data decoding device extracts individual information corresponding to the specified single frame index from the decoded individual information, decodes the extracted individual information, and thereby restores the point group data of the frame corresponding to the specified frame index (S5433).

[0390] On the other hand, when decoding a plurality of frames ("No" in S5432), the three-dimensional data decoding device extracts individual information corresponding to the frame index of the specified plurality of frames (or all frames), decodes the extracted individual information, and thereby restores the point cloud data of the specified plurality of frames (S5434). Next, the three-dimensional data decoding device divides the decoded point cloud data (individual information) based on the frame index (S5435). That is, the three-dimensional data decoding device divides the decoded point cloud data into a plurality of frames.

[0391] Furthermore, the three-dimensional data decoding device may decode the data of all the combined frames at once and divide the decoded data into individual frames, or may decode any part of all the combined frames at once and divide the decoded data into individual frames. Furthermore, the three-dimensional data decoding device may also individually decode a predetermined unit frame composed of a plurality of frames.

[0392] Hereinafter, the configuration of the three-dimensional data encoding device according to this embodiment will be described. Fig.42 1 is a block diagram showing the structure of the encoding unit 5410 included in the three-dimensional data encoding device of this embodiment. The encoding unit 5410 generates encoded data (encoded stream) by encoding point group data (point cloud). The encoding unit 5410 includes a segmentation unit 5411, multiple position information encoding units 5412, multiple attribute information encoding units 5413, an additional information encoding unit 5414, and a multiplexing unit 5415.

[0393] The segmentation unit 5411 generates a plurality of segmentation data of a plurality of frames by segmenting the point group data of a plurality of frames. Specifically, the segmentation unit 5411 generates a plurality of segmentation data by segmenting the space of the point group data of each frame into a plurality of subspaces. Here, the subspace refers to one of a tile and a slice or a combination of a tile and a slice. More specifically, the point group data includes position information, attribute information (color or reflectivity, etc.) and additional information. In addition, the frame number is input to the segmentation unit 5411. The segmentation unit 5411 segments the position information of each frame into a plurality of segmentation position information, and segments the attribute information of each frame into a plurality of segmentation attribute information. In addition, the segmentation unit 5411 generates additional information related to the segmentation.

[0394] For example, the segmentation unit 5411 first segments the point group into tiles. Next, the segmentation unit 5411 further segments the obtained tiles into slices.

[0395] The multiple position information encoding unit 5412 generates multiple encoded position information by encoding multiple segmented position information. For example, the position information encoding unit 5412 uses an N-ary tree structure such as an octree to encode the segmented position information. Specifically, in the octree, the object space is divided into 8 nodes (subspaces), and 8 bits of information (occupancy code) are generated to indicate whether each node contains a point group. In addition, the node containing the point group is also divided into 8 nodes, and 8 bits of information are generated to indicate whether each of the 8 nodes contains a point group. This process is repeated until it becomes below the threshold of the number of point groups contained in a predetermined layer or node. For example, the multiple position information encoding unit 5412 processes multiple segmented position information in parallel.

[0396] The attribute information encoding unit 4632 generates the encoded attribute information as the encoded data by encoding using the structure information generated by the position information encoding unit 4631. For example, the attribute information encoding unit 4632 determines the reference point (reference node) to be referred to in the encoding of the object point (object node) of the processing object based on the octree structure generated by the position information encoding unit 4631. For example, the attribute information encoding unit 4632 refers to a node whose parent node in the octree is the same as the parent node of the object node among the surrounding nodes or adjacent nodes. In addition, the method of determining the reference relationship is not limited to this.

[0397] In addition, the encoding process of the position information or attribute information may also include at least one of a quantization process, a prediction process, and an arithmetic coding process. In this case, the reference means that the reference node is used in the calculation of the predicted value of the attribute information, or the state of the reference node is used in the determination of the encoding parameter (for example, indicating whether the occupancy information of the point group is included in the reference node). For example, the encoding parameter refers to the quantization parameter in the quantization process, or the context in the arithmetic coding, etc.

[0398] The multiple attribute information encoding unit 5413 generates multiple coded attribute information by encoding multiple pieces of segmented attribute information. For example, the multiple attribute information encoding unit 5413 processes multiple pieces of segmented attribute information in parallel.

[0399] The additional information encoding unit 5414 generates encoded additional information by encoding the additional information included in the point cloud data and the additional information related to data division generated when the data is divided by the dividing unit 5411.

[0400] The multiplexing unit 5415 generates coded data (coded stream) by multiplexing a plurality of coding position information, a plurality of coding attribute information, and coded additional information of a plurality of frames, and transmits the generated coded data. The coded additional information is used at the time of decoding.

[0401] Fig.435411 is a block diagram of the partitioning unit 5411. The partitioning unit 5411 includes a tile partitioning unit 5421 and a slice partitioning unit 5422.

[0402] The tile segmentation unit 5421 generates a plurality of tile position information by segmenting the position information (Position (Geometry)) of a plurality of frames into tiles. In addition, the tile segmentation unit 5421 generates a plurality of tile attribute information by segmenting the attribute information (Attribute) of a plurality of frames into tiles. In addition, the tile segmentation unit 5421 outputs tile additional information (Tile MetaData) including information on tile segmentation and information generated in tile segmentation.

[0403] The slice segmentation unit 5422 generates a plurality of segmentation position information (a plurality of slice position information) by segmenting a plurality of tile position information into slices. In addition, the slice segmentation unit 5422 generates a plurality of segmentation attribute information (a plurality of slice attribute information) by segmenting a plurality of tile attribute information into slices. In addition, the slice segmentation unit 5422 outputs slice additional information (Slice MetaData) including information on the slice segmentation and information generated in the slice segmentation.

[0404] In addition, the division unit 5411 uses a frame number (frame index) in order to indicate the origin coordinates and attribute information, etc. during the division process.

[0405] Fig.44 54 is a block diagram of the position information encoding unit 5412. The position information encoding unit 5412 includes a frame index generation unit 5431 and an entropy encoding unit 5432.

[0406] The frame index generation unit 5431 determines the value of the frame index based on the frame number, and adds the determined frame index to the position information. The entropy coding unit 5432 generates the coded position information by entropy coding the segmentation position information to which the frame index is added.

[0407] Fig.45 5413 is a block diagram of the attribute information encoding unit 5413. The attribute information encoding unit 5413 includes a frame index generation unit 5441 and an entropy encoding unit 5442.

[0408] The frame index generation unit 5441 determines the value of the frame index based on the frame number, and adds the determined frame index to the attribute information. The entropy coding unit 5442 generates coded attribute information by entropy coding the segmentation attribute information to which the frame index is added.

[0409] Next, the flow of the encoding process and the decoding process of the point cloud data according to the present embodiment will be described. Fig.46 This is a flowchart of the encoding process of the point cloud data according to the present embodiment.

[0410] First, the three-dimensional data encoding device determines the segmentation method to be used (S5441). The segmentation method includes whether to perform slice segmentation or tile segmentation. In addition, the segmentation method may also include the number of segmentations when performing slice segmentation or tile segmentation, and the type of segmentation.

[0411] When tile segmentation is performed (Yes in S5442), the three-dimensional data encoding device generates a plurality of tile position information and a plurality of tile attribute information by segmenting the position information and the attribute information (S5443). In addition, the three-dimensional data encoding device generates tile additional information of the tile segmentation.

[0412] When slice segmentation is performed ("Yes" in S5444), the three-dimensional data encoding device generates a plurality of segmentation position information and a plurality of segmentation attribute information (S5445) by segmenting a plurality of tile position information and a plurality of tile attribute information (or position information and attribute information). In addition, the three-dimensional data encoding device generates slice additional information of the slice segmentation.

[0413] Next, the three-dimensional data encoding device generates a plurality of encoding position information and a plurality of encoding attribute information by encoding the plurality of segmentation position information and the plurality of segmentation attribute information with the frame index respectively (S5446). In addition, the three-dimensional data encoding device generates dependency relationship information.

[0414] Next, the three-dimensional data encoding device generates encoded data (encoded stream) by NAL unitizing (multiplexing) the plurality of encoding position information, the plurality of encoding attribute information, and the additional information (S5447). In addition, the three-dimensional data encoding device transmits the generated encoded data.

[0415] Fig.47 1 is a flowchart of the encoding process (S5446). First, the three-dimensional data encoding device encodes the segmentation position information (S5451). Next, the three-dimensional data encoding device encodes the frame index for the segmentation position information (S5452).

[0416] When the segmentation attribute information exists ("Yes" in S5453), the three-dimensional data encoding device encodes the segmentation attribute information (S5454) and encodes the frame index used for the segmentation attribute information (S5455). On the other hand, when the segmentation attribute information does not exist ("No" in S5453), the three-dimensional data encoding device does not encode the segmentation attribute information and the frame index used for the segmentation attribute information. In addition, the frame index may be stored in either or both of the segmentation position information and the segmentation attribute information.

[0417] In addition, the 3D data encoding device may encode the attribute information using a frame index or without using a frame index. That is, the 3D data encoding device may use a frame index to identify the frame to which each point belongs and encode each frame, or may encode points belonging to all frames without identifying the frame.

[0418] The following describes the configuration of the three-dimensional data decoding device according to this embodiment. Fig.48 5450 is a block diagram showing the structure of the decoding unit 5450. The decoding unit 5450 restores the point group data by decoding the coded data (coded stream) generated by encoding the point group data. The decoding unit 5450 includes a demultiplexing unit 5451, a plurality of position information decoding units 5452, a plurality of attribute information decoding units 5453, an additional information decoding unit 5454, and a combining unit 5455.

[0419] The demultiplexing unit 5451 generates a plurality of encoding position information, a plurality of encoding attribute information, and encoding additional information by demultiplexing the encoded data (encoded stream).

[0420] The plurality of position information decoding units 5452 generate a plurality of divided position information by decoding a plurality of coded position information. For example, the plurality of position information decoding units 5452 processes a plurality of coded position information in parallel.

[0421] The multiple attribute information decoding unit 5453 generates multiple pieces of split attribute information by decoding multiple pieces of coded attribute information. For example, the multiple attribute information decoding unit 5453 processes multiple pieces of coded attribute information in parallel.

[0422] The plurality of additional information decoding units 5454 generate additional information by decoding the encoded additional information.

[0423] The combining unit 5455 generates the position information by combining the plurality of divided position information using the additional information. The combining unit 5455 generates the attribute information by combining the plurality of divided attribute information using the additional information. In addition, the combining unit 5455 divides the position information and the attribute information into the position information of a plurality of frames and the attribute information of a plurality of frames using the frame index.

[0424] Fig.49 : is a block diagram of the position information decoding unit 5452. The position information decoding unit 5452 includes an entropy decoding unit 5461 and a frame index obtaining unit 5462. The entropy decoding unit 5461 generates segmentation position information by entropy decoding the encoding position information. The frame index obtaining unit 5462 obtains a frame index from the segmentation position information.

[0425] Fig.5054 is a block diagram of the attribute information decoding unit 5453. The attribute information decoding unit 5453 includes an entropy decoding unit 5471 and a frame index obtaining unit 5472. The entropy decoding unit 5471 generates segmentation attribute information by entropy decoding the encoded attribute information. The frame index obtaining unit 5472 obtains a frame index from the segmentation attribute information.

[0426] Fig.51 5455 is a diagram showing the structure of the combining unit 5455. The combining unit 5455 generates position information by combining a plurality of divided position information. The combining unit 5455 generates attribute information by combining a plurality of divided attribute information. In addition, the combining unit 5455 divides the position information and the attribute information into position information of a plurality of frames and attribute information of a plurality of frames using a frame index.

[0427] Fig.52 Flowchart of the decoding process of the point cloud data of the present embodiment. First, the three-dimensional data decoding device determines the segmentation method (S5461) by parsing the additional information (slice additional information and tile additional information) of the segmentation method contained in the coded data (coded stream). The segmentation method includes whether to perform slice segmentation and whether to perform tile segmentation. In addition, the segmentation method may also include the number of segmentations when performing slice segmentation or tile segmentation, and the type of segmentation.

[0428] Next, the three-dimensional data decoding apparatus decodes the plurality of pieces of encoding position information and the plurality of pieces of encoding attribute information contained in the encoded data using the dependency information contained in the encoded data, thereby generating segmentation position information and segmentation attribute information (S5462).

[0429] When the additional information indicates that the slice segmentation is performed ("Yes" in S5463), the three-dimensional data decoding device generates a plurality of tile position information by combining a plurality of segmentation position information and generates a plurality of tile attribute information by combining a plurality of segmentation attribute information based on the slice additional information (S5464). Here, the plurality of segmentation position information, the plurality of segmentation attribute information, the plurality of tile position information, and the plurality of tile attribute information include a frame index.

[0430] When the additional information indicates that tile segmentation has been performed ("Yes" in S5465), the three-dimensional data decoding device generates position information by combining a plurality of tile position information (a plurality of segmentation position information) based on the tile additional information, and generates attribute information by combining a plurality of tile attribute information (a plurality of segmentation attribute information) (S5466). Here, the plurality of tile position information, the plurality of tile attribute information, the position information, and the attribute information include a frame index.

[0431] Fig.531 is a flowchart of the decoding process (S5464 or S5466). First, the three-dimensional data decoding device decodes the segmentation position information (slice position information) (S5471). Next, the three-dimensional data decoding device decodes the frame index used for the segmentation position information (S5472).

[0432] If the segmentation attribute information exists ("Yes" in S5473), the three-dimensional data decoding device decodes the segmentation attribute information (S5474) and decodes the frame index used for the segmentation attribute information (S5475). On the other hand, if the segmentation attribute information does not exist ("No" in S5473), the three-dimensional data decoding device does not decode the segmentation attribute information and the frame index used for the segmentation attribute information.

[0433] Furthermore, the three-dimensional data decoding apparatus may decode the attribute information using a frame index or may decode the attribute information without using a frame index.

[0434] The following describes the coding unit in frame combination. Fig.54 2 is a diagram showing an example of a frame combination pattern. The example in this diagram is an example of a case where PCC frames are time-series and data generation and encoding are performed in real time.

[0435] Fig.54 (a) shows a case where four frames are fixedly combined. The three-dimensional data encoding device generates the encoded data after waiting for the generation of data for four frames.

[0436] Fig.54 (b) shows a case where the number of frames is adaptively changed. For example, the three-dimensional data encoding device changes the number of combined frames in order to adjust the code amount of the encoded data in rate control.

[0437] Furthermore, the three-dimensional data encoding device may not combine frames when there is a possibility that there will be no effect due to combining frames. Furthermore, the three-dimensional data encoding device may switch between combining frames and not combining frames.

[0438] Fig.54 (c) is an example of a case where a part of a plurality of frames to be combined overlaps a part of a plurality of frames to be combined next. This example is useful when real-time performance or low delay is required, such as sequential transmission from frames that can be encoded.

[0439] Fig.55 3D data encoding apparatus may also be configured to include at least a data unit capable of decoding the combined frame individually. Fig.55As shown in (a) of FIG. 1 , when all PCC frames are intra-coded and the PCC frames can be decoded individually, any of the above-mentioned patterns can be applied.

[0440] In addition, if Fig.55 As shown in (b), when a random access unit such as GOF (group of frame) is set when inter-frame prediction is applied, the three-dimensional data encoding device can also combine data using the GOF unit as the minimum unit.

[0441] Furthermore, the three-dimensional data encoding device may encode the common information and the individual information together or separately. Furthermore, the three-dimensional data encoding device may use a common data structure for the common information and the individual information or use different data structures.

[0442] Alternatively, after generating the occupancy code for each frame, the three-dimensional data encoding device may compare the occupancy codes of the plurality of frames, for example, determine whether there are many common parts between the occupancy codes of the plurality of frames based on a predetermined reference, and generate common information if there are many common parts. Alternatively, the three-dimensional data encoding device may determine whether to combine frames, which frames to combine, or the number of frames to combine based on whether there are many common parts.

[0443] Next, the structure of the encoded position information is described. Fig.56 This is a diagram showing the structure of the coded position information. The coded position information includes a header and a payload.

[0444] Fig.57 The diagram shows an example of the syntax of the header (Geometry_header) of the coded position information. The header of the coded position information includes a GPS index (gps_idx), offset information (offset), other information (other_geometry_information), a frame combine flag (combine_frame_flag), and the number of combined frames (number_of_combine_frame).

[0445] The GPS index represents the identifier (ID) of the parameter set (GPS) corresponding to the encoded location information. GPS is a parameter set for encoding location information for one or more frames. In addition, in the case where a parameter set exists in each frame, the identifiers of multiple parameter sets can also be represented in the header.

[0446] The offset information indicates an offset position for obtaining the combined data. The other information indicates other information about the position information (for example, a differential value of a quantization parameter (QPdelta), etc.). The frame combination flag is a flag indicating whether the encoded data is frame combined. The combined frame number indicates the number of combined frames.

[0447] In addition, part or all of the above information may be recorded in SPS or GPS. In addition, SPS is a parameter set in sequence (a plurality of frames) units, and is a parameter set commonly used in encoding position information and encoding attribute information.

[0448] Fig.58 This is a diagram showing a syntax example of the payload (Geometry_data) of the coded position information. The payload of the coded position information includes common information and leaf node information.

[0449] The common information is data obtained by combining more than one frame, and includes occupancy code (occupancy_Code) and the like.

[0450] The leaf node information (combine_information) is information of each leaf node. The leaf node information may be displayed for each frame as a loop of the number of frames.

[0451] As a method of expressing the frame index of a point included in a leaf node, either method 1 or method 2 can be used. Fig.59 This is a diagram showing an example of leaf node information in the case of method 1. Fig.59 The leaf node information shown includes a three-dimensional point number (NumberOfPoints) indicating the number of points included in the node and a frame index (FrameIndex) of each point.

[0452] Fig.60 This is a diagram showing an example of leaf node information in the case of method 2. Fig.60 In the example shown, the leaf node information includes bitmap information (bitmapIsFramePointsFlag) indicating frame indexes of a plurality of points through bitmap. Fig.61 is a diagram showing an example of bitmap information. In this example, a leaf node including three-dimensional points of frame indexes 1, 3, and 5 is represented by a bitmap.

[0453] In addition, when the quantization resolution is low, there may be duplicate points in the same frame. In this case, the 3D point number (NumberOfPoints) may be shared to show the number of 3D points in each frame and the total number of 3D points in multiple frames.

[0454] In addition, when irreversible compression is used, the three-dimensional data encoding device can also delete duplicate points to reduce the amount of information. The three-dimensional data encoding device can delete duplicate points before combining frames or after combining frames.

[0455] Next, the structure of the encoded attribute information is explained. Fig.62This is a diagram showing the structure of the encoding attribute information. The encoding attribute information includes a header and a payload.

[0456] Fig.63 This is a diagram showing a syntax example of a header (Attribute_header) of coding attribute information. The header of coding attribute information includes an APS index (aps_idx), offset information (offset), other information (other_attribute_information), a frame combine flag (combine_frame_flag), and the number of combined frames (number_of_combine_frame).

[0457] The APS index represents the identifier (ID) of the parameter set (APS) corresponding to the coding attribute information. APS is a parameter set of coding attribute information for one or more frames. In addition, when a parameter set exists in each frame, the identifiers of multiple parameter sets can also be represented in the header.

[0458] The offset information indicates the offset position for obtaining the combined data. The other information indicates other information related to the attribute information (for example, the differential value of the quantization parameter (QPdelta), etc.). The frame combination flag is a flag indicating whether the encoded data is frame combined. The combined frame number indicates the number of combined frames.

[0459] In addition, part or all of the above information may be recorded in the SPS or APS.

[0460] Fig.64 This is a diagram showing a syntax example of a payload (Attribute_data) of coded attribute information. The payload of coded attribute information includes leaf node information (combine_information). For example, the structure of the leaf node information is the same as the leaf node information included in the payload of coded position information. That is, the leaf node information (frame index) can also be included in the attribute information.

[0461] In addition, the leaf node information (frame index) may be stored in one of the coding position information and the coding attribute information, but not in the other. In this case, the leaf node information (frame index) stored in one of the coding position information and the coding attribute information is referenced when the information of the other is decoded. In addition, the information indicating the reference target may also be included in the coding position information or the coding attribute information.

[0462] Next, an example of the transmission order and decoding order of encoded data is described. Fig.65 A diagram showing the structure of encoded data. The encoded data includes a header and a payload.

[0463] Figure 66 to Figure 68This is a diagram showing the order in which data is sent and the reference relationship between the data. In this diagram, G(1) and the like represent the encoded position information, GPS(1) and the like represent the parameter set of the encoded position information, and SPS represents the parameter set of the sequence (multiple frames). In addition, the numbers in () represent the values ​​of the frame index. In addition, the three-dimensional data encoding device can also send data in the decoding order.

[0464] Fig.66 This is a diagram showing an example of the transmission order when frames are not combined. Fig.67 This is a diagram showing an example of adding metadata (parameter set) to each PCC frame when combining frames. Fig.68 This is a diagram showing an example of adding metadata (parameter set) in units of combining when combining frames.

[0465] In the header of the frame-combined data, in order to obtain the metadata of the frame, an identifier of the metadata of the reference target is stored. Fig.68 As shown, metadata of each of the plurality of frames may be aggregated. Parameters common to the plurality of frames that have been combined may also be aggregated into one. Parameters that are not common to the frames represent values ​​for each frame.

[0466] The information of each frame (parameters not common to the frames) refers to, for example, a timestamp indicating the time when the frame data was generated, encoded, or decoded. In addition, the information of each frame may also include information of the sensor that obtained the frame data (sensor speed, acceleration, position information, sensor orientation, and other sensor information).

[0467] Fig.69 It means in Fig.67 FIG. 1 is a diagram showing an example of decoding a portion of a frame in the example shown. Fig.69 As shown, if there is no dependency between frames in the frame combination data, the three-dimensional data decoding device can decode each data independently.

[0468] In the case where the point group data has attribute information, the three-dimensional data encoding device can also perform frame combination on the attribute information. The attribute information is encoded and decoded with reference to the position information. The referenced position information can be the position information before the frame combination is performed, or it can be the position information after the frame combination is performed. The number of frames combined with the position information and the number of frames combined with the attribute information can be the same (the same) or independent (different).

[0469] Figure 70 to Figure 73 This is a diagram showing the order in which data is sent and the reference relationship between the data. Fig.70 as well as Fig.71 This shows an example of combining position information and attribute information in 4 frames. Fig.70In , metadata (parameter set) is added to each PCC frame. Fig.71 In the figure, metadata (parameter set) is added in a combined unit. In this figure, A(1) and the like represent encoding attribute information, and APS(1) and the like represent parameter sets of encoding attribute information. In addition, the numbers in () represent the values ​​of frame indexes.

[0470] Fig.72 This shows an example of combining position information with 4 frames without combining attribute information. Fig.72 As shown, the position information may be frame-bound, but the attribute information may not be frame-bound.

[0471] Fig.73 This shows an example of combining frame merging and tile segmentation. Fig.73 As shown, when tile segmentation is performed, the header of each tile position information includes information such as GPS index (gps_idx) and number of combined frames (number_of_combine_frame). In addition, the header of each tile position information includes a tile index (tile_idx) for identifying the tile.

[0472] As described above, the three-dimensional data encoding device of this embodiment performs Fig.74 First, the three-dimensional data encoding device generates the third point group data by combining the first point group data with the second point group data (S5481). Next, the three-dimensional data encoding device generates the encoded data by encoding the third point group data (S5482). In addition, the encoded data includes identification information (e.g., frame index) indicating to which of the first point group data and the second point group data the plurality of three-dimensional points included in the third point group data belong respectively.

[0473] Therefore, the three-dimensional data encoding device can improve the encoding efficiency by collectively encoding a plurality of point group data.

[0474] For example, the first point group data and the second point group data are point group data (eg, PCC frames) at different times. For example, the first point group data and the second point group data are point group data (eg, PCC frames) at different times of the same object.

[0475] The encoded data includes position information and attribute information of each of the plurality of three-dimensional points included in the third point group data, and the identification information is included in the attribute information.

[0476] For example, the encoded data includes position information (eg, occupancy rate code) representing the position of each of the plurality of three-dimensional points included in the third point group data using an N-ary tree (N is an integer greater than or equal to 2).

[0477] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.

[0478] In addition, the three-dimensional data decoding device of this embodiment performs Fig.75 First, the three-dimensional data decoding device decodes the encoded data to obtain the third point group data generated by combining the first point group data and the second point group data, and identification information indicating to which of the first point group data and the second point group data the plurality of three-dimensional points included in the third point group data respectively belong (S5491). Next, the three-dimensional data decoding device uses the identification information to separate the first point group data and the second point group data from the third point group data (S5492).

[0479] Thus, the three-dimensional data decoding device can decode the encoded data with improved encoding efficiency by collectively encoding a plurality of point group data.

[0480] For example, the first point group data and the second point group data are point group data (eg, PCC frames) at different times. For example, the first point group data and the second point group data are point group data (eg, PCC frames) at different times of the same object.

[0481] The encoded data includes position information and attribute information of each of the plurality of three-dimensional points included in the third point group data, and the identification information is included in the attribute information.

[0482] For example, the encoded data includes position information (eg, occupancy rate code) representing the position of each of a plurality of three-dimensional points included in the third point cloud data using an N-ary tree (N is an integer greater than or equal to 2).

[0483] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.

[0484] (Implementation 6)

[0485] The information of the three-dimensional point group includes position information (geometry) and attribute information (attribute). The position information includes coordinates (x coordinate, y coordinate, z coordinate) based on a certain point. When encoding the position information, instead of directly encoding the coordinates of each three-dimensional point, a method is used to reduce the amount of encoding by using an octree to represent the position of each three-dimensional point and encoding the information of the octree.

[0486] On the other hand, the attribute information includes information indicating color information (RGB, YUV, etc.), reflectivity, normal vector, etc. of each 3D point. For example, the 3D data encoding device can encode the attribute information using a different encoding method from that of the position information.

[0487] In the present embodiment, a method for encoding attribute information when encoding position information in combination with a plurality of point group data of a plurality of frames is described. In addition, in the present embodiment, integer values ​​are used as the values ​​of the attribute information for description. For example, when each color component of the color information RGB or YUV has an 8-bit precision, each color component takes an integer value of 0 to 255. When the value of the reflectivity has a 10-bit precision, the value of the reflectivity takes an integer value of 0 to 1023. In addition, when the bit precision of the attribute information is a decimal precision, the three-dimensional data encoding device may also multiply the value by a scaling value and round it to an integer value so that the value of the attribute information becomes an integer value. In addition, the three-dimensional data encoding device may also attach the scaling value to the header of the bit stream, etc.

[0488] As a method for encoding attribute information when encoding each position information of a three-dimensional point group in combination with point group data of multiple frames, for example, it is considered to encode the attribute information corresponding to each position information using the combined position information. Here, the combined position information may also include the position information of the three-dimensional point group and the frame_index (frame index) to which the three-dimensional point group belongs. In addition, when encoding the attribute information of the first three-dimensional point in the three-dimensional point group, not only the position information or attribute information of the three-dimensional point group contained in the frame to which the first three-dimensional point belongs, but also the position information or attribute information of the three-dimensional point group contained in a frame different from the frame to which the first three-dimensional point belongs may be used.

[0489] Multiple frames respectively contain point group data. The first point group data belonging to the first frame and the second point group data belonging to the second frame in the multiple frames are point group data at different times. In addition, the first point group data and the second point group data are, for example, point group data of the same object at different times. The first point group data includes a frame index indicating that the three-dimensional point group included in the first point group data belongs to the first point group data. The second point group data includes a frame index indicating that the three-dimensional point group included in the second point group data belongs to the second point group data. The frame index is identification information indicating to which point group data the three-dimensional point group included in the combined point group data that is a combination of multiple point group data belonging to different frames belongs. In addition, the three-dimensional point group is also referred to as multiple three-dimensional points.

[0490] As a method for encoding attribute information of a three-dimensional point, it is considered to calculate the predicted value of the attribute information of the three-dimensional point and encode the difference (prediction residual) between the original attribute information value and the predicted value. For example, when the value of the attribute information of the three-dimensional point p is Ap and the predicted value is Pp, the three-dimensional data encoding device encodes its differential absolute value Diffp = |Ap-Pp|. In this case, if the predicted value Pp can be generated with high precision, the value of the differential absolute value Diffp becomes smaller. Therefore, for example, by entropy encoding the differential absolute value Diffp using a coding table that generates a smaller number of bits as the value is smaller, the amount of coding can be reduced.

[0491] As a method for generating a prediction value of attribute information, it is considered to use the attribute information of other three-dimensional points located around the object three-dimensional point of the encoding object, that is, the reference three-dimensional point. In this way, the three-dimensional data encoding device can also encode the attribute information of the first three-dimensional point using the attribute information of the surrounding three-dimensional points. Here, another surrounding three-dimensional point located around the three-dimensional point of the encoding object can exist in the frame to which the object three-dimensional point of the encoding object belongs, and can also exist in a frame different from the frame to which the object three-dimensional point of the encoding object belongs. That is, it is also possible that the attribute information of the object three-dimensional point includes a first frame index indicating that the object three-dimensional point belongs to the first point group data, and the attribute information of the surrounding three-dimensional points includes a second frame index indicating that the surrounding three-dimensional point belongs to the second point group data. Thus, by also referring to the attribute information of the three-dimensional points outside the frame to which the three-dimensional point of the encoding object belongs, a high-precision prediction value Pp can be generated, and the encoding efficiency can be improved.

[0492] Here, the reference 3D point refers to a 3D point within a predetermined distance range from the target 3D point. For example, when there is a target 3D point p = (x1, y1, z1) and a 3D point q = (x2, y2, z2), the 3D data encoding device calculates the Euclidean distance d(p, q) between the 3D point p and the 3D point q shown in (Formula H1).

[0493]

Formula 1

[0494]

[0495] When the Euclidean distance d(p, q) is less than a predetermined threshold value THd, the three-dimensional data encoding device determines that the position of the three-dimensional point q is close to the position of the target three-dimensional point p, and determines that the value of the attribute information of the three-dimensional point q is used in the generation of the predicted value of the attribute information of the target three-dimensional point p. In addition, the distance calculation method may also be other methods, for example, the Mahalanobis distance may also be used. In addition, the three-dimensional data encoding device may also determine not to use the three-dimensional points outside the predetermined distance range from the target three-dimensional point for prediction processing. For example, when there is a three-dimensional point r and the distance d(p, r) between the target three-dimensional point p and the three-dimensional point r is greater than the threshold value THd, the three-dimensional data encoding device may also determine not to use the three-dimensional point r for prediction. In addition, the three-dimensional data encoding device may also attach information indicating the threshold value THd to the header of the bit stream, etc. In addition, when the three-dimensional data encoding device combines the position information of the three-dimensional point group with the point group data of multiple frames for encoding, it may also calculate the distance between each three-dimensional point based on the combined three-dimensional point group. That is, the three-dimensional data encoding device may calculate the distance between two three-dimensional points belonging to different frames, or may calculate the distance between two three-dimensional points belonging to the same frame.

[0496] Fig.763D points are shown in the figure. In this example, the distance d(p, q) between the target 3D point p and the 3D point q is less than the threshold value THd. Therefore, the 3D data encoding device determines that the 3D point q is a reference 3D point of the target 3D point p, and determines that the value of the attribute information Aq of the 3D point q is used in the generation of the predicted value Pp of the attribute information Ap of the target 3D point p.

[0497] On the other hand, the distance d(p, r) between the object 3D point p and the 3D point r is greater than the threshold value THd. Therefore, the 3D data encoding device determines that the 3D point r is not a reference 3D point of the object 3D point p, and determines that the value of the attribute information Ar of the 3D point r is not used in the generation of the predicted value Pp of the attribute information Ap of the object 3D point p.

[0498] Here, the 3D point p belongs to the frame indicated by the frame index (frame_idx=0), the 3D point q belongs to the frame indicated by the frame index (frame_idx=1), and the 3D point r belongs to the frame indicated by the frame index (frame_idx=0). The 3D encoding device can calculate the distance between the 3D point p and the 3D point r that belong to the same frame indicated by the frame index, and can also calculate the distance between the 3D point p and the 3D point q that belong to different frames indicated by the frame index.

[0499] Furthermore, when encoding the attribute information of the target 3D point using the prediction value, the 3D data encoding device uses the 3D point whose attribute information has been encoded and decoded as a reference 3D point. Similarly, when decoding the attribute information of the target 3D point of the decoding target using the prediction value, the 3D data decoding device uses the 3D point whose attribute information has been decoded as a reference 3D point. Thus, the same prediction value can be generated during encoding and decoding, so that the bit stream of the 3D point generated by encoding can be correctly decoded on the decoding side.

[0500] In addition, another surrounding 3D point around the object 3D point of the encoding object may exist in the frame to which the object 3D point of the encoding object belongs, or may exist in a frame different from the frame to which the object 3D point of the encoding object belongs, but is not necessarily limited to this. For example, it may be determined that the 3D point existing in the frame different from the frame to which the object 3D point of the encoding object belongs does not exist around the object 3D point of the encoding object and is not used as a prediction value. Thus, the 3D data encoding device performs frame combination and encoding on the position information, and predictively encodes the attribute information using the attribute information of other surrounding 3D points in the same frame, for example, when the attribute information of each 3D point of a plurality of frames of the combined object is greatly different, thereby improving the encoding efficiency. In addition, the 3D data encoding device may also attach information indicating whether the attribute information of the object 3D point is encoded using the attribute information of the surrounding 3D points of the same frame or the attribute information of the surrounding 3D points outside the same frame is encoded to the header of the encoded data and switch. Therefore, the three-dimensional data decoding device can determine whether to use the same frame or the attribute information of the surrounding three-dimensional points in the same frame and outside the same frame to decode the attribute information of the object three-dimensional point when decoding the frame-based encoded data by decoding the header, and switch which decoding to perform, so that the bit stream can be decoded appropriately.

[0501] In addition, when encoding the attribute information of 3D points after combining multiple frames, it is considered to classify each 3D point into multiple levels using the position information of the 3D points belonging to the same frame or different frames and then encode them. Here, each level after classification is called LoD (Level of Detail). Fig.77 The method of generating LoD is described.

[0502] First, the three-dimensional data encoding device selects an initial point a0 from the combined three-dimensional point group and assigns it to LoD0. Next, the three-dimensional data encoding device extracts a point a1 whose distance from point a0 is greater than the threshold Thres_LoD[0] of LoD0 and assigns it to LoD0. Next, the three-dimensional data encoding device extracts a point a2 whose distance from point a1 is greater than the threshold Thres_LoD[0] of LoD0 and assigns it to LoD0. In this way, the three-dimensional data encoding device constructs LoD0 in such a way that the distance between each point in LoD0 is greater than the threshold Thres_LoD[0]. In addition, the three-dimensional data encoding device can also calculate the distance between the three-dimensional points of the two points by the same process regardless of whether they belong to the same frame or different frames. For example, point a0 and point a1 can belong to the same frame or different frames. Therefore, the distance between point a0 and point a1 is calculated by the same process regardless of whether they belong to the same frame or different frames.

[0503] Next, the 3D data encoding device selects a point b0 that has not been assigned a LoD, and assigns it to LoD1. Next, the 3D data encoding device extracts a point b1 that has not been assigned a LoD and whose distance from point b0 is greater than the threshold Thres_LoD[1] of LoD1, and which has not been assigned a LoD, and assigns it to LoD1. Next, the 3D data encoding device extracts a point b2 that has not been assigned a LoD and whose distance from point b1 is greater than the threshold Thres_LoD[1] of LoD1, and which has not been assigned a LoD, and assigns it to LoD1. In this way, the 3D data encoding device constructs LoD1 in such a way that the distance between each point in LoD1 is greater than the threshold Thres_LoD[1].

[0504] Next, the three-dimensional data encoding device selects point c0 that has not been assigned a LoD, and assigns it to LoD2. Next, the three-dimensional data encoding device extracts point c1 that has not been assigned a LoD and whose distance from point c0 is greater than the threshold Thres_LoD[2] of LoD2, and which has not been assigned a LoD, and assigns it to LoD2. Next, the three-dimensional data encoding device extracts point c2 that has not been assigned a LoD and whose distance from point c1 is greater than the threshold Thres_LoD[2] of LoD2, and which has not been assigned a LoD, and assigns it to LoD2. In this way, the three-dimensional data encoding device constructs LoD2 in such a way that the distance between each point in LoD2 is greater than the threshold Thres_LoD[2]. For example, Fig.78 As shown, the threshold values ​​of each LoD, Thres_LoD[0], Thres_LoD[1], and Thres_LoD[2] are set.

[0505] In addition, the three-dimensional data encoding device may also add information indicating the threshold value of each LoD to the header of the bit stream. Fig.78 In the illustrated example, the three-dimensional data encoding device may add threshold values ​​Thres_LoD[0], Thres_LoD[1], and Thres_LoD[2] to the header.

[0506] Alternatively, the 3D data encoding device may also assign all 3D points that are not assigned LoD to the lowest layer of LoD. In this case, the 3D data encoding device can reduce the amount of coding in the head by not adding the threshold of the lowest layer of LoD to the head. Fig.78 In the example shown, the 3D data encoding device adds the thresholds Thres_LoD[0] and Thres_LoD[1] to the header, and does not add Thres_LoD[2] to the header. In this case, the 3D data decoding device may also estimate the value of Thres_LoD[2] to be 0. In addition, the 3D data encoding device may also add the number of layers of LoD to the header. Thus, the 3D data decoding device can use the number of layers of LoD to determine the LoD of the lowest layer.

[0507] In addition, if Fig.78 As shown in FIG. 1 , the threshold value of each LoD layer is set to be larger as it is closer to the upper layer, so that the upper layer (the layer closer to LoD0) becomes a sparse point group with a long distance between three-dimensional points, and the lower layer becomes a dense point group with a short distance between three-dimensional points. Fig.78 In the example shown, LoD0 is the highest layer.

[0508] In addition, the method of selecting the initial three-dimensional point when setting each LoD may also depend on the encoding order when encoding the position information. For example, the three-dimensional data encoding device selects the three-dimensional point that is first encoded when encoding the position information as the initial point a0 of LoD0, and selects points a1 and a2 to constitute LoD0 based on the initial point a0. Moreover, the three-dimensional data encoding device may also select the three-dimensional point whose position information is encoded earliest among the three-dimensional points that do not belong to LoD0 as the initial point b0 of LoD1. That is, the three-dimensional data encoding device may also select the three-dimensional point whose position information is encoded earliest among the three-dimensional points that do not belong to the upper layer (LoD0 to LoDn-1) of LoDn as the initial point n0 of LoDn. Thus, the three-dimensional data decoding device can construct the same LoD as that during encoding by using the same initial point selection method during decoding, and thus can properly decode the bit stream. Specifically, the three-dimensional data decoding device selects the three-dimensional point whose position information is decoded earliest among the three-dimensional points that do not belong to the upper layer of LoDn as the initial point n0 of LoDn.

[0509] The following describes a method for generating predicted values ​​of attribute information of three-dimensional points using LoD information. For example, when encoding three-dimensional points included in LoD0 in sequence, the three-dimensional data encoding device uses the encoded and decoded (hereinafter, also referred to as "encoded") attribute information included in LoD0 and LoD1 to generate the object three-dimensional point included in LoD1. In this way, the three-dimensional data encoding device uses the encoded attribute information included in LoDn' (n'<=n) to generate predicted values ​​of attribute information of three-dimensional points included in LoDn. That is, the three-dimensional data encoding device does not use the attribute information of three-dimensional points included in the lower layer of LoDn in the calculation of the predicted values ​​of the attribute information of the three-dimensional points included in LoDn.

[0510] For example, the three-dimensional data encoding device generates a predicted value of the attribute information of the three-dimensional point by calculating the average value of the attribute values ​​of less than N three-dimensional points among the encoded three-dimensional points around the object three-dimensional point of the encoding object. In addition, the three-dimensional data encoding device can add the value of N to the header of the bit stream, etc. In addition, the three-dimensional data encoding device can also change the value of N for each three-dimensional point and add the value of N to each three-dimensional point. Thus, it is possible to select an appropriate N for each three-dimensional point, thereby improving the accuracy of the predicted value. Therefore, the prediction residual can be reduced. In addition, the three-dimensional data encoding device can also add the value of N to the header of the bit stream and fix the value of N in the bit stream. Thus, it is not necessary to encode or decode the value of N for each three-dimensional point, thereby reducing the amount of processing. In addition, the three-dimensional data encoding device can also encode the value of N for each LoD separately. Thus, by selecting an appropriate N for each LoD, the coding efficiency can be improved.

[0511] Alternatively, the 3D data encoding device may also calculate the predicted value of the attribute information of the 3D point by taking the weighted average of the attribute information of the surrounding N 3D points that have been encoded. For example, the 3D data encoding device calculates the weight using the distance information of the target 3D point and the surrounding N 3D points.

[0512] When the three-dimensional data encoding device encodes the value of N for each LoD, for example, the higher the LoD layer, the larger the value of N is set, and the lower the LoD layer, the smaller the value of N is set. In the upper layer of the LoD, the distance between the three-dimensional points belonging to the layer is far, so it is possible to set the value of N to a large value and select a plurality of surrounding three-dimensional points for averaging, thereby improving the prediction accuracy. In addition, since the distance between the three-dimensional points belonging to the layer in the lower layer of the LoD is close, it is possible to set the value of N to a small value to suppress the processing amount of averaging while performing efficient prediction.

[0513] Fig.79 is a diagram showing an example of attribute information used in the prediction value. As described above, the predicted value of the point P included in LoDN' (N' <= N) is generated using the encoded surrounding points P' included in LoDN'. Here, the surrounding points P' are selected based on the distance from the point P. For example, the attribute information of points a0, a1, a2, b0, and b1 is used to generate Fig.79 The predicted value of the attribute information of point b2 is shown.

[0514] The selected surrounding points change according to the value of N. For example, when N=5, a0, a1, a2, b0, and b1 are selected as the surrounding points of point b2. When N=4, points a0, a1, a2, and b1 are selected based on the distance information.

[0515] The prediction is calculated by weighted averaging depending on the distance. For example, Fig.79 In the example shown, the predicted value a2p of point a2 is calculated by taking the weighted average of the attribute information of points a0 and a1 as shown in (Formula H2) and (Formula H3). i It is the value of the attribute information of point ai.

[0516]

Formula 2

[0517]

[0518]

[0519] In addition, the predicted value b2p of point b2 is calculated by taking the weighted average of the attribute information of points a0, a1, a2, b0, and b1 as shown in (Equations H4) to (Equations H6). i It is the value of the attribute information of point bi.

[0520]

Formula 3

[0521]

[0522]

[0523]

[0524] In addition, the three-dimensional data encoding device can also calculate the difference between the value of the attribute information of the three-dimensional point and the predicted value generated from the surrounding points (prediction residual), and quantize the calculated prediction residual. For example, the three-dimensional data encoding device quantizes the prediction residual by dividing it by a quantization scale (also called a quantization step size). In this case, the smaller the quantization scale, the smaller the error (quantization error) caused by quantization. On the contrary, the larger the quantization scale, the larger the quantization error.

[0525] In addition, the three-dimensional data encoding device may also change the quantization scale used for each LoD. For example, the higher the level of the three-dimensional data encoding device, the smaller the quantization scale, and the lower the level of the three-dimensional data encoding device, the larger the quantization scale. The value of the attribute information of the three-dimensional point belonging to the upper layer may be used as the predicted value of the attribute information of the three-dimensional point belonging to the lower layer. Therefore, the quantization scale of the upper layer can be reduced to suppress the quantization error generated in the upper layer, and the coding efficiency can be improved by improving the accuracy of the predicted value. In addition, the three-dimensional data encoding device may also attach the quantization scale used for each LoD to the header, etc. As a result, the three-dimensional data decoding device can correctly decode the quantization scale, and thus can properly decode the bit stream.

[0526] In addition, the three-dimensional data encoding device may also transform the signed integer value (signed quantized value) of the quantized prediction residual into an unsigned integer value (unsigned quantized value). Thus, when entropy encoding is performed on the prediction residual, there is no need to consider the generation of negative integers. In addition, the three-dimensional data encoding device does not necessarily need to transform the signed integer value into an unsigned integer value, and for example, the sign bit may be entropy encoded separately.

[0527] The prediction residual is calculated by subtracting the predicted value from the original value. For example, as shown in (Equation H7), the prediction residual a2r of point a2 is obtained by subtracting the value A of the attribute information of point a2 from 2 The prediction residual b2r of point b2 is calculated by subtracting the predicted value a2p of point a2. As shown in (Equation H8), the prediction residual b2r of point b2 is obtained by calculating the value B of the attribute information of point b2. 2 Calculated by subtracting the predicted value b2p of point b2.

[0528] a2r=A 2 -a2p…(Formula H7)

[0529] b2r=B 2 -b2p…(Formula H8)

[0530] In addition, the prediction residual is quantized by dividing by QS (Quantization Step). For example, the quantization value a2q of point a2 is calculated by (Formula H9). The quantization value b2q of point b2 is calculated by (Formula H10). Here, QS_LoD0 is the QS for LoD0, and QS_LoD1 is the QS for LoD1. That is, QS can be changed according to LoD.

[0531] a2q=a2r / QS_LoD0…(Formula H9)

[0532] b2q=b2r / QS_LoD1…(Formula H10)

[0533] In addition, as described below, the three-dimensional data encoding device converts the signed integer value as the quantized value into an unsigned integer value. When the signed integer value a2q is less than 0, the three-dimensional data encoding device sets the unsigned integer value a2u to -1-(2×a2q). When the signed integer value a2q is greater than 0, the three-dimensional data encoding device sets the unsigned integer value a2u to 2×a2q.

[0534] Similarly, when the signed integer value b2q is less than 0, the three-dimensional data encoding device sets the unsigned integer value b2u to -1-(2×b2q). When the signed integer value b2q is greater than or equal to 0, the three-dimensional data encoding device sets the unsigned integer value b2u to 2×b2q.

[0535] Furthermore, the three-dimensional data encoding device may encode the quantized prediction residual (unsigned integer value) by entropy encoding. For example, after binarizing the unsigned integer value, binary arithmetic coding may be applied.

[0536] In addition, in this case, the three-dimensional data encoding device may also switch the binarization method according to the value of the prediction residual. For example, when the prediction residual pu is less than the threshold R_TH, the three-dimensional data encoding device binarizes the prediction residual pu with the required fixed number of bits to represent the threshold R_TH. In addition, when the prediction residual pu is greater than the threshold R_TH, the three-dimensional data encoding device binarizes the binarized data of the threshold R_TH and the value of (pu-R_TH) using Exponential-Golomb or the like.

[0537] For example, when the threshold R_TH is 63 and the prediction residual pu is less than 63, the three-dimensional data encoding device binarizes the prediction residual pu with 6 bits. In addition, when the prediction residual pu is greater than 63, the three-dimensional data encoding device binarizes the binary data (111111) and (pu-63) of the threshold R_TH using Exponential Golomb, thereby performing arithmetic coding.

[0538] In a more specific example, when the prediction residual pu is 32, the three-dimensional data encoding device generates 6-bit binary data (100000) and performs arithmetic coding on the bit string. In addition, when the prediction residual pu is 66, the three-dimensional data encoding device generates binary data (111111) representing the threshold value R_TH using Exponential Golomb and a bit string (00100) of value 3 (66-63), and performs arithmetic coding on the bit string (111111+00100).

[0539] Thus, the 3D data encoding device switches the binarization method according to the size of the prediction residual, thereby being able to perform encoding while suppressing a sharp increase in the number of binarization bits when the prediction residual becomes larger. In addition, the 3D data encoding device may also add the threshold R_TH to the header of the bitstream, etc.

[0540] For example, in the case of encoding at a high bit rate, that is, when the quantization scale is small, the quantization error becomes smaller, the prediction accuracy becomes higher, and the prediction residual may not become larger as a result. Therefore, in this case, the three-dimensional data encoding device sets the threshold R_TH to be large. As a result, the possibility of encoding the binary data of the threshold R_TH becomes lower, and the encoding efficiency is improved. On the contrary, in the case of encoding at a low bit rate, that is, when the quantization scale is large, the quantization error becomes larger, the prediction accuracy becomes worse, and the prediction residual may become larger as a result. Therefore, in this case, the three-dimensional data encoding device sets the threshold R_TH to be small. As a result, it is possible to prevent the sharp increase in the bit length of the binary data.

[0541] In addition, the three-dimensional data encoding device may also switch the threshold R_TH for each LoD, and attach the threshold R_TH of each LoD to the header, etc. That is, the three-dimensional data encoding device may also switch the binarization method for each LoD. For example, in the upper layer, due to the long distance between the three-dimensional points, the prediction accuracy deteriorates, and the prediction residual may become larger as a result. Therefore, the three-dimensional data encoding device prevents the sharp increase in the bit length of the binary data by setting the threshold R_TH to be small for the upper layer. In addition, in the lower layer, due to the short distance between the three-dimensional points, the prediction accuracy becomes high, and the prediction residual may become smaller as a result. Therefore, the three-dimensional data encoding device improves the encoding efficiency by setting the threshold R_TH to be large for the hierarchy.

[0542] Fig.80 is a diagram showing an example of an Exp Golomb code, and is a diagram showing the relationship between the value before binarization (multi-value) and the bit after binarization (code). Fig.80 The 0s and 1s shown are reversed.

[0543] In addition, the three-dimensional data encoding device applies arithmetic coding to the binary data of the prediction residual. Thus, the coding efficiency can be improved. In addition, when arithmetic coding is applied, in the binary data, the tendency of the occurrence probability of 0 and 1 of each bit may be different in the part binarized with n bits, that is, the n-bit code (n-bit code) and the part binarized using the exponential Golomb, that is, the remaining code (remaining code). Therefore, the three-dimensional data encoding device can also switch the application method of arithmetic coding through n-bit coding and remaining coding.

[0544] For example, for n-bit coding, the three-dimensional data coding device uses a different coding table (probability table) to perform arithmetic coding on each bit. At this time, the three-dimensional data coding device can also change the number of coding tables used for each bit. For example, the three-dimensional data coding device uses 1 coding table to perform arithmetic coding on the leading bit b0 of the n-bit coding. In addition, the three-dimensional data coding device uses 2 coding tables for the next bit b1. In addition, the three-dimensional data coding device switches the coding table used in the arithmetic coding of the bit b1 according to the value of b0 (0 or 1). Similarly, the three-dimensional data coding device also uses 4 coding tables for the next bit b2. In addition, the three-dimensional data coding device switches the coding table used in the arithmetic coding of the bit b2 according to the values ​​of b0 and b1 (0 to 3).

[0545] Thus, the three-dimensional data encoding device uses 2 when performing arithmetic coding on each bit bn-1 of the n-bit code. n-1 In addition, the three-dimensional data encoding device switches the encoding table to be used according to the value (occurrence pattern) of the bit before bn-1. As a result, the three-dimensional data encoding device can use an appropriate encoding table for each bit, thereby improving the encoding efficiency.

[0546] In addition, the three-dimensional data encoding device may also reduce the number of encoding tables used for each bit. For example, when performing arithmetic coding on each bit bn-1, the three-dimensional data encoding device may also switch between two encoding tables according to the value (generation pattern) of the m bits (m<n-1) before bn-1. m The three-dimensional data encoding device can also update the probability of occurrence of 0 and 1 in each coding table according to the value of the binary data actually generated. In addition, the three-dimensional data encoding device can also fix the probability of occurrence of 0 and 1 in the coding table of a part of the bits. Thus, the number of updates of the probability of occurrence can be suppressed, so the amount of processing can be reduced.

[0547] For example, when the n-bit code is b0b1b2…bn-1, there is one coding table for b0 (CTb0). There are two coding tables for b1 (CTb10, CTb11). In addition, the coding table used is switched according to the value of b0 (0 to 1). There are four coding tables for b2 (CTb20, CTb21, CTb22, CTb23). In addition, the coding table used is switched according to the values ​​of b0 and b1 (0 to 3). There are 2 coding tables for bn-1. n-1 (CTbn0, CTbn1, ..., CTbn(2 n-1 -1)). In addition, according to the value of b0b1…bn-2 (0~2 n-1 -1) to switch the encoding table used.

[0548] In addition, the three-dimensional data encoding device may also set 0 to 2 instead of binarizing the n-bit code. n m-ary arithmetic coding of the value -1 (m=2 n ). In addition, when the three-dimensional data encoding device performs arithmetic encoding on the n-bit code using m-ary, the three-dimensional data decoding device can also restore the n-bit code by arithmetic decoding of m-ary.

[0549] Fig.81 2 is a diagram for explaining the processing when the residual code is an exponential Golomb code, for example. Fig.81 As shown, the portion binarized using the exponential Golomb, i.e., the remaining code, includes the prefix part and the suffix part. For example, the three-dimensional data encoding device switches the coding table between the prefix part and the suffix part. That is, the three-dimensional data encoding device uses the coding table for the prefix to perform arithmetic coding on each bit included in the prefix part, and uses the coding table for the suffix to perform arithmetic coding on each bit included in the suffix part.

[0550] In addition, the three-dimensional data encoding device may also update the occurrence probabilities of 0 and 1 in each coding table according to the value of the binary data actually generated. Alternatively, the three-dimensional data encoding device may also fix the occurrence probabilities of 0 and 1 in a certain coding table. Thus, the number of updates of the occurrence probability can be suppressed, thereby reducing the amount of processing. For example, the three-dimensional data encoding device may also update the occurrence probability for the prefix part and fix the occurrence probability for the suffix part.

[0551] In addition, the three-dimensional data encoding device decodes the quantized prediction residual by inverse quantization and reconstruction, and uses the decoded prediction residual, i.e., the decoded value, for subsequent prediction of the three-dimensional point of the encoding object. Specifically, the three-dimensional data encoding device calculates the inverse quantization value by multiplying the quantized prediction residual (quantization value) by the quantization scale, and adds the inverse quantization value and the prediction value to obtain the decoded value (reconstruction value).

[0552] For example, the inverse quantization value a2iq of point a2 is calculated by (Formula H11) using the quantization value a2q of point a2. The inverse quantization value b2iq of point b2 is calculated by (Formula H12) using the quantization value b2q of point b2. Here, QS_LoD0 is the QS for LoD0, and QS_LoD1 is the QS for LoD1. That is, QS can be changed according to LoD.

[0553] a2iq=a2q×QS_LoD0…(Formula H11)

[0554] b2iq=b2q×QS_LoD1…(Formula H12)

[0555] For example, as shown in (Formula H13), the decoded value a2rec of point a2 is calculated by adding the inverse quantized value a2iq of point a2 to the predicted value a2p of point a2. As shown in (Formula H14), the decoded value b2rec of point b2 is calculated by adding the inverse quantized value b2iq of point b2 to the predicted value b2p of point b2.

[0556] a2rec=a2iq+a2p…(Formula H13)

[0557] b2rec=b2iq+b2p…(Formula H14)

[0558] Hereinafter, an example of syntax of a bit stream according to the present embodiment will be described. Fig.82 1 is a diagram showing an example of the syntax of the attribute header (attribute_header) of this embodiment. The attribute header is the header information of the attribute information. Fig.82 As shown, the attribute header includes the number of layers information (NumLoD), three-dimensional point information (NumOfPoint[i]), layer threshold (Thres_Lod[i]), surrounding point information (NumNeighborPoint[i]), prediction threshold (THd[i]), quantization scale (QS[i]), and binarization threshold (R_TH[i]).

[0559] The number of layers information (NumLoD) indicates the number of layers of LoD used.

[0560] The three-dimensional point number information (NumOfPoint[i]) indicates the number of three-dimensional points belonging to layer i. In addition, the three-dimensional data encoding device may also attach the three-dimensional point total number information (AllNumOfPoint) indicating the total number of three-dimensional points to other headers. In this case, the three-dimensional data encoding device may not attach NumOfPoint[NumLoD-1] indicating the number of three-dimensional points belonging to the bottom layer to the header. In this case, the three-dimensional data decoding device can calculate NumOfPoint[NumLoD-1] by (Formula H15). In this way, the amount of coding in the header can be reduced.

[0561]

Formula 4

[0562]

[0563] The hierarchical threshold (Thres_Lod[i]) is used to set the threshold for level i. The 3D data encoding device and the 3D data decoding device form LoDi such that the distance between each point within LoDi is greater than the threshold Thres_LoD[i]. Additionally, the 3D data encoding device may not attach the value of Thres_Lod[NumLoD - 1] (the bottommost level) to the header. In this case, the 3D data decoding device estimates the value of Thres_Lod[NumLoD - 1] as 0. Thereby, the encoding amount of the header can be reduced.

[0564] The number of surrounding points information (NumNeighborPoint[i]) represents the upper limit value of the number of surrounding points used in the generation of the predicted value of the 3D points belonging to level i. When the number of surrounding points M is less than NumNeighborPoint[i] (M < NumNeighborPoint[i]), the 3D data encoding device may also use M surrounding points to calculate the predicted value. Additionally, when it is not necessary to separate the value of NumNeighborPoint[i] in each LoD, the 3D data encoding device may also attach 1 number of surrounding points information (NumNeighborPoint) used in all LoDs to the header.

[0565] The prediction threshold (THd[i]) represents the upper limit value of the distance between the surrounding 3D points and the object 3D point used in the prediction of the object 3D point to be encoded or decoded at level i. The 3D data encoding device and the 3D data decoding device do not use 3D points whose distance from the object 3D point is farther than THd[i] for prediction. Additionally, when it is not necessary to separate the value of THd[i] in each LoD, the 3D data encoding device may also attach 1 prediction threshold (THd) used in all LoDs to the header.

[0566] The quantization scale (QS[i]) represents the quantization scale used in the quantization and inverse quantization at level i.

[0567] The binarization threshold (R_TH[i]) is a threshold used to switch the binarization method of the prediction residual of the 3D points belonging to level i. For example, when the prediction residual is less than the threshold R_TH, the 3D data encoding device binarizes the prediction residual pu with a fixed number of bits, and when the prediction residual is greater than or equal to the threshold R_TH, it uses exponential Golomb to binarize the binarized data of the threshold R_TH and the value of (pu - R_TH). Additionally, when it is not necessary to switch the value of R_TH[i] in each LoD, the 3D data encoding device may also attach 1 binarization threshold (R_TH) used in all LoDs to the header.

[0568] In addition, R_TH[i] may also be a maximum value represented by nbit. For example, in 6bit, R_TH is 63, and in 8bit, R_TH is 255. In addition, the three-dimensional data encoding device may encode the number of bits instead of encoding the maximum value represented by nbit as the binarization threshold. For example, the three-dimensional data encoding device may append the value 6 to the header when R_TH[i]=63, and append the value 8 to the header when R_TH[i]=255. In addition, the three-dimensional data encoding device may also define a minimum value (minimum number of bits) for the number of bits representing R_TH[i], and append the relative number of bits according to the minimum value to the header. For example, the three-dimensional data encoding device may append the value 0 to the header when R_TH[i]=63 and the minimum number of bits is 6, and append the value 2 to the header when R_TH[i]=255 and the minimum number of bits is 6.

[0569] In addition, the three-dimensional data encoding device may entropy encode at least one of NumLoD, Thres_Lod[i], NumNeighborPoint[i], THd[i], QS[i], and R_TH[i] and append it to the header. For example, the three-dimensional data encoding device may binarize each value and perform arithmetic encoding. In addition, in order to reduce the amount of processing, the three-dimensional data encoding device may encode each value with a fixed length.

[0570] In addition, the three-dimensional data encoding device may not add at least one of NumLoD, Thres_Lod[i], NumNeighborPoint[i], THd[i], QS[i], and R_TH[i] to the header. For example, the value of at least one of them may also be specified by a profile or level of a standard. In this way, the bit amount of the header can be reduced.

[0571] Fig.83 1 is a diagram showing a syntax example of attribute data (attribute_data) according to the present embodiment. The attribute data includes encoded data of attribute information of a plurality of three-dimensional points. Fig.83 As shown, the attribute data includes an n-bit code and a remaining code.

[0572] The n-bit code is the coded data of the prediction residual of the value of the attribute information or a part thereof. The bit length of the n-bit code depends on the value of R_TH[i]. For example, when the value shown in R_TH[i] is 63, the n-bit code is 6 bits, and when the value shown in R_TH[i] is 255, the n-bit code is 8 bits.

[0573] The remaining code is the coded data after exponential Golomb coding in the coded data of the prediction residual of the value of the attribute information. When the n-bit code is the same as R_TH[i], the remaining code is encoded or decoded. In addition, the three-dimensional data decoding device adds the value of the n-bit code and the value of the remaining code to decode the prediction residual. In addition, when the n-bit code is not the same value as R_TH[i], the remaining code may not be encoded or decoded.

[0574] The following describes the flow of processing in the three-dimensional data encoding device. Fig.84 This is a flowchart of a three-dimensional data encoding process performed by a three-dimensional data encoding device.

[0575] First, the three-dimensional data encoding device combines multiple frames (S5601). For example, the three-dimensional data encoding device combines multiple three-dimensional point groups belonging to multiple frames input into one three-dimensional point group. In addition, when combining, the three-dimensional data encoding device adds a frame index indicating the frame to which each three-dimensional point group belongs to each three-dimensional point group.

[0576] Next, the three-dimensional data encoding device encodes the position information (geometry) after the frame combination (S5602). For example, the three-dimensional data encoding device uses octree representation to perform encoding.

[0577] After encoding the position information, the three-dimensional data encoding device reallocates the original three-dimensional point's attribute information to the changed three-dimensional point when the position of the three-dimensional point changes due to quantization or the like (S5603). For example, the three-dimensional data encoding device reallocates the attribute information by interpolating the value of the attribute information according to the amount of change in the position. For example, the three-dimensional data encoding device detects N three-dimensional points before the change that are close to the changed three-dimensional position, and performs weighted averaging on the values ​​of the attribute information of the N three-dimensional points. For example, in the weighted averaging, the three-dimensional data encoding device determines the weight based on the distance from the changed three-dimensional position to each of the N three-dimensional points. Then, the three-dimensional data encoding device determines the value obtained by the weighted averaging as the value of the attribute information of the changed three-dimensional point. In addition, when two or more three-dimensional points change to the same three-dimensional position due to quantization or the like, the three-dimensional data encoding device may also allocate the average value of the attribute information of the two or more three-dimensional points before the change as the value of the attribute information of the changed three-dimensional point.

[0578] Next, the three-dimensional data encoding device encodes the reallocated attribute information (Attribute) (S5604). For example, here, the three-dimensional data encoding device encodes the frame index of each of the multiple three-dimensional points as the attribute information of the three-dimensional point. In addition, for example, in the case of encoding multiple attribute information, the three-dimensional data encoding device may also encode the multiple attribute information in sequence. For example, in the case of encoding color, reflectivity, and frame index as attribute information, the three-dimensional data encoding device may also generate the following bit stream: the encoding result of reflectivity is attached after the encoding result of color, and the encoding result of frame index is attached after the encoding result of reflectivity. In addition, the order of multiple encoding results of attribute information attached to the bit stream is not limited to this order, and can be any order. In addition, the three-dimensional data encoding device encodes the frame index as attribute information in the same data form as other attribute information such as color or reflectivity that is different from the frame index. Therefore, the encoded data contains the frame index in the same data form as other attribute information different from the frame index.

[0579] In addition, the three-dimensional data encoding device may also attach information indicating the starting position of the encoded data of each attribute information in the bit stream to the header, etc. Thus, the three-dimensional data decoding device can selectively decode the attribute information that needs to be decoded, thereby omitting the decoding process of the attribute information that does not need to be decoded. Therefore, the processing amount of the three-dimensional data decoding device can be reduced. In addition, the three-dimensional data encoding device may also encode multiple attribute information in parallel and integrate the encoding results into one bit stream. Thus, the three-dimensional data encoding device can encode multiple attribute information at high speed.

[0580] Fig.85 4 is a flowchart of the attribute information encoding process (S5604). First, the three-dimensional data encoding device sets the LoD (S5611). That is, the three-dimensional data encoding device assigns each three-dimensional point to any one of a plurality of LoDs.

[0581] Next, the three-dimensional data encoding device starts a loop in LoD units (S5612). That is, the three-dimensional data encoding device repeatedly performs the processing of steps S5613 to S5621 for each LoD.

[0582] Next, the three-dimensional data encoding device starts a loop in three-dimensional point units (S5613). That is, the three-dimensional data encoding device repeatedly performs the processing of steps S5614 to S5620 for each three-dimensional point.

[0583] First, the three-dimensional data encoding device searches for three-dimensional points that are used in calculating the predicted value of the object three-dimensional point of the processing object, that is, multiple surrounding points (S5614). Next, the three-dimensional data encoding device calculates the weighted average of the values ​​of the attribute information of the multiple surrounding points, and sets the obtained value as the predicted value P (S5615). Next, the three-dimensional data encoding device calculates the difference between the attribute information of the object three-dimensional point and the predicted value, that is, the prediction residual (S5616). Next, the three-dimensional data encoding device calculates a quantized value by quantizing the prediction residual (S5617). Next, the three-dimensional data encoding device performs arithmetic coding on the quantized value (S5618).

[0584] In addition, the three-dimensional data encoding device calculates an inverse quantization value by inverse quantizing the quantization value (S5619). Next, the three-dimensional data encoding device generates a decoded value by adding a prediction value to the inverse quantization value (S5620). Next, the three-dimensional data encoding device ends the loop of the three-dimensional point unit (S5621). In addition, the three-dimensional data encoding device ends the loop of the LoD unit (S5622).

[0585] Hereinafter, a three-dimensional data decoding process in a three-dimensional data decoding device for decoding a bit stream generated by the three-dimensional data encoding device described above will be described.

[0586] The three-dimensional data decoding device generates decoded binary data by performing arithmetic decoding on the binary data of the attribute information in the bit stream generated by the three-dimensional data encoding device in the same method as the three-dimensional data encoding device. In addition, in the three-dimensional data encoding device, when the application method of arithmetic coding is switched between the part binarized by n bits (n-bit coding) and the part binarized by exponential Golomb (residual coding), the three-dimensional data decoding device performs decoding in accordance with the arithmetic decoding when the arithmetic decoding is applied.

[0587] For example, in an arithmetic decoding method for n-bit coding, a three-dimensional data decoding device uses a different coding table (decoding table) to perform arithmetic decoding on each bit. At this time, the three-dimensional data decoding device may also change the number of coding tables used for each bit. For example, one coding table is used to perform arithmetic decoding on the leading bit b0 of the n-bit coding. In addition, the three-dimensional data decoding device uses two coding tables for the next bit b1. In addition, the three-dimensional data decoding device switches the coding table used in the arithmetic decoding of the bit b1 according to the value of b0 (0 or 1). Similarly, the three-dimensional data decoding device further uses four coding tables for the next bit b2. In addition, the three-dimensional data decoding device switches the coding table used in the arithmetic decoding of the bit b2 according to the values ​​of b0 and b1 (0 to 3).

[0588] Thus, the three-dimensional data decoding device uses 2 when performing arithmetic decoding on each bit bn-1 of the n-bit code. n-1 In addition, the three-dimensional data decoding device switches the coding table to be used according to the value (occurrence pattern) of the bit before bn-1. As a result, the three-dimensional data decoding device can use an appropriate coding table for each bit to appropriately decode the bit stream with improved coding efficiency.

[0589] In addition, the three-dimensional data decoding device may also reduce the number of coding tables used for each bit. For example, when performing arithmetic decoding on each bit bn-1, the three-dimensional data decoding device may switch between two encoding tables according to the value (occurrence pattern) of the m bits (m<n-1) before bn-1. m The three-dimensional data decoding device can appropriately decode the bit stream with improved coding efficiency while suppressing the number of coding tables used in each bit. In addition, the three-dimensional data decoding device can also update the occurrence probability of 0 and 1 in each coding table according to the value of the binary data actually generated. In addition, the three-dimensional data decoding device can also fix the occurrence probability of 0 and 1 in the coding table of a part of the bits. In this way, the number of updates of the occurrence probability can be suppressed, so the processing amount can be reduced.

[0590] For example, when the n-bit code is b0b1b2…bn-1, there is one coding table for b0 (CTb0). There are two coding tables for b1 (CTb10, CTb11). In addition, the coding table is switched according to the value of b0 (0 to 1). There are four coding tables for b2 (CTb20, CTb21, CTb22, CTb23). In addition, the coding table is switched according to the values ​​of b0 and b1 (0 to 3). There are 2 coding tables for bn-1. n-1 (CTbn0, CTbn1, ..., CTbn(2 n-1 -1)). In addition, according to the value of b0b1…bn-2 (0~2 n-1 -1) to switch the encoding table.

[0591] For example, Fig.86 2 is a diagram for explaining the processing when the residual code is an exponential Golomb code. Fig.86 As shown, the part (remaining code) that the three-dimensional data encoding device binarizes using the exponential Golomb and encodes includes the prefix part and the suffix part. For example, the three-dimensional data decoding device switches the encoding table between the prefix part and the suffix part. That is, the three-dimensional data decoding device uses the encoding table for the prefix to perform arithmetic decoding on each bit included in the prefix part, and uses the encoding table for the suffix to perform arithmetic decoding on each bit included in the suffix part.

[0592] In addition, the three-dimensional data decoding device may also update the occurrence probabilities of 0 and 1 in each coding table according to the value of the binary data generated during decoding. Alternatively, the three-dimensional data decoding device may also fix the occurrence probabilities of 0 and 1 in a certain coding table. Thus, the number of updates of the occurrence probabilities can be suppressed, thereby reducing the amount of processing. For example, the three-dimensional data decoding device may also update the occurrence probability for the prefix part and fix the occurrence probability for the suffix part.

[0593] In addition, the three-dimensional data decoding device converts the binary data of the prediction residual obtained by arithmetic decoding into multiple values ​​in accordance with the encoding method used in the three-dimensional data encoding device, thereby decoding the quantized prediction residual (unsigned integer value). The three-dimensional data decoding device first calculates the value of the n-bit code decoded by arithmetic decoding the binary data of the n-bit code. Then, the three-dimensional data decoding device compares the value of the n-bit code with the value of R_TH.

[0594] When the value of the n-bit code is consistent with the value of R_TH, the three-dimensional data decoding device determines that there is a bit encoded by Exponential Golomb next, and performs arithmetic decoding on the binary data encoded by Exponential Golomb, that is, the residual code. Then, the three-dimensional data decoding device calculates the value of the residual code based on the decoded residual code using a back-calculation table indicating the relationship between the residual code and the value. Fig.87 : is a diagram showing an example of a back-estimation table showing the relationship between the residual code and its value. Next, the three-dimensional data decoding apparatus adds the obtained residual code value to R_TH to obtain a multi-valued quantized prediction residual.

[0595] On the other hand, when the value of the n-bit code is inconsistent with the value of R_TH (the value is smaller than R_TH), the three-dimensional data decoding device directly determines the value of the n-bit code as the prediction residual after quantization. Thus, the three-dimensional data decoding device can appropriately decode the bit stream generated by switching the binarization method according to the value of the prediction residual in the three-dimensional data encoding device.

[0596] Furthermore, when the threshold R_TH is added to the header of the bit stream, the three-dimensional data decoding device may decode the value of the threshold R_TH from the header and use the decoded value of the threshold R_TH to switch the decoding method. Furthermore, when the threshold R_TH is added to the header for each LoD, the three-dimensional data decoding device may switch the decoding method using the threshold R_TH decoded for each LoD.

[0597] For example, when the threshold R_TH is 63 and the decoded n-bit code value is 63, the three-dimensional data decoding device decodes the remaining code using the Exponential Golomb method to obtain the value of the remaining code. Fig.87 In the example shown, the remaining code is 00100, and the remaining code value is 3. Next, the three-dimensional data decoding device adds the value 63 of the threshold value R_TH and the value 3 of the remaining code to obtain a prediction residual value 66.

[0598] Furthermore, when the decoded n-bit code value is 32, the three-dimensional data decoding apparatus sets the n-bit code value 32 as the value of the prediction residual.

[0599] In addition, the three-dimensional data decoding device converts the decoded quantized prediction residual from an unsigned integer value to a signed integer value by, for example, processing opposite to the processing in the three-dimensional data encoding device. Thus, the three-dimensional data decoding device can appropriately decode a bit stream generated without considering the generation of negative integers when entropy encoding the prediction residual. In addition, the three-dimensional data decoding device does not necessarily need to convert the unsigned integer value to a signed integer value, and can also decode the sign bit when decoding a bit stream generated by separately entropy encoding the sign bit.

[0600] The three-dimensional data decoding device decodes the quantized prediction residual transformed into a signed integer value through inverse quantization and reconstruction, thereby generating a decoded value. In addition, the three-dimensional data decoding device uses the generated decoded value for subsequent prediction of the three-dimensional point of the decoding object. Specifically, the three-dimensional data decoding device calculates the inverse quantization value by multiplying the quantized prediction residual by the decoded quantization scale, and adds the inverse quantization value and the prediction value to obtain the decoded value.

[0601] The decoded unsigned integer value (unsigned quantized value) is converted into a signed integer value by the following processing. When the LSB (least significant bit) of the decoded unsigned integer value a2u is 1, the three-dimensional data decoding device sets the signed integer value a2q to -((a2u+1)>>1). When the LSB of the unsigned integer value a2u is not 1, the three-dimensional data decoding device sets the signed integer value a2q to (a2u>>1).

[0602] Similarly, when the LSB of the decoded unsigned integer value b2u is 1, the three-dimensional data decoding device sets the signed integer value b2q to -((b2u+1)>>1). When the LSB of the unsigned integer value n2u is not 1, the three-dimensional data decoding device sets the signed integer value b2q to (b2u>>1).

[0603] Note that details of the inverse quantization and reconstruction processing performed by the three-dimensional data decoding device are the same as those of the inverse quantization and reconstruction processing in the three-dimensional data encoding device.

[0604] The following describes the flow of processing in the three-dimensional data decoding device. Fig.88 3D data decoding processing performed by a 3D data decoding device. First, the 3D data decoding device decodes position information (geometry) from a bit stream (S5631). For example, the 3D data decoding device performs decoding using an octree representation.

[0605] Next, the three-dimensional data decoding device decodes attribute information (Attribute) from the bit stream (S5632). For example, in the case of decoding multiple types of attribute information, the three-dimensional data decoding device may also decode the multiple types of attribute information in sequence. For example, in the case of decoding color, reflectivity, and frame index as attribute information, the three-dimensional data decoding device decodes the encoding result of color, the encoding result of reflectivity, and the encoding result of frame index in the order in which they are attached to the bit stream. For example, in the case where the encoding result of reflectivity is attached after the encoding result of color in the bit stream, the three-dimensional data decoding device decodes the encoding result of color and then decodes the encoding result of reflectivity. In addition, in the case where the encoding result of frame index is attached after the encoding result of reflectivity in the bit stream, the three-dimensional data decoding device decodes the encoding result of frame index after the decoding of the encoding result of reflectivity. In addition, the three-dimensional data decoding device may decode the encoding results of attribute information attached to the bit stream in any order.

[0606] In addition, the three-dimensional data decoding device can also obtain information indicating the starting position of the coded data of each attribute information in the bit stream by decoding the header. As a result, the three-dimensional data decoding device can selectively decode the attribute information that needs to be decoded, so that the decoding process of the attribute information that does not need to be decoded can be omitted. Therefore, the processing amount of the three-dimensional data decoding device can be reduced. In addition, the three-dimensional data decoding device can also decode multiple attribute information in parallel and integrate the decoding results into one three-dimensional point group. As a result, the three-dimensional data decoding device can decode multiple attribute information at high speed.

[0607] Next, the 3D data decoding device divides the decoded 3D point group into a plurality of frames based on the value of the frame index decoded together with the position information of each 3D point (S5633). For example, when the frame index of the decoded 3D point a is 0, the 3D data decoding device adds the position information and attribute information of the 3D point a to frame 0, and when the frame index of the decoded 3D point b is 1, the 3D data decoding device adds the position information and attribute information of the 3D point b to frame 1, thereby dividing the 3D point group obtained by decoding into a plurality of 3D point groups belonging to different frames.

[0608] Fig.89 4 is a flowchart of the attribute information decoding process (S5632). First, the three-dimensional data decoding device sets the LoD (S5641). That is, the three-dimensional data decoding device assigns a plurality of three-dimensional points having decoded position information to any one of a plurality of LoDs. For example, the assignment method is the same method as the assignment method used in the three-dimensional data encoding device.

[0609] Next, the three-dimensional data decoding device starts a loop in LoD units (S5642). That is, the three-dimensional data decoding device repeatedly performs the processing of steps S5643 to S5649 for each LoD.

[0610] Next, the three-dimensional data decoding device starts a three-dimensional point unit loop (S5643). That is, the three-dimensional data decoding device repeatedly performs the processing of steps S5644 to S5648 for each three-dimensional point.

[0611] First, the three-dimensional data decoding device searches for three-dimensional points that are used in calculating the predicted value of the object three-dimensional point of the processing object, that is, multiple surrounding points (S5644). Next, the three-dimensional data decoding device calculates the weighted average of the values ​​of the attribute information of the multiple surrounding points and sets the obtained value as the predicted value P (S5645). In addition, these processes are the same as those in the three-dimensional data encoding device.

[0612] Next, the three-dimensional data decoding device performs arithmetic decoding on the quantized value from the bit stream (S5646). In addition, the three-dimensional data decoding device calculates an inverse quantized value by inverse quantizing the decoded quantized value (S5647). Next, the three-dimensional data decoding device generates a decoded value by adding a predicted value to the inverse quantized value (S5648). Next, the three-dimensional data decoding device ends the loop of the three-dimensional point unit (S5649). In addition, the three-dimensional data decoding device ends the loop of the LoD unit (S5650).

[0613] Next, the configurations of the three-dimensional data encoding device and the three-dimensional data decoding device according to the present embodiment will be described. Fig.903D data encoding device 5600 of this embodiment is a block diagram showing a structure of the 3D data encoding device 5600. The 3D data encoding device 5600 includes a frame combining unit 5601, a position information encoding unit 5602, an attribute information reallocation unit 5603, and an attribute information encoding unit 5604.

[0614] The frame combining unit 5601 combines a plurality of frames. The position information encoding unit 5602 encodes the position information (geometry) of a plurality of three-dimensional points included in the input point group. The attribute information reallocation unit 5603 reallocates the values ​​of the attribute information of the plurality of three-dimensional points included in the input point group using the encoding and decoding results of the position information. The attribute information encoding unit 5604 encodes the reallocated attribute information (attribute). In addition, the three-dimensional data encoding device 5600 generates a bit stream including the encoded position information and the encoded attribute information.

[0615] Fig.91 3D data decoding device 5610 is a block diagram showing a structure of the present embodiment. The 3D data decoding device 5610 includes a position information decoding unit 5611 , a property information decoding unit 5612 , and a frame division unit 5613 .

[0616] The position information decoding unit 5611 decodes the position information (geometry) of the plurality of three-dimensional points from the bit stream. The attribute information decoding unit 5612 decodes the attribute information (attribute) of the plurality of three-dimensional points from the bit stream. The frame segmentation unit 5613 segments the decoded three-dimensional point group into a plurality of frames based on the value of the frame index decoded together with the position information of each three-dimensional point. In addition, the three-dimensional data decoding device 5610 generates an output point group by combining the decoded position information with the decoded attribute information.

[0617] Fig.92 It is a diagram showing the structure of attribute information. Fig.92 (a) is a diagram showing the structure of compressed attribute information. Fig.92 (b) is a diagram showing an example of the syntax of the header of attribute information. Fig.92 (c) is a diagram showing an example of the syntax of the payload (data) of attribute information.

[0618] like Fig.92As shown in (b), the syntax of the header of the attribute information is explained. apx_idx represents the ID of the corresponding parameter set. In apx_idx, multiple IDs can also be represented when there is a parameter set for each frame. offset represents the offset position used to obtain the combined data. other_attribute_information represents other attribute data such as QPΔ (Japanese: デルタ) representing the differential value of the quantization parameter. combine_frame_flag is a flag indicating whether the encoded data is combined by frames. number_of_combine_frame represents the number N of combined frames. number_of_combine_frame can be included in SPS or APS.

[0619] refer_different_frame is a flag indicating whether to use the same frame, or the attribute information of the surrounding 3D points belonging to the same frame and outside the same frame to encode / decode the attribute information of the object 3D point of the encoding / decoding object. For example, consider the allocation of values ​​as described below. When refer_different_frame is 0, the 3D data encoding device or the 3D data decoding device encodes / decodes the attribute information of the object 3D point using the attribute information of the surrounding 3D points in the same frame as the object 3D point. In this case, the 3D data encoding device or the 3D data decoding device does not encode / decode the attribute information of the object 3D point using the attribute information of the surrounding 3D points in a frame different from the object 3D point.

[0620] On the other hand, when refer_different_frame is 1, the 3D data encoding device or the 3D data decoding device encodes / decodes the attribute information of the object 3D point using the attribute information of the surrounding 3D points in the same frame as the frame to which the object 3D point belongs and outside the same frame. That is, the 3D data encoding device or the 3D data decoding device encodes / decodes the attribute information of the object 3D point using the attribute information of the surrounding 3D points regardless of whether the frame to which the object 3D point belongs belongs to the same frame.

[0621] As the attribute information of the object 3D point, an example of encoding color information or reflectivity information using the attribute information of the surrounding 3D points is shown, but the frame index of the object 3D point can also be encoded using the frame index of the surrounding 3D points. The 3D data encoding device can also, for example, set the frame index attached to each 3D point when combining multiple frames as the attribute information of each 3D point, and encode it using the prediction encoding method described in the present disclosure. For example, the 3D data encoding device can also calculate the predicted value of the frame index of the 3D point A based on the values ​​of the frame indexes of the surrounding 3D points B, C, and D of the 3D point A, and encode the prediction residual. As a result, the 3D data encoding device can reduce the amount of bits used to encode the frame index, and can improve the encoding efficiency.

[0622] Fig.93 This is a diagram for explaining encoded data.

[0623] When the point group data includes attribute information, the attribute information may be frame-bound. The attribute information is encoded or decoded with reference to the position information. The referenced position information may be the position information before the frame-bound or the position information after the frame-bound. The number of frames bound to the position information may be the same as the number of frames bound to the attribute information, or may be independent and different.

[0624] Fig.93 The numerical value in the brackets in represents the frame. For example, in the case of 1, it indicates the information of frame 1, and in the case of 1-4, it indicates the information of the combined frames 1 to 4. In addition, G indicates the position information, and A indicates the attribute information. Frame_idx1 is the frame index of frame 1.

[0625] Fig.93 (a) represents an example of a case where refer_different_frame is 1. When refer_different_frame is 1, the three-dimensional data encoding device or the three-dimensional data decoding device encodes or decodes A(1-4) based on the information of G(1-4). During decoding, the three-dimensional data decoding device uses Frame_idx1-4 decoded together with G(1-4) to divide G(1-4) and A(1-4) into Frame1-4. In addition, when encoding or decoding A(1-4), the three-dimensional data encoding device or the three-dimensional data decoding device may also refer to other attribute information of A(1-4). That is, when encoding or decoding A(1), the three-dimensional data encoding device or the three-dimensional data decoding device may refer to other A(1) or A(2-4). In addition, arrows represent the reference source and reference target of information, the direction of the arrow represents the reference source, and the direction of the arrow represents the reference target.

[0626] Fig.93(b) shows an example of a case where refer_different_frame is 0. When refer_different_frame is 0, unlike when refer_different_frame is 1, the 3D data encoding device or the 3D data decoding device does not refer to attribute information of different frames. That is, when encoding or decoding A(1), the 3D data encoding device or the 3D data decoding device refers to other A(1) and does not refer to A(2-4).

[0627] Fig.93 (c) shows another example in which refer_different_frame is 0. In this case, the position information is encoded in the combined frames, but the attribute information is encoded for each frame. Therefore, when encoding or decoding A(1), the three-dimensional data encoding device or the three-dimensional data decoding device refers to other A(1). Similarly, when encoding or decoding the attribute information, other attribute information belonging to the same frame is referred to. In addition, A(1-4) can also attach each APS to the header.

[0628] As described above, the three-dimensional data encoding device of this embodiment performs Fig.94 The processing shown. The three-dimensional data encoding device obtains the third point group data, which combines the first point group data and the second point group data, and includes the position information of each of the multiple three-dimensional points included in the third point group data, and identification information indicating to which of the first point group data and the second point group data the multiple three-dimensional points belong respectively (S5661). Next, the three-dimensional data encoding device generates encoded data by encoding the obtained third point group data (S5662). In generating the encoded data, the three-dimensional data encoding device encodes the identification information of each of the multiple three-dimensional points as the attribute information of the three-dimensional point.

[0629] Therefore, the three-dimensional data encoding method can improve the encoding efficiency by collectively encoding multiple point group data.

[0630] For example, in the generation of the encoded data (S5662), the attribute information of the first three-dimensional point is encoded using the attribute information of the second three-dimensional points around the first three-dimensional point among the plurality of three-dimensional points.

[0631] For example, the attribute information of the first three-dimensional point includes first identification information indicating that the first three-dimensional point belongs to the first point group data, and the attribute information of the second three-dimensional point includes second identification information indicating that the second three-dimensional point belongs to the second point group data.

[0632] For example, in the generation of encoded data (S5662), the attribute information of the second three-dimensional point is used to calculate the predicted value of the attribute information of the first three-dimensional point, and the difference between the attribute information of the first three-dimensional point and the predicted value, i.e., the prediction residual, is calculated to generate encoded data including the prediction residual.

[0633] For example, in the acquisition (S5661), the third point group data is generated by combining the first point group data and the second point group data, thereby acquiring the third point group data.

[0634] For example, the encoded data contains the identification information in the same data form as other attribute information different from the identification information.

[0635] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.

[0636] In addition, the three-dimensional data decoding device of this embodiment performs Fig.95 The three-dimensional data decoding device obtains the encoded data (S5671). Next, the three-dimensional data decoding device obtains the position information and attribute information of each of the plurality of three-dimensional points included in the third point group data combining the first point group data and the second point group data by decoding the encoded data (S5672). In addition, the attribute information includes identification information indicating which of the first point group data and the second point group data the three-dimensional point corresponding to the attribute information belongs to.

[0637] Thus, the three-dimensional data decoding device can decode the encoded data with improved encoding efficiency by collectively encoding a plurality of point group data.

[0638] For example, in the obtaining (S5671), the attribute information of the first three-dimensional point is decoded using the attribute information of the second three-dimensional points around the first three-dimensional point among the plurality of three-dimensional points.

[0639] For example, the attribute information of the first three-dimensional point includes first identification information indicating that the first three-dimensional point belongs to the first point group data, and the attribute information of the second three-dimensional point includes second identification information indicating that the second three-dimensional point belongs to the second point group data.

[0640] For example, the encoded data includes a prediction residual. Then, in decoding (S5672) of the encoded data, the predicted value of the attribute information of the first three-dimensional point is calculated using the attribute information of the second three-dimensional point, and the predicted value and the prediction residual are added to thereby calculate the attribute information of the first three-dimensional point.

[0641] For example, the three-dimensional data decoding device further uses the identification information to divide the third three-dimensional point group data into the first point group data and the second point group data.

[0642] For example, the encoded data contains the identification information in the same data form as other attribute information different from the identification information.

[0643] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.

[0644] As mentioned above, the three-dimensional data encoding device and the three-dimensional data decoding device and the like according to the embodiments of the present disclosure have been described, but the present disclosure is not limited to these embodiments.

[0645] Furthermore, each processing unit included in the three-dimensional data encoding device and three-dimensional data decoding device of the above-mentioned embodiment can be typically implemented as an LSI of an integrated circuit. These can be made into one chip separately, or part or all of them can be made into one chip.

[0646] Furthermore, integrated circuits are not limited to LSIs, and can be implemented by dedicated circuits or general-purpose processors. FPGAs (Field Programmable Gate Arrays) that are programmable after LSI manufacturing, or reconfigurable processors that can reconfigure the connections or settings of circuits within LSIs can also be used.

[0647] Furthermore, in each of the above-mentioned embodiments, each component may be formed by dedicated hardware, or may be implemented by executing a software program suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or a processor reading and executing a software program recorded on a recording medium such as a hard disk or a semiconductor memory.

[0648] Furthermore, the present disclosure can be implemented as a three-dimensional data encoding method or a three-dimensional data decoding method, etc., which are executed by a three-dimensional data encoding device or a three-dimensional data decoding device, etc.

[0649] Furthermore, the division of the functional blocks in the block diagram is an example, and multiple functional blocks can be implemented as one functional block, and one functional block can also be divided into multiple blocks, and a part of the functions can also be moved to other functional blocks. Furthermore, the functions of multiple functional blocks with similar functions can also be processed in parallel or in time division by a single hardware or software.

[0650] Furthermore, the execution order of each step in the flowchart is an example given for the purpose of specifically describing the present disclosure, and may be an order other than the above. Furthermore, part of the above steps may be executed simultaneously (in parallel) with other steps.

[0651] The above describes one or more forms of a three-dimensional data encoding device and a three-dimensional data decoding device based on the embodiments, but the present disclosure is not limited to these embodiments. Without departing from the scope of the present disclosure, forms obtained by performing various modifications that can be thought of by those skilled in the art on the present embodiment, and forms obtained by combining constituent elements in different embodiments are all included in the scope of one or more forms.

[0652] Industrial Applicability

[0653] The present disclosure is applicable to a three-dimensional data encoding device and a three-dimensional data decoding device.

[0654] Description of Reference Numerals

[0655] 4601 Three-dimensional data encoding system

[0656] 4602 3D data decoding system

[0657] 4603 Sensor Terminal

[0658] 4604 External connection

[0659] 4611 Point Group Data Generation System

[0660] 4612 Prompt Department

[0661] 4613 Coding Department

[0662] 4614 Multiplexing Department

[0663] 4615 Input and Output

[0664] 4616 Control Department

[0665] 4617 Sensor Information Acquisition Department

[0666] 4618 Point Group Data Generation Department

[0667] 4621 Sensor Information Acquisition Department

[0668] 4622 Input and Output

[0669] 4623 Inverse Multiplexing Department

[0670] 4624 Decoding Department

[0671] 4625 Prompt Department

[0672] 4626 User Interface

[0673] 4627 Control Department

[0674] 4630 1st Coding Department

[0675] 4631 Position Information Coding Unit

[0676] 4632 Attribute Information Coding Unit

[0677] 4633 Additional Information Coding Department

[0678] 4634 Multiplexing Department

[0679] 4640 Decoding Unit 1

[0680] 4641 Inverse Multiplexing Department

[0681] 4642 Position information decoding unit

[0682] 4643 Attribute information decoding unit

[0683] 4644 Additional information decoding unit

[0684] 4650 No. 2 Coding Department

[0685] 4651 Additional Information Generation Department

[0686] 4652 Position Image Generation Unit

[0687] 4653 Attribute Image Generation Unit

[0688] 4654 Video Coding Department

[0689] 4655 Additional Information Coding Department

[0690] 4656 Multiplexing Department

[0691] 4660 Decoding Unit 2

[0692] 4661 Inverse Multiplexing Department

[0693] 4662 Image Decoding Department

[0694] 4663 Additional information decoding unit

[0695] 4664 Position Information Generation Unit

[0696] 4665 Attribute Information Generation Unit

[0697] 4801 Coding Department

[0698] 4802 Multiplexing Department

[0699] 4910 1st Coding Department

[0700] 4911 Division

[0701] 4912 Position Information Coding Unit

[0702] 4913 Attribute Information Coding Unit

[0703] 4914 Additional Information Coding Department

[0704] 4915 Multiplexing Department

[0705] 4920 Decoding Unit 1

[0706] 4921 Inverse Multiplexing Department

[0707] 4922 Position information decoding unit

[0708] 4923 Attribute Information Decoding Unit

[0709] 4924 Additional information decoding unit

[0710] 4925 Joint

[0711] 4931 Slice Division

[0712] 4932 Position Information Tile Segmentation Unit

[0713] 4933 Attribute Information Tile Division

[0714] 4941 Position information tile joint

[0715] 4942 Attribute information tile joint

[0716] 4943 Slice joint

[0717] 5410 Coding Department

[0718] 5411 Division

[0719] 5412 Position Information Coding Unit

[0720] 5413 Attribute Information Coding Unit

[0721] 5414 Additional Information Coding Department

[0722] 5415 Multiplexing Department

[0723] 5421 Tile Division

[0724] 5422 Slice Division

[0725] 5431, 5441 frame index generation unit

[0726] 5432, 5442 Entropy coding unit

[0727] 5450 Decoding Department

[0728] 5451 Inverse Multiplexing Unit

[0729] 5452 Position information decoding unit

[0730] 5453 Attribute information decoding unit

[0731] 5454 Additional information decoding unit

[0732] 5455 Joint

[0733] 5461, 5471 Entropy decoding unit

[0734] 5462, 5472 Frame index acquisition unit

[0735] 5600 3D Data Encoding Device

[0736] 5601 Frame combination unit

[0737] 5602 Position Information Coding Unit

[0738] 5603 Attribute Information Redistribution Department

[0739] 5604 Attribute Information Coding Unit

[0740] 5610 3D data decoding device

[0741] 5611 Position information decoding unit

[0742] 5612 Attribute information decoding unit

[0743] 5613 Frame segmentation unit

Claims

1. A three-dimensional data encoding method, in, obtaining third point group data in a third frame, wherein the third point group data includes geometric information of each of a plurality of three-dimensional points and a frame index, wherein the frame index indicates how to divide the plurality of three-dimensional points included in the third point group data into first point group data in a first frame and second point group data in a second frame, By encoding the obtained third point group data to generate encoded data, The frame index is encoded as attribute information associated with each of the plurality of three-dimensional points.

2. The three-dimensional data encoding method according to claim 1, in, In the generation, attribute information of the first three-dimensional point is encoded using attribute information of second three-dimensional points around the first three-dimensional point, and the first three-dimensional point and the second three-dimensional point are included in the plurality of three-dimensional points.

3. The three-dimensional data encoding method according to claim 2, in, The attribute information of the first three-dimensional point includes first identification information indicating that the first three-dimensional point belongs to the first point group data. The attribute information of the second three-dimensional point includes second identification information indicating that the second three-dimensional point belongs to the second point group data.

4. The three-dimensional data encoding method according to claim 2 or 3, in, In the generation, using the attribute information of the second three-dimensional point to calculate a predicted value of the attribute information of the first three-dimensional point, Calculating the difference between the attribute information of the first three-dimensional point and the predicted value, that is, the prediction residual, Generate coded data including the prediction residual.

5. A three-dimensional data decoding method, in, Get the encoded data, By decoding the encoded data, third point group data in the third frame is obtained, the third point group data includes geometric information of each of the plurality of three-dimensional points and a frame index, the frame index indicates how to divide the plurality of three-dimensional points included in the third point group data into the first point group data in the first frame and the second point group data in the second frame, The frame index is decoded into attribute information associated with each of the plurality of three-dimensional points.

6. The three-dimensional data decoding method according to claim 5, in, In the obtaining, attribute information of the first three-dimensional point is decoded using attribute information of a second three-dimensional point around the first three-dimensional point, and the first three-dimensional point and the second three-dimensional point are included in the plurality of three-dimensional points.

7. The three-dimensional data decoding method according to claim 6, in, The attribute information of the first three-dimensional point includes first identification information indicating that the first three-dimensional point belongs to the first point group data. The attribute information of the second three-dimensional point includes second identification information indicating that the second three-dimensional point belongs to the second point group data.

8. The three-dimensional data decoding method according to claim 6 or 7, in, The coded data comprises a prediction residual, In decoding the encoded data, using the attribute information of the second three-dimensional point to calculate a predicted value of the attribute information of the first three-dimensional point, The attribute information of the first three-dimensional point is calculated by adding the predicted value to the prediction residual.

9. The three-dimensional data decoding method according to any one of claims 5 to 7, in, Further, The third point group data is divided into the first point group data and the second point group data using the attribute information.

10. A three-dimensional data encoding device, in, have: Processor; and Memory, The processor uses the memory and is configured as follows: obtaining third point group data in a third frame, wherein the third point group data includes geometric information of each of a plurality of three-dimensional points and a frame index, wherein the frame index indicates how to divide the plurality of three-dimensional points included in the third point group data into first point group data in a first frame and second point group data in a second frame, By encoding the obtained third point group data to generate encoded data, The frame index is encoded as attribute information associated with each of the plurality of three-dimensional points.

11. A three-dimensional data decoding device, in, have: Processor; and Memory, The processor uses the memory and is configured as follows: Get the encoded data, By decoding the encoded data, third point group data in the third frame is obtained, the third point group data includes geometric information of each of the plurality of three-dimensional points and a frame index, the frame index indicates how to divide the plurality of three-dimensional points included in the third point group data into the first point group data in the first frame and the second point group data in the second frame, The frame index is decoded into attribute information associated with each of the plurality of three-dimensional points.

Citation Information

Patent Citations

  • Map display device

    WO2014020663A1

  • Region-adaptive hierarchical transform and entropy coding for point cloud compression, and corresponding decompression

    US20170347100A1

  • Method for creating three-dimensional data and device for creating three-dimensional data

    WO2018051746A1