Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device

By calculating predicted values using position and angle information, the method improves encoding efficiency and supports mixed codecs, addressing the inefficiencies in existing three-dimensional data encoding and decoding methods.

JP2025108616AActive Publication Date: 2025-07-23PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025067955
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-11-13
Filing Date
2025-04-17
Publication Date
2025-07-23
Estimated Expiration
2040-11-13

AI Technical Summary

Technical Problem

Existing three-dimensional data encoding and decoding methods lack efficiency in compressing and transmitting large amounts of point cloud data, particularly in applications requiring mixed encoding formats and network bandwidth optimization.

Method used

A method that calculates predicted values using position and angle information of reference points in point cloud data to improve encoding efficiency, incorporating scanning angles for enhanced prediction accuracy and encoding methods that support mixed codecs like PCC.

Benefits of technology

Enhances encoding efficiency by improving prediction accuracy and supporting mixed encoding formats, facilitating efficient transmission and decoding of three-dimensional data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025108616000001_ABST
    Figure 2025108616000001_ABST
Patent Text Reader

Abstract

To improve coding efficiency.SOLUTION: A three-dimensional data encoding method includes: calculating, by a processor, a first predicted value of position information of a first three-dimensional point included in point cloud data, using position information of a reference three-dimensional point included in the point cloud data and angle information (S10061); calculating, by the processor, a first difference value between the position information of the first three-dimensional point and the first predicted value (S10062); calculating, by the processor, a second predicted value of position information of a second three-dimensional point included in the point cloud data, using the position information of the first three-dimensional point and the angle information; and calculating, by the processor, a second difference value between the position information of the second three-dimensional point and the second predicted value.SELECTED DRAWING: Figure 113
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding apparatus, and a three-dimensional data decoding apparatus.

Background Art

[0002] In the future, the spread of devices or services using three-dimensional data is expected in a wide range of fields such as computer vision, map information, monitoring, infrastructure inspection, or video distribution for autonomous operation of automobiles or robots. Three-dimensional data is acquired by various methods such as a distance sensor such as a range finder, a stereo camera, or a combination of a plurality of monocular cameras.

[0003] As one of the methods for expressing three-dimensional data, there is a method called point cloud that represents the shape of a three-dimensional structure by a point group in a three-dimensional space. In a point cloud, the position and color of the point group are stored. Although point cloud is expected to become mainstream as a method for expressing three-dimensional data, the amount of data of the point group is very large. Therefore, in the accumulation or transmission of three-dimensional data, as with two-dimensional moving images (for example, MPEG-4 AVC or HEVC standardized by MPEG), compression of the amount of data by encoding is essential.

[0004] Also, regarding the compression of point cloud, it is partially supported by a public library (Point Cloud Library) that performs point cloud-related processing.

[0005] Also, a technique is known for searching for and displaying facilities located around a vehicle using three-dimensional map data (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0006]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0007] In the encoding process and decoding process of three-dimensional data, it is desired to improve the encoding efficiency.

[0008] The present disclosure aims to provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device capable of improving the encoding efficiency.

Means for Solving the Problems

[0009] A three-dimensional data encoding method according to an aspect of the present disclosure calculates, by a processor, a first predicted value of position information of a first three-dimensional point included in point cloud data using position information and angle information of a reference three-dimensional point included in the point cloud data, calculates, by the processor, a first difference value between the position information of the first three-dimensional point and the first predicted value, calculates, by the processor, a second predicted value of position information of a second three-dimensional point included in the point cloud data using the position information of the first three-dimensional point and the angle information, and calculates, by the processor, a second difference value between the position information of the second three-dimensional point and the second predicted value.

[0010] A three-dimensional data decoding method according to an aspect of the present disclosure calculates, by a processor, a first predicted value using position information and angle information of a reference three-dimensional point included in point cloud data, calculates, by the processor, the position information of a first three-dimensional point included in the point cloud data by adding the first predicted value and a first difference value, calculates, by the processor, a second predicted value using the position information of the first three-dimensional point and the angle information, and calculates, by the processor, the position information of a second three-dimensional point included in the point cloud data by adding the second predicted value and a second difference value.

Advantages of the Invention

[0011] The present disclosure can provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding apparatus, or a three-dimensional data decoding apparatus that can improve encoding efficiency.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Figure 36

Figure 37

Figure 38

Figure 39

Figure 40

Figure 41

Figure 42

Figure 43

Figure 44

Figure 45

Figure 46

Figure 47

Figure 48

Figure 49

Figure 50

Figure 51

Figure 52

Figure 53

Figure 54

Figure 55

Figure 56

Figure 57

Figure 58

Figure 59

Figure 60

Figure 61

Figure 62

Figure 63

Figure 64

Figure 65

Figure 66

Figure 67

Figure 68

Figure 69

Figure 70

Figure 71

Figure 72

Figure 73

Figure 74

Figure 75

Figure 76

Figure 77

Figure 78

Figure 79

Figure 80

Figure 81

Figure 82

Figure 83

Figure 84

Figure 85

Figure 86

Figure 87

Figure 88

Figure 89

Figure 90

Figure 91

Figure 92

Figure 93

Figure 94

Figure 95

Figure 96

Figure 97

Figure 98

Figure 99

Figure 100

Figure 101

Figure 102

Figure 103

Figure 104

Figure 105

Figure 106

Figure 107

Figure 108

Figure 109

Figure 110

Figure 111

Figure 112

Figure 113

Figure 114

Figure 115

Figure 116

Figure 117

Figure 118

Figure 119

Figure 120

Figure 121

Figure 122

Figure 123

Figure 124

Figure 125

Figure 126

Figure 127

Figure 128

Figure 129

Figure 130

Figure 131

Figure 132

Figure 133

Figure 134

Figure 135

Figure 136

Figure 137

Figure 138

Embodiments for Carrying Out the Invention

[0013] In the three-dimensional data encoding method according to one aspect of the present disclosure, a first predicted value of the position information of a first three-dimensional point included in point cloud data obtained by a sensor is generated using the position information of a reference three-dimensional point included in the point cloud data and the scanning angle of the sensor, a first difference value between the position information of the first three-dimensional point and the first predicted value is calculated, and a bit stream including the first difference value and information indicating the scanning angle is generated.

[0014] According to this, since the three-dimensional data encoding method can improve the prediction accuracy by generating a predicted value using the scanning angle of the sensor that generated the point cloud data, the encoding efficiency can be improved.

[0015] For example, the position information of the first three-dimensional point may be represented by xyz coordinates, and the scanning angle may be a scanning angle in the xy plane.

[0016] For example, the sensor may be a radar that performs rotational scanning for each scanning angle.

[0017] For example, in the generation of the first predicted value, the first predicted value may be generated by rotating and moving the position information of the reference three-dimensional point immediately before the first three-dimensional point in the scanning order included in the point cloud data by the scanning angle around a reference point corresponding to the position of the sensor.

[0018] For example, a second predicted value may be generated using a value obtained by multiplying a reference predicted value by n (n is an integer of 1 or more), a second difference value between the position information of the first three-dimensional point and the second predicted value may be calculated, and the bit stream may include the second difference value and information indicating the n.

[0019] According to this, since the three-dimensional data encoding method can improve the prediction accuracy, the encoding efficiency can be improved.

[0020] For example, the reference predicted value may correspond to the difference between the position information of two reference three-dimensional points before the first three-dimensional point in the scanning order included in the point cloud data.

[0021] For example, the reference predicted value may be generated using the scanning angle.

[0022] A three-dimensional data decoding method according to an aspect of the present disclosure acquires, from a bit stream, a first difference value between the position information of a first three-dimensional point included in point cloud data obtained by a sensor and a first predicted value, and information indicating the scanning angle of the sensor, generates the first predicted value using the position information of a reference three-dimensional point included in the point cloud data and the scanning angle of the sensor, and calculates the position information of the first three-dimensional point by adding the first predicted value and the first difference value.

[0023] According to this, since the three-dimensional data decoding method can improve the prediction accuracy by generating a predicted value using the scanning angle of the sensor that generated the point cloud data, the encoding efficiency can be improved.

[0024] For example, the position information of the first three-dimensional point may be represented by xyz coordinates, and the scanning angle may be a scanning angle in the xy plane.

[0025] For example, the sensor may be a radar that performs rotational scanning for each scanning angle.

[0026] For example, in the generation of the first predicted value, the position information of the reference three-dimensional point immediately before the first three-dimensional point in the scanning order included in the point cloud data may be rotationally moved by the scanning angle around a reference point corresponding to the position of the sensor to generate the first predicted value.

[0027] For example, the three-dimensional data decoding method may further obtain, from the bit stream, a second difference value between the position information of the second three-dimensional point included in the point cloud data and the second predicted value, and information indicating n (where n is an integer of 1 or more), generate the second predicted value using a value obtained by multiplying the reference predicted value by n, and calculate the position information of the second three-dimensional point by adding the second predicted value and the second difference value.

[0028] According to this, since the three-dimensional data decoding method can improve the prediction accuracy, the coding efficiency can be improved.

[0029] For example, the reference predicted value may correspond to the difference between the position information of the two reference three-dimensional points before the second three-dimensional point in the scanning order included in the point cloud data.

[0030] For example, the reference predicted value may be generated using the scanning angle.

[0031] Further, a three-dimensional data encoding apparatus according to an aspect of the present disclosure includes a processor and a memory. The processor uses the memory to generate a first predicted value of the position information of a first three-dimensional point included in point cloud data obtained by a sensor, using the position information of a reference three-dimensional point included in the point cloud data and the scanning angle of the sensor, calculates a first difference value between the position information of the first three-dimensional point and the first predicted value, and generates a bit stream including the first difference value and information indicating the scanning angle.

[0032] According to this, since the three-dimensional data encoding apparatus can improve the prediction accuracy by generating a predicted value using the scanning angle of the sensor that generated the point cloud data, the coding efficiency can be improved.

[0033] In addition, a three-dimensional data decoding apparatus according to an aspect of the present disclosure includes a processor and a memory. The processor uses the memory to obtain, from a bitstream, a first difference value between the position information of a first three-dimensional point included in point cloud data obtained by a sensor and a first predicted value, and information indicating the scanning angle of the sensor. The first predicted value is generated using the position information of a reference three-dimensional point included in the point cloud data and the scanning angle of the sensor, and the position information of the first three-dimensional point is calculated by adding the first predicted value and the first difference value.

[0034] According to this, the three-dimensional data decoding apparatus can improve the prediction accuracy by generating a predicted value using the scanning angle of the sensor that generated the point cloud data, and thus can improve the encoding efficiency.

[0035] These general or specific aspects may be implemented by a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.

[0036] Hereinafter, embodiments will be specifically described with reference to the drawings. Note that each of the embodiments described below shows a specific example of the present disclosure. The numerical values, shapes, materials, components, arrangement positions and connection forms of the components, steps, order of steps, etc. shown in the following embodiments are merely examples and are not intended to limit the present disclosure. In addition, among the components in the following embodiments, the components not described in the independent claims are described as optional components.

[0037] (Embodiment 1) When using the encoded data of the point cloud in an actual device or service, it is desirable to transmit and receive the necessary information according to the application in order to suppress the network bandwidth. However, until now, such a function has not existed in the encoding structure of three-dimensional data, and there has also been no encoding method therefor.

[0038] In this embodiment, a three-dimensional data encoding method, a three-dimensional data encoding device for providing a function of transmitting and receiving necessary information according to the application in the encoded data of three-dimensional point cloud, a three-dimensional data decoding method and a three-dimensional data decoding device for decoding the encoded data, a three-dimensional data multiplexing method for multiplexing the encoded data, and a three-dimensional data transmission method for transmitting the encoded data will be described.

[0039] In particular, currently, as an encoding method (encoding format) of point cloud data, a first encoding method and a second encoding method are being studied, but the configuration of the encoded data and the method of storing the encoded data in the system format are not defined, and there is a problem that MUX processing (multiplexing) in the encoding unit, or transmission or storage cannot be performed as it is.

[0040] Also, there has been no method to support a format in which two codecs, the first encoding method and the second encoding method, are mixed like PCC (Point Cloud Compression).

[0041] In this embodiment, the configuration of PCC encoded data in which two codecs, the first encoding method and the second encoding method, are mixed and the method of storing the encoded data in the system format will be described.

[0042] First, the configuration of the three-dimensional data (point cloud data) encoding / decoding system according to this embodiment will be described. FIG. 1 is a diagram showing a configuration example of the three-dimensional data encoding / decoding system according to this embodiment. As shown in FIG. 1, the three-dimensional data encoding / decoding system includes a three-dimensional data encoding system 4601, a three-dimensional data decoding system 4602, a sensor terminal 4603, and an external connection unit 4604.

[0043] The three-dimensional data encoding system 4601 generates encoded data or multiplexed data by encoding point cloud data which is three-dimensional data. Note that the three-dimensional data encoding system 4601 may be a three-dimensional data encoding device realized by a single device, or may be a system realized by a plurality of devices. Further, the three-dimensional data encoding device may include a part of a plurality of processing units included in the three-dimensional data encoding system 4601.

[0044] The three-dimensional data encoding system 4601 includes a point cloud data generation system 4611, a presentation unit 4612, an encoding unit 4613, a multiplexing unit 4614, an input / output unit 4615, and a control unit 4616. The point cloud data generation system 4611 includes a sensor information acquisition unit 4617 and a point cloud data generation unit 4618.

[0045] The sensor information acquisition unit 4617 acquires sensor information from the sensor terminal 4603 and outputs the sensor information to the point cloud data generation unit 4618. The point cloud data generation unit 4618 generates point cloud data from the sensor information and outputs the point cloud data to the encoding unit 4613.

[0046] The presentation unit 4612 presents the sensor information or the point cloud data to the user. For example, the presentation unit 4612 displays information or an image based on the sensor information or the point cloud data.

[0047] The encoding unit 4613 encodes (compresses) the point cloud data and outputs the obtained encoded data, control information obtained in the encoding process, and other additional information to the multiplexing unit 4614. The additional information includes, for example, sensor information.

[0048] The multiplexing unit 4614 generates multiplexed data by multiplexing the encoded data, control information, and additional information input from the encoding unit 4613. The format of the multiplexed data is, for example, a file format for storage or a packet format for transmission.

[0049] The input / output unit 4615 (e.g., a communication unit or an interface) outputs the multiplexed data to the outside. Alternatively, the multiplexed data is stored in a storage unit such as an internal memory. The control unit 4616 (or the application execution unit) controls each processing unit. That is, the control unit 4616 performs controls such as encoding and multiplexing.

[0050] Note that the sensor information may be input to the encoding unit 4613 or the multiplexing unit 4614. Also, the input / output unit 4615 may output the point cloud data or the encoded data to the outside as it is.

[0051] The transmission signal (multiplexed data) output from the three-dimensional data encoding system 4601 is input to the three-dimensional data decoding system 4602 via the external connection unit 4604.

[0052] The three-dimensional data decoding system 4602 generates point cloud data, which is three-dimensional data, by decoding the encoded data or the multiplexed data. Note that the three-dimensional data decoding system 4602 may be a three-dimensional data decoding device realized by a single device, or may be a system realized by a plurality of devices. Also, the three-dimensional data decoding device may include a part of a plurality of processing units included in the three-dimensional data decoding system 4602.

[0053] The three-dimensional data decoding system 4602 includes a sensor information acquisition unit 4621, an input / output unit 4622, a demultiplexing unit 4623, a decoding unit 4624, a presentation unit 4625, a user interface 4626, and a control unit 4627.

[0054] The sensor information acquisition unit 4621 acquires sensor information from the sensor terminal 4603.

[0055] The input / output unit 4622 acquires the transmission signal, decodes the multiplexed data (file format or packet) from the transmission signal, and outputs the multiplexed data to the demultiplexing unit 4623.

[0056] The inverse multiplexing unit 4623 acquires encoded data, control information, and additional information from the multiplexed data, and outputs the encoded data, control information, and additional information to the decoding unit 4624.

[0057] The decoding unit 4624 reconstructs the point cloud data by decoding the encoded data.

[0058] The presentation unit 4625 presents the point cloud data to the user. For example, the presentation unit 4625 displays information or an image based on the point cloud data. The user interface 4626 acquires an instruction based on the user's operation. The control unit 4627 (or the application execution unit) controls each processing unit. That is, the control unit 4627 performs controls such as inverse multiplexing, decoding, and presentation.

[0059] Note that the input / output unit 4622 may directly acquire point cloud data or encoded data from the outside. Also, the presentation unit 4625 may acquire additional information such as sensor information and present information based on the additional information. Further, the presentation unit 4625 may perform the presentation based on the user's instruction acquired by the user interface 4626.

[0060] The sensor terminal 4603 generates sensor information, which is information obtained by a sensor. The sensor terminal 4603 is a terminal equipped with a sensor or a camera, and examples include a moving body such as an automobile, a flying object such as an airplane, a mobile terminal, or a camera.

[0061] The sensor information that can be acquired by the sensor terminal 4603 includes, for example, (1) the distance between the sensor terminal 4603 and an object or the reflectivity of the object obtained from a LIDAR, millimeter-wave radar, or infrared sensor, (2) the distance between the camera and the object or the reflectivity of the object obtained from a plurality of monocular camera images or stereo camera images. Also, the sensor information may include the attitude, orientation, gyro (angular velocity), position (GPS information or altitude), speed, or acceleration of the sensor. Further, the sensor information may include temperature, atmospheric pressure, humidity, or magnetism.

[0062] The external connection unit 4604 is realized by an integrated circuit (LSI or IC), an external storage unit, communication with a cloud server via the Internet, or broadcasting or the like.

[0063] Next, the point cloud data will be described. FIG. 2 is a diagram showing the configuration of the point cloud data. FIG. 3 is a diagram showing a configuration example of a data file in which the information of the point cloud data is described.

[0064] The point cloud data includes data of a plurality of points. The data of each point includes position information (three-dimensional coordinates) and attribute information for the position information. A collection of these points is called a point cloud. For example, the point cloud shows the three-dimensional shape of an object.

[0065] Position information such as three-dimensional coordinates (Position) may also be called geometry. Also, the data of each point may include attribute information (attribute) of a plurality of attribute types. The attribute types are, for example, color or reflectance.

[0066] One piece of attribute information may be associated with one piece of position information, or attribute information having a plurality of different attribute types may be associated with one piece of position information. Also, a plurality of pieces of attribute information of the same attribute type may be associated with one piece of position information.

[0067] The configuration example of the data file shown in FIG. 3 is an example in the case where the position information and the attribute information correspond one-to-one, and shows the position information and the attribute information of N points constituting the point cloud data.

[0068] The position information is, for example, information on three axes of x, y, and z. The attribute information is, for example, RGB color information. A typical data file is a ply file or the like.

[0069] Next, the types of the point cloud data will be described. FIG. 4 is a diagram showing the types of the point cloud data. As shown in FIG. 4, the point cloud data includes static objects and dynamic objects.

[0070] The static object is three-dimensional point cloud data at any time (a certain moment). The dynamic object is three-dimensional point cloud data that changes over time. Hereinafter, the three-dimensional point cloud data at a certain moment is referred to as a PCC frame or a frame.

[0071] The object may be a point cloud with a limited area like normal video data, or a large-scale point cloud without a limited area like map information.

[0072] Also, there is point cloud data with various densities, and there may be sparse point cloud data and dense point cloud data.

[0073] Hereinafter, the details of each processing unit will be described. The sensor information is obtained by various methods such as a distance sensor such as a LIDAR or a range finder, a stereo camera, or a combination of a plurality of monocular cameras. The point cloud data generation unit 4618 generates point cloud data based on the sensor information obtained by the sensor information acquisition unit 4617. The point cloud data generation unit 4618 generates position information as the point cloud data, and adds attribute information for the position information to the position information.

[0074] The point cloud data generation unit 4618 may process the point cloud data when generating the position information or adding the attribute information. For example, the point cloud data generation unit 4618 may reduce the data amount by deleting the point cloud with overlapping positions. Also, the point cloud data generation unit 4618 may convert the position information (such as position shift, rotation, or normalization), or may render the attribute information.

[0075] Note that in FIG. 1, the point cloud data generation system 4611 is included in the three-dimensional data encoding system 4601, but may be provided independently outside the three-dimensional data encoding system 4601.

[0076] The symbolization unit 4613 generates encoded data by encoding the point cloud data based on a predefined encoding method. There are roughly two types of encoding methods as follows. The first is an encoding method using position information, which will be described as the first encoding method hereinafter. The second is an encoding method using a video codec, which will be described as the second encoding method hereinafter.

[0077] The decoding unit 4624 decodes the point cloud data by decoding the encoded data based on a predefined encoding method.

[0078] The multiplexing unit 4614 generates multiplexed data by multiplexing the encoded data using an existing multiplexing method. The generated multiplexed data is transmitted or stored. In addition to the PCC encoded data, the multiplexing unit 4614 multiplexes other media such as video, audio, subtitles, applications, files, etc., or reference time information. Further, the multiplexing unit 4614 may further multiplex sensor information or attribute information related to the point cloud data.

[0079] Examples of the multiplexing method or file format include ISOBMFF, MPEG-DASH which is an ISOBMFF-based transmission method, MMT, MPEG-2 TS Systems, RMP, etc.

[0080] The demultiplexing unit 4623 extracts the PCC encoded data, other media, time information, etc. from the multiplexed data.

[0081] The input / output unit 4615 transmits the multiplexed data using a method suitable for the medium for transmission such as broadcasting or communication or the medium for storage. The input / output unit 4615 may communicate with other devices via the Internet or communicate with a storage unit such as a cloud server.

[0082] Examples of the communication protocol used include http, ftp, TCP, or UDP. A PULL-type communication method may be used, or a PUSH-type communication method may be used.

[0083] Either wired transmission or wireless transmission may be used. As the wired transmission, Ethernet (registered trademark), USB, RS-232C, HDMI (registered trademark), or a coaxial cable, etc. are used. As the wireless transmission, wireless LAN, Wi-Fi (registered trademark), Bluetooth (registered trademark), or millimeter wave, etc. are used.

[0084] Also, as the broadcast method, for example, DVB-T2, DVB-S2, DVB-C2, ATSC 3.0, or ISDB-S3, etc. are used.

[0085] FIG. 5 is a diagram showing the configuration of a first encoding unit 4630 which is an example of an encoding unit 4613 that performs encoding according to a first encoding method. FIG. 6 is a block diagram of the first encoding unit 4630. The first encoding unit 4630 generates encoded data (encoded stream) by encoding point cloud data according to the first encoding method. This first encoding unit 4630 includes a position information encoding unit 4631, an attribute information encoding unit 4632, an additional information encoding unit 4633, and a multiplexing unit 4634.

[0086] The first encoding unit 4630 is characterized in that encoding is performed while being aware of the three-dimensional structure. Also, the first encoding unit 4630 is characterized in that the attribute information encoding unit 4632 performs encoding using the information obtained from the position information encoding unit 4631. The first encoding method is also called GPCC (Geometry based PCC).

[0087] The point cloud data is PCC point cloud data such as a PLY file, or PCC point cloud data generated from sensor information, and includes position information (Position), attribute information (Attribute), and other additional information (MetaData). The position information is input to the position information encoding unit 4631, the attribute information is input to the attribute information encoding unit 4632, and the additional information is input to the additional information encoding unit 4633.

[0088] The position information encoding unit 4631 generates encoded position information (Compressed Geometry), which is encoded data, by encoding the position information. For example, the position information encoding unit 4631 encodes the position information using an N-ary tree structure such as an octree. Specifically, in an octree, the target space is divided into eight nodes (sub-spaces), and 8-bit information (occupancy code) indicating whether or not each node contains a point cloud is generated. Further, a node containing a point cloud is further divided into eight nodes, and 8-bit information indicating whether or not each of the eight nodes contains a point cloud is generated. This process is repeated until the number of point clouds contained in a predetermined hierarchy or node falls below a threshold value.

[0089] The attribute information encoding unit 4632 generates encoded attribute information (Compressed Attribute), which is encoded data, by encoding using the configuration information generated by the position information encoding unit 4631. For example, the attribute information encoding unit 4632 determines a reference point (reference node) to be referred to in the encoding of the target point (target node) to be processed based on the octree structure generated by the position information encoding unit 4631. For example, the attribute information encoding unit 4632 refers to a node in which the parent node in the octree is the same as the target node among the surrounding nodes or adjacent nodes. Note that the method of determining the reference relationship is not limited to this.

[0090] Further, the encoding process of the attribute information may include at least one of quantization processing, prediction processing, and arithmetic encoding processing. In this case, reference means using the reference node to calculate the predicted value of the attribute information, or using the state of the reference node (for example, occupancy information indicating whether or not the reference node contains a point cloud) to determine the encoding parameter. For example, the encoding parameter is a quantization parameter in quantization processing or a context in arithmetic encoding.

[0091] The additional information encoding unit 4633 generates encoded additional information (Compressed MetaData), which is encoded data, by encoding compressible data among the additional information.

[0092] The multiplexing unit 4634 generates an encoded stream (Compressed Stream), which is encoded data, by multiplexing the encoded position information, encoded attribute information, encoded additional information, and other additional information. The generated encoded stream is output to a processing unit in a system layer (not shown).

[0093] Next, a first decoding unit 4640, which is an example of a decoding unit 4624 that decodes the first encoding method, will be described. FIG. 7 is a diagram showing the configuration of the first decoding unit 4640. FIG. 8 is a block diagram of the first decoding unit 4640. The first decoding unit 4640 generates point cloud data by decoding the encoded data (encoded stream) encoded by the first encoding method using the first encoding method. The first decoding unit 4640 includes a demultiplexing unit 4641, a position information decoding unit 4642, an attribute information decoding unit 4643, and an additional information decoding unit 4644.

[0094] An encoded stream (Compressed Stream), which is encoded data, is input from a processing unit in a system layer (not shown) to the first decoding unit 4640.

[0095] The demultiplexing unit 4641 separates the encoded position information (Compressed Geometry), encoded attribute information (Compressed Attribute), encoded additional information (Compressed MetaData), and other additional information from the encoded data.

[0096] The position information decoding unit 4642 generates position information by decoding the encoded position information. For example, the position information decoding unit 4642 restores the position information of the point cloud represented by three-dimensional coordinates from the encoded position information represented by an N-ary tree structure such as an octree.

[0097] The attribute information decoding unit 4643 decodes the encoded attribute information based on the configuration information generated by the position information decoding unit 4642. For example, the attribute information decoding unit 4643 determines a reference point (reference node) to be referred to in decoding a target point (target node) to be processed based on the octree structure obtained by the position information decoding unit 4642. For example, the attribute information decoding unit 4643 refers to a node among peripheral nodes or adjacent nodes whose parent node in the octree is the same as the target node. Note that the method for determining the reference relationship is not limited to this.

[0098] Also, the decoding process of the attribute information may include at least one of inverse quantization processing, prediction processing, and arithmetic decoding processing. In this case, reference means using the reference node to calculate the predicted value of the attribute information, or using the state of the reference node (for example, occupancy information indicating whether the reference node contains a point cloud) to determine the decoding parameter. For example, the decoding parameter is a quantization parameter in the inverse quantization process or a context in arithmetic decoding.

[0099] The additional information decoding unit 4644 generates additional information by decoding the encoded additional information. Also, the first decoding unit 4640 uses the additional information necessary for the decoding processes of the position information and the attribute information during decoding and outputs the additional information necessary for the application to the outside.

[0100] Next, a configuration example of the position information encoding unit will be described. FIG. 9 is a block diagram of the position information encoding unit 2700 according to the present embodiment. The position information encoding unit 2700 includes an octree generation unit 2701, a geometric information calculation unit 2702, an encoding table selection unit 2703, and an entropy encoding unit 2704.

[0101] The octree generation unit 2701 generates, for example, an octree from the input position information and generates occupancy codes for each node of the octree. The geometric information calculation unit 2702 acquires information indicating whether an adjacent node of the target node is an occupied node. For example, the geometric information calculation unit 2702 calculates occupancy information of an adjacent node (information indicating whether the adjacent node is an occupied node) from the occupancy code of the parent node to which the target node belongs. Further, the geometric information calculation unit 2702 may save the encoded nodes in a list and search for adjacent nodes from within that list. Note that the geometric information calculation unit 2702 may switch adjacent nodes according to the position within the parent node of the target node.

[0102] The encoding table selection unit 2703 selects an encoding table to be used for entropy encoding of the target node by using the occupancy information of the adjacent node calculated by the geometric information calculation unit 2702. For example, the encoding table selection unit 2703 may generate a bit string by using the occupancy information of the adjacent node and select an encoding table with an index number generated from that bit string.

[0103] The entropy encoding unit 2704 generates encoded position information and metadata by performing entropy encoding on the occupancy code of the target node by using the encoding table with the selected index number. The entropy encoding unit 2704 may add information indicating the selected encoding table to the encoded position information.

[0104] Hereinafter, the octree representation and the scan order of the position information will be described. The position information (position data) is encoded after being converted (octreed) into an octree structure. The octree structure is composed of nodes and leaves. Each node has eight nodes or leaves, and each leaf has voxel (VXL) information. FIG. 10 is a diagram showing an example of the structure of position information including a plurality of voxels. FIG. 11 is a diagram showing an example of converting the position information shown in FIG. 10 into an octree structure. Here, among the leaves shown in FIG. 11, leaves 1, 2, and 3 represent voxels VXL1, VXL2, and VXL3 shown in FIG. 10, respectively, and represent voxels containing point clouds (hereinafter, valid VXLs).

[0105] Specifically, Node 1 corresponds to the entire space that includes the position information in FIG. 10. The entire space corresponding to Node 1 is divided into eight nodes. Among the eight nodes, the nodes containing valid VXLs are further divided into eight nodes or leaves, and this process is repeated hierarchically in a tree structure. Here, each node corresponds to a sub-space and has information (occupancy code) indicating which position of the divided space has the next node or leaf as node information. Also, the bottommost block is set as a leaf, and information such as the number of point groups included in the leaf is retained as leaf information.

[0106] Next, a configuration example of the position information decoding unit will be described. FIG. 12 is a block diagram of the position information decoding unit 2710 according to the present embodiment. The position information decoding unit 2710 includes an octree generation unit 2711, a geometric information calculation unit 2712, a coding table selection unit 2713, and an entropy decoding unit 2714.

[0107] The octree generation unit 2711 generates an octree of a certain space (node) using the header information or metadata of the bitstream. For example, the octree generation unit 2711 generates a large space (root node) using the sizes in the x-axis, y-axis, and z-axis directions of a certain space added to the header information, and generates an octree by dividing that space into two in each of the x-axis, y-axis, and z-axis directions to generate eight small spaces A (nodes A0 to A7). Also, nodes A0 to A7 are sequentially set as the target nodes.

[0108] The geometric information calculation unit 2712 acquires occupancy information indicating whether an adjacent node of the target node is an occupied node. For example, the geometric information calculation unit 2712 calculates the occupancy information of the adjacent node from the occupancy code of the parent node to which the target node belongs. Also, the geometric information calculation unit 2712 may save the decoded nodes in a list and search for adjacent nodes from within that list. Note that the geometric information calculation unit 2712 may switch the adjacent nodes according to the position within the parent node of the target node.

[0109] The symbolic table selection unit 2713 selects a symbolic table (decoding table) to be used for entropy decoding of the target node by using the occupancy information of the adjacent nodes calculated by the geometric information calculation unit 2712. For example, the symbolic table selection unit 2713 may generate a bit string by using the occupancy information of the adjacent nodes, and select the symbolic table of the index number generated from the bit string.

[0110] The entropy decoding unit 2714 generates position information by entropy decoding the occupancy code of the target node by using the selected symbolic table. Note that the entropy decoding unit 2714 may decode and acquire the information of the selected symbolic table from the bit stream, and entropy decode the occupancy code of the target node by using the symbolic table indicated by the information.

[0111] Hereinafter, the configurations of the attribute information encoding unit and the attribute information decoding unit will be described. FIG. 13 is a block diagram showing a configuration example of the attribute information encoding unit A100. The attribute information encoding unit may include a plurality of encoding units that execute different encoding methods. For example, the attribute information encoding unit may switch between and use the following two methods according to the use case.

[0112] The attribute information encoding unit A100 includes a LoD attribute information encoding unit A101 and a transformed attribute information encoding unit A102. The LoD attribute information encoding unit A101 classifies each three-dimensional point into a plurality of layers by using the position information of the three-dimensional points, predicts the attribute information of the three-dimensional points belonging to each layer, and encodes the prediction residual. Here, each classified layer is called LoD (Level of Detail).

[0113] The transformed attribute information encoding unit A102 encodes the attribute information by using RAHT (Region Adaptive Hierarchical Transform). Specifically, the transformed attribute information encoding unit A102 generates high-frequency components and low-frequency components of each layer by applying RAHT or Haar transform to each attribute information based on the position information of the three-dimensional points, and encodes those values by using quantization and entropy encoding or the like.

[0114] FIG. 14 is a block diagram showing a configuration example of the attribute information decoding unit A110. The attribute information decoding unit may include a plurality of decoding units that execute different decoding methods. For example, the attribute information decoding unit may switch and decode based on the information included in the following two methods based on the information included in the header or metadata.

[0115] The attribute information decoding unit A110 includes a LoD attribute information decoding unit A111 and a conversion attribute information decoding unit A112. The LoD attribute information decoding unit A111 classifies each three-dimensional point into a plurality of layers using the position information of the three-dimensional points, and decodes the attribute value while predicting the attribute information of the three-dimensional points belonging to each layer.

[0116] The conversion attribute information decoding unit A112 decodes the attribute information using RAHT (Region Adaptive Hierarchical Transform). Specifically, the conversion attribute information decoding unit A112 decodes the attribute value by applying inverse RAHT or inverse Haar transform to the high-frequency component and the low-frequency component of each attribute value based on the position information of the three-dimensional points.

[0117] FIG. 15 is a block diagram showing the configuration of an attribute information encoding unit 3140, which is an example of the LoD attribute information encoding unit A101.

[0118] The attribute information encoding unit 3140 includes a LoD generation unit 3141, a surrounding search unit 3142, a prediction unit 3143, a prediction residual calculation unit 3144, a quantization unit 3145, an arithmetic encoding unit 3146, an inverse quantization unit 3147, a decoded value generation unit 3148, and a memory 3149.

[0119] The LoD generation unit 3141 generates LoD using the position information of the three-dimensional points.

[0120] The surrounding search unit 3142 searches for neighboring three-dimensional points adjacent to each three-dimensional point using the LoD generation result by the LoD generation unit 3141 and distance information indicating the distance between each three-dimensional point.

[0121] The prediction unit 3143 generates a predicted value of the attribute information of the target three-dimensional points to be encoded.

[0122] The prediction residual calculation unit 3144 calculates (generates) the prediction residual of the predicted value of the attribute information generated by the prediction unit 3143.

[0123] The quantization unit 3145 quantizes the prediction residual of the attribute information calculated by the prediction residual calculation unit 3144.

[0124] The arithmetic coding unit 3146 arithmetically encodes the prediction residual after quantization by the quantization unit 3145. The arithmetic coding unit 3146 outputs a bit stream including the arithmetically encoded prediction residual to, for example, a three-dimensional data decoding device.

[0125] Note that the prediction residual may be binarized by, for example, the quantization unit 3145 before being arithmetically encoded by the arithmetic coding unit 3146.

[0126] Also, for example, the arithmetic coding unit 3146 may initialize the coding table used for arithmetic coding before arithmetic coding. The arithmetic coding unit 3146 may initialize the coding table for each layer. Further, the arithmetic coding unit 3146 may output information indicating the position of the layer for which the coding table is initialized, included in the bit stream.

[0127] The inverse quantization unit 3147 inversely quantizes the prediction residual after quantization by the quantization unit 3145.

[0128] The decoded value generation unit 3148 generates a decoded value by adding the predicted value of the attribute information generated by the prediction unit 3143 and the prediction residual after inverse quantization by the inverse quantization unit 3147.

[0129] Memory 3149 is a memory that stores the decoded values of the attribute information of each three-dimensional point decoded by the decoded value generation unit 3148. For example, when the prediction unit 3143 generates a predicted value of a three-dimensional point that has not yet been encoded, it uses the decoded values of the attribute information of each three-dimensional point stored in the memory 3149 to generate a predicted value.

[0130] FIG. 16 is a block diagram of an attribute information encoding unit 6600, which is an example of the conversion attribute information encoding unit A102. The attribute information encoding unit 6600 includes a sorting unit 6601, a Haar transform unit 6602, a quantization unit 6603, an inverse quantization unit 6604, an inverse Haar transform unit 6605, a memory 6606, and an arithmetic encoding unit 6607.

[0131] The sorting unit 6601 generates a Morton code using the position information of the three-dimensional points and sorts the plurality of three-dimensional points in Morton code order. The Haar transform unit 6602 generates encoded coefficients by applying a Haar transform to the attribute information. The quantization unit 6603 quantizes the encoded coefficients of the attribute information.

[0132] The inverse quantization unit 6604 inverse-quantizes the quantized encoded coefficients. The inverse Haar transform unit 6605 applies an inverse Haar transform to the encoded coefficients. The memory 6606 stores the values of the attribute information of the plurality of decoded three-dimensional points. For example, the decoded attribute information of the three-dimensional points stored in the memory 6606 may be used for prediction of unencoded three-dimensional points and the like.

[0133] The arithmetic encoding unit 6607 calculates ZeroCnt from the quantized encoded coefficients and arithmetic-encodes ZeroCnt. Further, the arithmetic encoding unit 6607 arithmetic-encodes the non-zero quantized encoded coefficients. The arithmetic encoding unit 6607 may binarize the encoded coefficients before arithmetic encoding. Further, the arithmetic encoding unit 6607 may generate and encode various header information.

[0134] FIG. 17 is a block diagram showing the configuration of an attribute information decoding unit 3150, which is an example of the LoD attribute information decoding unit A111.

[0135] The attribute information decoding unit 3150 includes a LoD generation unit 3151, a surrounding search unit 3152, a prediction unit 3153, an arithmetic decoding unit 3154, an inverse quantization unit 3155, a decoded value generation unit 3156, and a memory 3157.

[0136] The LoD generation unit 3151 generates LoD using the position information of the three-dimensional points decoded by a position information decoding unit (not shown in FIG. 17).

[0137] The surrounding search unit 3152 searches for neighboring three-dimensional points adjacent to each three-dimensional point using the generation result of LoD by the LoD generation unit 3151 and distance information indicating the distance between each three-dimensional point.

[0138] The prediction unit 3153 generates a predicted value of the attribute information of the target three-dimensional point to be decoded.

[0139] The arithmetic decoding unit 3154 arithmetically decodes the prediction residual in the bit stream obtained from the attribute information encoding unit 3140 shown in FIG. 15. Note that the arithmetic decoding unit 3154 may initialize the decoding table used for arithmetic decoding. The arithmetic decoding unit 3154 initializes the decoding table used for arithmetic decoding for the layer on which the arithmetic encoding unit 3146 shown in FIG. 15 performed the encoding process. The arithmetic decoding unit 3154 may initialize the decoding table for each layer. Further, the arithmetic decoding unit 3154 may initialize the decoding table based on the information indicating the position of the layer in which the encoding table was initialized, included in the bit stream.

[0140] The inverse quantization unit 3155 inverse quantizes the prediction residual arithmetically decoded by the arithmetic decoding unit 3154.

[0141] The decoded value generation unit 3156 generates a decoded value by adding the predicted value generated by the prediction unit 3153 and the prediction residual after being inverse quantized by the inverse quantization unit 3155. The decoded value generation unit 3156 outputs the decoded attribute information data to other devices.

[0142] The memory 3157 is a memory that stores the decoded values of the attribute information of each three-dimensional point decoded by the decoded value generation unit 3156. For example, when the prediction unit 3153 generates a predicted value of a three-dimensional point that has not yet been decoded, it generates a predicted value using the decoded values of the attribute information of each three-dimensional point stored in the memory 3157.

[0143] FIG. 18 is a block diagram of an attribute information decoding unit 6610 which is an example of the conversion attribute information decoding unit A112. The attribute information decoding unit 6610 includes an arithmetic decoding unit 6611, an inverse quantization unit 6612, an inverse Haar transform unit 6613, and a memory 6614.

[0144] The arithmetic decoding unit 6611 arithmetically decodes the ZeroCnt and the encoding coefficients included in the bit stream. Note that the arithmetic decoding unit 6611 may decode various header information.

[0145] The inverse quantization unit 6612 inverse quantizes the arithmetically decoded encoding coefficients. The inverse Haar transform unit 6613 applies an inverse Haar transform to the encoding coefficients after inverse quantization. The memory 6614 stores the values of the attribute information of a plurality of decoded three-dimensional points. For example, the decoded attribute information of the three-dimensional points stored in the memory 6614 may be used for predicting three-dimensional points that have not been decoded.

[0146] Next, a second encoding unit 4650 which is an example of an encoding unit 4613 that performs encoding using the second encoding method will be described. FIG. 19 is a diagram showing the configuration of the second encoding unit 4650. FIG. 20 is a block diagram of the second encoding unit 4650.

[0147] The second encoding unit 4650 generates encoded data (encoded stream) by encoding the point cloud data using the second encoding method. This second encoding unit 4650 includes an additional information generation unit 4651, a position image generation unit 4652, an attribute image generation unit 4653, a video encoding unit 4654, an additional information encoding unit 4655, and a multiplexing unit 4656.

[0148] The second encoding unit 4650 is characterized by generating a position image and an attribute image by projecting a three-dimensional structure onto a two-dimensional image, and encoding the generated position image and attribute image using an existing video encoding method. The second encoding method is also called VPCC (Video based PCC).

[0149] The point cloud data is PCC point cloud data such as a PLY file, or PCC point cloud data generated from sensor information, and includes position information (Position), attribute information (Attribute), and other additional information (MetaData).

[0150] The additional information generation unit 4651 generates map information of a plurality of two-dimensional images by projecting a three-dimensional structure onto a two-dimensional image.

[0151] The position image generation unit 4652 generates a position image (Geometry Image) based on the position information and the map information generated by the additional information generation unit 4651. This position image is, for example, a depth image in which the distance (Depth) is shown as a pixel value. Note that this depth image may be an image of a plurality of point clouds viewed from one viewpoint (an image obtained by projecting a plurality of point clouds onto one two-dimensional plane), or a plurality of images of a plurality of point clouds viewed from a plurality of viewpoints, or a single image obtained by integrating these plurality of images.

[0152] The attribute image generation unit 4653 generates an attribute image based on the attribute information and the map information generated by the additional information generation unit 4651. This attribute image is, for example, an image in which attribute information (e.g., color (RGB)) is shown as a pixel value. Note that this image may be an image of a plurality of point clouds viewed from one viewpoint (an image obtained by projecting a plurality of point clouds onto one two-dimensional plane), or a plurality of images of a plurality of point clouds viewed from a plurality of viewpoints, or a single image obtained by integrating these plurality of images.

[0153] The video encoding unit 4654 generates an encoded position image (Compressed Geometry Image) and an encoded attribute image (Compressed Attribute Image), which are encoded data, by encoding the position image and the attribute image using a video encoding method. Note that any known encoding method may be used as the video encoding method. For example, the video encoding method is AVC or HEVC, etc.

[0154] The additional information encoding unit 4655 generates encoded additional information (Compressed MetaData) by encoding the additional information included in the point cloud data and map information, etc.

[0155] The multiplexing unit 4656 generates an encoded stream (Compressed Stream), which is encoded data, by multiplexing the encoded position image, the encoded attribute image, the encoded additional information, and other additional information. The generated encoded stream is output to a processing unit of a system layer (not shown).

[0156] Next, a second decoding unit 4660, which is an example of a decoding unit 4624 that decodes the second encoding method, will be described. FIG. 21 is a diagram showing the configuration of the second decoding unit 4660. FIG. 22 is a block diagram of the second decoding unit 4660. The second decoding unit 4660 generates point cloud data by decoding the encoded data (encoded stream) encoded by the second encoding method using the second encoding method. This second decoding unit 4660 includes a demultiplexing unit 4661, a video decoding unit 4662, an additional information decoding unit 4663, a position information generation unit 4664, and an attribute information generation unit 4665.

[0157] An encoded stream (Compressed Stream), which is encoded data, is input from a processing unit of a system layer (not shown) to the second decoding unit 4660.

[0158] The inverse multiplexing unit 4661 separates, from the encoded data, an encoded position image (Compressed Geometry Image), an encoded attribute image (Compressed Attribute Image), encoded additional information (Compressed MetaData), and other additional information.

[0159] The video decoding unit 4662 generates a position image and an attribute image by decoding the encoded position image and the encoded attribute image using a video encoding method. Note that any known encoding method may be used as the video encoding method. For example, the video encoding method is AVC or HEVC or the like.

[0160] The additional information decoding unit 4663 generates additional information including map information and the like by decoding the encoded additional information.

[0161] The position information generation unit 4664 generates position information using the position image and the map information. The attribute information generation unit 4665 generates attribute information using the attribute image and the map information.

[0162] The second decoding unit 4660 uses, at the time of decoding, additional information necessary for decoding and outputs, to the outside, additional information necessary for the application.

[0163] Hereinafter, problems in the PCC encoding method will be described. FIG. 23 is a diagram showing a protocol stack related to PCC encoded data. FIG. 23 shows an example in which data of other media such as video (for example, HEVC) or audio is multiplexed with the PCC encoded data and transmitted or stored.

[0164] The multiplexing method and the file format have functions for multiplexing various encoded data and transmitting or storing it. In order to transmit or store the encoded data, the encoded data must be converted into the format of the multiplexing method. For example, in HEVC, a technique is defined in which the encoded data is stored in a data structure called an NAL unit, and the NAL unit is stored in ISOBMFF.

[0165] On the one hand, currently, as encoding methods for point cloud data, a first encoding method (Codec1) and a second encoding method (Codec2) are being considered. However, the configuration of the encoded data and the method of storing the encoded data into the system format are not defined, and there is a problem that MUX processing (multiplexing), transmission, and storage in the encoding unit cannot be performed as it is.

[0166] Note that hereinafter, unless otherwise specified for a particular encoding method, it shall indicate either the first encoding method or the second encoding method.

[0167] (Embodiment 2) In this embodiment, the types of encoded data (Geometry (position information), Attribute (attribute information), Metadata (additional information)) generated by the above-described first encoding unit 4630 or second encoding unit 4650, the method of generating additional information (metadata), and the multiplexing process in the multiplexing unit will be described. Note that the additional information (metadata) may also be referred to as a parameter set or control information.

[0168] In this embodiment, the dynamic object (three-dimensional point cloud data that changes over time) described in FIG. 4 will be used as an example for explanation, but the same method may also be used for a static object (three-dimensional point cloud data at an arbitrary time).

[0169] FIG. 24 is a diagram showing the configuration of an encoding unit 4801 and a multiplexing unit 4802 included in the three-dimensional data encoding apparatus according to this embodiment. The encoding unit 4801 corresponds to, for example, the above-described first encoding unit 4630 or second encoding unit 4650. The multiplexing unit 4802 corresponds to the above-described multiplexing unit 4634 or 4656.

[0170] The encoding unit 4801 encodes the point cloud data of a plurality of PCC (Point Cloud Compression) frames and generates encoded data (Multiple Compressed Data) of a plurality of position information, attribute information, and additional information.

[0171] The multiplexing unit 4802 converts the data into a data configuration considering data access in the decoding device by NAL unitizing the data of multiple data types (position information, attribute information, and additional information).

[0172] FIG. 25 is a diagram showing a configuration example of the encoded data generated by the encoding unit 4801. The arrows in the figure indicate the dependency relationships related to the decoding of the encoded data, and the origin of the arrow depends on the data at the tip of the arrow. That is, the decoding device decodes the data at the tip of the arrow and uses the decoded data to decode the data at the origin of the arrow. In other words, being dependent means that the data at the dependent destination is referenced (used) in the processing (encoding or decoding, etc.) of the data at the dependent source.

[0173] First, the generation process of the encoded data of the position information will be described. The encoding unit 4801 generates encoded position data (Compressed Geometry Data) for each frame by encoding the position information of each frame. Also, the encoded position data is represented by G(i). Here, i indicates the frame number, or the time of the frame, etc.

[0174] In addition, the encoding unit 4801 generates a position parameter set (GPS(i)) corresponding to each frame. The position parameter set includes parameters that can be used for decoding the encoded position data. Also, the encoded position data for each frame depends on the corresponding position parameter set.

[0175] Also, the encoded position data consisting of multiple frames is defined as a geometry sequence. The encoding unit 4801 generates a geometry sequence parameter set (Geometry Sequence PS: also denoted as position SPS) that stores parameters commonly used for the decoding process of multiple frames within the geometry sequence. The geometry sequence depends on the position SPS.

[0176] Next, the generation process of the encoded data of the attribute information will be described. The encoding unit 4801 generates encoded attribute data (Compressed Attribute Data) for each frame by encoding the attribute information of each frame. Also, the encoded attribute data is represented by A(i). Further, FIG. 25 shows an example in which there are attribute X and attribute Y, the encoded attribute data of attribute X is represented by AX(i), and the encoded attribute data of attribute Y is represented by AY(i).

[0177] Also, the encoding unit 4801 generates an attribute parameter set (APS(i)) corresponding to each frame. Also, the attribute parameter set of attribute X is represented by AXPS(i), and the attribute parameter set of attribute Y is represented by AYPS(i). The attribute parameter set includes parameters that can be used for decoding the encoded attribute information. The encoded attribute data depends on the corresponding attribute parameter set.

[0178] Also, the encoded attribute data composed of a plurality of frames is defined as an attribute sequence. The encoding unit 4801 generates an attribute sequence parameter set (Attribute Sequence PS: also denoted as attribute SPS) that stores parameters commonly used for the decoding process for a plurality of frames in the attribute sequence. The attribute sequence depends on the attribute SPS.

[0179] Also, in the first encoding method, the encoded attribute data depends on the encoding position data.

[0180] Also, FIG. 25 shows an example in the case where there are two types of attribute information (attribute X and attribute Y). When there are two types of attribute information, for example, two encoding units generate respective data and metadata. Also, for example, an attribute sequence is defined for each type of attribute information, and an attribute SPS is generated for each type of attribute information.

[0181] Note that in Fig. 25, an example is shown where there is one type of position information and two types of attribute information. However, this is not the only case. The attribute information may be one type or three or more types. Also in this case, the encoded data can be generated in the same way. Further, in the case of point cloud data without attribute information, the attribute information may not be present. In that case, the encoding unit 4801 may not generate a parameter set related to the attribute information.

[0182] Next, the generation process of additional information (metadata) will be described. The encoding unit 4801 generates a PCC stream PS (also referred to as stream PS), which is a parameter set for the entire PCC stream. The encoding unit 4801 stores in the stream PS parameters that can be commonly used for the decoding process of one or more position sequences and one or more attribute sequences. For example, the stream PS includes identification information indicating the codec of the point cloud data, information indicating the algorithm used for encoding, and the like. The position sequence and the attribute sequence depend on the stream PS.

[0183] Next, the access unit and the GOF will be described. In the present embodiment, the concepts of a new access unit (Access Unit: AU) and a GOF (Group of Frame) are introduced.

[0184] The access unit is a basic unit for accessing data during decoding and is composed of one or more data and one or more metadata. For example, the access unit is composed of position information at the same time and one or more pieces of attribute information. The GOF is a random access unit and is composed of one or more access units.

[0185] The symbolization unit 4801 generates an access unit header (AU Header) as identification information indicating the start of an access unit. The symbolization unit 4801 stores parameters related to the access unit in the access unit header. For example, the access unit header includes the configuration or information of the encoded data included in the access unit. Also, the access unit header includes parameters commonly used for the data included in the access unit, such as parameters related to the decoding of the encoded data.

[0186] Note that the symbolization unit 4801 may generate an access unit delimiter that does not include parameters related to the access unit instead of the access unit header. This access unit delimiter is used as identification information indicating the start of the access unit. The decoding device identifies the start of the access unit by detecting the access unit header or the access unit delimiter.

[0187] Next, the generation of the identification information at the start of a GOP will be described. The symbolization unit 4801 generates a GOP header (GOP Header) as identification information indicating the start of a GOP. The symbolization unit 4801 stores parameters related to the GOP in the GOP header. For example, the GOP header includes the configuration or information of the encoded data included in the GOP. Also, the GOP header includes parameters commonly used for the data included in the GOP, such as parameters related to the decoding of the encoded data.

[0188] Note that the symbolization unit 4801 may generate a GOP delimiter that does not include parameters related to the GOP instead of the GOP header. This GOP delimiter is used as identification information indicating the start of the GOP. The decoding device identifies the start of the GOP by detecting the GOP header or the GOP delimiter.

[0189] In the PCC encoded data, for example, an access unit is defined in units of PCC frames. The decoding device accesses the PCC frame based on the identification information at the start of the access unit.

[0190] Also, for example, a GOF is defined as one random access unit. The decoding device accesses the random access unit based on the identification information at the head of the GOF. For example, if PCC frames are independent of each other and can be decoded alone, the PCC frames may be defined as random access units.

[0191] Note that two or more PCC frames may be assigned to one access unit, or a plurality of random access units may be assigned to one GOF.

[0192] Also, the encoding unit 4801 may define and generate parameter sets or metadata other than the above. For example, the encoding unit 4801 may generate an SEI (Supplemental Enhancement Information) that stores parameters (optional parameters) that may not necessarily be used during decoding.

[0193] Next, the configuration of the encoded data and the method of storing the encoded data in the NAL unit will be described.

[0194] For example, a data format is defined for each type of encoded data. FIG. 26 is a diagram showing an example of encoded data and NAL units.

[0195] For example, as shown in FIG. 26, the encoded data includes a header and a payload. Note that the encoded data may include length information indicating the length (data amount) of the encoded data, the header, or the payload. Also, the encoded data may not include a header.

[0196] The header includes, for example, identification information for specifying the data. This identification information indicates, for example, the data type or the frame number.

[0197] The header includes, for example, identification information indicating a reference relationship. This identification information is stored in the header, for example, when there is a dependency between data, and is information for referring from a reference source to a reference destination. For example, the header of the reference destination includes identification information for specifying the data. The header of the reference source includes identification information indicating the reference destination.

[0198] Note that when the reference destination or the reference source can be identified or derived from other information, the identification information for specifying the data or the identification information indicating the reference relationship may be omitted.

[0199] The multiplexing unit 4802 stores the encoded data in the payload of the NAL unit. The NAL unit header includes pcc_nal_unit_type, which is identification information of the encoded data. FIG. 27 is a diagram showing an example of the semantics of pcc_nal_unit_type.

[0200] As shown in FIG. 27, when pcc_codec_type is Codec1 (the first encoding method), the values 0 to 10 of pcc_nal_unit_type are assigned to the encoded position data (Geometry), the encoded attribute X data (AttributeX), the encoded attribute Y data (AttributeY), the position PS (Geom.PS), the attribute XPS (AttrX.PS), the attribute YPS (AttrX.PS), the position SPS (Geometry Sequence PS), the attribute XSPS (AttributeX Sequence PS), the attribute YSPS (AttributeY Sequence PS), the AU header (AU Header), and the GOF header (GOF Header) in Codec1. Further, the values 11 and later are assigned for future use in Codec1.

[0201] When pcc_codec_type is Codec2 (the second encoding method), values 0 to 2 of pcc_nal_unit_type are assigned to DataA, MetaDataA, and MetaDataB of the codec. Also, values 3 and later are assigned to the reserve of Codec2.

[0202] Next, the data transmission order will be described. Hereinafter, the constraints on the transmission order of NAL units will be described.

[0203] The multiplexing unit 4802 transmits NAL units in units of GOF or AU. The multiplexing unit 4802 places a GOF header at the head of the GOF and an AU header at the head of the AU.

[0204] Even if data is lost due to packet loss or the like, the multiplexing unit 4802 may place the sequence parameter set (SPS) for each AU so that the decoding device can decode from the next AU.

[0205] When there is a dependency relationship related to decoding in the encoded data, the decoding device decodes the reference destination data first and then decodes the reference source data. In the decoding device, in order to be able to decode in the received order without rearranging the data, the multiplexing unit 4802 transmits the reference destination data first.

[0206] FIG. 28 is a diagram showing an example of the transmission order of NAL units. FIG. 28 shows three examples: position information priority, parameter priority, and data integration.

[0207] The transmission order of position information priority is an example of transmitting information related to position information and information related to attribute information together. In this transmission order, the transmission of information related to position information is completed earlier than the transmission of information related to attribute information.

[0208] For example, by using this transmission order, a decoder that does not decrypt the attribute information may be able to provide time for non - processing by ignoring the decryption of the attribute information. Also, for example, in the case of a decoder that wants to quickly decrypt the position information, it may be possible to more quickly decrypt the position information by obtaining the encoded data of the position information earlier.

[0209] Note that in FIG. 28, the attribute XSPS and the attribute YSPS are integrated and described as the attribute SPS, but the attribute XSPS and the attribute YSPS may be arranged individually.

[0210] In the transmission order with parameter set priority, the parameter set is transmitted first and the data is transmitted later.

[0211] As described above, according to the constraints of the NAL unit transmission order, the multiplexing unit 4802 may transmit the NAL units in any order. For example, order identification information is defined, and the multiplexing unit 4802 may have a function of transmitting the NAL units in a plurality of pattern orders. For example, the order identification information of the NAL units is stored in the stream PS.

[0212] The three - dimensional data decoder may perform decoding based on the order identification information. A desired transmission order is instructed from the three - dimensional data decoder to the three - dimensional data encoder, and the three - dimensional data encoder (multiplexing unit 4802) may control the transmission order according to the instructed transmission order.

[0213] Note that the multiplexing unit 4802 may generate encoded data by merging a plurality of functions as long as it is within the range of following the constraints of the transmission order, such as the transmission order of data integration. For example, as shown in FIG. 28, the GOF header and the AU header may be integrated, or the AXPS and the AYPS may be integrated. In this case, an identifier indicating that the data has a plurality of functions is defined for pcc_nal_unit_type.

[0214] Hereinafter, a modification example of the present embodiment will be described. PS has levels such as frame-level PS, sequence-level PS, and PCC sequence-level PS. When the PCC sequence level is the upper level and the frame level is the lower level, the following method may be used for storing parameters.

[0215] The value of the default PS is indicated by the higher-level PS. Also, when the value of the lower-level PS is different from the value of the higher-level PS, the value of the PS is indicated by the lower-level PS. Or, the value of the PS is not described at the higher level, and the value of the PS is described at the lower-level PS. Or, information indicating whether the value of the PS is indicated by the lower-level PS, the higher-level PS, or both is indicated in either or both of the lower-level PS and the higher-level PS. Or, the lower-level PS may be merged into the higher-level PS. Or, when the lower-level PS and the higher-level PS overlap, the multiplexing unit 4802 may omit the transmission of either one.

[0216] Note that the encoding unit 4801 or the multiplexing unit 4802 may divide the data into slices or tiles, etc., and transmit the divided data. The divided data includes information for identifying the divided data, and the parameters used for decoding the divided data are included in the parameter set. In this case, the pcc_nal_unit_type is defined as an identifier indicating that it is data for storing data or parameters related to tiles or slices.

[0217] Hereinafter, the process related to the order identification information will be described. FIG. 29 is a flowchart of the process by the three-dimensional data encoding device (encoding unit 4801 and multiplexing unit 4802) related to the transmission order of the NAL unit.

[0218] First, the three-dimensional data encoding device determines the transmission order (position information priority or parameter set priority) of the NAL unit (S4801). For example, the three-dimensional data encoding device determines the transmission order based on a designation from a user or an external device (e.g., a three-dimensional data decoding device).

[0219] When the determined transmission order is position information priority (position information priority in S4802), the three-dimensional data encoding device sets the order identification information included in the stream PS to position information priority (S4803). That is, in this case, the order identification information indicates that the NAL units are transmitted in the order of position information priority. Then, the three-dimensional data encoding device transmits the NAL units in the order of position information priority (S4804).

[0220] On the other hand, when the determined transmission order is parameter set priority (parameter set priority in S4802), the three-dimensional data encoding device sets the order identification information included in the stream PS to parameter set priority (S4805). That is, in this case, the order identification information indicates that the NAL units are transmitted in the order of parameter set priority. Then, the three-dimensional data encoding device transmits the NAL units in the order of parameter set priority (S4806).

[0221] Figure 30 is a flowchart of the processing by the three-dimensional data decoding device regarding the transmission order of the NAL units. First, the three-dimensional data decoding device analyzes the order identification information included in the stream PS (S4811).

[0222] When the transmission order indicated by the order identification information is position information priority (position information priority in S4812), the three-dimensional data decoding device decodes the NAL units assuming that the transmission order of the NAL units is position information priority (S4813).

[0223] On the other hand, when the transmission order indicated by the order identification information is parameter set priority (parameter set priority in S4812), the three-dimensional data decoding device decodes the NAL units assuming that the transmission order of the NAL units is parameter set priority (S4814).

[0224] For example, when the three-dimensional data decoding device does not decode the attribute information, in step S4813, instead of acquiring all NAL units, it may acquire the NAL units related to the position information and decode the position information from the acquired NAL units.

[0225] Next, the processing related to the generation of AUs and GOPs will be described. FIG. 31 is a flowchart of the processing by the three-dimensional data encoding device (multiplexing unit 4802) related to the generation of AUs and GOPs in the multiplexing of NAL units.

[0226] First, the three-dimensional data encoding device determines the type of the encoded data (S4821). Specifically, the three-dimensional data encoding device determines whether the encoded data to be processed is data at the head of an AU, data at the head of a GOP, or other data.

[0227] When the encoded data is data at the head of a GOP (GOP head in S4822), the three-dimensional data encoding device generates an NAL unit by arranging the GOP header and the AU header at the head of the encoded data belonging to the GOP (S4823).

[0228] When the encoded data is data at the head of an AU (AU head in S4822), the three-dimensional data encoding device generates an NAL unit by arranging the AU header at the head of the encoded data belonging to the AU (S4824).

[0229] When the encoded data is neither at the head of a GOP nor at the head of an AU (neither GOP head nor AU head in S4822), the three-dimensional data encoding device generates an NAL unit by arranging the encoded data after the AU header of the AU to which the encoded data belongs (S4825).

[0230] Next, the processing related to access to AUs and GOPs will be described. FIG. 32 is a flowchart of the processing of the three-dimensional data decoding device related to access to AUs and GOPs in the demultiplexing of NAL units.

[0231] First, the three-dimensional data decoding device determines the type of encoded data included in the NAL unit by analyzing the nal_unit_type included in the NAL unit (S4831). Specifically, the three-dimensional data decoding device determines whether the encoded data included in the NAL unit is data at the beginning of an AU, data at the beginning of a GOP, or other data.

[0232] When the encoded data included in the NAL unit is data at the beginning of a GOP (GOP start in S4832), the three-dimensional data decoding device determines that the NAL unit is the start position of random access, accesses the NAL unit, and starts the decoding process (S4833).

[0233] On the other hand, when the encoded data included in the NAL unit is data at the beginning of an AU (AU start in S4832), the three-dimensional data decoding device determines that the NAL unit is at the beginning of an AU, accesses the data included in the NAL unit, and decodes the AU (S4834).

[0234] On the other hand, when the encoded data included in the NAL unit is neither at the beginning of a GOP nor at the beginning of an AU (other than GOP start and AU start in S4832), the three-dimensional data decoding device does not process the NAL unit.

[0235] (Embodiment 3) In this embodiment, a method for representing three-dimensional points (point clouds) in the encoding of three-dimensional data will be described.

[0236] FIG. 33 is a block diagram showing the configuration of a three-dimensional data distribution system according to this embodiment. The distribution system shown in FIG. 33 includes a server 1501 and a plurality of clients 1502.

[0237] The server 1501 includes a storage unit 1511 and a control unit 1512. The storage unit 1511 stores an encoded three-dimensional map 1513, which is encoded three-dimensional data.

[0238] FIG. 34 is a diagram showing a configuration example of the bit stream of the encoded three-dimensional map 1513. The three-dimensional map is divided into a plurality of sub-maps, and each sub-map is encoded. A random access header (RA) including sub-coordinate information is added to each sub-map. The sub-coordinate information is used to improve the encoding efficiency of the sub-map. This sub-coordinate information indicates the sub-coordinate of the sub-map. The sub-coordinate is the coordinate of the sub-map with respect to the reference coordinate. Note that a three-dimensional map including a plurality of sub-maps is called an overall map. Also, the coordinate (for example, the origin) serving as a reference in the overall map is called the reference coordinate. That is, the sub-coordinate is the coordinate of the sub-map in the coordinate system of the overall map. In other words, the sub-coordinate indicates the offset between the coordinate system of the overall map and the coordinate system of the sub-map. Also, the coordinate in the coordinate system of the overall map with respect to the reference coordinate is called the overall coordinate. The coordinate in the coordinate system of the sub-map with respect to the sub-coordinate is called the differential coordinate.

[0239] The client 1502 transmits a message to the server 1501. This message includes the position information of the client 1502. The control unit 1512 included in the server 1501 acquires the bit stream of the sub-map at the position closest to the position of the client 1502 based on the position information included in the received message. The bit stream of the sub-map includes sub-coordinate information and is transmitted to the client 1502. The decoder 1521 included in the client 1502 obtains the overall coordinate of the sub-map with respect to the reference coordinate using this sub-coordinate information. The application 1522 included in the client 1502 executes an application related to its own position using the obtained overall coordinate of the sub-map.

[0240] Also, the sub-map indicates a partial area of the entire map. The sub-coordinates are the coordinates where the sub-map is located in the reference coordinate space of the entire map. For example, assume that in the entire map of A, there are a sub-map A of AA and a sub-map B of AB. When the vehicle wants to refer to the map of AA, it starts decoding from sub-map A, and when it wants to refer to the map of AB, it starts decoding from sub-map B. Here, the sub-map is a random access point. Specifically, A is Osaka Prefecture, AA is Osaka City, AB is Takatsuki City, etc.

[0241] Each sub-map is transmitted to the client together with sub-coordinate information. The sub-coordinate information is included in the header information of each sub-map, or in the transmission packet, etc.

[0242] The reference coordinates that serve as the reference coordinates for the sub-coordinate information of each sub-map may be added to the header information of the space higher than the sub-map, such as the header information of the entire map.

[0243] The sub-map may be composed of one space (SPC). Also, the sub-map may be composed of multiple SPCs.

[0244] Also, the sub-map may include a GOS (Group of Space). Also, the sub-map may be composed of a world. For example, when there are multiple objects in the sub-map, if the multiple objects are assigned to separate SPCs, the sub-map is composed of multiple SPCs. Also, if the multiple objects are assigned to one SPC, the sub-map is composed of one SPC.

[0245] Next, the improvement effect of the encoding efficiency when using sub-coordinate information will be described. FIG. 35 is a diagram for explaining this effect. For example, in order to encode the three-dimensional point A at a position far from the reference coordinate shown in FIG. 35, a large number of bits are required. Here, the distance between the sub-coordinate and the three-dimensional point A is shorter than the distance between the reference coordinate and the three-dimensional point A. Therefore, the encoding efficiency can be improved by encoding the coordinates of the three-dimensional point A with respect to the sub-coordinate rather than encoding the coordinates of the three-dimensional point A with respect to the reference coordinate. Also, the bit stream of the sub-map includes sub-coordinate information. By sending the bit stream of the sub-map and the reference coordinate to the decoding side (client), the overall coordinates of the sub-map can be restored on the decoding side.

[0246] FIG. 36 is a flowchart of the processing by the server 1501 which is the transmission side of the sub-map.

[0247] First, the server 1501 receives a message including the position information of the client 1502 from the client 1502 (S1501). The control unit 1512 acquires the encoded bit stream of the sub-map based on the position information of the client from the storage unit 1511 (S1502). Then, the server 1501 transmits the encoded bit stream of the sub-map and the reference coordinate to the client 1502 (S1503).

[0248] FIG. 37 is a flowchart of the processing by the client 1502 which is the reception side of the sub-map.

[0249] First, the client 1502 receives the encoded bit stream of the sub-map and the reference coordinate transmitted from the server 1501 (S1511). Next, the client 1502 acquires the sub-map and the sub-coordinate information by decoding the encoded bit stream (S1512). Next, the client 1502 restores the differential coordinates in the sub-map to the overall coordinates using the reference coordinate and the sub-coordinate (S1513).

[0250] Next, a syntax example of information regarding the sub-map will be described. In the encoding of the sub-map, the three-dimensional data encoding device calculates differential coordinates by subtracting the sub-coordinates from the coordinates of each point cloud (three-dimensional points). Then, the three-dimensional data encoding device encodes the differential coordinates into a bit stream as the value of each point cloud. Further, the encoding device encodes sub-coordinate information indicating the sub-coordinates as header information of the bit stream. Thereby, the three-dimensional data decoding device can obtain the overall coordinates of each point cloud. For example, the three-dimensional data encoding device is included in the server 1501, and the three-dimensional data decoding device is included in the client 1502.

[0251] FIG. 38 is a diagram showing a syntax example of the sub-map. NumOfPoint shown in FIG. 38 indicates the number of point clouds included in the sub-map. sub_coordinate_x, sub_coordinate_y, and sub_coordinate_z are sub-coordinate information. sub_coordinate_x indicates the x coordinate of the sub-coordinates. sub_coordinate_y indicates the y coordinate of the sub-coordinates. sub_coordinate_z indicates the z coordinate of the sub-coordinates.

[0252] Also, diff_x[i], diff_y[i], and diff_z[i] are the differential coordinates of the i-th point cloud in the sub-map. diff_x[i] indicates the difference value between the x coordinate of the i-th point cloud in the sub-map and the x coordinate of the sub-coordinates. diff_y[i] indicates the difference value between the y coordinate of the i-th point cloud in the sub-map and the y coordinate of the sub-coordinates. diff_z[i] indicates the difference value between the z coordinate of the i-th point cloud in the sub-map and the z coordinate of the sub-coordinates.

[0253] The three-dimensional data decoding device decodes point_cloud[i]_x, point_cloud[i]_y, and point_cloud[i]_z, which are the global coordinates of the i-th point cloud, using the following equations. point_cloud[i]_x is the x-coordinate of the global coordinates of the i-th point cloud. point_cloud[i]_y is the y-coordinate of the global coordinates of the i-th point cloud. point_cloud[i]_z is the z-coordinate of the global coordinates of the i-th point cloud.

[0254] point_cloud[i]_x = sub_coordinate_x + diff_x[i] point_cloud[i]_y = sub_coordinate_y + diff_y[i] point_cloud[i]_z = sub_coordinate_z + diff_z[i]

[0255] Next, the switching process for applying octree encoding will be described. The three-dimensional data encoding device selects whether to use octree encoding (hereinafter referred to as octree encoding) to encode each point cloud in an octree representation or to encode the difference value from the sub-coordinates (hereinafter referred to as non-octree encoding) during sub-map encoding. FIG. 39 is a diagram schematically showing this operation. For example, when the number of point clouds in the sub-map is equal to or greater than a predetermined threshold, the three-dimensional data encoding device applies octree encoding to the sub-map. When the number of point clouds in the sub-map is less than the above threshold, the three-dimensional data encoding device applies non-octree encoding to the sub-map. Thereby, the three-dimensional data encoding device can appropriately select whether to use octree encoding or non-octree encoding according to the shape and density of the objects included in the sub-map, so that the encoding efficiency can be improved.

[0256] In addition, the three-dimensional data encoding device adds information indicating whether octree encoding or non-octree encoding is applied to the sub-map (hereinafter referred to as octree encoding application information) to the header of the sub-map or the like. As a result, the three-dimensional data decoding device can determine whether the bitstream is a bitstream obtained by octree encoding of the sub-map or a bitstream obtained by non-octree encoding of the sub-map.

[0257] In addition, the three-dimensional data encoding device calculates the encoding efficiency when applying octree encoding and non-octree encoding respectively to the same point cloud, and may apply the encoding method with good encoding efficiency to the sub-map.

[0258] FIG. 40 is a diagram showing a syntax example of a sub-map when this switching is performed. The coding_type shown in FIG. 40 is information indicating the encoding type and is the above-mentioned octree encoding application information. coding_type = 00 indicates that octree encoding has been applied. coding_type = 01 indicates that non-octree encoding has been applied. coding_type = 10 or 11 indicates that another encoding method other than the above has been applied.

[0259] When the encoding type is non-octree encoding (non_octree), the sub-map includes NumOfPoint and sub-coordinate information (sub_coordinate_x, sub_coordinate_y, and sub_coordinate_z).

[0260] When the encoding type is octree encoding (octree), the sub-map includes octree_info. octree_info is information necessary for octree encoding and includes, for example, depth information.

[0261] When the encoding type is non-octree encoding (non_octree), the sub-map includes differential coordinates (diff_x[i], diff_y[i], and diff_z[i]).

[0262] When the symbolization type is octree encoding, the sub-map includes octree_data, which is the encoding data related to octree encoding.

[0263] Here, an example using the xyz coordinate system as the coordinate system of the point cloud is shown, but the polar coordinate system may also be used.

[0264] Figure 41 is a flowchart of the three-dimensional data encoding process by the three-dimensional data encoding device. First, the three-dimensional data encoding device calculates the number of point clouds in the target sub-map, which is the sub-map to be processed (S1521). Next, the three-dimensional data encoding device determines whether the calculated number of point clouds is greater than or equal to a predetermined threshold (S1522).

[0265] If the number of point clouds is greater than or equal to the threshold (Yes in S1522), the three-dimensional data encoding device applies octree encoding to the target sub-map (S1523). Also, the three-dimensional point data encoding device adds octree encoding application information indicating that octree encoding has been applied to the target sub-map to the header of the bit stream (S1525).

[0266] On the other hand, if the number of point clouds is less than the threshold (No in S1522), the three-dimensional data encoding device applies non-octree encoding to the target sub-map (S1524). Also, the three-dimensional point data encoding device adds octree encoding application information indicating that non-octree encoding has been applied to the target sub-map to the header of the bit stream (S1525).

[0267] Figure 42 is a flowchart of the three-dimensional data decoding process by the three-dimensional data decoding device. First, the three-dimensional data decoding device decodes the octree encoding application information from the header of the bit stream (S1531). Next, based on the decoded octree encoding application information, the three-dimensional data decoding device determines whether the encoding type applied to the target sub-map is octree encoding (S1532).

[0268] When the encoding type indicated by the octree encoding application information is octree encoding (Yes in S1532), the three-dimensional data decoding device decodes the target sub-map by octree decoding (S1533). On the other hand, when the encoding type indicated by the octree encoding application information is non-octree encoding (No in S1532), the three-dimensional data decoding device decodes the target sub-map by non-octree decoding (S1534).

[0269] Hereinafter, a modification example of the present embodiment will be described. FIGS. 43 to 45 are diagrams schematically showing the operation of a modification example of the encoding type switching process.

[0270] As shown in FIG. 43, the three-dimensional data encoding device may select whether to apply octree encoding or non-octree encoding for each space. In this case, the three-dimensional data encoding device adds octree encoding application information to the header of the space. Thereby, the three-dimensional data decoding device can determine whether octree encoding has been applied for each space. Also, in this case, the three-dimensional data encoding device sets sub-coordinates for each space and encodes the difference value obtained by subtracting the value of the sub-coordinates from the coordinates of each point cloud in the space.

[0271] Thereby, the three-dimensional data encoding device can appropriately switch whether to apply octree encoding according to the shape of the object or the number of point clouds in the space, so that the encoding efficiency can be improved.

[0272] Also, as shown in FIG. 44, the three-dimensional data encoding device may select whether to apply octree encoding or non-octree encoding for each volume. In this case, the three-dimensional data encoding device adds octree encoding application information to the header of the volume. Thereby, the three-dimensional data decoding device can determine whether octree encoding has been applied for each volume. Also, in this case, the three-dimensional data encoding device sets sub-coordinates for each volume and encodes the difference value obtained by subtracting the value of the sub-coordinates from the coordinates of each point cloud in the volume.

[0273] As a result, the three-dimensional data encoding device can appropriately switch whether to apply octree encoding according to the shape of the object in the volume or the number of point clouds, so that the encoding efficiency can be improved.

[0274] Also, in the above description, as an example of non-octree encoding, an example of encoding the difference obtained by subtracting the sub-coordinates from the coordinates of each point cloud is shown. However, it is not necessarily limited to this, and any encoding method other than octree encoding may be used for encoding. For example, as shown in FIG. 45, as non-octree encoding, the three-dimensional data encoding device may use a method (hereinafter referred to as original coordinate encoding) of encoding not the difference from the sub-coordinates but the value of the point cloud itself in the sub-map, space, or volume.

[0275] In that case, the three-dimensional data encoding device stores information indicating that the original coordinate encoding has been applied to the target space (sub-map, space, or volume) in the header. Thereby, the three-dimensional data decoding device can determine whether the original coordinate encoding has been applied to the target space.

[0276] Also, when applying the original coordinate encoding, the three-dimensional data encoding device may perform encoding without applying quantization and arithmetic encoding to the original coordinates. Also, the three-dimensional data encoding device may encode the original coordinates with a predetermined fixed bit length. Thereby, the three-dimensional data encoding device can generate a stream of a certain bit length at a certain timing.

[0277] Also, in the above description, as an example of non-octree encoding, an example of encoding the difference obtained by subtracting the sub-coordinates from the coordinates of each point cloud is shown, but it is not necessarily limited to this.

[0278] For example, the three-dimensional data encoding device may sequentially encode the difference values between the coordinates of each point cloud. FIG. 46 is a diagram for explaining the operation in this case. For example, in the example shown in FIG. 46, when encoding the point cloud PA, the three-dimensional data encoding device uses the sub-coordinates as the predicted coordinates and encodes the difference value between the coordinates of the point cloud PA and the predicted coordinates. Also, when encoding the point cloud PB, the three-dimensional data encoding device uses the coordinates of the point cloud PA as the predicted coordinates and encodes the difference value between the point cloud PB and the predicted coordinates. Further, when encoding the point cloud PC, the three-dimensional data encoding device uses the point cloud PB as the predicted coordinates and encodes the difference value between the point cloud PB and the predicted coordinates. In this way, the three-dimensional data encoding device may set the scan order for a plurality of point clouds and encode the difference value between the coordinates of the target point cloud to be processed and the coordinates of the immediately preceding point cloud in the scan order with respect to the target point cloud.

[0279] Also, in the above description, the sub-coordinates were the coordinates of the lower left front corner of the sub-map, but the position of the sub-coordinates is not limited to this. FIGS. 47 to 49 are diagrams showing another example of the position of the sub-coordinates. The sub-coordinates may be set at any coordinates within the target space (sub-map, space, or volume). That is, the sub-coordinates may be the coordinates of the lower left front corner of the target space as described above. As shown in FIG. 47, the sub-coordinates may be the coordinates of the center of the target space. As shown in FIG. 48, the sub-coordinates may be the coordinates of the upper right back corner of the target space. Also, the sub-coordinates are not limited to the coordinates of the lower left front or upper right back corner of the target space, and may be the coordinates of any corner of the target space.

[0280] Also, the set position of the sub-coordinates may be the same as the coordinates of a certain point cloud within the target space (sub-map, space, or volume). For example, in the example shown in FIG. 49, the coordinates of the sub-coordinates match the coordinates of the point cloud PD.

[0281] In addition, in this embodiment, an example of switching between applying octree encoding and non-octree encoding is shown, but it is not necessarily limited to this. For example, the three-dimensional data encoding device may switch between applying another tree structure other than the octree and applying a non-tree structure other than the tree structure. For example, another tree structure is a kd-tree that performs division using a plane perpendicular to one of the coordinate axes. Note that any method may be used as another tree structure.

[0282] In addition, in this embodiment, an example of encoding the coordinate information of the point cloud is shown, but it is not necessarily limited to this. The three-dimensional data encoding device may encode, for example, color information, three-dimensional feature amounts, or feature amounts of visible light in the same way as the coordinate information. For example, the three-dimensional data encoding device may set the average value of the color information of each point cloud in the submap as sub-color information, and encode the difference between the color information of each point cloud and the sub-color information.

[0283] In addition, in this embodiment, an example of selecting an encoding method (octree encoding or non-octree encoding) with good encoding efficiency according to the number of point clouds, etc. is shown, but it is not necessarily limited to this. For example, the three-dimensional data encoding device on the server side holds the bitstream of the point cloud encoded by octree encoding, the bitstream of the point cloud encoded by non-octree encoding, and the bitstream of the point cloud encoded by both, and switches the bitstream to be transmitted to the three-dimensional data decoding device according to the communication environment or the processing ability of the three-dimensional data decoding device.

[0284] FIG. 50 is a diagram showing a syntax example of volume when switching the application of octree encoding. The syntax shown in FIG. 50 is basically the same as the syntax shown in FIG. 40, but the difference is that each piece of information is information in volume units. Specifically, NumOfPoint indicates the number of point clouds included in the volume. sub_coordinate_x, sub_coordinate_y, and sub_coordinate_z are sub-coordinate information of the volume.

[0285] Also, diff_x[i], diff_y[i], and diff_z[i] are the differential coordinates of the i-th point cloud in the volume. diff_x[i] indicates the difference value between the x coordinate of the i-th point cloud in the volume and the x coordinate of the sub-coordinate. diff_y[i] indicates the difference value between the y coordinate of the i-th point cloud in the volume and the y coordinate of the sub-coordinate. diff_z[i] indicates the difference value between the z coordinate of the i-th point cloud in the volume and the z coordinate of the sub-coordinate.

[0286] Note that when the relative position of the volume in the space can be calculated, the three-dimensional data encoding device may not include the sub-coordinate information in the header of the volume. That is, the three-dimensional data encoding device may calculate the relative position of the volume in the space without including the sub-coordinate information in the header, and use the calculated position as the sub-coordinate of each volume.

[0287] As described above, the three-dimensional data encoding device according to the present embodiment determines whether to encode a target space unit among a plurality of space units (for example, submaps, spaces, or volumes) included in the three-dimensional data in an octree structure (for example, S1522 in FIG. 41). For example, when the number of three-dimensional points included in the target space unit is more than a predetermined threshold, the three-dimensional data encoding device determines to encode the target space unit in an octree structure. Also, when the number of three-dimensional points included in the target space unit is less than or equal to the above threshold, the three-dimensional data encoding device determines not to encode the target space unit in an octree structure.

[0288] When it is determined that the target space unit is to be encoded in an octree structure (Yes in S1522), the three-dimensional data encoder encodes the target space unit using the octree structure (S1523). Also, when it is determined that the target space unit is not to be encoded in an octree structure (No in S1522), the three-dimensional data encoder encodes the target space unit in a manner different from the octree structure (S1524). For example, in a different manner, the three-dimensional data encoder encodes the coordinates of the three-dimensional points included in the target space unit. Specifically, in a different manner, the three-dimensional data encoder encodes the difference between the reference coordinates of the target space unit and the coordinates of the three-dimensional points included in the target space unit.

[0289] Next, the three-dimensional data encoder adds information indicating whether the target space unit has been encoded in an octree structure to the bitstream (S1525).

[0290] According to this, the three-dimensional data encoder can reduce the data amount of the encoded signal, so the encoding efficiency can be improved.

[0291] For example, the three-dimensional data encoder includes a processor and a memory, and the processor performs the above processing using the memory.

[0292] Also, the three-dimensional data decoder according to the present embodiment decodes from the bitstream information indicating whether to decode a target space unit (for example, a submap, a space, or a volume) included in the three-dimensional data in an octree structure (for example, S1531 in FIG. 42). When it is indicated by the above information that the target space unit is to be decoded in an octree structure (Yes in S1532), the three-dimensional data decoder decodes the target space unit using the octree structure (S1533).

[0293] When it is shown that the target space unit is not decoded in an octree structure based on the above information (No in S1532), the three-dimensional data decoding device decodes the target space unit in a method different from the octree structure (S1534). For example, the three-dimensional data decoding device decodes the coordinates of the three-dimensional points included in the target space unit in a different method. Specifically, the three-dimensional data decoding device decodes the difference between the reference coordinates of the target space unit and the coordinates of the three-dimensional points included in the target space unit in a different method.

[0294] According to this, since the three-dimensional data decoding device can reduce the data amount of the encoded signal, the encoding efficiency can be improved.

[0295] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor performs the above processing using the memory.

[0296] (Embodiment 4) In the three-dimensional data encoding method according to Embodiment 4, the position information of a plurality of three-dimensional points is encoded using a prediction tree generated based on the position information.

[0297] FIG. 51 is a diagram showing an example of a prediction tree used in the three-dimensional data encoding method according to Embodiment 4. FIG. 52 is a flowchart showing an example of the three-dimensional data encoding method according to Embodiment 4. FIG. 53 is a flowchart showing an example of the three-dimensional data decoding method according to Embodiment 4.

[0298] As shown in FIGS. 51 and 52, in the three-dimensional data encoding method, a prediction tree is generated using a plurality of three-dimensional points, and then the node information included in each node of the prediction tree is encoded. Thereby, a bit stream including the encoded node information is obtained. Each node information is, for example, information regarding one node of the prediction tree. Each node information includes, for example, the position information of one node, the index of the one node, the number of child nodes that the one node has, the prediction mode used to encode the position information of the one node, and the prediction residual.

[0299] Also, as shown in FIGS. 51 and 53, in the three-dimensional data decoding method, each encoded node information included in the bit stream is decoded, and then the position information is decoded while generating a prediction tree.

[0300] Next, a method for generating a prediction tree will be described with reference to FIG. 54.

[0301] FIG. 54 is a diagram for explaining a method for generating a prediction tree according to Embodiment 4.

[0302] In the method for generating a prediction tree, as shown in FIG. 54(a), the three-dimensional data encoding device first adds point 0 as an initial point of the prediction tree. The position information of point 0 is represented by coordinates including three elements of (x0, y0, z0). The position information of point 0 may be represented by coordinates in a three-axis orthogonal coordinate system or coordinates in a polar coordinate system.

[0303] child_count is incremented by 1 each time one child node is added to the node for which child_count is set. The child_count of each node after completion of generating the prediction tree indicates the number of child nodes each node has, and is added to the bit stream. pred_mode indicates a prediction mode for predicting the value of the position information of each node. Details of the prediction mode will be described later.

[0304] Next, as shown in FIG. 54(b), the three-dimensional data encoding device adds point 1 to the prediction tree. At this time, the three-dimensional data encoding device may search for the nearest neighbor point of point 1 from the point group already added to the prediction tree and add point 1 as a child node of the nearest neighbor point. The position information of point 1 is represented by coordinates including three elements of (x1, y1, z1). The position information of point 1 may be represented by coordinates in a three-axis orthogonal coordinate system or coordinates in a polar coordinate system. In the case of FIG. 54, point 0 becomes the nearest neighbor point of point 1, and point 1 is added as a child node of point 0. Then, the three-dimensional data encoding device increments by 1 the value indicated by the child_count of point 0.

[0305] Note that the predicted value of the position information of each node may be calculated when a node is added to the prediction tree. For example, in the case of (b) in FIG. 54, the three-dimensional data encoding device may add point 1 as a child node of point 0 and calculate the position information of point 0 as the predicted value. In that case, pred_mode may be set to 1. pred_mode is prediction mode information (prediction mode value) indicating the prediction mode. Also, after calculating the predicted value, the three-dimensional data encoding device may calculate the residual_value (prediction residual) of point 1. Here, residual_value is the difference value obtained by subtracting the predicted value calculated in the prediction mode indicated by pred_mode from the position information of each node. Thus, in the three-dimensional data encoding method, the encoding efficiency can be improved by encoding not the position information itself but the difference value from the predicted value.

[0306] Next, as shown in (c) of FIG. 54, the three-dimensional data encoding device adds point 2 to the prediction tree. At this time, the three-dimensional data encoding device may search for the nearest neighbor point of point 2 from the point group already added to the prediction tree and add point 2 as a child node of the nearest neighbor point. The position information of point 2 is represented by coordinates including three elements of (x2, y2, z2). The position information of point 2 may be represented by coordinates in a three-axis orthogonal coordinate system or by coordinates in a polar coordinate system. In the case of FIG. 54, point 1 becomes the nearest neighbor point of point 2, and point 2 is added as a child node of point 1. Then, the three-dimensional data encoding device increments by 1 the value indicated by child_count of point 1.

[0307] Next, as shown in (d) of FIG. 54, the three-dimensional data encoding device adds point 3 to the prediction tree. At this time, the three-dimensional data encoding device may search for the nearest neighbor point of point 3 from the point group already added to the prediction tree and add point 3 as a child node of the nearest neighbor point. The position information of point 3 is represented by coordinates including three elements of (x3, y3, z3). The position information of point 3 may be represented by coordinates in a three-axis orthogonal coordinate system or by coordinates in a polar coordinate system. In the case of FIG. 54, point 0 becomes the nearest neighbor point of point 3, and point 3 is added as a child node of point 0. Then, the three-dimensional data encoding device increments by 1 the value indicated by child_count of point 0.

[0308] In this way, the three-dimensional data encoding device adds all the points to the prediction tree and completes the generation of the prediction tree. When the generation of the prediction tree is completed, the nodes with finally child_count = 0 become the leaves of the prediction tree. After the generation of the prediction tree is completed, the three-dimensional data encoding device encodes the child_count, pred_mode, and residual_value of each node selected in depth-first order from the root node. That is, when the three-dimensional data encoding device selects nodes in depth-first order, as the next node of the selected node, it selects an unselected child node among one or more child nodes of the selected node. When there is no child node for the selected node, the three-dimensional data encoding device selects an unselected other child node of the parent node of the selected node.

[0309] Note that the encoding order is not limited to depth-first order, and for example, it may be in breadth-first order. When the three-dimensional data encoding device selects nodes in breadth-first order, as the next node of the selected node, it selects an unselected node among one or more nodes at the same depth (hierarchy) as the selected node. When there is no node at the same depth as the selected node, the three-dimensional data encoding device selects an unselected node among one or more nodes at the next depth.

[0310] Note that points 0 to 3 are an example of a plurality of three-dimensional points.

[0311] Note that in the above three-dimensional data encoding method, it is assumed that child_count, pred_mode, and residual_value are calculated when each point is added to the prediction tree, but it is not necessarily limited to this. For example, they may be calculated after the generation of the prediction tree is completed.

[0312] The input order of the three-dimensional data of a plurality of three-dimensional points to the three-dimensional data encoding device may be to sort the input three-dimensional points in ascending or descending order of Morton order and process them in order from the first three-dimensional point. Thereby, the three-dimensional data encoding device can efficiently search for the nearest neighbor points of the three-dimensional points to be processed and improve the encoding efficiency. Also, the three-dimensional data encoding device may process the input three-dimensional points in the order in which they are input without sorting them. For example, the three-dimensional data encoding device may generate a prediction tree without branches in the input order of a plurality of three-dimensional points. Specifically, the three-dimensional data encoding device may add the three-dimensional point input next to the input three-dimensional point as a child node of a predetermined three-dimensional point in the input order of the plurality of three-dimensional points.

[0313] Next, a first example of the prediction mode will be described with reference to FIG. 55. FIG. 55 is a diagram for explaining a first example of the prediction mode according to Embodiment 4. FIG. 55 is a diagram showing a part of a prediction tree.

[0314] The prediction mode may be set to eight as shown below. For example, as shown in FIG. 55, the case of calculating the predicted value of point c will be described as an example. In the prediction tree, it is shown that the parent node of point c is point p0, the grandparent node of point c is point p1, and the great-grandparent node of point c is point p2. Note that point c, point p0, point p1, and point p2 are examples of a plurality of three-dimensional points.

[0315] The prediction mode with a prediction mode value of 0 (hereinafter referred to as prediction mode 0) may be set without prediction. That is, the three-dimensional data encoding device may calculate the position information of the input point c as the predicted value of the point c in prediction mode 0.

[0316] Also, the prediction mode with a prediction mode value of 1 (hereinafter referred to as prediction mode 1) may be set to differential prediction with point p0. That is, the three-dimensional data encoding device may calculate the position information of point p0, which is the parent node of point c, as the predicted value of the point c.

[0317] Also, the prediction mode with a prediction mode value of 2 (hereinafter referred to as prediction mode 2) may be set to linear prediction using point p0 and point p1. That is, the three-dimensional data encoding device may calculate, as the predicted value of point c, the prediction result by linear prediction using the position information of point p0, which is the parent node of point c, and the position information of point p1, which is the grandparent node of point c. Specifically, the three-dimensional data encoding device calculates the predicted value of point c in prediction mode 2 using the following formula T1.

[0318] Predicted value = 2 × p0 - p1 (Formula T1)

[0319] In formula T1, p0 represents the position information of point p0, and p1 represents the position information of point p1.

[0320] Also, the prediction mode with a prediction mode value of 3 (hereinafter referred to as prediction mode 3) may be set to Parallelogram prediction using point p0, point p1, and point p2. That is, the three-dimensional data encoding device may calculate, as the predicted value of point c, the prediction result by Parallelogram prediction using the position information of point p0, which is the parent node of point c, the position information of point p1, which is the grandparent node of point c, and the position information of point p2, which is the great-grandparent node of point c. Specifically, the three-dimensional data encoding device calculates the predicted value of point c in prediction mode 3 using the following formula T2.

[0321] Predicted value = p0 + p1 - p2 (Formula T2)

[0322] In formula T2, p0 represents the position information of point p0, p1 represents the position information of point p1, and p2 represents the position information of point p2.

[0323] Also, the prediction mode with a prediction mode value of 4 (hereinafter referred to as prediction mode 4) may be set to differential prediction with point p1. That is, the three-dimensional data encoding device may calculate the position information of point p1, which is the grandparent node of point c, as the predicted value of point c.

[0324] Also, the prediction mode with a prediction mode value of 5 (hereinafter referred to as prediction mode 5) may be set to differential prediction with respect to point p2. That is, the three-dimensional data encoding device may calculate the position information of point p2, which is the great-grandparent node of point c, as the predicted value of point c.

[0325] Also, the prediction mode with a prediction mode value of 6 (hereinafter referred to as prediction mode 6) may be set to the average of the position information of any two or more of point p0, point p1, and point p2. That is, the three-dimensional data encoding device may calculate the average value of the position information of two or more of the position information of point p0, which is the parent node of point c, the position information of point p1, which is the grandparent node of point c, and the position information of point p2, which is the great-grandparent node of point c, as the predicted value of point c. For example, when the three-dimensional data encoding device uses the position information of point p0 and the position information of point p1 to calculate the predicted value, the predicted value of point c in prediction mode 6 is calculated using the following formula T3.

[0326] Predicted value = (p0 + p1) / 2 (Formula T3)

[0327] In formula T3, p0 represents the position information of point p0, and p1 represents the position information of point p1.

[0328] Also, the prediction mode with a prediction mode value of 7 (hereinafter referred to as prediction mode 7) may be set to non-linear prediction using the distance d0 between point p0 and point p1 and the distance d1 between point p2 and point p1. That is, the three-dimensional data encoding device may calculate the prediction result by non-linear prediction using the distance d0 and the distance d1 as the predicted value of point c.

[0329] Note that the prediction method assigned to each prediction mode is not limited to the above example. Also, the above eight prediction modes and the above eight prediction methods do not have to be in the above combination, and any combination may be used. For example, when encoding a prediction mode using entropy encoding such as arithmetic coding, a prediction method with a high usage frequency may be assigned to prediction mode 0. This can improve the encoding efficiency. Further, the three-dimensional data encoding device may improve the encoding efficiency by dynamically changing the assignment of the prediction mode according to the usage frequency of the prediction mode while proceeding with the encoding process. For example, the three-dimensional data encoding device may count the usage frequency of each prediction mode during encoding and assign a prediction mode indicated by a smaller value to a prediction method with a higher usage frequency. This can improve the encoding efficiency. Note that M is the number of prediction modes indicating the number of prediction modes, and in the case of the above example, since there are eight prediction modes from prediction mode 0 to 7, M = 8.

[0330] The three-dimensional data encoding device may calculate a predicted value used for calculating the position information of the three-dimensional point to be encoded as the predicted values (px, py, pz) of the position information (x, y, z) of the three-dimensional point using the position information of the three-dimensional points near the three-dimensional point to be encoded among the three-dimensional points around the three-dimensional point to be encoded. Further, the three-dimensional data encoding device may add prediction mode information (pred_mode) for each three-dimensional point so that a predicted value calculated according to the prediction mode can be selected.

[0331] For example, in the prediction modes with a total number of M, it is conceivable to assign the position information of the nearest neighbor three-dimensional point p0 to prediction mode 0, ···, the position information of three-dimensional point p2 to prediction mode M - 1, and add the prediction mode used for prediction to the bit stream for each three-dimensional point.

[0332] Note that the number of prediction modes M may be added to the bit stream. Also, the number of prediction modes M may be defined by values in the profile, level, etc. of the standard without being added to the bit stream. Further, the number of prediction modes M may be a value calculated from the three-dimensional point number N used for prediction. For example, the number of prediction modes M may be calculated by M = N + 1.

[0333] FIG. 56 is a diagram showing a second example of a table indicating predicted values calculated in each prediction mode according to Embodiment 4.

[0334] The table shown in FIG. 56 is an example in the case where the three-dimensional point number N = 4 used for prediction and the number of prediction modes M = 5.

[0335] In the second example, the predicted value of the position information of point c is calculated using the position information of at least one of point p0, point p1, and point p2. The prediction mode is added for each three-dimensional point to be encoded. The predicted value is calculated to be a value corresponding to the added prediction mode.

[0336] FIG. 57 is a diagram showing a specific example of the second example of the table indicating predicted values calculated in each prediction mode according to Embodiment 4.

[0337] The three-dimensional data encoding device may, for example, select prediction mode 1 and encode the position information (x, y, z) of the three-dimensional point to be encoded using the predicted values (p0x, p0y, p0z), respectively. In this case, "1", which is the prediction mode value indicating the selected prediction mode 1, is added to the bit stream.

[0338] In this way, the three-dimensional data encoding device may select a common prediction mode for the three elements in selecting a prediction mode, as one prediction mode for calculating the predicted value of each of the three elements included in the position information of the three-dimensional point to be encoded.

[0339] FIG. 58 is a diagram showing a third example of a table indicating predicted values calculated in each prediction mode according to Embodiment 4.

[0340] The table shown in FIG. 58 is an example in the case where the three-dimensional number of points N = 2 used for prediction and the number of prediction modes M = 5.

[0341] In the third example, the predicted value of the position information of point c is calculated using the position information of at least one of point p0 and point p1. The prediction mode is added for each three-dimensional point to be encoded. The predicted value is calculated to be a value corresponding to the added prediction mode.

[0342] Note that, as in the third example, when the number of points (number of adjacent points) around point c is less than 3, the prediction mode for which the predicted value is unassigned may be set to not available. Also, when a prediction mode with not available set occurs, another prediction method may be assigned to that prediction mode. For example, the position information of point p2 may be assigned as the predicted value to that prediction mode. Also, the predicted value assigned to another prediction mode may be assigned to that prediction mode. For example, the position information of point p1 assigned to prediction mode 4 may be assigned to prediction mode 3 for which not available is set. In that case, the position information of point p2 may be newly assigned to prediction mode 4. In this way, when a prediction mode with not available set occurs, the encoding efficiency can be improved by assigning a new prediction method to the prediction mode.

[0343] In addition, when the position information has three elements such as a three-axis orthogonal coordinate system or a polar coordinate system, the predicted values may be calculated in a prediction mode divided for each of the three elements. For example, when the three elements are represented by x, y, and z of the coordinates (x, y, z) in a three-axis orthogonal coordinate system, the predicted values of the three elements may be calculated in the prediction mode selected for each element. For example, a prediction mode pred_mode_x for calculating the predicted value of element x (i.e., the x coordinate), a prediction mode pred_mode_y for calculating the predicted value of element y (i.e., the y coordinate), and a prediction mode pred_mode_z for calculating the predicted value of element z (i.e., the z coordinate), the prediction mode values may be selected respectively. In this case, as the prediction mode values indicating the prediction mode of each element, the values in the tables of FIGS. 59 to 61 described later are used, and these prediction mode values may be added to the bit stream respectively. Note that in the above, as an example of the position information, the coordinates of the three-axis orthogonal coordinate system have been described, but the same can be applied to the coordinates of the polar coordinate system.

[0344] As described above, in the selection of the prediction mode, the three-dimensional data encoding device may select an independent prediction mode for each of the three elements as one prediction mode for calculating the predicted value of each element included in the position information of the three-dimensional point to be encoded.

[0345] In addition, the predicted values including two or more elements among the plurality of elements of the position information may be calculated in a common prediction mode. For example, when the three elements are represented by x, y, and z of the coordinates (x, y, z) in a three-axis orthogonal coordinate system, the prediction mode values may be selected respectively in a prediction mode pred_mode_x for calculating the predicted value using element x and a prediction mode pred_mode_yz for calculating the predicted value using element y and element z. In this case, as the prediction mode values indicating the prediction mode of each component, the values in the tables of FIGS. 59 and 62 described later are used, and these prediction mode values may be added to the bit stream respectively.

[0346] In this way, in the selection of the prediction mode, the three-dimensional data encoding device may select, as one prediction mode for calculating the predicted value of each of the three elements included in the position information of the three-dimensional point to be encoded, a common prediction mode for two of the three elements, and select a prediction mode independent of the two elements for the remaining one element.

[0347] FIG. 59 is a diagram showing a fourth example of a table indicating predicted values calculated in each prediction mode. Specifically, the fourth example is an example in the case where the position information used for the predicted value is the value of the element x of the position information of the surrounding three-dimensional points.

[0348] As shown in FIG. 59, the predicted value calculated in the prediction mode pred_mode_x indicated by the prediction mode value "0" is 0. Also, the predicted value calculated in the prediction mode pred_mode_x indicated by the prediction mode value "1" is the x coordinate of the point p0, which is p0x. Also, the predicted value calculated in the prediction mode pred_mode_x indicated by the prediction mode value "2" is the prediction result of linear prediction based on the x coordinate of the point p0 and the x coordinate of the point p1, which is (2 × p0x - p1x). Also, the predicted value calculated in the prediction mode pred_mode_x indicated by the prediction mode value "3" is the prediction result of parallelogram prediction based on the x coordinate of the point p0, the x coordinate of the point p1, and the x coordinate of the point p2, which is (p0x + p1x - p2x). Also, the predicted value calculated in the prediction mode pred_mode_x indicated by the prediction mode value "4" is the x coordinate of the point p1, which is p1x.

[0349] For example, when the prediction mode pred_mode_x indicated by the prediction mode value "1" in the table of FIG. 59 is selected, the x coordinate of the position information of the three-dimensional point to be encoded may be encoded using the predicted value p0x. In this case, "1" as the prediction mode value is added to the bit stream.

[0350] FIG. 60 is a diagram showing a fifth example of a table indicating predicted values calculated in each prediction mode. Specifically, the fifth example is an example where the position information used for the predicted value is the value of the element y of the position information of the surrounding three-dimensional points.

[0351] As shown in FIG. 60, the predicted value calculated in the prediction mode pred_mode_y indicated by the prediction mode value "0" is 0. Also, the predicted value calculated in the prediction mode pred_mode_y indicated by the prediction mode value "1" is the y-coordinate of point p0, which is p0y. Also, the predicted value calculated in the prediction mode pred_mode_y indicated by the prediction mode value "2" is the prediction result of linear prediction based on the y-coordinate of point p0 and the y-coordinate of point p1, which is (2×p0y - p1y). Also, the predicted value calculated in the prediction mode pred_mode_y indicated by the prediction mode value "3" is the prediction result of Parallelogram prediction based on the y-coordinate of point p0, the y-coordinate of point p1, and the y-coordinate of point p2, which is (p0y + p1y - p2y). Also, the predicted value calculated in the prediction mode pred_mode_y indicated by the prediction mode value "4" is the y-coordinate of point p1, which is p1y.

[0352] For example, when the prediction mode pred_mode_y indicated by the prediction mode value "1" in the table of FIG. 60 is selected, the y-coordinate of the position information of the three-dimensional point to be encoded may be encoded using the predicted value p0y. In this case, "1" as the prediction mode value is added to the bit stream.

[0353] FIG. 61 is a diagram showing a sixth example of a table indicating predicted values calculated in each prediction mode. Specifically, the sixth example is an example where the position information used for the predicted value is the value of the element z of the position information of the surrounding three-dimensional points.

[0354] As shown in FIG. 61, the predicted value calculated in the prediction mode pred_mode_z where the prediction mode value is indicated by "0" is 0. Also, the predicted value calculated in the prediction mode pred_mode_z where the prediction mode value is indicated by "1" is the z coordinate of point p0, which is p0z. Further, the predicted value calculated in the prediction mode pred_mode_z where the prediction mode value is indicated by "2" is the prediction result of linear prediction based on the z coordinate of point p0 and the z coordinate of point p1, which is (2×p0z - p1z). Also, the predicted value calculated in the prediction mode pred_mode_z where the prediction mode value is indicated by "3" is the prediction result of Parallelogram prediction based on the z coordinate of point p0, the z coordinate of point p1, and the z coordinate of point p2, which is (p0z + p1z - p2z). Also, the predicted value calculated in the prediction mode pred_mode_z where the prediction mode value is indicated by "4" is the z coordinate of point p1, which is p1z.

[0355] Note that, for example, when the prediction mode pred_mode_z where the prediction mode value is indicated by "1" is selected in the table of FIG. 61, the z coordinate of the position information of the three-dimensional point to be encoded may be encoded using the predicted value p0z. In this case, "1" as the prediction mode value is added to the bit stream.

[0356] FIG. 62 is a diagram showing a seventh example of a table indicating the predicted values calculated in each prediction mode. Specifically, the seventh example is an example where the position information used for the predicted value is the values of element y and element z of the position information of surrounding three-dimensional points.

[0357] As shown in FIG. 62, in the prediction mode pred_mode_yz where the prediction mode value is indicated by "0", the predicted value calculated is 0. Also, in the prediction mode pred_mode_yz where the prediction mode value is indicated by "1", the predicted values are the y-coordinate and z-coordinate of point p0, which are (p0y, p0z). Further, in the prediction mode pred_mode_yz where the prediction mode value is indicated by "2", the predicted value calculated is the prediction result of linear prediction based on the y-coordinate and z-coordinate of point p0 and the y-coordinate and z-coordinate of point p1, which is (2×p0y - p1y, 2×p0z - p1z). Moreover, in the prediction mode pred_mode_yz where the prediction mode value is indicated by "3", the predicted value calculated is the prediction result of Parallelogram prediction based on the y-coordinate and z-coordinate of point p0, the y-coordinate and z-coordinate of point p1, and the y-coordinate and z-coordinate of point p2, which is (p0y + p1y - p2y, p0z + p1z - p2z). Also, in the prediction mode pred_mode_yz where the prediction mode value is indicated by "4", the predicted values are the y-coordinate and z-coordinate of point p1, which are (p1y, p1z).

[0358] Note that, for example, when the prediction mode pred_mode_yz where the prediction mode value is indicated by "1" is selected in the table of FIG. 62, the y-coordinate and z-coordinate of the position information of the three-dimensional point to be encoded may be encoded using the predicted values (p0y, p0z). In this case, "1" as the prediction mode value is added to the bitstream.

[0359] In the tables in the 4th to 7th examples, the correspondence between the prediction mode and the prediction method of the calculated predicted values is the same as the above correspondence in the table of the 2nd example.

[0360] The prediction mode during symbolization may be selected by RD optimization. For example, it is conceivable to calculate the cost cost(P) when a certain prediction mode P is selected, and select the prediction mode P that minimizes cost(P). As the cost cost(P), for example, it may be calculated by Equation D1 using the prediction residual residual_value(P) when the predicted value of the prediction mode P is used, the number of bits bit(P) required to encode the prediction mode P, and the adjustment parameter λ value.

[0361] cost(P)=abs(residual(P))+λ×bit(P)···(Equation D1)

[0362] abs(x) indicates the absolute value of x.

[0363] Instead of abs(x), the squared value of x may be used.

[0364] By using the above Equation D1, it becomes possible to select a prediction mode that takes into account the balance between the magnitude of the prediction residual and the number of bits required to encode the prediction mode. Note that the adjustment parameter λ may be set to different values according to the value of the quantization scale. For example, when the quantization scale is small (at high bit rate), by making the λ value small, a prediction mode with a small prediction residual residual_value(P) is selected to improve the prediction accuracy as much as possible. When the quantization scale is large (at low bit rate), by making the λ value large, an appropriate prediction mode may be selected while considering the number of bits bit(P) required to encode the prediction mode P.

[0365] Note that the case where the quantization scale is small means, for example, the case where it is smaller than the first quantization scale. The case where the quantization scale is large means, for example, the case where it is larger than the second quantization scale that is equal to or greater than the first quantization scale. Also, the smaller the quantization scale, the smaller the λ value may be set.

[0366] The prediction residual residual_value(P) is calculated by subtracting the predicted value of the prediction mode P from the position information of the three-dimensional point to be encoded. Note that instead of the prediction residual residual_value(P) at the time of cost calculation, the prediction residual residual_value(P) may be quantized, dequantized, added to the predicted value to obtain a decoded value, and the difference (encoding error) between the decoded value when using the position information of the original three-dimensional point and the prediction mode P may be reflected in the cost value. Thereby, it becomes possible to select a prediction mode with a small encoding error.

[0367] The number of bits bit(P) required to encode the prediction mode P may be, for example, the number of bits after binarization when the prediction mode is binarized and encoded.

[0368] For example, when the number of prediction modes M = 5, as shown in FIG. 63, the prediction mode value indicating the prediction mode may be binarized with a truncated unary code having a maximum value of 5 using the number of prediction modes M. In this case, when the prediction mode value is "0", it is 1 bit, when the prediction mode value is "1", it is 2 bits, when the prediction mode value is "2", it is 3 bits, and when the prediction mode values are "3" and "4", it is 4 bits, which are used as the number of bits bit(P) required for encoding each prediction mode value. By using the truncated unary code, the smaller the value of the prediction mode, the fewer the number of bits. Therefore, 0 calculated as the predicted value when the prediction mode value is "0", or the position information of the three-dimensional point p0 calculated as the predicted value when the prediction mode value is "1", that is, the position information of the three-dimensional point close to the three-dimensional point to be encoded, for example, it is easy to select a prediction mode that calculates a predicted value for which cost(P) is likely to be minimized, and the amount of code for the prediction mode value indicating the prediction mode can be reduced.

[0369] In this way, the three-dimensional data encoding device may encode the prediction mode value indicating the selected prediction mode using the number of prediction modes. Specifically, the three-dimensional data encoding device may encode the prediction mode value with a truncated unary code having the number of prediction modes as the maximum value.

[0370] Also, when the maximum value of the number of prediction modes is not determined, as shown in FIG. 64, the prediction mode value indicating the prediction mode may be binarized with a unary code. Also, when the occurrence probabilities of each prediction mode are close, as shown in FIG. 65, the prediction mode value indicating the prediction mode may be binarized with a fixed code to reduce the amount of code.

[0371] Note that, as the number of bits bit(P) required to encode the prediction mode value indicating the prediction mode P, the binary data of the prediction mode value indicating the prediction mode P may be arithmetic-coded, and the amount of code after arithmetic coding may be used as the value of bit(P). Thereby, since the cost can be calculated using the more accurate required number of bits bit(P), it becomes possible to select a more appropriate prediction mode.

[0372] Note that FIG. 63 is a diagram showing a first example of a binarization table in the case of binarizing and encoding the prediction mode value according to Embodiment 4. Specifically, the first example is an example of binarizing the prediction mode value with a truncated unary code when the number of prediction modes M = 5.

[0373] Also, FIG. 64 is a diagram showing a second example of a binarization table in the case of binarizing and encoding the prediction mode value according to Embodiment 4. Specifically, the second example is an example of binarizing the prediction mode value with a unary code when the number of prediction modes M = 5.

[0374] Also, FIG. 65 is a diagram showing a third example of a binarization table in the case of binarizing and encoding the prediction mode value according to Embodiment 4. Specifically, the third example is an example of binarizing the prediction mode value with a fixed code when the number of prediction modes M = 5.

[0375] The prediction mode value indicating the prediction mode (pred_mode) may be arithmetic-coded after binarization and added to the bitstream. As described above, the prediction mode value may be binarized, for example, with a truncated unary code using the value of the number of prediction modes M. In this case, the maximum number of bits after binarization of the prediction mode value is M - 1.

[0376] Also, the binarized binary data may be arithmetic-coded using a coding table. In this case, for example, the coding efficiency may be improved by switching the coding table for each bit of the binary data for coding. Also, in order to suppress the number of coding tables, among the binary data, the first bit one bit may be coded using the coding table A for one bit, and each bit of the remaining bits remaining bit may be coded using the coding table B for the remaining bit. For example, when coding the binary data "1110" with the prediction mode value of "3" shown in FIG. 66, the first bit one bit "1" may be coded using the coding table A, and each bit of the remaining bits remaining bit "110" may be coded using the coding table B.

[0377] Note that FIG. 66 is a diagram for explaining an example of coding binary data in a binarization table when binarizing and coding the prediction mode according to Embodiment 4. The binarization table in FIG. 66 is an example of binarizing the prediction mode value with a truncated unary code when the number of prediction modes M = 5.

[0378] Thereby, while suppressing the number of coding tables, the coding efficiency can be improved by switching the coding table according to the bit position of the binary data. Note that when coding the Remaining bit, furthermore, the coding table may be switched for each bit for arithmetic coding, or the coding table may be switched according to the result of arithmetic coding for decoding.

[0379] When encoding the prediction mode value by binarizing it with a truncated unary code using the number of prediction modes M, the number of prediction modes M used for the truncated unary code may be added to the header of the bit stream or the like so that the prediction mode can be specified from the binary data decoded on the decoding side. The header of the bit stream is, for example, a sequence parameter set (SPS), a position parameter set (GPS), a slice header, or the like. Also, the possible value MaxM of the number of prediction modes may be defined by a standard or the like, and the value of MaxM - M (M <= MaxM) may be added to the header. Also, the number of prediction modes M may be defined by a profile or level of a standard or the like without being added to the stream.

[0380] Note that the prediction mode value binarized using the truncated unary code can be arithmetically encoded by switching the encoding table between the one bit part and the remaining part as described above. Note that the occurrence probabilities of 0 and 1 in each encoding table may be updated according to the values of the actually generated binary data. Also, the occurrence probabilities of 0 and 1 in either encoding table may be fixed. Thereby, the number of update times of the occurrence probability may be suppressed to reduce the processing amount. For example, the occurrence probability of the one bit part may be updated and the occurrence probability of the remaining bit part may be fixed.

[0381] FIG. 67 is a flowchart showing an example of encoding of a prediction mode value according to Embodiment 4. FIG. 68 is a flowchart showing an example of decoding of a prediction mode value according to Embodiment 4.

[0382] As shown in FIG. 67, in the encoding of the prediction mode value, first, the prediction mode value is binarized with a truncated unary code using the number of prediction modes M (S9701).

[0383] Next, the binary data of the truncated unary code is arithmetically encoded (S9702). Thereby, the binary data is included as the prediction mode in the bit stream.

[0384] Also, as shown in FIG. 68, in the decoding of the prediction mode value, first, the bit stream is arithmetically decoded using the number of prediction modes M to generate binary data of truncated unary code (S9711).

[0385] Next, the prediction mode value is calculated from the binary data of the truncated unary code (S9712).

[0386] As an example of the method for binarizing the prediction mode value indicating the prediction mode (pred_mode), an example of binarization using truncated unary code with the value of the number of prediction modes M is shown, but it is not necessarily limited to this. For example, the prediction mode value may be binarized using truncated unary code with the number L (L <= M) of prediction values assigned to the prediction mode. For example, when the number of prediction modes M = 5 and there is 1 surrounding three-dimensional point available for predicting a certain coded three-dimensional point, as shown in FIG. 69, there is a case where 2 prediction modes become available and the remaining 3 prediction modes become not available. For example, as shown in FIG. 69, when the number of prediction modes M = 5 and the number of three-dimensional points available for prediction around the coded three-dimensional point is 1, there may be a case where no prediction value is assigned to the prediction modes indicating "2", "3", and "4" for the prediction mode value.

[0387] In this case, as shown in FIG. 70, by binarizing the prediction mode value using truncated unary code with the value L assigned to the prediction mode as the maximum value, it may be possible to reduce the number of bits after binarization compared to the case of using truncated unary code with the number of prediction modes M. For example, in this case L = 3, so the number of bits can be reduced by binarizing with truncated unary code with the maximum value of 3. In this way, by binarizing with truncated unary code with the number L of prediction values assigned to the prediction mode as the maximum value, the number of bits after binarization of the prediction mode value may be reduced.

[0388] The binary data after binarization may be arithmetic-coded using a coding table. In this case, for example, the coding efficiency may be improved by switching the coding table for each bit of the binary data and performing coding. Also, in order to suppress the number of coding tables, among the binary data, for the first bit one bit, coding may be performed using the coding table A for one bit, and for each bit of the remaining bits remaining bit, coding may be performed using the coding table B for the remaining bit. For example, when coding the binary data "1" whose predicted mode value shown in FIG. 70 is "1", the coding table A is used to code the "1" of the first bit one bit. Since there are no remaining bits remaining bit, it may not be necessary to perform coding. If there are remaining bits remaining bit, the remaining bits remaining bit may be coded using the coding table B.

[0389] Note that FIG. 70 is a diagram for explaining an example of coding the binary data of the binarization table in the case of binarizing and coding the prediction mode according to Embodiment 4. The binarization table in FIG. 70 is an example of binarizing the prediction mode value with a truncated unary code when the number L to which the prediction value is assigned in the prediction mode is 2.

[0390] Thereby, while suppressing the number of coding tables, the coding efficiency can be improved by switching the coding table according to the bit position of the binary data. Note that at the time of coding the Remaining bit, further, the coding table may be switched for each bit and arithmetic coding may be performed, or the coding table may be switched according to the result of arithmetic coding and decoding may be performed.

[0391] When binarizing and coding the prediction mode value with a truncated unary code using the number L to which the prediction value is assigned, in order to be able to specify the prediction mode from the binary data decoded on the decoding side, the prediction value is assigned to the prediction mode in the same method as at the time of coding to calculate the number L, and the calculated L may be used to decode the prediction mode.

[0392] Note that for the predicted mode value binarized using the truncated unary code, it is conceivable to perform arithmetic coding by switching the coding table between the one-bit part and the remaining part as described above. Note that the occurrence probabilities of 0 and 1 in each coding table may be updated according to the value of the actually generated binary data. Also, the occurrence probabilities of 0 and 1 in either coding table may be fixed. Thereby, the number of updates of the occurrence probability may be suppressed to reduce the processing amount. For example, the occurrence probability of the one-bit part may be updated and the occurrence probability of the remaining bit part may be fixed.

[0393] FIG. 71 is a flowchart showing another example of the coding of the predicted mode value according to Embodiment 4. FIG. 72 is a flowchart showing another example of the decoding of the predicted mode value according to Embodiment 4.

[0394] As shown in FIG. 71, in the coding of the predicted mode value, first, the number L to which the predicted value is assigned to the prediction mode is calculated (S9721).

[0395] Next, the predicted mode value is binarized with the truncated unary code using the number L (S9722).

[0396] Next, the binary data of the truncated unary code is arithmetically coded (S9723).

[0397] Also, as shown in FIG. 72, in the decoding of the predicted mode value, first, the number L to which the predicted value is assigned to the prediction mode is calculated (S9731).

[0398] Next, the bit stream is arithmetically decoded using the number L to generate the binary data of the truncated unary code (S9732).

[0399] Next, the predicted mode value is calculated from the binary data of the truncated unary code (S9733).

[0400] The prediction mode value does not have to be added for each piece of position information. For example, if a certain condition is satisfied, the prediction mode may be fixed and the prediction mode value may not be added to the bit stream. If a certain condition is not satisfied, the prediction mode may be selected and the prediction mode value may be added to the bit stream. For example, if condition A is satisfied, the prediction mode value may be fixed at "2" and the predicted value may be calculated from the linear prediction of the surrounding three-dimensional points. If condition A is not satisfied, one prediction mode may be selected from a plurality of prediction modes and the prediction mode value indicating the selected prediction mode may be added to the bit stream.

[0401] As a certain condition A, for example, the distance d0 between point p1 and point P0 and the distance d1 between point p2 and point p1 are calculated, and the absolute value of the difference distdiff = |d0 - d1| is less than the threshold Thfix. When the absolute value of the difference is less than the threshold Thfix, the three-dimensional data encoding device determines that the difference between the predicted value by linear prediction and the position information of the point to be processed is small, fixes the prediction mode value at "2", and does not encode the prediction mode value, thereby reducing the amount of code for encoding the prediction mode and generating an appropriate predicted value. Note that when the absolute value of the difference is greater than or equal to the threshold Thfix, the three-dimensional data encoding device may select a prediction mode and encode the prediction mode value indicating the selected prediction mode.

[0402] Note that the threshold Thfix may be added to the header of the bitstream or the like, and the encoder may be able to change the value of the threshold Thfix for encoding. For example, when encoding at a high bitrate, the encoder may add the value of the threshold Thfix to the header with a smaller value than that at a low bitrate, and increase the cases of encoding by selecting the prediction mode so as to make the prediction residual as small as possible. Also, when encoding at a low bitrate, the encoder may add the value of the threshold Thfix to the header with a larger value than that at a high bitrate, and encode with the prediction mode fixed. In this way, by increasing the cases where the prediction mode is fixed and encoded at a low bitrate, it is possible to improve the encoding efficiency while suppressing the amount of bits for encoding the prediction mode. Also, the threshold Thfix may be defined by the profile or level of the standard without being added to the bitstream.

[0403] The N three-dimensional points around the three-dimensional point to be encoded used for prediction are N encoded and decoded three-dimensional points whose distance from the three-dimensional point to be encoded is smaller than the threshold THd. The maximum value of N may be added to the bitstream as NumNeighborPoint. The value of N does not necessarily have to always match the value of NumNeighborPoint, such as when the number of encoded and decoded three-dimensional points around is less than the value of NumNeighborPoint.

[0404] Although an example is shown in which the prediction mode value is fixed to "2" if the absolute difference in prediction distdiff is smaller than the threshold Thfix[i], it is not necessarily limited to this, and the prediction mode value may be fixed to any one of "0" to "M - 1". Also, the fixed prediction mode value may be added to the bitstream.

[0405] FIG. 73 is a flowchart showing an example of a process for determining whether to fix a prediction mode value according to condition A during encoding according to Embodiment 4. FIG. 74 is a flowchart showing an example of a process for determining whether to set a fixed value or decode a prediction mode value according to condition A during decoding according to Embodiment 4.

[0406] As shown in FIG. 73, first, the three-dimensional data encoding device calculates the distance d0 between point p1 and point p0 and the distance d1 between point p2 and point p1, and calculates the absolute difference distdiff = |d0 - d1| (S9741).

[0407] Next, the three-dimensional data encoding device determines whether the absolute difference distdiff is less than the threshold value Thfix (S9742). Note that the threshold value Thfix may be encoded and added to the header of the stream or the like.

[0408] When the absolute difference distdiff is less than the threshold value Thfix (Yes in S9742), the three-dimensional data encoding device determines the prediction mode value as "2" (S9743).

[0409] On the other hand, when the absolute difference distdiff is greater than or equal to the threshold value Thfix (No in S9742), the three-dimensional data encoding device sets one prediction mode among a plurality of prediction modes (S9744).

[0410] Then, the three-dimensional data encoding device arithmetic-encodes the prediction mode value indicating the set prediction mode (S9745). Specifically, the three-dimensional data encoding device arithmetic-encodes the prediction mode value by executing steps S9701 and S9702 described in FIG. 67. Note that the three-dimensional data encoding device may binarize and arithmetic-encode the prediction mode pred_mode using the number of prediction modes to which the prediction value is assigned in truncated unary code. That is, the three-dimensional data encoding device may arithmetic-encode the prediction mode value by executing steps S9721 to S9723 described in FIG. 71.

[0411] The three-dimensional data encoding device calculates the predicted value of the prediction mode determined in step S9743 or the prediction mode set in step S9745, and outputs the calculated predicted value (S9746). When using the prediction mode value determined in step S9743, the three-dimensional data encoding device calculates the predicted value of the prediction mode indicated by the prediction mode value "2" by linear prediction of the position information of N surrounding three-dimensional points.

[0412] Also, as shown in FIG. 74, first, the three-dimensional data decoding device calculates the distance d0 between point p1 and point p0 and the distance d1 between point p2 and point p1, and calculates the absolute value of the difference distdiff = |d0 - d1| (S9751).

[0413] Next, the three-dimensional data decoding device determines whether the absolute value of the difference distdiff is less than the threshold Thfix (S9752). Note that the threshold Thfix may be set after decoding the header of the stream or the like.

[0414] When the absolute value of the difference distdiff is less than the threshold Thfix (Yes in S9752), the three-dimensional data decoding device determines the prediction mode value as "2" (S9753).

[0415] On the other hand, when the absolute value of the difference distdiff is greater than or equal to the threshold Thfix (No in S9752), the three-dimensional data decoding device decodes the prediction mode value from the bit stream (S9754).

[0416] The three-dimensional data decoding device calculates the predicted value of the prediction mode indicated by the prediction mode value determined in step S9753 or the prediction mode value decoded in step S9754, and outputs the calculated predicted value (S9755). When using the prediction mode value determined in step S9753, the three-dimensional data decoding device calculates the predicted value of the prediction mode indicated by the prediction mode value "2" by linear prediction of the position information of N surrounding three-dimensional points.

[0417] FIG. 75 is a diagram showing an example of the syntax of the header of position information. NumNeighborPoint, NumPredMode, Thfix, QP, and unique_point_per_leaf in the syntax of FIG. 75 will be described in order.

[0418] NumNeighborPoint indicates the upper limit value of the number of surrounding points used to generate the predicted value of the position information of the three-dimensional points. When the number of surrounding points M is less than NumNeighborPoint (M < NumNeighborPoint), in the calculation process of the predicted value, the predicted value may be calculated using M surrounding points.

[0419] NumPredMode indicates the total number M of prediction modes used for predicting position information. Note that the maximum value MaxM of the possible values of the number of prediction modes may be defined by a standard or the like. The three-dimensional data encoding device may add the value of (MaxM - M) (0 < M <= MaxM) as NumPredMode to the header, and binarize and encode (MaxM - 1) with a truncated unary code. Also, the number of prediction modes NumPredMode may not be added to the bit stream, and may be defined by a profile or level of a standard or the like. Also, the number of prediction modes may be defined by NumNeighborPoint + NumPredMode.

[0420] Thfix is a threshold value for determining whether to fix the prediction mode. The distance d0 between the points p1 and p0 used for prediction and the distance d1 between the points p2 and p1 are calculated, and if the absolute value of the difference distdiff = |d0 - d1| is smaller than the threshold value Thfix[i], the prediction mode is fixed to α. α is a prediction mode for calculating a predicted value using linear prediction, and is "2" in the above embodiment. Note that Thfix may not be added to the bit stream, and may be defined by a profile or level of a standard or the like.

[0421] QP indicates the quantization parameter used when quantizing position information. The three-dimensional data encoder may calculate a quantization step from the quantization parameter and quantize the position information using the calculated quantization step.

[0422] unique_point_per_leaf is information indicating whether the bitstream contains a duplicated point (points with the same position information). unique_point_per_leaf = 1 indicates that there is no duplicated point in the bitstream. unique_point_per_leaf = 0 indicates that there is one or more duplicated points in the bitstream.

[0423] In this embodiment, the determination of whether to fix the prediction mode is made using the absolute value of the difference between distance d0 and distance d1. However, it is not necessarily limited to this, and any method of determination may be used. For example, this determination calculates the distance d0 between point p1 and point p0. If the distance d0 is greater than the threshold, it is determined that point p1 cannot be used for prediction, and the prediction mode value is fixed to "1" (predicted value p0). Otherwise, the prediction mode may be set. This can improve the coding efficiency while suppressing overhead.

[0424] The above NumNeighborPoint, NumPredMode, Thfix, and unique_point_per_leaf may be entropy-coded and added to the header. For example, each value may be binarized and then calculated and coded. Also, each value may be coded with a fixed length to reduce the processing amount.

[0425] FIG. 76 is a diagram showing an example of the syntax of position information. The NumOfPoint, child_count, pred_mode, and residual_value[j] in the syntax of FIG. 76 will be described in order.

[0426] NumOfPoint indicates the total number of three-dimensional points included in the bitstream.

[0427] child_count indicates the number of child nodes that the i-th three-dimensional point (node[i]) has.

[0428] pred_mode indicates the prediction mode for encoding or decoding the position information of the i-th three-dimensional point. pred_mode takes values from 0 to M-1 (M is the total number of prediction modes). When pred_mode is not in the bitstream (when the condition distdiff >= Thfix[i] && NumPredMode > 1 is not satisfied), pred_mode may be estimated as a fixed value α. α is a prediction mode for calculating a predicted value using linear prediction, and in the above embodiment, it is "2". Note that α is not limited to "2", and any value from 0 to M-1 may be set as the estimated value. Also, the estimated value when pred_mode is not in the bitstream may be added separately to a header or the like. Further, pred_mode may be binarized using a truncated unary code with the number of prediction modes to which the predicted value is assigned and then arithmetic-coded.

[0429] Note that when NumPredMode = 1, that is, when the number of prediction modes is 1, the three-dimensional data encoding device may generate a bitstream that does not include the prediction mode value without encoding the prediction mode value indicating the prediction mode. Also, when the three-dimensional data decoding device obtains a bitstream that does not include the prediction mode value, in calculating the predicted value, it may calculate the predicted value of a specific prediction mode. The specific prediction mode is a predefined prediction mode.

[0430] residual_value[j] indicates the encoded data of the prediction residual between the predicted value of the position information. residual_value[0] may indicate the element x of the position information, residual_value[1] may indicate the element y of the position information, and residual_value[2] may indicate the element z of the position information.

[0431] FIG. 77 is a diagram showing another example of the syntax of the position information. The example of FIG. 77 is a modification of the example of FIG. 76.

[0432] As shown in FIG. 77, pred_mode may indicate the prediction mode for each of the three elements of the position information (x, y, z). That is, pred_mode[0] indicates the prediction mode of element x, pred_mode[1] indicates the prediction mode of element y, and pred_mode[2] indicates the prediction mode of element z. pred_mode[0], pred_mode[1], and pred_mode[2] may be added to the bit stream.

[0433] (Embodiment 5) FIG. 78 is a diagram showing an example of a prediction tree used in the three-dimensional data encoding method according to Embodiment 5.

[0434] In Embodiment 5, compared with Embodiment 4, in the method of generating the prediction tree, when generating the prediction tree, the depth of each node may be calculated.

[0435] For example, the root of the prediction tree may be set to depth = 0, the child node of the root may be set to depth = 1, and the child node of that child node may be set to depth = 2. At this time, the value that pred_mode can take may be changed according to the value of depth. That is, in setting the prediction mode, the three-dimensional data encoding device may set the prediction mode for predicting the three-dimensional point based on the depth of the hierarchical structure of each three-dimensional point. For example, pred_mode may be limited to a value less than or equal to the value of depth. That is, the set prediction mode value may be set to be less than or equal to the value of the depth of the hierarchical structure of each three-dimensional point.

[0436] Also, when pred_mode is binarized with a truncated unary code and arithmetic coded according to the number of prediction modes, it may be binarized with a truncated unary code with the number of prediction modes = min(depth, number of prediction modes M). This can reduce the bit length of the binary data of pred_mode when depth < M, and improve the coding efficiency.

[0437] In the method for generating a prediction tree, when adding a three-dimensional point A to the prediction tree, an example was shown in which the nearest neighbor point B was searched and the three-dimensional point A was added to the child node of the three-dimensional point B. Here, any method may be used for the method of searching for the nearest neighbor point. For example, the nearest neighbor point may be searched using the kd-tree method. This can efficiently search for the nearest neighbor point and improve the coding efficiency.

[0438] Also, the nearest neighbor point may be searched using the nearest neighbor method. This can search for the nearest neighbor point while suppressing the processing load, and balance the processing amount and the coding efficiency. Also, when searching for the nearest neighbor point using the nearest neighbor method, a search range may be set. This can reduce the processing amount.

[0439] In addition, the three-dimensional data encoding device may quantize and encode the prediction residual residual_value. For example, the three-dimensional data encoding device may add a quantization parameter QP to a header such as a slice, quantize the residual_value using Qstep calculated from QP, binarize the quantization value, and perform arithmetic coding. In this case, the three-dimensional data decoding device may apply inverse quantization to the quantization value of the residual_value using the same Qstep and add it to the predicted value to decode the position information. In that case, the decoded position information may be added to the prediction tree. As a result, even when quantization is applied, the three-dimensional data encoding device or the three-dimensional data decoding device can calculate the predicted value using the decoded position information, so the three-dimensional data encoding device can generate a bitstream that the three-dimensional data decoding device can correctly decode. Although an example of searching for the nearest neighbor point of the three-dimensional point and adding it to the prediction tree at the time of generating the prediction tree has been shown, it is not necessarily limited to this, and the prediction tree may be generated by any method or order. For example, when the input three-dimensional points are data acquired by lidar, the three-dimensional points may be added in the order scanned by lidar to generate a prediction tree. Thereby, the prediction accuracy can be improved and the encoding efficiency can be improved.

[0440] FIG. 79 is a diagram showing another example of the syntax of the position information. The residual_is_zero, residual_sign, residual_bitcount_minus1, and residual_bit[k] in the syntax of FIG. 79 will be described in order.

[0441] residual_is_zero is information indicating whether the residual_value is zero or not. For example, residual_is_zero = 1 indicates that residual_value is zero, and residual_is_zero = 0 indicates that residual_value is not zero. Note that when pred_mode = 0 (no prediction, predicted value is 0), since it is less likely for residual_value to be zero, it may not be necessary to encode residual_is_zero and add it to the bitstream. When pred_mode = 0, the three-dimensional data decoding device may estimate that residual_is_zero = 0 without decoding residual_is_zero from the bitstream.

[0442] residual_sign is positive / negative information (sign bit) indicating whether residual_value is positive or negative. For example, residual_sign = 1 indicates that residual_value is negative, and residual_sign = 0 indicates that residual_value is positive.

[0443] Note that when pred_mode = 0, since the predicted value is 0, residual_value must be positive or 0. Therefore, the three-dimensional data encoding device may not be necessary to encode residual_sign and add it to the bitstream. That is, when the three-dimensional data encoding device is set to a prediction mode in which the predicted value is calculated to be 0, it may generate a bitstream that does not include positive / negative information without encoding the positive / negative information indicating whether the prediction residual is positive or negative. When pred_mode = 0, the three-dimensional data decoding device may estimate that residual_sign = 0 without decoding residual_sign from the bitstream. That is, when the three-dimensional data decoding device obtains a bitstream that does not include positive / negative information indicating whether the prediction residual is positive or negative, it may handle the prediction residual as 0 or a positive number.

[0444] residual_bitcount_minus1 indicates the number obtained by subtracting 1 from the number of bits of residual_bit. That is, residual_bitcount is equal to the number obtained by adding 1 to residual_bitcount_minus1.

[0445] residual_bit[k] indicates the k-th bit information when the absolute value of residual_value is binarized with a fixed length according to the value of residual_bitcount.

[0446] Note that when condition A is defined as "when using the position information of any one of the points p0, p1, and p2 as the direct prediction value as in prediction mode 1, unique_point_per_leaf = 1 (no duplicated point)", since it is impossible for residual_is_zero[0] of element x, residual_is_zero[1] of element y, and residual_is_zero[2] of element z to all become 0 at the same time, it is not necessary to add residual_is_zero of any one element to the bit stream.

[0447] For example, when the three-dimensional data encoding device has condition A being true and residual_is_zero[0] and residual_is_zero[1] being 0, it is not necessary to add residual_is_zero[2] to the bit stream. Also, in this case, the three-dimensional data decoding device may estimate that residual_is_zero[2] not added to the bit stream is 1.

[0448] (Variant example) In this embodiment, an example was shown in which a prediction tree is generated using the position information (x, y, z) of three-dimensional points, and the position information is encoded and decoded. However, it is not necessarily limited to this. For example, predictive encoding using a prediction tree may be applied to the encoding of the attribute information (color, reflectance, etc.) of three-dimensional points. Also, the prediction tree generated in the encoding of the position information may be used in the encoding of the attribute information. As a result, it is not necessary to generate a prediction tree during the encoding of the attribute information, and the processing amount can be reduced.

[0449] FIG. 80 is a diagram showing an example of the configuration of a prediction tree commonly used for the encoding of position information and attribute information.

[0450] As shown in FIG. 80, each node of this prediction tree includes child_count, g_pred_mode, g_residual_value, a_pred_mode, and a_residual_value. g_pred_mode indicates the prediction mode of the position information. g_residual_value indicates the prediction residual of the position information. a_pred_mode indicates the prediction mode of the attribute information. a_residual_value indicates the prediction mode of the attribute information.

[0451] Here, child_count may be shared by the position information and the attribute information. Thereby, overhead can be suppressed and encoding efficiency can be improved.

[0452] Note that child_count may be added independently for the position information and the attribute information. Thereby, the three-dimensional data decoding device can decode the position information and the attribute information independently. For example, the three-dimensional data decoding device can also decode only the attribute information.

[0453] Note that the three-dimensional data encoding device may generate separate prediction trees for the position information and the attribute information. Thereby, the three-dimensional data encoding device can generate a prediction tree suitable for each of the position information and the attribute information, and can improve the encoding efficiency. In this case, the three-dimensional data encoding device may add information (such as child_count) necessary for the three-dimensional data decoding device to reconstruct each of the prediction trees for the position information and the attribute information to the bit stream. Note that the three-dimensional data encoding device may add identification information indicating whether to share the prediction tree for the position information and the attribute information to a header or the like. Thereby, it is possible to adaptively switch whether to share the prediction tree for the position information and the attribute information, and to control the balance between the encoding efficiency and the low processing quantization.

[0454] FIG. 81 is a flowchart showing an example of a three-dimensional data encoding method according to a modification of Embodiment 5.

[0455] The three-dimensional data encoding device generates a prediction tree using the position information of a plurality of three-dimensional points (S9761).

[0456] Next, the three-dimensional data encoding device encodes the node information included in each node of the prediction tree and the prediction residual of the position information. Specifically, the three-dimensional data encoding device calculates a predicted value for predicting the position information of each node, calculates a prediction residual that is the difference between the calculated predicted value and the position information of the node, and encodes the node information and the prediction residual of the position information.

[0457] Next, the three-dimensional data encoding device encodes the node information included in each node of the prediction tree and the prediction residual of the attribute information (S9763). Specifically, the three-dimensional data encoding device calculates a predicted value for predicting the attribute information of each node, calculates a prediction residual that is the difference between the calculated predicted value and the attribute information of the node, and encodes the node information and the prediction residual of the attribute information.

[0458] FIG. 82 is a flowchart showing an example of a three-dimensional data decoding method according to a modification of Embodiment 5.

[0459] The three-dimensional data decoding device decodes the node information to reconstruct the prediction tree (S9771).

[0460] Next, the three-dimensional data decoding device decodes the position information of the nodes (S9772). Specifically, the three-dimensional data decoding device calculates the predicted value of the position information of each node, and adds the calculated predicted value and the obtained prediction residual to decode the position information.

[0461] Next, the three-dimensional data decoding device decodes the attribute information of the nodes (S9773). Specifically, the three-dimensional data decoding device calculates the predicted value of the attribute information of each node, and adds the calculated predicted value and the obtained prediction residual to decode the position information.

[0462] Next, the three-dimensional data decoding device determines whether the decoding of all nodes is completed (S9774). If the decoding of all nodes is completed, the three-dimensional data decoding method ends. If the decoding of all nodes is not completed, steps S9771 to S9773 are executed for the unprocessed nodes.

[0463] FIG. 83 is a diagram showing an example of the syntax of the header of the attribute information. The NumNeighborPoint, NumPredMode, Thfix, QP, and unique_point_per_leaf in the syntax of FIG. 83 will be described in order.

[0464] NumNeighborPoint indicates the upper limit value of the number of surrounding points used for generating the predicted value of the attribute information of the three-dimensional points. When the number of surrounding points M is less than NumNeighborPoint (M < NumNeighborPoint), in the calculation process of the predicted value, the predicted value may be calculated using M surrounding points.

[0465] NumPredMode indicates the total number M of prediction modes used for predicting attribute information. Note that the maximum value MaxM of the possible values of the number of prediction modes may be specified by a standard or the like. The three-dimensional data encoding device may add the value of (MaxM - M) (0 < M <= MaxM) as NumPredMode to the header and binarize and encode (MaxM - 1) with a truncated unary code. Also, the number of prediction modes NumPredMode may not be added to the bitstream and may be specified by a profile or level in a standard or the like. Also, the number of prediction modes may be defined by NumNeighborPoint + NumPredMode.

[0466] Thfix is a threshold for determining whether to fix the prediction mode. The distance d0 between the points p1 and p0 used for prediction and the distance d1 between the points p2 and p1 are calculated, and if the absolute difference distdiff = |d0 - d1| is smaller than the threshold Thfix[i], the prediction mode is fixed to α. α is a prediction mode for calculating a predicted value using linear prediction and is "2" in the above embodiment. Note that Thfix may not be added to the bitstream and may be specified by a profile or level in a standard or the like.

[0467] QP indicates the quantization parameter used when quantizing attribute information. The three-dimensional data encoding device may calculate a quantization step from the quantization parameter and use the calculated quantization step to quantize the attribute information.

[0468] unique_point_per_leaf is information indicating whether the bitstream contains a duplicated point (points with the same position information). unique_point_per_leaf = 1 indicates that there are no duplicated points in the bitstream. unique_point_per_leaf = 0 indicates that there is one or more duplicated points in the bitstream.

[0469] In addition, in this embodiment, it is assumed that the determination of whether to fix the prediction mode is performed using the absolute value of the difference between the distance d0 and the distance d1. However, it is not necessarily limited to this, and any method of determination may be used. For example, this determination calculates the distance d0 between the point p1 and the point p0. If the distance d0 is greater than the threshold value, it is determined that the point p1 cannot be used for prediction, and the prediction mode value is fixed to "1" (predicted value p0). Otherwise, the prediction mode may be set. Thereby, while suppressing overhead, the coding efficiency can be improved.

[0470] The above NumNeighborPoint, NumPredMode, Thfix, or unique_point_per_leaf may be shared with the position information and not added to the attribute_header. Thereby, overhead can be reduced.

[0471] The above NumNeighborPoint, NumPredMode, Thfix, unique_point_per_leaf may be entropy-coded and added to the header. For example, each value may be binarized and calculated and coded. Also, each value may be coded with a fixed length in order to suppress the processing amount.

[0472] FIG. 84 is a diagram showing another example of the syntax of the attribute information. The NumOfPoint, child_count, pred_mode, dimension, residual_is_zero, residual_sign, residual_bitcount_minus1, and residual_bit[k] in the syntax of FIG. 84 will be described in order.

[0473] NumOfPoint indicates the total number of three-dimensional points included in the bit stream. NumOfPoint may be shared with the NumOfPoint of the position information.

[0474] child_count indicates the number of child nodes that the i-th three-dimensional point (node[i]) has. Note that child_count may be shared with the child_count of the position information. When child_count is shared with the child_count of the position information, child_count may not be added to the attribute_data. This can reduce the overhead.

[0475] pred_mode indicates the prediction mode for encoding or decoding the position information of the i-th three-dimensional point. pred_mode takes values from 0 to M - 1 (M is the total number of prediction modes). When pred_mode is not in the bit stream (when the condition distdiff >= Thfix[i] && NumPredMode > 1 is not satisfied), pred_mode may be estimated as a fixed value α. α is a prediction mode for calculating a predicted value using linear prediction, and in the above embodiment, it is "2". Note that α is not limited to "2", and any value from 0 to M - 1 may be set as the estimated value. Also, the estimated value when pred_mode is not in the bit stream may be added separately to a header or the like. Further, pred_mode may be binarized with a truncated unary code using the number of prediction modes to which predicted values are assigned and then arithmetic-coded.

[0476] dimension is information indicating the dimension of the attribute information. dimension may be added to a header such as SPS. For example, when the attribute information is color, dimension may be set to "3", and when it is reflectance, dimension may be set to "1".

[0477] residual_is_zero is information indicating whether the residual_value is 0. For example, residual_is_zero = 1 indicates that residual_value is 0, and residual_is_zero = 0 indicates that residual_value is not 0. Note that when pred_mode = 0 (no prediction, predicted value 0), since the probability of residual_value becoming 0 is low, it may not be necessary to encode residual_is_zero and add it to the bitstream. When pred_mode = 0, the three-dimensional data decoding device may estimate that residual_is_zero = 0 without decoding residual_is_zero from the bitstream.

[0478] residual_sign is positive / negative information (sign bit) indicating whether residual_value is positive or negative. For example, residual_sign = 1 indicates that residual_value is negative, and residual_sign = 0 indicates that residual_value is positive.

[0479] Note that when pred_mode = 0 (no prediction, predicted value 0), since residual_value is positive, the three-dimensional data encoding device may not be necessary to encode residual_sign and add it to the bitstream. That is, when the prediction residual is positive, the three-dimensional data encoding device may generate a bitstream without positive / negative information without encoding the positive / negative information indicating whether the prediction residual is positive or negative, and when the prediction residual is negative, it may generate a bitstream including positive / negative information. When pred_mode = 0, the three-dimensional data decoding device may estimate that residual_sign = 0 without decoding residual_sign from the bitstream. That is, when the three-dimensional data decoding device obtains a bitstream without positive / negative information indicating whether the prediction residual is positive or negative, it may handle the prediction residual as a positive number, and when it obtains a bitstream including positive / negative information, it may handle the prediction residual as a negative number.

[0480] residual_bitcount_minus1 indicates the number obtained by subtracting 1 from the number of bits of residual_bit. That is, residual_bitcount is equal to the number obtained by adding 1 to residual_bitcount_minus1.

[0481] residual_bit[k] indicates the k-th bit information when the absolute value of residual_value is binarized with a fixed length according to the value of residual_bitcount.

[0482] Note that when condition A is defined as "when using the attribute information of any one of point p0, point p1, and point p2 as the direct prediction value as in prediction mode 1, unique_point_per_leaf = 1 (no duplicated point)", since it is impossible for residual_is_zero[0] of element x, residual_is_zero[1] of element y, and residual_is_zero[2] of element z to all become 0 at the same time, it is not necessary to add residual_is_zero of any one element to the bit stream.

[0483] For example, when the three-dimensional data encoding device has condition A being true and residual_is_zero[0] and residual_is_zero[1] being 0, it is not necessary to add residual_is_zero[2] to the bit stream. Also, in this case, the three-dimensional data decoding device may estimate that residual_is_zero[2] not added to the bit stream is 1.

[0484] FIG. 85 is a diagram showing an example of the syntax of position information and attribute information.

[0485] As shown in FIG. 85, encoded information of position information and attribute information may be stored in one data unit. Here, g_* indicates encoded information related to geometry, and a_* indicates encoded information related to attribute information. Thereby, the position information and the attribute information can be decoded simultaneously.

[0486] As described above, the three-dimensional data encoding apparatus according to one aspect of the present embodiment performs the process shown in FIG. 86. The three-dimensional data encoding apparatus executes a three-dimensional data encoding method for encoding a plurality of three-dimensional points having a hierarchical structure. The three-dimensional data encoding apparatus sets one prediction mode out of two or more prediction modes for calculating a predicted value of the first position information of the first three-dimensional point using the second position information of one or more second three-dimensional points around the first three-dimensional point (S9781). Next, the three-dimensional data encoding apparatus calculates a predicted value of the set prediction mode (S9782). Next, the three-dimensional data encoding apparatus calculates a prediction residual, which is a difference between the first position information and the calculated predicted value (S9783). Next, the three-dimensional data encoding apparatus generates a first bit stream including the set prediction mode and the prediction residual (S9784). In the setting (S9781), the prediction mode is set based on the depth of the hierarchical structure of the first three-dimensional point.

[0487] According to this, since the position information can be encoded using the predicted value of one prediction mode set based on the depth of the hierarchical structure among two or more prediction modes, the encoding efficiency of the position information can be improved.

[0488] For example, in the setting (S9784), the three-dimensional data encoding apparatus sets a prediction mode value equal to or less than the value of the depth of the hierarchical structure of the first three-dimensional point. The prediction mode value indicates the prediction mode.

[0489] For example, the first bit stream further includes a prediction mode number indicating the number of the two or more prediction modes.

[0490] For example, in the generation (S9784), the prediction mode value indicating the set prediction mode is encoded using the number of prediction modes. The first bitstream includes the encoded prediction mode value as the set prediction mode.

[0491] For example, in the generation (S9784), the prediction mode value is encoded with a truncated unary code having the maximum value of the number of prediction modes. Therefore, the amount of code for the prediction mode value can be reduced.

[0492] For example, each of the first position information and the second position information includes three elements. In the setting (S9781), the three-dimensional data encoding device sets a common prediction mode for the three elements as the one prediction mode for calculating the predicted value of each element included in the first position information. Therefore, the amount of code for the prediction mode value can be reduced.

[0493] For example, each of the first position information and the second position information includes three elements. In the setting, the three-dimensional data encoding device sets an independent prediction mode for each of the three elements as the one prediction mode for calculating the predicted value of each element included in the first position information. Therefore, the three-dimensional data decoding device can decode each element independently.

[0494] For example, each of the first position information and the second position information includes three elements. In the setting, the three-dimensional data encoding device sets a common prediction mode for two of the three elements as the one prediction mode for calculating the predicted value of each element included in the first position information, and sets an independent prediction mode for the remaining one element with respect to the two elements. Therefore, the amount of code for the prediction mode values for the two elements can be reduced. Also, the three-dimensional data decoding device can decode the remaining one element independently.

[0495] For example, in the generation by the three-dimensional data encoding device, when the number of prediction modes is 1, a second bit stream that does not include the prediction mode value is generated without encoding the prediction mode value indicating the prediction mode. Therefore, the amount of code of the bit stream can be reduced.

[0496] For example, in the generation by the three-dimensional data encoding device, when a prediction mode in which the predicted value calculated in the calculation is 0 is set, a third bit stream that does not include the positive / negative information indicating whether the prediction residual is positive or negative is generated without encoding the positive / negative information. Therefore, the amount of code of the bit stream can be reduced.

[0497] For example, a three-dimensional data encoding device includes a processor and a memory, and the processor performs the above processing using the memory.

[0498] In addition, the three-dimensional data decoding device according to one aspect of the present embodiment performs the processing shown in FIG. 87. The three-dimensional data decoding device executes a three-dimensional data decoding method for decoding a plurality of three-dimensional points having a hierarchical structure. The three-dimensional data decoding device acquires a first bit stream including the prediction mode of the first three-dimensional point among the plurality of encoded three-dimensional points and the encoded prediction residual (S9791). Next, the three-dimensional data decoding device decodes the prediction mode value indicating the encoded prediction mode and the encoded prediction residual (S9792). Next, the three-dimensional data decoding device calculates a predicted value of the prediction mode indicated by the prediction mode value obtained by decoding (S9793). Next, the three-dimensional data decoding device calculates first position information of the first three-dimensional point by adding the predicted value and the prediction residual obtained by decoding (S9794). The encoded prediction mode included in the first bit stream is a prediction mode set based on the depth of the hierarchical structure of the first three-dimensional point.

[0499] According to this, among two or more prediction modes, the position information encoded using the prediction value of one prediction mode set based on the depth of the hierarchical structure can be appropriately decoded.

[0500] For example, the prediction mode value indicating the encoded prediction mode included in the first bit stream is less than or equal to the value of the depth of the hierarchical structure of the first three-dimensional point.

[0501] For example, the first bit stream includes a prediction mode number indicating the number of the two or more prediction modes.

[0502] For example, in the decoding (S9792), the three-dimensional data decoding device decodes the encoded prediction mode value using a truncated unary code with the prediction mode number as the maximum value.

[0503] For example, each of the first position information and the second position information of one or more second three-dimensional points around the first three-dimensional point includes three elements. The prediction mode is used to calculate the prediction value of each element of the three elements included in the first position information, and is commonly set for the three elements.

[0504] For example, each of the first position information and the second position information of one or more second three-dimensional points around the first three-dimensional point includes three elements. The prediction mode is used to calculate the prediction value of each element of the three elements included in the first position information, and is independently set for each of the three elements.

[0505] For example, each of the first position information and the second position information of one or more second three-dimensional points around the first three-dimensional point includes three elements. The prediction mode is used to calculate the prediction value of each element of the three elements included in the first position information, and is commonly set for two of the three elements, and is set independently of the two elements for the remaining one element.

[0506] For example, when the three-dimensional data decoding device acquires, in the acquisition (S9791), a second bit stream that does not include the prediction mode value, in the calculation of the prediction value, a prediction value of a specific prediction mode is calculated.

[0507] For example, when the three-dimensional data decoding device acquires, in the acquisition (S9791), a third bit stream that does not include positive / negative information indicating whether the prediction residual is positive or negative, in the calculation of the first position information (S9794), the prediction residual is treated as 0 or a positive number.

[0508] For example, a three-dimensional data decoding device includes a processor and a memory, and the processor performs the above processing using the memory.

[0509] (Embodiment 6) FIG. 88 is a diagram showing an example of point cloud data obtained by LiDAR. The point cloud (point cloud) obtained by scanning with LiDAR is usually sparse, and has specific characteristics according to the specifications of LiDAR such as 16 layers, 32 layers, 64 layers, 128 layers, etc. The three-dimensional data encoding device compresses the position (geometry) information of the point cloud using point cloud compression using an octree. The three-dimensional data encoding device uses point cloud compression using prediction for sparse point clouds. In addition, the three-dimensional data encoding device uses the characteristics of sparse point clouds (such as information obtained from LiDAR) to construct a prediction tree.

[0510] FIG. 89 is a block diagram showing the configuration of the three-dimensional data encoding device according to the present embodiment. The three-dimensional data encoding device includes a rearrangement unit 10001, a prediction tree generation unit 10002, a prediction conversion unit 10003, a quantization unit 10004, and an entropy encoding unit 10005.

[0511] The rearrangement unit 10001 rearranges a plurality of position information of the input point cloud based on a predetermined standard. The prediction tree generation unit 10002 generates a prediction tree for the plurality of position information. Here, the prediction tree indicates the reference relationship in the prediction of the plurality of position information. Specifically, the prediction tree generation unit 10002 generates a prediction tree using a setting file. The setting file is information on the hardware specifications such as LiDAR that generates the point cloud, and indicates, for example, horizontal and vertical angle information (angle resolution), the number of horizontal scans, and the total number of nodes.

[0512] The prediction conversion unit 10003 generates a predicted value of a target point, which is a three-dimensional point to be processed, using the prediction tree, and calculates a prediction residual, which is the difference between the position information of the target point and the predicted value. For example, the prediction conversion unit 10003 generates a predicted value using a parent node, a grandparent node, a great-grandparent node, or a combination thereof indicated by the prediction tree.

[0513] The quantization unit 10004 quantizes the prediction residual generated by the prediction conversion unit 10003. The entropy encoding unit 10005 generates an encoded stream by entropy encoding (arithmetic encoding) the quantized prediction residual.

[0514] The three-dimensional data encoding device constructs a prediction tree having a good correlation with other adjacent point clouds. Thereby, the position information of the sparse point cloud obtained by LiDAR or the like can be encoded with high efficiency and a high compression gain.

[0515] The sparse point cloud is generated via a rotating laser such as LiDAR. The three-dimensional data encoding device generates a prediction tree using the information indicated by the setting file or the product specifications, such as the angle resolution of the vertical and horizontal axes and the range (for example, the number of horizontal layers). Also, the three-dimensional data encoding device may use the field of view, the range, and the relative distance between points. The three-dimensional data encoding device generates a prediction tree for the point cloud using these information.

[0516] Note that the three-dimensional data encoding device may perform conversion processing on the prediction residual and perform quantization on the converted prediction residual. This conversion is, for example, DCT (Discrete Cosine Transform) or Haar transform. Note that other methods such as a method of referring to a conversion table may be used for this conversion. Also, for the generation of the predicted value, a selected prediction method may be used from a plurality of prediction methods (prediction modes). For example, the plurality of prediction methods differ in the nodes to be referred to, the number of nodes, or the calculation method.

[0517] Also, the setting file is generated by the user, for example. Thereby, a bitstream according to the user's specification is generated.

[0518] FIG. 90 is a block diagram showing another configuration example of the three-dimensional data encoding device. The three-dimensional data encoding device shown in FIG. 90 includes a sensor detection unit 10006 in addition to the configuration of the three-dimensional data encoding device shown in FIG. 89.

[0519] The sensor detection unit 10006 determines the characteristics of the LiDAR hardware by comparing adjacent point cloud data and acquiring information such as the number of layers, resolution, and total number of point clouds per frame.

[0520] When it is known that the input point cloud data is a three-dimensional point cloud obtained by a specific LiDAR, or when the user executes a test and confirms that it is a three-dimensional point cloud obtained by a specific LiDAR, the sensor detection unit 10006 can start preprocessing after the sorting process.

[0521] Here, the sorting process is a process for arranging nodes based on a specific arrangement format. For example, based on the z-axis or the horizontal layer value, for each layer of the rotation scan layer of the LiDAR, the x coordinate and the y coordinate are arranged in ascending order from the minimum to the maximum. As another method, the absolute value of the distance from a certain point to each point may be calculated, and the point cloud may be sorted using the obtained absolute value of the distance.

[0522] In addition, rearrangement processing may be performed using the known hardware operation characteristics of LiDAR. Further, when using the characteristics of LiDAR, by applying normalization processing, the accuracy of determining horizontal or vertical angle information can be improved. Further, the extracted information is used for the generation of a prediction tree.

[0523] When the input point cloud is a point cloud obtained by scanning with a plurality of layers of LiDAR, the three-dimensional data encoding device may individually scan the point cloud of each LiDAR to form a prediction tree. For example, when the input point cloud includes point clouds generated by LiDAR1 and LiDAR2, a prediction tree 1 for the point cloud 1 obtained by LiDAR1 and a prediction tree 2 for the point cloud 2 obtained by LiDAR2 are generated. Further, the three-dimensional data encoding device may individually encode the point cloud 1 and the point 2. Similarly, the three-dimensional data decoding device may individually decode the point cloud 1 and the point cloud 2. Thereby, since the three-dimensional data encoding device can generate an appropriate prediction tree for each point cloud, the encoding efficiency can be improved.

[0524] In addition, the three-dimensional data encoding device may store the encoded data encoded using the prediction tree 1 and the encoded data encoded using the prediction tree 2 in one bit stream. At this time, the three-dimensional data encoding device may add at least one of the start position of the encoded data regarding the prediction tree 1 in the bit stream and the start position of the encoded data regarding the prediction tree 2 to the header of the bit stream or the like. Thereby, since the three-dimensional data decoding device can decode the encoded data of the prediction tree 1 and the encoded data of the prediction tree 2 in parallel, the decoding processing time can be shortened. Further, the three-dimensional data decoding device can selectively decode the encoded data regarding the prediction tree 2, for example, by reading the start position of the encoded data of each prediction tree from the header.

[0525] Hereinafter, a specific example of a prediction tree will be described. A sparse point cloud can be projected onto a specific view (plane). FIG. 91 is a top view of a LiDAR point cloud projected onto a plane. The three-dimensional data encoding device generates a prediction tree using the characteristics of the projected point cloud. Hereinafter, each three-dimensional point included in the point cloud is also referred to as a node.

[0526] FIG. 92 is a diagram showing the parent-child relationship of a plurality of points. The three-dimensional data encoding device generates a prediction tree using the parent-child relationship. Here, the child node (p0) is the three-dimensional point to be processed. Well, the parent node (p1) and the grandparent node (p2) are already encoded three-dimensional points. Also, the relationship between this child node, parent node, and grandparent node is shown in the prediction tree.

[0527] The three-dimensional data encoding device calculates a predicted value from the position information of the parent node or the grandparent node. For example, the predicted value may be calculated using the difference (p1 - p2) between the position information of the parent node and the position information of the grandparent node. The three-dimensional data encoding device calculates a prediction residual, which is the difference between the position information of the child node and the predicted value. The smaller the prediction residual, the more the encoding efficiency can be improved.

[0528] Next, a method for generating a prediction tree using angle information will be described. For example, a prediction model can be generated on a plane using the characteristics of the rotating laser of LiDAR. FIG. 93 is a top view of the point cloud. FIG. 94 is a diagram showing the relationship of the point cloud. As shown in FIG. 94, the coordinates of the grandparent node are (x, y), the coordinates of the parent node are (x', y'), and the coordinates of the child node are (x'', y''). Also, the coordinates of the center position corresponding to the sensor position (LiDAR position) are (Cx, Cy). Note that these coordinates are two-dimensional coordinates in the xy plane. Also, for example, the angle information θ indicating the angular resolution in the rotational scanning of LiDAR is obtained from, for example, a setting file. Also, the vector from the center to the grandparent node is r, the vector from the center to the parent node is r', and the vector from the center to the child node is r''.

[0529] The relationship between the direction of r and the direction of r' is defined by θ. The relationship between the direction of r' and the direction of r'' is defined by θ. The magnitude of r' (the distance between the center and the parent node) is expressed as the sum of the magnitude of r (the distance between the center and the grandparent node) and the residual. The magnitude of r'' (the distance between the center and the child node) is expressed as the sum of the magnitude of r' (the distance between the center and the parent node) and the residual. Here, the residual is the error between the predicted value and the actual value, and the residual becomes smaller if the prediction accuracy is high.

[0530] Here, the following relationships (Equations U1) to (U3) hold.

[0531] x 2 +y 2 =r 2 ···(Equation U1) x’ = Cx + r×sinθ ···(Equation U2) y’ = Cy + r×cosθ ···(Equation U3)

[0532] Also, when Cx and Cy are 0, the following (Equations U4) and (U5) hold.

[0533] x’’ = x’×cosθ - y’×sinθ ···(Equation U4) y’’ = x’×sinθ + y’×cosθ ···(Equation U5)

[0534] In this way, the three-dimensional data encoding device can generate the predicted value of the child node from the position information of the parent node and the grandparent node by also using the cosine function, sine function, and the angle information θ. Note that although the magnitudes of r, r’, and r’’ may be different, each node exists in the direction of the corresponding vector.

[0535] Also, the residual is a positive or negative value, and this residual may be encoded. Also, the angle information θ is stored in the SPS (Sequence Parameter Set), GPS (Geometry Parameter Set), or slice header, etc.

[0536] Also, here, an example of calculating a predicted value using (Equation U1) to (Equation U5) has been shown, but the method of calculating the predicted value is not necessarily limited to this. Any equation may be used as long as it is a method of generating a predicted value using the angle information θ of LiDAR. For example, as in (Equation U4) and (Equation U5), a predicted value may be generated using θ as the rotation angle. That is, the three-dimensional data encoding device may calculate the predicted value of the child node by rotating the position information of the previous point (for example, the parent node) by an angle θ with a radius r' around the sensor position.

[0537] Note that the angle information θ is, for example, a value based on the specifications of LiDAR, and the angle between two actual consecutive points (for example, the parent node and the child node) includes variations and does not necessarily match θ.

[0538] Also, the three-dimensional data encoding device may select a prediction method to be used from a plurality of prediction methods (prediction modes) including the prediction method using the angle information θ of LiDAR described here. Here, the plurality of prediction methods may include any known method. For example, the plurality of prediction methods include a mode in which the position information of any one of the parent node, the grandparent node, and the great-grandparent node is used as the predicted value as it is, and a method of calculating the predicted value using the position information of at least two of the parent node, the grandparent node, and the great-grandparent node. Further, the three-dimensional data encoding device may select a prediction method that minimizes the generated code amount from the plurality of prediction methods. Thereby, the encoding efficiency can be improved.

[0539] Also, in the above, the method of generating a predicted value using the angle information θ has been described, but the three-dimensional data encoding device may use the same method for the horizontal angle information α from the base layer to the upper layer. That is, in the above, an example of using the angle information in the xy plane has been described, but the same method may be used for the z direction. FIG. 95 is a top view of the point cloud of LiDAR projected onto a plane in this case. FIG. 96 is a diagram showing the relationship between a plurality of points.

[0540] As shown in FIG. 96, the coordinates of the grandparent node are (x, y), the coordinates of the parent node are (x', y'), and the coordinates of the child node are (x'', y''). Also, the coordinates of the center position corresponding to the sensor position (LiDAR position) are (Cx, Cy). Also, for example, the angle information α is obtained from, for example, a setting file. Also, the vector from the center to the grandparent node is r, the vector from the center to the parent node is r', and the vector from the center to the child node is r''.

[0541] The relationship between the direction of r and the direction of r' is defined by α. The relationship between the direction of r' and the direction of r'' is defined by α. The magnitude of r' (the distance between the center and the parent node) is represented by the sum of the magnitude of r (the distance between the center and the grandparent node) and the residual. The magnitude of r'' (the distance between the center and the child node) is represented by the sum of the magnitude of r' (the distance between the center and the parent node) and the residual.

[0542] Here, the following relationships (Equations U6) to (Equation U8) hold.

[0543] x 2 +y 2 =r 2 ···(Equation U6) x' = Cx + r × sin α ···(Equation U7) y' = Cy + r × cos α ···(Equation U8)

[0544] Also, when Cx and Cy are 0, the following (Equation U9) and (Equation U10) hold.

[0545] x'' = x' × cos α - y' × sin α ···(Equation U9) y'' = x' × sin α + y' × cos α ···(Equation U10)

[0546] In this way, the three-dimensional data encoding device can generate a predicted value of the child node from the position information of the parent node and the grandparent node by using the cosine function and the sine function in combination with the angle information α. Note that a combination of prediction in the same plane and prediction in the horizontal plane may be used. Thereby, the prediction accuracy can be improved.

[0547] Also, the residual is a positive or negative value, and this residual may be encoded. Also, the angle information α is stored in an SPS (Sequence Parameter Set), a GPS (Position Information Parameter Set), a slice header, or the like.

[0548] Also, here, an example of calculating a predicted value using (Equation U6) to (Equation U10) has been shown, but the method of calculating the predicted value is not necessarily limited to this. Any formula may be used as long as it is a method of generating a predicted value using the horizontal angle information α of LiDAR. For example, as in (Equation U9) and (Equation U10), the predicted value may be generated using α as the rotation angle. That is, that is, the three-dimensional data encoding device may calculate the predicted value of the child node by rotating the position information of the previous point (for example, the parent node) around the sensor position by an angle α with a radius r'.

[0549] Also, the three-dimensional data encoding device may select a prediction method to be used from a plurality of prediction methods (prediction modes) including the prediction method using the horizontal angle information α of LiDAR described here. Here, the plurality of prediction methods may include any known method. For example, the plurality of prediction methods include a mode in which the position information of any one of the parent node, the grandparent node, and the great-grandparent node is used as the predicted value as it is, and a method of calculating the predicted value using the position information of at least two of the parent node, the grandparent node, and the great-grandparent node. Also, the three-dimensional data encoding device may select a prediction method that minimizes the generated code amount from the plurality of prediction methods. Thereby, the encoding efficiency can be improved.

[0550] Note that here, an example in which the point cloud data is the point cloud data obtained by LiDAR has been described, but the point cloud data may be the point cloud data obtained by a rotating laser other than LiDAR. Also, the point cloud data may be the point cloud data obtained by a sensor other than the rotating laser (for example, a TOF (Time Of Flight) sensor). In this case, the angle information θ is, for example, the scanning angle of the sensor.

[0551] Next, a prediction method using virtual nodes will be described. In sparse three-dimensional point clouds or in certain cases, the position information of nodes may be significantly separated from adjacent points.

[0552] A virtual node is a virtual node inserted to improve the correlation in a prediction tree. FIG. 97 is a diagram showing an example of a virtual node in point cloud data. FIG. 98 is a diagram showing an example of a virtual node and a prediction residual. As shown in FIG. 98, when virtual nodes are not used, the prediction residual is 142. That is, the prediction residual becomes a large value.

[0553] In this example, two virtual nodes are added. Also, by adding two split indicators, two virtual nodes are generated. Also, a virtual node is an empty node and has no prediction residual. By adding virtual nodes, the prediction residual of the target node p0 becomes 2, and the prediction residual can be reduced.

[0554] Also, the three-dimensional data encoding device determines whether to add a virtual node using a threshold VRth. Specifically, the three-dimensional data encoding device adds a virtual node when the prediction residual is greater than the threshold VRth. The threshold VRth may be set based on a configuration file or may be calculated during the evaluation process of LiDAR characteristics.

[0555] Also, the three-dimensional data encoding device may add information indicating this threshold VRth to the bitstream. Note that the bitstream may not include information indicating the threshold VRth. Also, the three-dimensional data decoding device does not necessarily require information indicating the threshold VRth for the decoding process of the position information, but this information can be used for reference or post-processing purposes.

[0556] An example of this determination process is shown below. First, the three-dimensional data encoding device determines whether p0 - p1 > VRth holds. That is, the three-dimensional data encoding device determines whether the difference between the position information of the target node p0 and the position information of the parent node p1 is greater than the threshold VRth. When p0 - p1 > VRth holds, the three-dimensional data encoding device calculates p1 - p2 = a. Then, the three-dimensional data encoding device calculates n such that p0 - p1 = n × a. Here, n is the number of virtual nodes to be inserted.

[0557] For example, when VRth = 60 and p0 - p1 = 142, p0 - p1 = 142 is greater than VRth = 60. Also, the difference value a = p1 - p2 = 70. Therefore, the number of virtual nodes to be inserted n = 142 / 70 = 2. That is, n is the quotient obtained by dividing the original prediction residual (p0 - p1) by the difference value a. Note that n may be determined such that the absolute value of the difference between (p0 - p1) and n × a is minimized. At this time, the prediction residual of node p0 is p0 - p1 - 2a = 2.

[0558] Note that since there is no prediction residual for virtual nodes, the prediction residual of virtual nodes is not included in the bit stream. Also, the three-dimensional data encoding device may add information indicating the number of inserted virtual nodes to the bit stream for each three-dimensional point. For example, in the example shown in FIG. 98, the three-dimensional data encoding device adds to the bit stream information indicating that the virtual nodes of p1 to p5 are each 0 and the virtual nodes of p0 are 2. Thereby, the three-dimensional data decoding device can appropriately generate a predicted value by correcting the predicted value calculated from the prediction mode according to the number of virtual nodes added for each three-dimensional point. Thereby, the three-dimensional data decoding device can appropriately decode the bit stream encoded using virtual nodes.

[0559] In addition, in this embodiment, an example was shown in which the predicted value of p0 is corrected using the number of virtual nodes and the difference value a between the position information of p2 and the position information of p1. However, it is not necessarily limited to this. The three-dimensional data encoding device may use any method as long as it is a method for correcting the predicted value according to the number of virtual nodes. For example, the three-dimensional data encoding device may correct the predicted value of p0 using a value obtained by multiplying the difference value between the position information of p3 and the position information of p2 by the number of virtual nodes. Further, when generating the predicted value of p0, the three-dimensional data encoding device determines whether there is a virtual node between p0 and p1. If there is a virtual node, the value of the predicted value calculated using at least one or more of p1 to p5 may be corrected according to the number of virtual nodes. Further, when there is no virtual node, the three-dimensional data encoding device may use the predicted value calculated using at least one or more of p1 to p5 as the predicted value of p0.

[0560] For example, if the difference between the position information of p0 and the position information of p1 is greater than the threshold value VRth, the three-dimensional data encoding device determines that there is a virtual node between p0 and p1, and for example, may calculate the number of virtual nodes by the method described above. For example, when there are two virtual nodes between p0 and p1, in the prediction mode where p1 is the predicted value of p0, the predicted value p1 is corrected by adding n times the difference value a between p2 and p1 (n is the number of virtual nodes between p0 and p1) to the predicted value p1. Specifically, the predicted value becomes p1 + a×2. As a result, the prediction residual of p0 becomes p0 - (p1 + a×2). Therefore, as shown in FIG. 98, when p2 = 10, p1 = 80, and p0 = 222, when not using virtual nodes, the prediction residual of p0 is p0 - p1 = 142, and when correcting the predicted value using virtual nodes, the prediction residual is p0 - (p1 + a×2) = 222 - (80 + 70×2) = 2. Therefore, the prediction residual becomes smaller and the encoding efficiency is improved.

[0561] Here, an example of generating a predicted value using the difference value a of the position information has been described, but a predicted value may also be generated using the angle information θ. For example, a first virtual node may be generated by rotating by an angle θ with a radius r' (r' is the distance between the center and p1) around the sensor position, and a second virtual node may be generated by rotating by an angle 2θ with a radius r' around the sensor position.

[0562] Also, the three-dimensional data encoding device may select a difference value appropriate for generating a predicted value from a plurality of difference values a x , a y , a z . FIG. 99 is a diagram showing a configuration example of the difference value and the bit stream in this case.

[0563] For example, the three-dimensional data encoding device acquires the immediately preceding difference value ax and the threshold value VRthx from a setting file, SPS, or GPS, etc.

[0564] For example, another difference value a y is the difference value of the position information of the parent node and the grandparent node. Also, the threshold value VRthy is set according to a y . For example, VRthy = 0.5 × a y . Also, a different formula or function can be used to calculate the difference value a z and the threshold value VRthz.

[0565] In this case, the three-dimensional data encoding device uses a plurality of difference values and the threshold value of the virtual node. Also, for different division indicators included in the bit stream, identifiers such as 001 or 002 are used, for example. For example, the value of the prediction residual is a value between 0 and 255.

[0566] Also, the above division indicator may be added to the SPS, GPS, or slice header in order to distinguish the difference value used.

[0567] Also, RDO (Rate Distortion Optimization) may be used to check whether the best predicted value is selected, or to compare the case of using virtual nodes with the case of directly encoding the prediction residuals from p1 to p0 without using virtual nodes.

[0568] Next, direct encoding will be described. In a sparse three-dimensional point cloud or in a specific case, the position information of a node may be significantly separated from its adjacent points.

[0569] When an appropriate predicted value cannot be generated, the three-dimensional data encoding device uses direct encoding to directly encode the target node without using the predicted value. Also, the three-dimensional data encoding device starts generating a new prediction tree starting from the target node.

[0570] FIG. 100 is a diagram showing an example of point cloud data in this case. FIG. 101 is a diagram showing the reference relationship of each node. For example, the three-dimensional data encoding device determines whether to use direct encoding using a threshold DCth. Specifically, when the prediction residual is greater than the threshold DCth, the three-dimensional data encoding device determines to use direct encoding.

[0571] For example, when p0 - p1 > DCth is satisfied, the three-dimensional data encoding device uses direct encoding. Also, the three-dimensional data encoding device may evaluate the effectiveness of direct encoding using RDO, similar to the case of using virtual nodes.

[0572] Next, a method for generating a prediction tree using line tracing will be described. One common method for constructing a prediction tree is to use the position information of the nearest neighbor point as the predicted value. In normal situations, in most cases, this prediction method generates a good correlation between the positions of the three-dimensional point cloud.

[0573] However, in some cases, undesirable prediction trees may be generated. FIG. 102 is a diagram showing an example of point cloud data and prediction relationships. As shown in FIG. 102, when using the position information of adjacent points, inappropriate predictions may be made. For example, when the prediction tree in the horizontal LiDAR scan intersects with another prediction tree in the vertical direction, points with few correlation relationships may be used for prediction. This reduces the coding efficiency.

[0574] Here, the horizontal prediction tree and the vertical prediction tree are approximately straight lines. Therefore, by selecting prediction values in the horizontal or vertical direction, it is possible to suppress the generation of zigzag prediction trees. This can improve the coding efficiency.

[0575] Next, the operation of the three-dimensional data encoding device will be described. FIG. 103 is a flowchart of the three-dimensional data encoding process by the three-dimensional data encoding device.

[0576] The three-dimensional data encoding device starts encoding the prediction tree position information using the configuration file (S10001). Next, the three-dimensional data encoding device determines whether it can obtain hardware information from the configuration file (S10002). The hardware information includes angle information (horizontal / vertical scanning angle information), resolution, and range, etc.

[0577] If the hardware information cannot be obtained from the configuration file (No in S10002), the three-dimensional data encoding device determines whether the characteristic evaluation of the input point cloud data has been successful (S10003).

[0578] If the characteristic evaluation of the input point cloud data has not been successful (No in S10003), that is, when angle information etc. cannot be obtained from either the configuration file or the characteristic evaluation, the three-dimensional data encoding device generates a normal (angle information - unused) prediction tree (S10004). Also, the three-dimensional data encoding device encodes the point cloud data using the generated prediction tree (S10005).

[0579] On one hand, when the angle information can be obtained through a configuration file or characteristic evaluation (Yes in S10002 or Yes in S10003), the three-dimensional data encoding device performs angle prediction using the angle information (such as in FIGS. 93 and 94, etc.) (S10006). In this case, the three-dimensional data encoding device stores information related to angle prediction, such as the angle information, in the SPS or GPS.

[0580] Also, when the three-dimensional data encoding device can use virtual nodes (Yes in S10007), it performs virtual node prediction (such as in FIGS. 97 and 98, etc.) (S10008). Also, when the three-dimensional data encoding device can use direct encoding (Yes in S10009), it performs direct encoding (such as in FIGS. 100 and 101, etc.) (S10010).

[0581] Also, the three-dimensional data encoding device may determine which prediction mode to use by using RDO (Rate Distortion Optimization). For example, the three-dimensional data encoding device encodes the position information or attribute information of the three-dimensional points in all prediction modes and selects the prediction mode that results in the minimum amount of generated bits.

[0582] Also, data corresponding to the used prediction mode is stored in the bitstream. Specifically, the bitstream includes information indicating the prediction mode, information about virtual nodes, a segmentation indicator, etc.

[0583] Next, the operation of the three-dimensional data decoding device will be described. FIG. 104 is a flowchart of the three-dimensional data decoding process by the three-dimensional data decoding device.

[0584] First, the three-dimensional data decoding device decodes the bitstream generated by the process shown in FIG. 103, for example, and obtains various information (S10021). Specifically, the three-dimensional data decoding device obtains angle information, etc. from the SPS or GPS.

[0585] Next, the three-dimensional data decoding device starts decoding the prediction tree position information (S10022). The three-dimensional data decoding device generates a prediction tree (S10023). Also, when the hardware information and the angle information can be used (Yes in S10024), the three-dimensional data decoding device performs angle prediction using the angle information (S10025). Also, when the virtual node can be used (Yes in S10026), the three-dimensional data decoding device performs virtual node prediction (S10027). Also, when direct encoding (direct decoding) can be used (Yes in S10028), the three-dimensional data decoding device performs direct encoding (S10029). For example, the three-dimensional data decoding device determines whether each prediction mode can be used by using the information included in the bit stream.

[0586] Next, the syntax structure of the bit stream will be described. FIG. 105 is a diagram showing a syntax example of Header1 included in the bit stream. Here, Header1 is, for example, SPS.

[0587] Header1 includes an angle information available flag (hardware_specification_available_flag), horizontal angle information (horizontal_angular, θ), and vertical angle information (vertical_angular, α) based on the hardware specifications of the LiDAR. The horizontal angle information and the vertical angle information correspond to the angle information θ and α described above. The angle information available flag is information indicating whether angle prediction using the angle information is possible. When the angle information available flag is on (angle prediction is possible), the horizontal angle information and the vertical angle information are included in Header1.

[0588] Thereby, prediction values of angle prediction having the same value can be generated by the three-dimensional data encoding device and the three-dimensional data decoding device. Also, by storing these information in SPS, which is control information in units of sequences (multiple frames), they are commonly used for the entire sequence. Note that these information are used for encoding the position information as described above, but may also be used for encoding the attribute information.

[0589] FIG. 106 is a diagram showing an example of the syntax of Header2 included in a bit stream. Here, Header2 is, for example, GPS or APS (Attribute Parameter Set).

[0590] Header2 includes a virtual node permission flag (pt_virtual_node_enable_flag), a direct coding permission flag (pt_direct_code_enable_flag), and a line tracking permission flag (pt_line_follow_enable_flag).

[0591] The virtual node permission flag (pt_virtual_node_enable_flag) is information indicating whether (or not) virtual node prediction can be used (has been used). The direct coding permission flag (pt_direct_code_enable_flag) is information indicating whether (or not) direct coding can be used (has been used). The line tracking permission flag (pt_line_follow_enable_flag) is information indicating whether a prediction tree has been generated using line tracing shown in FIG. 102 or the like.

[0592] Also, when the virtual node permission flag is on (virtual node prediction can be used), Header2 includes a virtual node threshold (pt_virtual_node_threshold). The virtual node threshold indicates the above-described threshold VRth. When the direct coding permission flag is on (direct coding can be used), Header2 includes a direct coding threshold (pt_direct_code_threshold). The direct coding threshold indicates the above-described threshold DCth.

[0593] Note that at least one of pt_virtual_node_enable_flag and pt_direct_code_enable_flag may be stored in the SPS. In this case, the same value is used throughout the sequence. Thereby, the code amount of GPS or APS can be reduced.

[0594] FIG. 107 is a diagram showing a configuration example of a bit stream. As shown in FIG. 107, the bit stream includes an SPS, a GPS, an APS, and slice data (or block data). The slice data (or block data) is encoded data in units of slices (or blocks), and includes encoded data of position information (Geom) and encoded data of attribute information (Attr(0) and Attr(1)).

[0595] FIG. 108 is a diagram showing a syntax example of the encoded data (geometry_data) of the position information. As shown in FIG. 108, the encoded data of the position information includes the number of virtual nodes (num_virtual_node) and a prediction residual (residual_value[j]).

[0596] The number of virtual nodes (num_virtual_node) indicates the number of virtual nodes of the i-th three-dimensional point. The three-dimensional data encoding device and the three-dimensional data decoding device may correct the predicted value calculated in the prediction mode specified by pred_mode using the value of num_virtual_node. Note that when pred_mode = 0 (no prediction), the three-dimensional data encoding device may not add num_virtual_node to the bit stream. Thereby, the amount of code in the case of pred_mode = 0 can be reduced. Further, when num_virtual_node is not added to the bit stream, the three-dimensional data decoding device may estimate its value as 0. Thereby, the three-dimensional data decoding device can appropriately decode the bit stream.

[0597] The three-dimensional data encoding device may add num_virtual_node to the header after entropy encoding. For example, the three-dimensional data encoding device binarizes each value and calculates and encodes it. Further, the three-dimensional data encoding device may use fixed-length encoding in order to suppress the processing amount.

[0598] FIG. 109 is a flowchart of the prediction tree generation process by the three-dimensional data encoding device.

[0599] First, the three-dimensional data encoding device calculates the number of virtual nodes of the target three-dimensional points to be encoded (S10041). Note that when the virtual nodes are not used (pt_virtual_node_enable_flag = 0), this process may be omitted.

[0600] Next, the three-dimensional data encoding device determines the prediction mode (pred_mode) of the target three-dimensional points (S10042). For example, the three-dimensional data encoding device may determine the prediction mode using RDO. Also, after determining the prediction mode, the three-dimensional data encoding device may generate a predicted value using the determined prediction mode and calculate a prediction residual, which is the difference between the original value of the position information or attribute information of the target three-dimensional points and the predicted value.

[0601] Next, the three-dimensional data encoding device adds the target three-dimensional points to the prediction tree (S10043). The three-dimensional data encoding device repeats these series of processes (S10041 to S10043) until all the three-dimensional points are added to the prediction tree (S10044).

[0602] FIG. 110 is a flowchart of the prediction tree decoding process by the three-dimensional data decoding device. First, the three-dimensional data decoding device decodes (acquires) the number of virtual nodes from the bit stream (S10051). Note that when the virtual nodes are not used (pt_virtual_node_enable_flag = 0), the three-dimensional data decoding device may omit this process.

[0603] Next, the three-dimensional data decoding device decodes (acquires) the prediction mode (pred_mode) of the target three-dimensional points to be decoded from the bit stream (S10052). Next, the three-dimensional data decoding device calculates a predicted value using the prediction mode indicated by pred_mode, and adds the predicted value to the decoded prediction residual to decode the position information or attribute information of the three-dimensional points (S10053). The three-dimensional data decoding device repeats these series of processes (S10051 to S10053) until the processing of all three-dimensional points is completed (S10054).

[0604] FIG. 111 is a diagram showing the reference relationship in the prediction process. FIG. 112 is a diagram showing an example of the prediction mode. As shown in FIG. 112, in addition to the angle prediction and virtual node prediction described above, the plurality of prediction modes may include no prediction, linear prediction, parallelogram prediction, and the like.

[0605] As described above, the three-dimensional data encoding device according to the present embodiment performs the processing shown in FIG. 113. First, the three-dimensional data encoding device generates a first predicted value of the position information of the first three-dimensional points included in the point cloud data obtained by the sensor, using the position information of the reference three-dimensional points included in the point cloud data and the scanning angle (for example, θ) of the sensor (S10061). The three-dimensional data encoding device calculates a first difference value between the position information of the first three-dimensional points and the first predicted value (S10062), and generates a bit stream including the first difference value (for example, residual_value[j]) and information indicating the scanning angle (for example, horizontal_angular, θ, or vertical_angular, α) (S10063).

[0606] According to this, the three-dimensional data encoding device can improve the prediction accuracy by generating a predicted value using the scanning angle of the sensor that generated the point cloud data, and thus can improve the encoding efficiency.

[0607] For example, the position information of the first three-dimensional points is represented by xyz coordinates, and the scanning angle is the scanning angle (for example, θ) in the xy plane. For example, the sensor is a radar (for example, LiDAR) that performs rotational scanning for each scanning angle.

[0608] For example, in generating the first predicted value, the three-dimensional data encoding device generates the first predicted value by rotationally moving the position information of the reference three-dimensional point (e.g., the parent node) immediately before the first three-dimensional point in the scanning order, which is included in the point cloud data, by the scanning angle around a reference point corresponding to the position of the sensor.

[0609] For example, when the first difference value is greater than a predetermined threshold value, the three-dimensional data encoding device generates a second predicted value using a value obtained by multiplying a reference predicted value (e.g., a) by n (n is an integer of 1 or more). For example, the three-dimensional data encoding device generates a second predicted value by adding a value obtained by multiplying the reference predicted value by n to the first predicted value (e.g., the position information of the parent node). A second difference value between the position information of the first three-dimensional point and the second predicted value is calculated, and the bit stream includes the second difference value (e.g., residual_value[j]) and information indicating n (e.g., num_virtual_node).

[0610] According to this, the three-dimensional data encoding device can improve the prediction accuracy, and thus can improve the encoding efficiency.

[0611] For example, the reference predicted value corresponds to the difference (e.g., a) in the position information of the two reference three-dimensional points before the first three-dimensional point in the scanning order, which is included in the point cloud data. For example, the reference predicted value is generated using the scanning angle (e.g., θ).

[0612] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor performs the above processing using the memory.

[0613] In addition, the three-dimensional data decoding device according to the present embodiment performs the process shown in FIG. 114. First, the three-dimensional data decoding device obtains, from the bit stream, a first difference value (for example, residual_value[j]) between the position information of the first three-dimensional point included in the point cloud data obtained by the sensor and the first predicted value, and information indicating the scanning angle of the sensor (for example, horizontal_angular, θ, or vertical_angular, α) (S10071). The three-dimensional data decoding device generates the first predicted value using the position information of the reference three-dimensional point included in the point cloud data and the scanning angle of the sensor (S10072), and calculates the position information of the first three-dimensional point by adding the first predicted value and the first difference value (S10073).

[0614] According to this, the three-dimensional data decoding device can improve the prediction accuracy by generating the predicted value using the scanning angle of the sensor that generated the point cloud data, so that the coding efficiency can be improved.

[0615] For example, the position information of the first three-dimensional point is represented by xyz coordinates, and the scanning angle is the scanning angle (for example, θ) in the xy plane. For example, the sensor is a radar (for example, LiDAR) that performs rotational scanning for each scanning angle.

[0616] For example, in generating the first predicted value, the three-dimensional data decoding device generates the first predicted value by rotating and moving the position information of the reference three-dimensional point (for example, the parent node) immediately before the first three-dimensional point in the scanning order included in the point cloud data by the scanning angle around a reference point corresponding to the position of the sensor.

[0617] For example, the three-dimensional data decoding device further obtains, from the bit stream, a second difference value (for example, residual_value[j]) between the position information of the second three-dimensional point included in the point cloud data and the second predicted value, and information indicating n (n is an integer of 1 or more) (for example, num_virtual_node), generates the second predicted value using a value obtained by multiplying the reference predicted value (for example, a) by n, and calculates the position information of the second three-dimensional point by adding the second predicted value and the second difference value.

[0618] According to this, the three-dimensional data decoding device can improve the prediction accuracy, so that the coding efficiency can be improved.

[0619] For example, the reference prediction value corresponds to the difference (e.g., a) in the position information of the two reference three-dimensional points before the first three-dimensional point in the scanning order included in the point cloud data. For example, the reference prediction value is generated using the scanning angle (e.g., θ).

[0620] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor performs the above processing using the memory.

[0621] (Embodiment 7) Next, the configuration of the three-dimensional data creation device 810 according to this embodiment will be described. FIG. 115 is a block diagram showing a configuration example of the three-dimensional data creation device 810 according to this embodiment. This three-dimensional data creation device 810 is mounted on a vehicle, for example. The three-dimensional data creation device 810 performs three-dimensional data transmission and reception with an external traffic monitoring cloud, a preceding vehicle, or a following vehicle, and creates and accumulates three-dimensional data.

[0622] The three-dimensional data creation device 810 includes a data reception unit 811, a communication unit 812, a reception control unit 813, a format conversion unit 814, a plurality of sensors 815, a three-dimensional data creation unit 816, a three-dimensional data synthesis unit 817, a three-dimensional data storage unit 818, a communication unit 819, a transmission control unit 820, a format conversion unit 821, and a data transmission unit 822.

[0623] The data reception unit 811 receives three-dimensional data 831 from the traffic monitoring cloud or the preceding vehicle. The three-dimensional data 831 includes information such as point cloud, visible light video, depth information, sensor position information, or speed information, including areas that cannot be detected by the sensors 815 of the host vehicle, for example.

[0624] The communication unit 812 communicates with the traffic monitoring cloud or the preceding vehicle and transmits a data transmission request or the like to the traffic monitoring cloud or the preceding vehicle.

[0625] The reception control unit 813 exchanges information such as the corresponding format with the communication destination via the communication unit 812 to establish communication with the communication destination.

[0626] The format conversion unit 814 generates the three-dimensional data 832 by performing format conversion or the like on the three-dimensional data 831 received by the data reception unit 811. Further, when the three-dimensional data 831 is compressed or encoded, the format conversion unit 814 performs decompression or decoding processing.

[0627] The plurality of sensors 815 is a group of sensors that acquire information outside the vehicle, such as a LiDAR, a visible light camera, or an infrared camera, and generates sensor information 833. For example, when the sensor 815 is a laser sensor such as a LiDAR, the sensor information 833 is three-dimensional data such as a point cloud (point group data). Note that the number of sensors 815 does not have to be plural.

[0628] The three-dimensional data creation unit 816 generates three-dimensional data 834 from the sensor information 833. The three-dimensional data 834 includes information such as, for example, a point cloud, a visible light video, depth information, sensor position information, or speed information.

[0629] The three-dimensional data synthesis unit 817 synthesizes the three-dimensional data 832 created by the traffic monitoring cloud or the preceding vehicle or the like with the three-dimensional data 834 created based on the sensor information 833 of the host vehicle to construct three-dimensional data 835 including the space in front of the preceding vehicle that cannot be detected by the sensors 815 of the host vehicle.

[0630] The three-dimensional data storage unit 818 stores the generated three-dimensional data 835 and the like.

[0631] The communication unit 819 communicates with the traffic monitoring cloud or the following vehicle and transmits a data transmission request or the like to the traffic monitoring cloud or the following vehicle.

[0632] The transmission control unit 820 exchanges information such as the corresponding format with the communication destination via the communication unit 819 and establishes communication with the communication destination. Further, the transmission control unit 820 determines a transmission area, which is the space of the three-dimensional data to be transmitted, based on the three-dimensional data construction information of the three-dimensional data 832 generated by the three-dimensional data synthesis unit 817 and the data transmission request from the communication destination.

[0633] Specifically, the transmission control unit 820 determines a transmission area including the space in front of the host vehicle that cannot be detected by the sensors of the following vehicle in response to a data transmission request from the traffic monitoring cloud or the following vehicle. Further, the transmission control unit 820 determines the transmission area by judging whether there is an update of the space that can be transmitted or the transmitted space based on the three-dimensional data construction information. For example, the transmission control unit 820 determines as the transmission area the area specified in the data transmission request and in which the corresponding three-dimensional data 835 exists. Then, the transmission control unit 820 notifies the format conversion unit 821 of the format corresponding to the communication destination and the transmission area.

[0634] The format conversion unit 821 generates three-dimensional data 837 by converting the three-dimensional data 836 in the transmission area among the three-dimensional data 835 stored in the three-dimensional data storage unit 818 into the format corresponding to the receiving side. Note that the format conversion unit 821 may reduce the data amount by compressing or encoding the three-dimensional data 837.

[0635] The data transmission unit 822 transmits the three-dimensional data 837 to the traffic monitoring cloud or the following vehicle. This three-dimensional data 837 includes information such as point cloud in front of the host vehicle, visible light video, depth information, or sensor position information, which includes areas that become blind spots for the following vehicle.

[0636] Here, an example in which format conversion and the like are performed by the format conversion units 814 and 821 has been described, but the format conversion may not be performed.

[0637] With such a configuration, the three-dimensional data creation device 810 acquires three-dimensional data 831 of an area that cannot be detected by the sensors 815 of the host vehicle from the outside, and generates three-dimensional data 835 by synthesizing the three-dimensional data 831 and three-dimensional data 834 based on the sensor information 833 detected by the sensors 815 of the host vehicle. Thereby, the three-dimensional data creation device 810 can generate three-dimensional data in a range that cannot be detected by the sensors 815 of the host vehicle.

[0638] In addition, the three-dimensional data creation device 810 can transmit three-dimensional data including the space in front of the host vehicle that cannot be detected by the sensors of the following vehicle to the traffic monitoring cloud or the following vehicle, etc., in response to a data transmission request from the traffic monitoring cloud or the following vehicle.

[0639] Next, the procedure for transmitting three-dimensional data to the following vehicle in the three-dimensional data creation device 810 will be described. FIG. 116 is a flowchart showing an example of the procedure for transmitting three-dimensional data to the traffic monitoring cloud or the following vehicle by the three-dimensional data creation device 810.

[0640] First, the three-dimensional data creation device 810 generates and updates three-dimensional data 835 of a space including the space on the road in front of the host vehicle (S801). Specifically, the three-dimensional data creation device 810 constructs three-dimensional data 835 including the space in front of the preceding vehicle that cannot be detected by the sensors 815 of the host vehicle by synthesizing the three-dimensional data 831 created by the traffic monitoring cloud or the preceding vehicle, etc., with the three-dimensional data 834 created based on the sensor information 833 of the host vehicle.

[0641] Next, the three-dimensional data creation device 810 determines whether the three-dimensional data 835 included in the transmitted space has changed (S802).

[0642] If the three-dimensional data 835 included in the transmitted space has changed due to a vehicle or a person entering from the outside (Yes in S802), the three-dimensional data creation device 810 transmits the three-dimensional data including the three-dimensional data 835 of the changed space to the traffic monitoring cloud or the following vehicle (S803).

[0643] Note that the three-dimensional data creation device 810 may transmit the three-dimensional data of the space in which the change has occurred in accordance with the transmission timing of the three-dimensional data transmitted at predetermined intervals, or may transmit it immediately after detecting the change. That is, the three-dimensional data creation device 810 may transmit the three-dimensional data of the space in which the change has occurred with higher priority than the three-dimensional data transmitted at predetermined intervals.

[0644] In addition, the three-dimensional data creation device 810 may transmit all of the three-dimensional data of the space in which the change has occurred as the three-dimensional data of the space in which the change has occurred, or may transmit only the difference of the three-dimensional data (for example, information on three-dimensional points that have appeared or disappeared, or displacement information of three-dimensional points, etc.).

[0645] In addition, the three-dimensional data creation device 810 may transmit metadata related to the danger avoidance operation of the host vehicle, such as an emergency braking warning, to the following vehicle prior to the three-dimensional data of the space in which the change has occurred. According to this, the following vehicle can recognize the sudden braking of the preceding vehicle earlier and can start a danger avoidance operation such as deceleration earlier.

[0646] If there is no change in the three-dimensional data 835 included in the space that has already been transmitted (No in S802), or after step S803, the three-dimensional data creation device 810 transmits the three-dimensional data included in a space of a predetermined shape at the front distance L of the host vehicle to the traffic monitoring cloud or the following vehicle (S804).

[0647] In addition, for example, the processes of steps S801 to S804 are repeatedly performed at predetermined time intervals.

[0648] In addition, when there is no difference between the three-dimensional data 835 of the currently targeted space and the three-dimensional map, the three-dimensional data creation device 810 may not transmit the three-dimensional data 837 of the space.

[0649] In the present embodiment, the client device transmits sensor information obtained by sensors to a server or another client device.

[0650] First, the configuration of the system according to this embodiment will be described. FIG. 117 is a diagram showing the configuration of a three-dimensional map and sensor information transmission / reception system according to this embodiment. This system includes a server 901 and client devices 902A and 902B. When the client devices 902A and 902B are not particularly distinguished, they are also referred to as the client device 902.

[0651] The client device 902 is, for example, an in-vehicle device mounted on a moving body such as a vehicle. The server 901 is, for example, a traffic monitoring cloud or the like and can communicate with a plurality of client devices 902.

[0652] The server 901 transmits a three-dimensional map composed of point clouds to the client device 902. Note that the configuration of the three-dimensional map is not limited to point clouds and may represent other three-dimensional data such as a mesh structure.

[0653] The client device 902 transmits sensor information acquired by the client device 902 to the server 901. The sensor information includes, for example, at least one of LiDAR acquisition information, visible light image, infrared image, depth image, sensor position information, and speed information.

[0654] The data transmitted and received between the server 901 and the client device 902 may be compressed for data reduction or may remain uncompressed to maintain the accuracy of the data. When compressing the data, for example, a three-dimensional compression method based on an octree structure can be used for the point cloud. Also, a two-dimensional image compression method can be used for visible light images, infrared images, and depth images. The two-dimensional image compression method is, for example, MPEG-4 AVC or HEVC standardized by MPEG.

[0655] In addition, the server 901 transmits the three-dimensional map managed by the server 901 to the client device 902 in response to a transmission request for the three-dimensional map from the client device 902. Note that the server 901 may transmit the three-dimensional map without waiting for a transmission request for the three-dimensional map from the client device 902. For example, the server 901 may broadcast the three-dimensional map to one or more client devices 902 located in a predetermined space. Further, the server 901 may transmit a three-dimensional map suitable for the position of the client device 902 to the client device 902 that has received a transmission request at regular intervals. Also, the server 901 may transmit the three-dimensional map to the client device 902 each time the three-dimensional map managed by the server 901 is updated.

[0656] The client device 902 issues a transmission request for the three-dimensional map to the server 901. For example, when the client device 902 wants to estimate its own position during travel, the client device 902 transmits a transmission request for the three-dimensional map to the server 901.

[0657] Note that in the following cases, the client device 902 may issue a transmission request for the three-dimensional map to the server 901. When the three-dimensional map held by the client device 902 is old, the client device 902 may issue a transmission request for the three-dimensional map to the server 901. For example, when a certain period of time has elapsed since the client device 902 acquired the three-dimensional map, the client device 902 may issue a transmission request for the three-dimensional map to the server 901.

[0658] Before a certain time when the client device 902 exits from the space shown in the three-dimensional map held by the client device 902, the client device 902 may send a transmission request for the three-dimensional map to the server 901. For example, when the client device 902 exists within a predetermined distance from the boundary of the space shown in the three-dimensional map held by the client device 902, the client device 902 may send a transmission request for the three-dimensional map to the server 901. Also, when the movement path and movement speed of the client device 902 can be grasped, based on these, the time when the client device 902 exits from the space shown in the three-dimensional map held by the client device 902 may be predicted.

[0659] When the error at the time of alignment between the three-dimensional data created by the client device 902 from the sensor information and the three-dimensional map is a certain value or more, the client device 902 may send a transmission request for the three-dimensional map to the server 901.

[0660] The client device 902 transmits the sensor information to the server 901 in response to the transmission request for the sensor information sent from the server 901. Note that the client device 902 may send the sensor information to the server 901 without waiting for the transmission request for the sensor information from the server 901. For example, when the client device 902 once obtains a transmission request for the sensor information from the server 901, the client device 902 may periodically transmit the sensor information to the server 901 for a certain period. Also, when the error at the time of alignment between the three-dimensional data created by the client device 902 based on the sensor information and the three-dimensional map obtained from the server 901 is a certain value or more, and it is determined that there may be a change in the three-dimensional map around the client device 902, the client device 902 may send the fact and the sensor information to the server 901.

[0661] Server 901 sends a request to the client device 902 to transmit sensor information. For example, server 901 receives the location information of client device 902 such as GPS from client device 902. If server 901 determines based on the location information of client device 902 that client device 902 is approaching a space with less information in the three-dimensional map managed by server 901, server 901 sends a request to client device 902 to transmit sensor information in order to generate a new three-dimensional map. In addition, when server 901 wants to update the three-dimensional map, when it wants to check the road conditions such as during snow accumulation or disasters, when it wants to check the traffic congestion situation, or the accident situation, etc., server 901 may send a request to transmit sensor information.

[0662] Also, client device 902 may set the amount of sensor information data to be transmitted to server 901 according to the communication state or bandwidth at the time of receiving the request to transmit sensor information received from server 901. Setting the amount of sensor information data to be transmitted to server 901 means, for example, increasing or decreasing the data itself, or appropriately selecting a compression method.

[0663] FIG. 118 is a block diagram showing a configuration example of client device 902. Client device 902 receives a three-dimensional map composed of point clouds, etc. from server 901, and estimates its own position from the three-dimensional data created based on the sensor information of client device 902. In addition, client device 902 transmits the acquired sensor information to server 901.

[0664] Client device 902 includes a data reception unit 1011, a communication unit 1012, a reception control unit 1013, a format conversion unit 1014, a plurality of sensors 1015, a three-dimensional data creation unit 1016, a three-dimensional image processing unit 1017, a three-dimensional data storage unit 1018, a format conversion unit 1019, a communication unit 1020, a transmission control unit 1021, and a data transmission unit 1022.

[0665] The data reception unit 1011 receives the three-dimensional map 1031 from the server 901. The three-dimensional map 1031 is data including point clouds such as WLD or SWLD. The three-dimensional map 1031 may include either compressed data or uncompressed data.

[0666] The communication unit 1012 communicates with the server 901 and transmits a data transmission request (for example, a request to transmit a three-dimensional map) to the server 901.

[0667] The reception control unit 1013 exchanges information such as a corresponding format with the communication destination via the communication unit 1012 and establishes communication with the communication destination.

[0668] The format conversion unit 1014 generates a three-dimensional map 1032 by performing format conversion or the like on the three-dimensional map 1031 received by the data reception unit 1011. Further, when the three-dimensional map 1031 is compressed or encoded, the format conversion unit 1014 performs decompression or decoding processing. Note that the format conversion unit 1014 does not perform decompression or decoding processing if the three-dimensional map 1031 is uncompressed data.

[0669] The plurality of sensors 1015 is a group of sensors that acquire information outside the vehicle in which the client device 902 is mounted, such as a LiDAR, a visible light camera, an infrared camera, or a depth sensor, and generates sensor information 1033. For example, when the sensor 1015 is a laser sensor such as a LiDAR, the sensor information 1033 is three-dimensional data such as a point cloud (point group data). Note that the number of sensors 1015 does not have to be plural.

[0670] The three-dimensional data creation unit 1016 creates three-dimensional data 1034 around the host vehicle based on the sensor information 1033. For example, the three-dimensional data creation unit 1016 creates point cloud data with color information around the host vehicle using the information acquired by the LiDAR and the visible light video obtained by the visible light camera.

[0671] The three-dimensional image processing unit 1017 performs self-position estimation processing of the host vehicle using the received three-dimensional map 1032 such as point cloud and the three-dimensional data 1034 of the surroundings of the host vehicle generated from the sensor information 1033. Note that the three-dimensional image processing unit 1017 may create the three-dimensional data 1035 of the surroundings of the host vehicle by synthesizing the three-dimensional map 1032 and the three-dimensional data 1034, and perform self-position estimation processing using the created three-dimensional data 1035.

[0672] The three-dimensional data storage unit 1018 stores the three-dimensional map 1032, the three-dimensional data 1034, the three-dimensional data 1035, etc.

[0673] The format conversion unit 1019 generates the sensor information 1037 by converting the sensor information 1033 into a format supported by the receiving side. Note that the format conversion unit 1019 may reduce the data amount by compressing or encoding the sensor information 1037. Also, the format conversion unit 1019 may omit the process when format conversion is not necessary. Further, the format conversion unit 1019 may control the data amount to be transmitted according to the specification of the transmission range.

[0674] The communication unit 1020 communicates with the server 901 and receives a data transmission request (a transmission request for sensor information) etc. from the server 901.

[0675] The transmission control unit 1021 exchanges information such as the corresponding format with the communication destination via the communication unit 1020 and establishes communication.

[0676] The data transmission unit 1022 transmits the sensor information 1037 to the server 901. The sensor information 1037 includes information acquired by a plurality of sensors 1015 such as information acquired by LiDAR, a luminance image acquired by a visible light camera, an infrared image acquired by an infrared camera, a depth image acquired by a depth sensor, sensor position information, and speed information.

[0677] Next, the configuration of the server 901 will be described. FIG. 119 is a block diagram showing a configuration example of the server 901. The server 901 receives sensor information transmitted from the client device 902, and creates three-dimensional data based on the received sensor information. The server 901 updates the three-dimensional map managed by the server 901 using the created three-dimensional data. Further, the server 901 transmits the updated three-dimensional map to the client device 902 in response to a transmission request for the three-dimensional map from the client device 902.

[0678] The server 901 includes a data reception unit 1111, a communication unit 1112, a reception control unit 1113, a format conversion unit 1114, a three-dimensional data creation unit 1116, a three-dimensional data synthesis unit 1117, a three-dimensional data storage unit 1118, a format conversion unit 1119, a communication unit 1120, a transmission control unit 1121, and a data transmission unit 1122.

[0679] The data reception unit 1111 receives sensor information 1037 from the client device 902. The sensor information 1037 includes, for example, information acquired by LiDAR, a luminance image acquired by a visible light camera, an infrared image acquired by an infrared camera, a depth image acquired by a depth sensor, sensor position information, and speed information.

[0680] The communication unit 1112 communicates with the client device 902 and transmits a data transmission request (for example, a transmission request for sensor information) to the client device 902.

[0681] The reception control unit 1113 exchanges information such as the corresponding format with the communication destination via the communication unit 1112 and establishes communication.

[0682] When the received sensor information 1037 is compressed or encoded, the format conversion unit 1114 generates sensor information 1132 by performing decompression or decoding processing. Note that if the sensor information 1037 is uncompressed data, the format conversion unit 1114 does not perform decompression or decoding processing.

[0683] The three-dimensional data creation unit 1116 creates three-dimensional data 1134 around the client device 902 based on the sensor information 1132. For example, the three-dimensional data creation unit 1116 creates point cloud data with color information around the client device 902 using the information obtained by LiDAR and the visible light video obtained by the visible light camera.

[0684] The three-dimensional data synthesis unit 1117 updates the three-dimensional map 1135 by synthesizing the three-dimensional data 1134 created based on the sensor information 1132 into the three-dimensional map 1135 managed by the server 901.

[0685] The three-dimensional data storage unit 1118 stores the three-dimensional map 1135 and the like.

[0686] The format conversion unit 1119 generates the three-dimensional map 1031 by converting the three-dimensional map 1135 into a format supported by the receiving side. Note that the format conversion unit 1119 may reduce the data volume by compressing or encoding the three-dimensional map 1135. Also, if format conversion is not necessary, the format conversion unit 1119 may omit the process. Further, the format conversion unit 1119 may control the amount of data to be transmitted according to the specified transmission range.

[0687] The communication unit 1120 communicates with the client device 902 and receives a data transmission request (a request to transmit a three-dimensional map) and the like from the client device 902.

[0688] The transmission control unit 1121 exchanges information such as the corresponding format with the communication destination via the communication unit ...

Claims

1. A processor calculates a first predicted value of the position information of a first three-dimensional point included in point cloud data using the position information and angle information of a reference three-dimensional point included in the point cloud data. The processor calculates a first difference value between the position information of the first three-dimensional point and the first predicted value. The processor calculates a second predicted value of the position information of a second three-dimensional point included in the point cloud data using the position information of the first three-dimensional point and the angle information. The processor calculates a second difference value between the position information of the second three-dimensional point and the second predicted value. A three-dimensional data encoding method.

2. The position information of the first three-dimensional point is represented in xyz coordinates, and the angle information includes an angle in the xy plane. The three-dimensional data encoding method according to Claim 1.

3. The angle information includes a scanning angle of a sensor that acquires the point cloud data. The sensor is a radar that performs rotational scanning for each scanning angle. The three-dimensional data encoding method according to Claim 1 or 2.

4. The angle information includes a scanning angle of a sensor that acquires the point cloud data. In the calculation of the first predicted value, the processor calculates the first predicted value by rotationally moving the position information of the reference three-dimensional point immediately before the first three-dimensional point in the scanning order included in the point cloud data by the scanning angle around a reference point corresponding to the position of the sensor. The three-dimensional data encoding method according to Claim 1 or 2.

5. The position information of the first three-dimensional point is represented in polar coordinates. The three-dimensional data encoding method according to Claim 1.

6. The angle information includes information regarding an angle between the first three-dimensional point and the reference three-dimensional point. The three-dimensional data encoding method according to Claim 1 or 2.

7. The angle information includes information regarding angles between a plurality of three-dimensional points included in the point cloud data. The three-dimensional data encoding method according to Claim 1 or 2.

8. A processor calculates a first predicted value using the position information and angle information of a reference three-dimensional point included in point cloud data. The processor calculates the position information of the first three-dimensional point included in the point cloud data by adding the first predicted value and the first difference value. The processor calculates a second predicted value using the position information of the first three-dimensional point and the angle information. The processor calculates the position information of the second three-dimensional point included in the point cloud data by adding the second predicted value and the second difference value. Three-dimensional data decoding method.

9. The position information of the first three-dimensional point is represented by xyz coordinates, and the angle information includes an angle in the xy plane The three-dimensional data decoding method according to claim 8.

10. The angle information includes a scanning angle of a sensor that acquires the point cloud data, The sensor is a radar that performs rotational scanning for each scanning angle The three-dimensional data decoding method according to claim 8 or 9.

11. The angle information includes a scanning angle of a sensor that acquires the point cloud data, In the calculation of the first predicted value, the processor rotates and moves the position information of the reference three-dimensional point immediately before the first three-dimensional point in the scanning order included in the point cloud data by the scanning angle around a reference point corresponding to the position of the sensor to calculate the first predicted value The three-dimensional data decoding method according to claim 8 or 9.

12. The position information of the first three-dimensional point is represented by polar coordinates The three-dimensional data decoding method according to claim 8.

13. The angle information includes information regarding an angle between the first three-dimensional point and the reference three-dimensional point The three-dimensional data decoding method according to claim 8 or 9.

14. The angle information includes information regarding an angle between a plurality of three-dimensional points included in the point cloud data The three-dimensional data decoding method according to claim 8 or 9.

15. A processor and, A memory, The processor uses the memory to, Calculate a first predicted value of the position information of a first three-dimensional point included in point cloud data using the position information and angle information of a reference three-dimensional point included in the point cloud data, Calculate a first difference value between the position information of the first three-dimensional point and the first predicted value, Calculate a second predicted value of the position information of a second three-dimensional point included in the point cloud data using the position information of the first three-dimensional point and the angle information, Calculate a second difference value between the position information of the second three-dimensional point and the second predicted value, Three-dimensional data encoding device.

16. A processor and, A memory, The processor uses the memory to, Calculate a first predicted value using the position information and angle information of a reference three-dimensional point included in point cloud data, Calculate the position information of the first three-dimensional point included in the point cloud data by adding the first predicted value and the first difference value, Calculate a second predicted value using the position information of the first three-dimensional point and the angle information, Calculate the position information of the second three-dimensional point included in the point cloud data by adding the second predicted value and the second difference value Three-dimensional data decoding device.

Citation Information

Patent Citations

  • Three-dimensional data coding method, three-dimensional data decoding method, three-dimensional data coding device, and three-dimensional data decoding device

    WO2019198636A1

  • Map display device

    WO2014020663A1